Implicit Regularization of SGD Reduces Shortcut Learning
Shows how the implicit regularization of SGD, controlled by learning rate and batch size, suppresses reliance on spurious features while preserving accuracy.
Paper2025–2026 · Recent work
Publications and preprints co-authored by Dr. Mahdieh Soleymani and at least one additional researcher listed on our People page.
8 worksRobust learning under distribution shifts, spurious correlations, and increasing reasoning complexity.
Shows how the implicit regularization of SGD, controlled by learning rate and batch size, suppresses reliance on spurious features while preserving accuracy.
PaperFrames reasoning as generalization to problems whose solutions require greater representational or computational complexity than the training examples.
PaperRefines class prototypes with explicit awareness of spurious features to make out-of-distribution detection more reliable under shortcut-driven shifts.
PaperStudies compositional inductive biases that help language-informed reinforcement learning agents generalize systematically to unseen task combinations.
PaperUses cross-modal auxiliary objectives and instruction tracking to improve sample efficiency and systematic generalization in language-guided reinforcement learning.
PaperArgues that System-2 behavior is best understood as generalization beyond the computational complexity represented in training examples.
PaperStudies how the implicit bias of optimization can improve worst-group performance even when group annotations are unavailable.
PaperUses information already encoded by trained models to reduce shortcut reliance without requiring group labels during training.
Paper9 worksImproving object coverage, spatial alignment, and compositional control in diffusion models.
Improves compositional text-to-image generation by optimizing and exploring initial noise with category-aware rewards at inference time.
PaperConnects missing entities in generated images to overlapping attention maps and develops an inference-time strategy for better object coverage.
PaperTakes a probabilistic view of spatial relationships to improve how text-to-image models place and align objects described in prompts.
PaperOptimizes generation at test time to preserve every requested concept and improve compositional prompt fidelity.
PaperCompares compositional alignment in visual autoregressive and diffusion models to expose their distinct strengths and failure modes.
PreprintSystematically evaluates whether popular metrics reliably measure object, attribute, and relationship alignment in generated images.
PreprintIntroduces a fine-grained metric for measuring whether generated images satisfy the individual components of complex prompts.
PaperRefines alignment signals and initial noise to improve compositional correctness without retraining the text-to-image model.
PreprintBenchmarks diffusion and autoregressive image generators and analyzes their relative ability to follow compositional prompts.
Paper7 worksMultilingual evaluation, multi-object representations, and structured visual reasoning.
Introduces a large bilingual Persian–English benchmark for evaluating vision-language models across educational, scientific, and cultural reasoning tasks.
PaperIntroduces VISER, which adds lightweight spatial structure and sequential-scanning prompts to improve counting, visual search, and spatial reasoning.
PaperProvides a controlled analysis of how CLIP represents multiple objects and reveals limitations that aggregate benchmarks can conceal.
PaperUses Bayesian probing to detect object hallucinations and guide mitigation in vision-language model responses.
PaperIntroduces a rebus-puzzle benchmark that tests whether multimodal models can integrate visual form, language, and abstract concepts.
PreprintStudies how external cues create grounding identifiers that help multimodal models bind objects to their properties and locations.
PreprintUses controlled high-resolution data to isolate CLIP's limitations when images contain multiple objects and attributes.
Paper4 worksMechanistic analysis and faithful attribution methods for language and vision transformers.
Reveals how a System-2 strategy decomposes large counting problems, transfers partial counts through dedicated attention heads, and aggregates them reliably.
PaperBalances gradient flow to produce more stable, faithful, and broadly effective attribution maps for vision transformer architectures.
PaperDevelops activation-patching methods for tracing causal contributions inside neural networks rather than relying on correlational explanations.
PreprintCompares internal counting strategies across language and vision-language models to reveal shared mechanisms and modality-specific failures.
Preprint7 worksReasoning confidence, hidden shortcuts, evaluation, and language models for scientific discovery.
Surveys the use of large language models for scientific idea generation through a creativity-centered framework and evaluation perspective.
PaperUses model confidence to improve reasoning reliability and help language models make better decisions during multi-step inference.
PaperEvaluates strategic deception through interactive games and probes whether models represent truthful and deceptive alternatives internally.
PreprintShows that language-model judges can exploit hidden shortcuts while producing explanations that fail to acknowledge those cues.
PreprintTests long-horizon finite-state-machine execution to separate genuine procedural reasoning from shallow pattern matching.
PreprintExamines cases where implicit hints improve answers even though model explanations do not acknowledge using them.
PreprintDemonstrates that LLM evaluators can rely on shortcut signals while presenting apparently independent justifications for their scores.
Preprint5 worksCollaborative research in reinforcement learning, adversarial bandits, emergent communication, and medical image segmentation.
Factorizes environment states and skill variables to learn richer, more diverse, and compositionally reusable behaviors without external rewards.
PaperDemonstrates how small targeted perturbations to reward models can hijack offline bandits and studies a practical partial defense.
PaperUses confidence-aware query labeling at test time to improve the accuracy and adaptability of three-dimensional few-shot segmentation.
PaperLearns three-dimensional medical image segmentation without dense labels by propagating masks through flow-guided self-supervision.
PaperUses vector quantization to help interacting agents develop a discrete symbolic language in emergent communication games.
PreprintComplete research record
For the complete research record, including earlier work, preprints, and citation information, visit Dr. Mahdieh Soleymani's Google Scholar profile.