Saltar al contenido
Wall·IA

Últimas en IA

Muro en vivo de noticias de inteligencia artificial.

2070 noticias
ResumenEN→ES

Emergent Sparsity in Frozen Random CNN Feature Extractors for Deep Reinforcement Learning

arXiv:2607.26059v1 Announce Type: new Abstract: We report a striking phenomenon: deep reinforcement learning agents trained with frozen, randomly initialized CNN feature extractors spontaneously develop extremely sparse fully-connected representations, without any…

arXiv cs.LG1 min
Investigación & PapersProductos & HerramientasCultura & Opinión
ResumenEN→ES

Sim2Win: A Team-Agnostic, Event-Based Pre-Match Outcome Prediction and Tactical Profiling System for Football

arXiv:2607.26061v1 Announce Type: new Abstract: Pre-match tactical decision-making in professional football relies heavily on subjective expert analysis and identity-based scouting systems that cannot generalize to unseen teams. This paper presents Sim2Win, a team-agnostic,…

arXiv cs.LG1 min
Investigación & PapersModelos & LanzamientosProductos & Herramientas
ResumenEN→ES

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback

arXiv:2607.26094v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard approach for aligning large language models with human preferences, but its quality is limited by static, task-agnostic reward models. This mismatch leads to sparse…

arXiv cs.LG1 min
Investigación & PapersModelos & LanzamientosProductos & Herramientas
ResumenEN→ES

Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation

arXiv:2607.26164v1 Announce Type: new Abstract: Automated molecular structure elucidation from infrared (IR) spectroscopy data has seen significant advancements in recent years, but its broad applicability is limited by a reliance on pre-determined chemical formulas provided as…

arXiv cs.LG1 min
Investigación & PapersModelos & LanzamientosProductos & Herramientas
ResumenEN→ES

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

arXiv:2607.26173v1 Announce Type: new Abstract: Alignment training, model organisms, and toy models are usually treated as separate research areas. But projects in all three frequently use supervised fine-tuning (SFT) to pursue the same underlying goals. When projects share a…

arXiv cs.LG1 min
Investigación & PapersModelos & LanzamientosPolítica & Regulación
ResumenEN→ES

Dynamic Parameterization Is Not Dynamic Inference

arXiv:2607.26192v1 Announce Type: new Abstract: Input-dependent controller coefficients are often treated as evidence of dynamic inference or computational savings. This interpretation conflates three properties: coefficient variation, dependence of a frozen model on how…

arXiv cs.LG1 min
Investigación & PapersModelos & LanzamientosProductos & Herramientas
ResumenEN→ES

Weak-to-Strong On-Policy Distillation

arXiv:2607.26246v1 Announce Type: new Abstract: On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on the student's own rollouts, is an effective paradigm for transferring capabilities across LLMs. Prevailing approaches assume a…

arXiv cs.LG1 min
Investigación & PapersModelos & LanzamientosProductos & Herramientas
ResumenEN→ES

Between Gradient and Natural Gradient: A Continuum of LoRA Initializations

arXiv:2607.26247v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) fine-tunes large pretrained models at a fraction of the cost of full fine-tuning, but its performance depends strongly on how the adapters are initialized. Recent schemes initialize the adapters from the…

arXiv cs.LG1 min
Investigación & PapersModelos & LanzamientosProductos & Herramientas
ResumenEN→ES

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR

arXiv:2607.26253v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by rollout generation, yet many sampled prompts produce saturated groups (all responses correct or all incorrect) whose zero reward variance yields no…

arXiv cs.LG1 min
Investigación & PapersModelos & LanzamientosProductos & Herramientas
ResumenEN→ES

Top-$k$ Pareto Bandits: Hypervolume Regret for Multi-Objective Slate Selection

arXiv:2607.26273v1 Announce Type: new Abstract: We consider a stochastic multi-objective bandit problem where, at each round, the agent selects a slate of $k$ arms and observes their $d$-dimensional reward vectors under semi-bandit feedback. We do not aim at identifying a…

arXiv cs.LG1 min
Investigación & PapersProductos & HerramientasPolítica & Regulación
ResumenEN→ES

FloDR: An invertible dimensionality reduction method based on a normalising flow

arXiv:2607.26278v1 Announce Type: new Abstract: It is common for two-dimensional embeddings of high-dimensional data to be read far beyond what they can support. Distances in and between clusters, the meaning behind empty spaces, and the amount of structure hidden at each point…

arXiv cs.LG1 min
Investigación & PapersModelos & LanzamientosProductos & Herramientas
ResumenEN→ES

Entity Resolution in Practice: Lessons from a Self-Serve Pipeline

arXiv:2607.26298v1 Announce Type: new Abstract: We built and evaluated a self-serve entity resolution (ER) system on six benchmarks spanning 864 to 5M records, and three lessons emerged that are absent from existing ER literature. (1) No single matching algorithm wins…

arXiv cs.LG1 min
Investigación & PapersModelos & Lanzamientos