Yifei Zhang

Publications

* indicates equal contribution

Discover more on Google Scholar

All CS Finance

2026

XYZ Agentic Team
XYZ AI LAB Technical Report, 2026 πŸš€ Highest reported scores on BrowseComp-ZH, LiveBrowseComp, and WideSearch among listed systems
We present AI4AI, a bounded-exploration, verification-gated framework for improving complete agentic systems, instantiated in XYZ-Aquila, a family of open-weight Deep Search agents. XYZ-Aquila-mini attains the highest score in every column of the seven-benchmark sub-40B open-weight table, while XYZ-Aquila-pro leads the sub-400B table.
AI4AI at Scale
Yifei Zhang*, Xu Yang*, Xiao Yang*, Bowen Xian, Qizheng Li, Shikai Fang, Jingyuan Li, Jian Wang, Mingrui Xu, Weiqing Liu, Jiang Bian
Findings of ACL 2026 πŸ“ˆ Gradient-based optimization scales better than tree search with stronger reasoning models
We introduce Gome, the MLE setting of RD-Agent, mapping diagnostic reasoning to gradient computation, success memory to momentum, and multi-trace execution to distributed optimization. Under a closed-world protocol, it reaches 35.1% any-medal rate on MLE-Bench with a 12-hour single-V100 budget, and scaling experiments show that gradient-based optimization increasingly outperforms gradient-free tree search as model reasoning capabilities improve.
Gome
Qizheng Li*, Yifei Zhang*, Xiao Yang*, Xu Yang*, Zhuo Wang, Weiqing Liu, Jiang Bian
ICML 2026 πŸ€– Benchmarking autonomous end-to-end LLM fine-tuning
We introduce FT-Dojo, the fine-tuning setting of RD-Agent, with 13 tasks across 5 domains for autonomous LLM fine-tuning, together with FT-Agent, a purpose-built system that achieves the best performance on 10 of 13 tasks and iteratively improves strategies from evaluation history.
FT-Dojo

2025

Xu Yang, Xiao Yang, Shikai Fang, Yifei Zhang, Jian Wang, et al.
arXiv preprint 2025 πŸ… Top-performing open-source framework on MLE-Bench
A top-performing open-source framework on MLE-Bench that organizes autonomous data science into two parts: Research for proposing ideas and Development for converting them into executable workflows and solutions.
R&D-Agent
Tianyu Fan, Yuhao Yang, Yangqin Jiang, Yifei Zhang, Yuxuan Chen, Chao Huang, et al.
arXiv preprint 2025 πŸš€ Fully-automated and data-uncontaminated benchmark for LLM agents in finance
A comprehensive benchmark for evaluating autonomous trading agents in real-time financial markets across NASDAQ 100, SSE 50, and cryptocurrency markets, enabling fair comparison of AI trading strategies with zero human intervention.
AI-Trader
Yuzhe Yang*, Yifei Zhang*, Minghao Wu*, Kaidi Zhang, Yunmiao Zhang, et al.
NeurIPS 2025 πŸ† Best Paper Award, ICLR 2025 Workshop on Advances in Financial AI (1/53)
A multi-agent simulation platform for financial markets that integrates behavioral economics and social network dynamics, enabling large-scale market experiments with diverse agent behaviors.
TwinMarket
Yuzhe Yang*, Yifei Zhang*, Yan Hu*, Yilin Guo, Ruoli Gan, et al.
Findings of NAACL 2025 πŸ€— #1 Paper of the day on Hugging Face
A comprehensive benchmark evaluating LLMs' financial expertise from a user-centric perspective, covering domain knowledge, application capabilities, and trustworthiness across 300+ real-world questions.
UCFE
Separating Skill from Luck: LLM-Based Belief Extraction and Investment Ability Measurement of Financial Influencers
Honghai Yu, Yunmiao Zhang, Haining Wang, Yifei Zhang
Chengdu Four-University Joint Finance Forum & The 7th Finance PhD Academic Forum, 2025 πŸ… Outstanding Paper Award (2/20)
Analyzes 10M+ posts from 60,000+ financial influencers on Xueqiu using LLMs to extract stock expectations, combined with finite mixture models to separate skill from luck. Findings reveal only 46.63% possess positive investment ability, while following top 10% yields 22.4 bps daily returns.
Abstract: This study leverages Large Language Models to extract stock expectation beliefs from millions of posts by 60,000+ financial influencers on Xueqiu, and employs finite mixture models with EM algorithm to filter out "luck noise" and scientifically measure actual investment abilities. Results show only 46.63% of influencers possess positive investment ability, with merely 2.22% achieving strong weekly returns of 1.25%. A long-short portfolio following top 10% influencers yields 22.4 bps daily returns and 4.94x cumulative returns from 2016-2025. Retail investors struggle to identify skilled influencers, while institutional investors respond more to high-ability KOLs. The methodology provides regulators and platforms with tools to identify and monitor financial social media content.
Separating Skill from Luck
Shared Fortunes and Risks: Stock Price Spillover Effects of Corporate Ties β€” Evidence from the Chinese Social Media Platform "Xueqiu"
Honghai Yu, Yunmiao Zhang, Zhuo Chen, Chang Zeng, Yifei Zhang
The 7th Conference of the Chinese Society of Optimization, Overall Planning, and Economic Mathematics, 2025 πŸŽ–οΈ Outstanding Paper Award (10/190)
Investigates stock price spillover effects through corporate network ties on social media, revealing how social connections influence market dynamics and investor behavior.
Abstract: This paper investigates the stock price spillover effects arising from corporate network ties on the Chinese social media platform "Xueqiu". By constructing a comprehensive dataset of corporate relationships and stock price movements, we analyze how information and sentiment propagate through social networks and impact stock valuations. Our findings reveal significant spillover effects where connected firms experience correlated price movements, particularly during periods of high market volatility. We identify key mechanisms through which social media amplifies these effects, including information cascades, herding behavior, and investor attention. The results provide important insights into the role of social media in modern financial markets and have implications for portfolio management, risk assessment, and regulatory policy.
Shared Fortunes

2024

Jimin Huang, Mengxi Xiao, Dong Li, Zihao Jiang, Yuzhe Yang, Yifei Zhang, et al.
arXiv preprint 2024 🎯 First open-source financial multimodal LLM: FinLLaVA-8B
A series of Financial LLMs including FinLLaMA (pre-trained on 52B tokens), FinLLaMA-instruct (573K instructions), and FinLLaVA (first open-source financial multimodal LLM) trained with 1.43M image-text instructions.
Open-FinLLMs
Do investors' actions speak louder than words?
Honghai Yu, Zhuo Chen, Yunmiao Zhang, Haining Wang, Yifei Zhang
The 21st Annual Conference on Financial Engineering and Risk Management, 2024 πŸ“Š Distinguishing noise from information through trading behavior
Examines whether posts on Chinese social media propagate noise or information, proposing that both coexist but can be distinguished by posters' trading behavior. Observing trading actions helps assess the reliability of expressed views.
Abstract: A large body of literature has examined whether posts on social media propagate noise or information. In this paper, we propose that both coexist on Chinese social media platforms but can be distinguished by posters' trading behavior. Individuals may post articles on social media that do not reflect their true opinions, often for impression management purposes, resulting in inconsistency between their words and subsequent actions. Additionally, observing a poster's trading behavior prior to posting can help assess the reliability of their expressed views.
Do Investors Speak Louder