Our team at Tongyi Lab is dedicated to pioneer advancements in AI search technologies.
43
Public repositories
25,446
Total stars
1,656
Followers
Alibaba-NLP, part of Tongyi Lab at Alibaba Group, is actively contributing to the open-source community on GitHub. The organization focuses on AI search technologies, with primary repositories developed in Python, including notable projects like DeepResearch and ZeroSearch, which address advanced research and search capabilities in AI.
Tongyi Deep Research, the Leading Open-source Deep Research Agent
ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Multimodal Retrieval-augmented Generation Framework Built by Tongyi Lab, Alibaba Group.
[EMNLP 2025] ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
Repo for Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent
[ACL-IJCNLP 2021] Automated Concatenation of Embeddings for Structured Prediction
Repo for NAACL 2025 Paper "Unfolding the Headline: Iterative Self-Questioning for News Retrieval and Timeline Summarization"
An Instruction-tuned Large Language Model for E-commerce
qqr is an RL training framework for open-ended agents.
Hierarchy-Aware Global Model for Hierarchical Text Classification
SeqGPT: An Out-of-the-box Large Language Model for Open Domain Sequence Understanding
[SIGIR 2022] Multi-CPR: A Multi Domain Chinese Dataset for Passage Retrieval
Winner system (DAMO-NLP) of SemEval 2022 MultiCoNER shared task over 10 out of 13 tracks.
Repo for "MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability"
[ACL-IJCNLP 2021] Improving Named Entity Recognition by External Context Retrieving and Cooperative Learning
[ACL 2020] Structure-Level Knowledge Distillation For Multilingual Sequence Labeling
E2Rank: Your Text Embedding can Also be an Effective and Efficient Listwise Reranker
The code for LaRA Benchmark
No description provided for this repository.
code for paper 《RankingGPT: Empowering Large Language Models in Text Ranking with Progressive Enhancement》
Code for 'Prototypical Representation Learning for Relation Extraction'.
[EMNLP 2021] MuVER: Improving First-Stage Entity Retrieval with Multi-View Entity Representations
[ICASSP 2022] AISHELL-NER: Named Entity Recognition from Chinese Speech
Coupling Distant Annotation and Adversarial Training for Cross-Domain Chinese Word Segmentation
[ACL 2023] MANNER: A Variational Memory-Augmented Model for Cross Domain Few-Shot Named Entity Recognition
Hybrid List Aware Transformer Reranking
Code for our EMNLP 2020 Paper "AIN: Fast and Accurate Sequence Labeling with Approximate Inference Network"
CDQA: Chinese Dynamic Question Answering Benchmark
Codes for the EMNLP'2020 paper "Predicting Clinical Trial Results by Implicit Evidence Integration".
[ACL-IJCNLP 2021] Structural Knowledge Distillation: Tractably Distilling Information for Structured Predictor
A new evaluation paradigm for deep search that identifies specific LLM failure sources, introduces challenging hint-free datasets with holistic evaluation, and offers a strong baseline incorporating memory and verification.
Source code of paper Improving "Retrieval Augmented Open-Domain Question-Answering with Vectorized Contexts
This is the official repository for the IBKD knowledge distillation method, as described in the paper .
No description provided for this repository.
[EMNLP 2025] Code for "Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based Inference"
No description provided for this repository.
Implementation of NeurIPS 20 paper: Latent Template Induction with Gumbel-CRFs
Implementation of AAAI 21 paper: Nested Named Entity Recognition with Partially Observed TreeCRFs
No description provided for this repository.
[ACL 2022 Findings] Fusing Heterogeneous Factors with Triaffine Mechanism for Nested Named Entity Recognition
[ACL 2022] Code Synonyms Do Matter: Multiple Synonyms Matching Network for Automatic ICD Coding
Codes and data for Alibaba's winning systems at the TREC Precision Medicine Track 2020.
Implementation of ICLR 21 paper: Probing BERT in Hyperbolic Spaces
Alibaba-NLP builds various tools and frameworks focused on AI search technologies. Key repositories include DeepResearch, which is an open-source deep research agent, and ZeroSearch, aimed at enhancing the search capabilities of large language models.
Alibaba-NLP primarily uses Python for its development work. This language is prevalent across their public repositories, allowing for efficient implementation of their AI-driven projects and frameworks.
Yes, Alibaba-NLP's repositories are public on GitHub. This openness allows collaboration and engagement with the broader development community, fostering advancements in AI search technologies and other related fields.
Monitor Tongyi Lab, Alibaba Group with RepoGuard and get alerted the moment a new public repository appears.
Monitor this account