1,442
Public repositories
492,319
Total stars
37,432
Followers
The facebookresearch organization on GitHub hosts a wide range of public repositories focused on advanced machine learning and AI research. Notable projects include segment-anything, detectron2, and fairseq, employing languages such as Python, Jupyter Notebook, and C++. This diverse array of tools contributes significantly to the research community.
The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
A library for efficient similarity search and clustering of dense vectors.
Detectron2 is a platform for object detection, segmentation and other visual recognition tasks.
Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
Library for fast text representation and classification.
FAIR's research platform for object detection research, implementing popular algorithms like Mask R-CNN and RetinaNet.
Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
End-to-End Object Detection with Transformers
[CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
PyTorch code and models for the DINOv2 self-supervised learning method.
The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
Reference PyTorch implementation and models for DINOv3
Hackable and optimized Transformers building blocks, supporting a composable construction.
Implementation of Nougat Neural Optical Understanding for Academic Documents
Fast, modular reference implementation of Instance Segmentation and Object Detection algorithms in PyTorch.
ImageBind One Embedding Space to Bind Them All
PySlowFast: video understanding codebase from FAIR for reproducing state-of-the-art video models.
SAM 3D Objects
Kats, a kit to analyze time series data, a lightweight, easy-to-use, generalizable, and extendable framework to perform time series analysis, from understanding the key statistics and characteristics, detecting change points and anomalies, to forecasting future trends.
High-resolution models for human tasks.
CoTracker is a model for tracking any point (pixel) on a video.
A PyTorch library for implementing flow matching algorithms, featuring continuous and discrete flow matching implementations. It includes practical examples for both text and image modalities.
PyTorch code and models for VJEPA2 self-supervised learning from video.
[CVPR 2026 Oral] VGGT Omega
Evolutionary Scale Modeling (esm): Pretrained language models for proteins
PyTorch code and models for V-JEPA self-supervised learning from video.
State-of-the-art deep learning based audio codec supporting both mono 24 kHz audio and stereo 48 kHz audio.
A flexible, high-performance 3D simulator for Embodied AI research.
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
Language-Agnostic SEntence Representations
The repository provides code for running inference with the Meta Segment Anything Audio Model (SAM-Audio), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
The repository provides code for running inference with the SAM 3D Body Model (3DB), links for downloading the trained model checkpoints and datasets, and example notebooks that show how to use the model.
This repository contains the code to train and evaluate TRIBE v2, a multimodal model for brain response prediction
A modular high-level library to train embodied AI agents across a variety of tasks and environments.
Omnilingual ASR Open-Source Multilingual SpeechRecognition for 1600+ Languages
Self-referential self-improving agents that can optimize for any computable task
Large Concept Models: Language modeling in a sentence representation space
Schedule-Free Optimization in PyTorch
FAIR Chemistry's library of machine learning methods for chemistry
Code for BLT research paper
TorchMultimodal is a PyTorch library for training state-of-the-art multimodal multi-task models at scale.
Training Large Language Model to Reason in a Continuous Latent Space
MobileLLM Optimizing Sub-billion Parameter Language Models for On-Device Use Cases. In ICML 2024.
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
The Replica Dataset v1 as published in https://arxiv.org/abs/1906.05797 .
Tooling for the Common Objects In 3D dataset.
Fast Differentiable Tensor Library in JavaScript and TypeScript with Bun + Flashlight
FAIR Sequence Modeling Toolkit 2
1K resolution vision transformers pretrained on 1B human images.
Non-invasive decoding of typed sentences from MEG and EEG brain recordings using a convolutional encoder, transformer, and character-level language model.
Can Language Models Rebuild Programs From Scratch?
Mobile vision models and code
Code for the ShapeR research paper
Code release for "Omni3D A Large Benchmark and Model for 3D Object Detection in the Wild"
Momentum Human Rig is an anatomically-inspired parametric full-body digital human model developed at Meta. It includes: A parametric body skeletal model; A realistic 3D mesh skinned to the skeleton with levels of detail;A body blendshape and pose corrective model; A facial blendshape model.Its design is friendly for both CG and CV communities.
projectaria_tools is an C++/Python open-source toolkit to interact with Project Aria data.
Open and efficient video and image watermarking
Unified automatic quality assessment for speech, music, and sound.
This is the official PyTorch implementation of the paper Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP.
Official PyTorch implementation of "InterHand2.6M: A Dataset and Baseline for 3D Interacting Hand Pose Estimation from a Single RGB Image", ECCV 2020
Efficient Universal Perception Encoder: a single on-device vision encoder with versatile representations that match or exceed specialized experts across multiple task domains.
An algorithm that generalizes the paradigm of self-play reinforcement learning and search to imperfect-information games.
SuperDex brings together a purpose-built physics engine, robotics authoring tools, and a scalable reinforcement learning interface in a unified simulation platform, with VR-based teleoperation and additional capabilities planned for future releases.
Code for the Boxer research paper
Fusion-in-Decoder
Code repository for supporting the paper "Atlas Few-shot Learning with Retrieval Augmented Language Models",(https//arxiv.org/abs/2208.03299)
A library to analyze PyTorch traces.
A Large-scale Video Action Dataset
Code, data and weights for the paper **What drives success in physical planning with Joint-Embedding Predictive World Models?**
Official implementation of paper "VLM³: Vision Language Models Are Native 3D Learners".
A library for human kinematic motion and numerical optimization solvers to apply human motion
Code for "LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding", ACL 2024
Python suite for neuroscience research across all modalities.
Comprehensive benchmark for RAG
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models. TMLR 2025.
This repository contains a command-line interface(CLI) that can detect and blur out faces and license plates(PII) from images and videos. The CLI takes an image or video file as input, runs an anonymization algorithm on it, and writes the blurred output to a specified path.
Code for exploring surface electromyography (sEMG) data and training models associated with Reality Labs' paper
Content Seal is a state-of-the-art framework for invisible, robust watermarking across all modalities audio, image, video, and text. This suite spans the entire generative lifecycle, from training data and inference to generated media.
Python Library to evaluate VLM models' robustness across diverse benchmarks
This repository contains the implementation of our novel approach associated with the paper "Don't Splat Your Gaussians" to modeling and rendering scattering and emissive media using volumetric primitives with the Mitsuba renderer.
Fillerbuster: Multi-View Scene Completion for Casual Captures
Autoform Bot
Code for ExploreTom
MobileLLM-R1
Evaluation codes and data for GenEval2
A holistic benchmark for LLM abstention
FlowWM stochastic world modeling via flow matching in DINOv3 feature space, with the FuturePerception (Waymo) benchmark.
Official implementation of the paper "Watermarking Autoregressive Image Generation" (NeurIPS'25)
LSRM is a SOTA, feed-forward 3D reconstruction model that generates high-fidelity, relightable 3D digital twins from sparse 2D views.
This repository contains the training code from paper "SpidR Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision". SpidR is a self-supervised speech representation model that efficiently learns strong representations for spoken language modeling.
Body motion estimation from monocular videos via two stage diffusion (CVPR 2026).
Automatic textbook formalization of Grinberg Algebraic Combinatorics
The repository provides code for running STyMo Fast and Controllable Few-Shot Motion Style Transfer.
WearableQA A Benchmark for Health Reasoningover Real-World Wearable Data
Official Claude Code marketplace for Project Aria — install ARK plugins for Aria SDK development, VRS data processing, MPS workflows, and device operations.
Training and evaluation code for our paper "MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning"
Code for simulating recommendation data and training early stage retrieval models collaborative filtering, mixture of experts, online and offline reinforcement learning using policy gradient methods.
This is a code base that implements the core algorithm of the paper Contrastive Activation Steering for Activation Training (CASAL).
schlepp is the Python package for the Schlepp synthetic dataset, which features multi-camera, multi-actor, multi-modal object manipulation sequences. It provides dense ground-truth optical flow, depth, segmentation, point tracks, oriented bounding boxes, object animation, and MHR body parameters.
facebookresearch builds a variety of tools and libraries related to machine learning and artificial intelligence. Key projects include segment-anything for segmentation tasks and detectron2 for object detection, showcasing their focus on research and development in these fields.
The primary programming languages used by facebookresearch include Python, Jupyter Notebook, C++, HTML, and TypeScript. This selection reflects their commitment to developing robust machine learning frameworks and research tools.
Yes, all of facebookresearch's repositories are public. This openness allows researchers and developers to access, contribute to, and utilize their projects, fostering collaboration and innovation within the AI community.
Monitor Meta Research with RepoGuard and get alerted the moment a new public repository appears.
Monitor this account