facebookresearch는 Meta Research의 GitHub 계정으로, Python, Jupyter Notebook, C++ 등 다양한 프로그래밍 언어로 여러 공개 리포지토리를 운영하고 있습니다. 주요 프로젝트로는 Segment Anything, Detectron2, fairseq 등이 있으며, 이러한 리포지토리는 객체 탐지, 오디오 처리 및 심층 학습과 관련된 연구에 중점을 두고 있습니다.
The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
A library for efficient similarity search and clustering of dense vectors.
Detectron2 is a platform for object detection, segmentation and other visual recognition tasks.
Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
Library for fast text representation and classification.
FAIR's research platform for object detection research, implementing popular algorithms like Mask R-CNN and RetinaNet.
Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
End-to-End Object Detection with Transformers
[CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
PyTorch code and models for the DINOv2 self-supervised learning method.
Code to accompany "A Method for Animating Children's Drawings of the Human Figure"
Foundational Models for State-of-the-Art Speech and Text Translation
The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
Reference PyTorch implementation and models for DINOv3
Code for the paper Hybrid Spectrogram and Waveform Source Separation
Implementation of Nougat Neural Optical Understanding for Academic Documents
Fast, modular reference implementation of Instance Segmentation and Object Detection algorithms in PyTorch.
ImageBind One Embedding Space to Bind Them All
PySlowFast: video understanding codebase from FAIR for reproducing state-of-the-art video models.
SAM 3D Objects
Kats, a kit to analyze time series data, a lightweight, easy-to-use, generalizable, and extendable framework to perform time series analysis, from understanding the key statistics and characteristics, detecting change points and anomalies, to forecasting future trends.
A natural language modeling framework based on PyTorch
High-resolution models for human tasks.
CoTracker is a model for tracking any point (pixel) on a video.
A PyTorch library for implementing flow matching algorithms, featuring continuous and discrete flow matching implementations. It includes practical examples for both text and image modalities.
PyTorch code and models for VJEPA2 self-supervised learning from video.
[CVPR 2026 Oral] VGGT Omega
A Python toolbox for performing gradient-free optimization
Evolutionary Scale Modeling (esm): Pretrained language models for proteins
PyTorch code and models for V-JEPA self-supervised learning from video.
State-of-the-art deep learning based audio codec supporting both mono 24 kHz audio and stereo 48 kHz audio.
A flexible, high-performance 3D simulator for Embodied AI research.
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
Language-Agnostic SEntence Representations
The repository provides code for running inference with the Meta Segment Anything Audio Model (SAM-Audio), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
The repository provides code for running inference with the SAM 3D Body Model (3DB), links for downloading the trained model checkpoints and datasets, and example notebooks that show how to use the model.
This repository contains the code to train and evaluate TRIBE v2, a multimodal model for brain response prediction
A modular high-level library to train embodied AI agents across a variety of tasks and environments.
Omnilingual ASR Open-Source Multilingual SpeechRecognition for 1600+ Languages
Self-referential self-improving agents that can optimize for any computable task
Large Concept Models: Language modeling in a sentence representation space
FAIR Chemistry's library of machine learning methods for chemistry
Code for BLT research paper
TorchMultimodal is a PyTorch library for training state-of-the-art multimodal multi-task models at scale.
Training Large Language Model to Reason in a Continuous Latent Space
MobileLLM Optimizing Sub-billion Parameter Language Models for On-Device Use Cases. In ICML 2024.
Tooling for the Common Objects In 3D dataset.
Fast Differentiable Tensor Library in JavaScript and TypeScript with Bun + Flashlight
FAIR Sequence Modeling Toolkit 2
[CVPR 2025] Official PyTorch implementation of "EdgeTAM: On-Device Track Anything Model"
1K resolution vision transformers pretrained on 1B human images.
Non-invasive decoding of typed sentences from MEG and EEG brain recordings using a convolutional encoder, transformer, and character-level language model.
Research code artifacts for Code World Model (CWM) including inference tools, reproducibility, and documentation.
Code for the ShapeR research paper
Momentum Human Rig is an anatomically-inspired parametric full-body digital human model developed at Meta. It includes: A parametric body skeletal model; A realistic 3D mesh skinned to the skeleton with levels of detail;A body blendshape and pose corrective model; A facial blendshape model.Its design is friendly for both CG and CV communities.
projectaria_tools is an C++/Python open-source toolkit to interact with Project Aria data.
Unified automatic quality assessment for speech, music, and sound.
This is the official PyTorch implementation of the paper Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP.
Official PyTorch implementation of "InterHand2.6M: A Dataset and Baseline for 3D Interacting Hand Pose Estimation from a Single RGB Image", ECCV 2020
Efficient Universal Perception Encoder: a single on-device vision encoder with versatile representations that match or exceed specialized experts across multiple task domains.
An algorithm that generalizes the paradigm of self-play reinforcement learning and search to imperfect-information games.
BenchMARL is a library for benchmarking Multi-Agent Reinforcement Learning (MARL). BenchMARL allows to quickly compare different MARL algorithms, tasks, and models while being systematically grounded in its two core tenets: reproducibility and standardization.
SuperDex brings together a purpose-built physics engine, robotics authoring tools, and a scalable reinforcement learning interface in a unified simulation platform, with VR-based teleoperation and additional capabilities planned for future releases.
Code for the Boxer research paper
Fusion-in-Decoder
Code repository for supporting the paper "Atlas Few-shot Learning with Retrieval Augmented Language Models",(https//arxiv.org/abs/2208.03299)
A library to analyze PyTorch traces.
Dr. Zero Self-Evolving Search Agents without Training Data
A Large-scale Video Action Dataset
Code, data and weights for the paper **What drives success in physical planning with Joint-Embedding Predictive World Models?**
VRS is a file format optimized to record & playback streams of sensor data, such as images, audio samples, and any other discrete sensors (IMU, temperature, etc), stored in per-device streams of timestamped records.
Official implementation of paper "VLM³: Vision Language Models Are Native 3D Learners".
🎬ActionMesh: A fast video to animated mesh model with unprecedented quality. Generate animated mesh seamlessly importable into any 3D software in less than a minute.
A library for human kinematic motion and numerical optimization solvers to apply human motion
Python suite for neuroscience research across all modalities.
Comprehensive benchmark for RAG
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models. TMLR 2025.
This repository contains a command-line interface(CLI) that can detect and blur out faces and license plates(PII) from images and videos. The CLI takes an image or video file as input, runs an anonymization algorithm on it, and writes the blurred output to a specified path.
Code for exploring surface electromyography (sEMG) data and training models associated with Reality Labs' paper
Python Library to evaluate VLM models' robustness across diverse benchmarks
Efficient baselines for autocurricula in JAX.
CoWTracker: Tracking by Warping instead of Correlation
This repository contains the implementation of our novel approach associated with the paper "Don't Splat Your Gaussians" to modeling and rendering scattering and emissive media using volumetric primitives with the Mitsuba renderer.
Fillerbuster: Multi-View Scene Completion for Casual Captures
Autoform Bot
Code for ExploreTom
MobileLLM-R1
Official implementation of the paper "Watermarking Autoregressive Image Generation" (NeurIPS'25)
LSRM is a SOTA, feed-forward 3D reconstruction model that generates high-fidelity, relightable 3D digital twins from sparse 2D views.
Body motion estimation from monocular videos via two stage diffusion (CVPR 2026).
WybeCoder Verified Generation of Imperative Code with LLMs
Automatic textbook formalization of Grinberg Algebraic Combinatorics
The repository provides code for running STyMo Fast and Controllable Few-Shot Motion Style Transfer.
WearableQA A Benchmark for Health Reasoningover Real-World Wearable Data
Official Claude Code marketplace for Project Aria — install ARK plugins for Aria SDK development, VRS data processing, MPS workflows, and device operations.
Training and evaluation code for our paper "MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning"
Code for simulating recommendation data and training early stage retrieval models collaborative filtering, mixture of experts, online and offline reinforcement learning using policy gradient methods.
This is a code base that implements the core algorithm of the paper Contrastive Activation Steering for Activation Training (CASAL).
schlepp is the Python package for the Schlepp synthetic dataset, which features multi-camera, multi-actor, multi-modal object manipulation sequences. It provides dense ground-truth optical flow, depth, segmentation, point tracks, oriented bounding boxes, object animation, and MHR body parameters.
facebookresearch는 객체 탐지, 텍스트 분류 및 오디오 처리와 같은 분야에서 연구를 진행하며, Segment Anything, Detectron2, fairseq 등의 여러 리포지토리를 개발하고 있습니다.
facebookresearch는 주로 Python, Jupyter Notebook, C++, HTML 및 C를 사용하여 다양한 리포지토리를 개발하고 있습니다. 이러한 언어들은 연구 및 개발 작업에 적합합니다.
예, facebookresearch의 모든 리포지토리는 공개되어 있어 누구나 접근할 수 있습니다. 이는 연구 결과와 코드를 공유하여 커뮤니티의 발전에 기여하고자 하는 의도를 반영합니다.