Basic.AI

Basic.AI

Share

All-in-One Smart Data Annotation Platform. Training Data Solutions.

With over 7 years experience in AI training data solutions, we exceed in delivering the best-quality data to our global clients, from data collection to data annotation.

Photos from Basic.AI's post 09/08/2026

Weโ€™re heading to Malmรถ, Sweden, for the Expo this month!

Held every two years, brings together the global research community and industry practitioners.

The main conference and Expo will take place September 10-12. Youโ€™ll find BasicAI at ๐๐จ๐จ๐ญ๐ก 14 at ๐Œ๐š๐ฅ๐ฆรถ๐ฆรค๐ฌ๐ฌ๐š๐ง, ๐Œรค๐ฌ๐ฌ๐ ๐š๐ญ๐š๐ง 6, ๐Œ๐š๐ฅ๐ฆรถ, ๐’๐ฐ๐ž๐๐ž๐ง.

https://hallerickson.ungerboeck.com/prod/app85.cshtml?aat=3939617841664b564d50586e734b667a4e4539696b3276326d53356f4a62726e686d724969344b4a784f6f3d&ExhibitorID=3065

Weโ€™re glad to share the Expo floor with Helsing, PAL Robotics, Niulinx, SafeAD, LINKERBOT, Milestone Systems, and DeepRoute.ai, fellow exhibitors helping move computer vision into real-world systems across robotics, mobility, security, and intelligent infrastructure.

Weโ€™re looking forward to meeting computer vision researchers, engineers, and teams to talk about what youโ€™re building, the data you work with, and the practical challenges you face. Feel free to stop by and say hello!

Weโ€™ll have a few small gifts waiting, too.

๐’๐ž๐ž ๐ฒ๐จ๐ฎ ๐ข๐ง ๐Œ๐š๐ฅ๐ฆรถ!

07/24/2026

3D multimodal models still face a clear capability gap when grounding answers in specific 3D regions and estimating real-world dimensions.

In multi-turn conversations, they must also maintain reference consistency, such as locating the chair closest to a table and then measuring โ€œitsโ€ backrest height.

๐†๐ซ๐จ๐ฎ๐ง๐3๐ƒ-๐‹๐Œ๐Œ, recently accepted at , puts those capabilities into one framework.

It takes a colored , an optional aligned RGB image, and a natural-language question. From that input, it can produce a text response, a point-level mask, and a measurement in meters or centimeters. Generating the special token triggers mask prediction, grounding a phrase such as โ€œthe chairโ€™s backrestโ€ in a specific 3D location.

๐Ÿ  ๐๐ซ๐จ๐ฃ๐ž๐œ๐ญ: https://amolharsh.github.io/ground3d-lmm/

The work also defines the 3D Grounded Measurement task, which requires a model to identify the correct 3D region and return the corresponding measurement. Its Grounded-Measurement Success Rate counts a prediction as successful only when the mask IoU and measurement error each meet their required threshold.

The accompanying Ground3D is built on and ScanNet++ and contains 2,475,307 question-answer pairs. It covers 8 task groups, including functional grounding, dimension estimation, depth relationships, size comparison, distance queries, and spatial relationships. The manually verified evaluation set contains 104,405 pairs.

On Ground3D-ScanNet++, however, the 3D-only model scores 24.56, below the image baseline at 27.80. Adding RGB lifts the result to 30.42. This suggests that appearance information still provides useful visual cues beyond point clouds alone.

Connecting conversational language, fine-grained 3D grounding, and real-world measurements through one interface could support research in , AR/VR, indoor , and assistive manipulation.

For now, it is better viewed as a 3D understanding research framework than a precision measurement system. Practical measurement applications still need to address occlusion, point cloud density, camera error, cross-environment generalization, and the quality limits of automatically generated .

07/17/2026

Next-token prediction helped move from task-specific systems toward that can adapt across tasks.

is still seeking for a similar general-purpose pretraining objective. Depth estimation, segmentation, camera pose, and human pose often rely on separate and models.

A recent 2026 paper suggests that large-scale text-to-video generation may offer one path forward. The team presents GenCeption, which turns geometry, motion, temporal consistency, and language semantics learned by a video generation model into general visual perception capabilities.

๐Ÿ“– Project: https://genception.github.io/

GenCeption builds on WAN 2.1. It replaces random noise with clean latents from the input video, fixes the diffusion timestep, and adjusts the DiT output, reducing roughly 50 diffusion steps to a single forward pass.

Text instructions select the task. Dense outputs such as depth, surface normals, segmentation, and camera information are encoded as RGB, while 2D and 3D keypoints are handled through learnable tokens that interact with video features.

Training relies mainly on synthetic human videos and several synthetic datasets. Referring-expression segmentation also uses real-world data.

For , the 14B model uses about 1.23 million frames of post-training data and reaches an average AbsRel of 0.071 on Sintel, KITTI, and ETH3D. D4RT uses about 86 million frames and achieves 0.082, while VGGT-ฮฉ uses about 600 million frames and reaches 0.067.

The paper estimates that GenCeption uses roughly 1/7 to 1/500 as much as leading models. It also reports competitive results on surface normals, camera pose, foreground segmentation, referring-expression segmentation, and 3D keypoints .

This work offers a clear path for adapting video generation to geometry, segmentation, pose, and motion perception with limited post-training data. A text-controlled visual system with shared architecture has direct research value for , , AR/VR, and video understanding .

Joint multi-task training still creates conflicts, especially between 3D keypoints and dense tasks. The 14B model also carries a high compute cost. The approach still needs broader validation on a larger scale.

07/10/2026

Data is often called the new oil. For research, high-quality, large-scale data with clear licenses and stable access is increasingly scarce.

-1K helped drive several generations of generative models. Yet its 1,000-category setup cannot cover the complex conditions required by modern text-to-image systems. Years of optimizing FID on a fixed have also pushed the benchmark toward saturation.

Recently, Fei-Fei Liโ€™s team at Stanford and collaborators have released the ๐†๐ข๐š๐ง๐ญ ๐๐ž๐ซ๐ฆ๐ข๐ฌ๐ฌ๐ข๐ฏ๐ž ๐ˆ๐ฆ๐š๐ ๐ž ๐‚๐จ๐ซ๐ฉ๐ฎ๐ฌ, or ๐†๐๐ˆ๐‚. It contains 100M training samples, 200K validation samples, and 1M test samples, totaling about 28T pixels and 12.9TB.

GPIC follows four design goals. ๐˜—๐˜ฆ๐˜ณ๐˜ฎ๐˜ช๐˜ด๐˜ด๐˜ช๐˜ท๐˜ฆ means licenses are open and traceable. ๐˜š๐˜ต๐˜ข๐˜ฃ๐˜ญ๐˜ฆ means the data files remain fixed. ๐˜“๐˜ข๐˜ณ๐˜จ๐˜ฆ refers to its scale and rich text conditions. ๐˜ˆ๐˜ค๐˜ค๐˜ฆ๐˜ด๐˜ด๐˜ช๐˜ฃ๐˜ญ๐˜ฆ means researchers can download the data directly for training.
๐Ÿ“– ๐๐š๐ฉ๐ž๐ซ: https://arxiv.org/abs/2605.30341

The images come from Flickr and Wikimedia. The pipeline includes quality filtering, safety filtering, and deduplication while retaining license and attribution metadata. The team used Qwen3-VL-4B-Instruct to generate captions at several lengths for text-conditioned generation research.

GPIC also provides a new evaluation protocol. Its main metric, FD-DINOv2, uses fixed 50K test captions and independent 1M test images. This separates from evaluation data while measuring generation quality and distributional differences.

The licensing boundary still matters. The Hugging Face page marks the dataset release layer as MIT, while the underlying images retain their own CC BY, CC0, public-domain, or other license terms.

GPIC brings data, evaluation tools, and a reference training setup into one reproducible framework. It offers a new experimental foundation for open research beyond ImageNet-1K.

Whether it becomes a widely adopted benchmark will depend on further training results and independent evaluation from the community.

06/26/2026

In AI , many questions look simple at first. Once a real project starts, the details appear quickly.

๐˜ž๐˜ฉ๐˜ข๐˜ต ๐˜ช๐˜ด ๐˜ฅ๐˜ข๐˜ต๐˜ข ๐˜ข๐˜ฏ๐˜ฏ๐˜ฐ๐˜ต๐˜ข๐˜ต๐˜ช๐˜ฐ๐˜ฏ? ๐˜๐˜ฐ๐˜ธ ๐˜ด๐˜ฉ๐˜ฐ๐˜ถ๐˜ญ๐˜ฅ ๐˜ข๐˜ฏ๐˜ฏ๐˜ฐ๐˜ต๐˜ข๐˜ต๐˜ช๐˜ฐ๐˜ฏ ๐˜ฒ๐˜ถ๐˜ข๐˜ญ๐˜ช๐˜ต๐˜บ ๐˜ฃ๐˜ฆ ๐˜ฆ๐˜ท๐˜ข๐˜ญ๐˜ถ๐˜ข๐˜ต๐˜ฆ๐˜ฅ? ๐˜๐˜ฐ๐˜ธ ๐˜ฅ๐˜ฐ ๐˜ธ๐˜ฐ๐˜ณ๐˜ฌ๐˜ง๐˜ญ๐˜ฐ๐˜ธ๐˜ด ๐˜ฅ๐˜ช๐˜ง๐˜ง๐˜ฆ๐˜ณ ๐˜ข๐˜ค๐˜ณ๐˜ฐ๐˜ด๐˜ด , ๐˜ช๐˜ฎ๐˜ข๐˜จ๐˜ฆ๐˜ด, ๐˜ข๐˜ฏ๐˜ฅ ๐˜“๐˜“๐˜” ๐˜ฅ๐˜ข๐˜ต๐˜ข? Many industry terms also shift in meaning depending on the project context.

To make these resources easier to read and search, we recently launched two new pages on the BasicAI website.

1๏ธโƒฃ ๐๐š๐ฌ๐ข๐œ๐€๐ˆ ๐€๐œ๐š๐๐ž๐ฆ๐ฒ

Weโ€™ve been sharing knowledge, best practices, and lessons from real projects across our blog, video channels, and social platforms.

Those resources now live in a more structured place. BasicAI Academy brings selected content together and organizes it into a clear learning path for the industry. If you are new to data annotation, or want a broader view of the field, this might be a good place to start.

๐€๐œ๐š๐๐ž๐ฆ๐ฒ: https://www.basic.ai/academy

2๏ธโƒฃ ๐๐š๐ฌ๐ข๐œ๐€๐ˆ ๐–๐ข๐ค๐ข

We also created BasicAI Wiki as a quick reference for common terms that appear in our articles and project discussions. Each entry includes a concise explanation and a visual example. This matters because the same term can carry different meanings in different contexts.

If you come across an unfamiliar term while reading, planning a project, or talking with a team, we hope the Wiki helps you find the answer faster.

๐–๐ข๐ค๐ข: https://www.basic.ai/wiki

We are actively maintaining both hubs and will keep adding new content. If there are topics or terms you would like us to cover, feel free to send us a message.

06/18/2026

Video (VPS) is one of the most annotation-heavy tasks in .

A VPS model has to classify every pixel, separate different instances of the same category, and keep each object identity consistent across frames.

๐•๐ข๐๐ž๐จ๐‚๐”๐๐’, a Highlight work, proposes the first method for VPS. It uses ordinary monocular videos, without human VPS labels or stereo video, to generate temporally consistent panoptic video pseudo-labels and train a VPS model from them.

The core idea is to build supervision from the video itself. VideoCUPS uses motion, depth, and self-supervised visual features to discover which pixels belong to the same object and which regions are semantically similar. It then aligns these signals over time to form consistent video-level panoptic pseudo-labels.

The model is trained on these pseudo-labels with Video DropLoss, which helps when the pseudo-labels are sparse or incomplete. This gives the model room to recover objects that the pseudo-labels miss, such as static objects, while still learning useful object tracks. The method also uses self-enhanced video copy-paste to improve small-object performance.

๐Ÿ  ๐๐ซ๐จ๐ฃ๐ž๐œ๐ญ: https://visinf.github.io/videocups/

On -VPS, VideoCUPS reaches 22.2 STQ, outperforming CUPS + SORT trained with stereo video at 20.6. It also outperforms all unsupervised baselines on KITTI-STEP, Waymo, and MOTS.

The method is also useful in low-label settings. With VideoCUPS initialization, fine-tuning on only 10% of Cityscapes-VPS labels reaches the same STQ as a randomly initialized model trained with 100% of the labels.

This work shows that ordinary monocular video already contains strong signals. Motion, depth, and visual features can be turned into trainable panoptic video pseudo-labels, moving VPS from dense manual toward scalable video self-learning.

06/05/2026

is moving from finding an object to executing an intent.

When the target is no longer just a class name such as car or road, but an anomaly, a cross-image consistent region, or an area defined by function, the model has to understand the concept behind the request.

A research team from NTU and collaborators proposed ๐‚๐จ๐ง๐œ๐ž๐ฉ๐ญ๐’๐ž๐ -๐‘1. Itโ€˜s a framework that formalizes generalized and takes an early step toward ๐’๐ž๐ ๐ฆ๐ž๐ง๐ญ ๐€๐ง๐ฒ ๐‚๐จ๐ง๐œ๐ž๐ฉ๐ญ.

But what is a โ€œconceptโ€ in tasks?

The paper organizes concepts into three levels: CI, CD, and CR.

CI (context-independent) concepts are mostly defined by the object itself, covering common classes, fine-grained classes, and long-tail categories. CD concepts depend on context. The target may be salient, transparent, or abnormal in an industrial scene. CR (context-reasoning) concepts require reasoning across visual and textual evidence.

ConceptSeg-R1 turns this into a rule-driven segmentation pipeline. It induces rules from reference images, checks them with proxy queries, translates the learned concept into the prompt space, and then produces pixel-level masks.

For efficiency, ConceptSeg-R1 also adds a Shortcut Router. Easier high-confidence CI tasks can use a faster SAM3 path, while harder CD and CR tasks go through the full reasoning pipeline.

The experiments cover all 3 tasks, along with external benchmarks such as and ReasonSeg. On zero-shot Cityscapes, ConceptSeg-R1-3B improves SAM 3 from 60.6 to 62.6 mean mIoU. On the ReasonSeg test set, ConceptSeg-R1-7B achieves 63.0 gIoU and 59.3 cIoU.

๐Ÿ  ๐๐ซ๐จ๐ฃ๐ž๐œ๐ญ: https://ntu-ai4x.github.io/ConceptSeg-R1/

segmentation moves beyond closed label set. Promptable segmentation made the interface more flexible through points, boxes, masks, and text.

ConceptSeg-R1 pushes this line further. It asks the model to infer the rule behind a target concept, validate it in context, and turn it into a pixel-level mask.

This becomes important when the target is hard to name cleanly. Lesions, anomalies, camouflaged regions, cross-image differences, and functional areas often start from user intent, beyond a fixed class label or a simple prompt.

05/21/2026

Urban infrastructure maintenance punishes delay. The later you find a problem, the more it costs to fix. U.S. infrastructure repairs alone are projected to run into trillions between 2024 and 2033.

Yet most bridge inspections still rely on human eyes or expensive specialty vehicles. Meanwhile, the forward-facing cameras already mounted on millions of modern cars sit largely unused for this purpose.

A new study from the CHE Lab at the University at Buffalo, SUNY takes a different route. It turns connected vehicles passing over roads and bridges every day into a distributed inspection fleet, using the cameras they already have.
๐Ÿ“– Paper: https://arxiv.org/abs/2604.24616

The challenge is that dashcam footage is noisy. It contains lane markings, shadows, road texture, glare, and perspective distortion. Running on the full image can lead to many false positives.

The teamโ€™s solution is to let the infrastructure tell the vehicle where to look. A roadside unit broadcasts the GPS coordinates of areas worth inspecting over C-V2X. The vehicle combines that with its own localization, camera intrinsics and extrinsics, and geometric projection to work out where the target region should appear in the image. The system then crops a tight window around that location and passes only the focused patch to the crack . The task shifts from hunting a thin crack across a cluttered scene to analyzing a pre-selected region of interest.

The system also goes beyond detection. A four-quadrant corner detection method recovers the actual physical length of each crack, closing the loop from spotting a defect to measuring it.

The team also released the first real-world crack captured from vehicle front-view cameras. It covers varied lighting, glare, motion blur, and road surface textures.

The dynamic cropping step alone lifts precision from 0.20 to 0.70 and improves AP by 822%. On their RT100 dataset, the final model reaches an ODS F1 of 0.5522, up from a 0.2923 baseline, an 89% relative gain.

This work points to a practical model for sensing. Infrastructure can actively assign perception tasks to vehicles already passing by, turning a biannual inspection cycle into continuous hourly monitoring at close to zero marginal cost.

04/24/2026

The SAM family has kept refining on interaction. The original used points and boxes for . SAM2 extended to .

introduced Promptable Concept Segmentation (PCS), locating all instances in an image that match a given noun phrase. But for longer, more complex natural language instructions, SAM3 has to route through an external to translate them into noun phrases first. That makes the system heavier, and fine-grained meaning can get lost along the way.

A recent multi-institution research team proposes SAM3-I (Segment Anything with Instructions), defining a new task, Promptable Instruction Segmentation (P*S). It gives the family a direct path to handle complex natural-language instructions, without routing through an LLM middle layer.

๐Ÿ“– ๐๐š๐ฉ๐ž๐ซ: https://arxiv.org/abs/2512.04585
๐Ÿ  ๐Ž๐ฉ๐ž๐ง-๐ฌ๐จ๐ฎ๐ซ๐œ๐ž๐ ๐จ๐ง: https://github.com/debby-0527/SAM3-I

SAM3-I organizes instructions by difficulty into Concept / Simple / Complex levels. On SAM3โ€™s text side, it inserts an Instruction-Aware Cascaded Adapter that learns progressively across these levels. The S-Adapter focuses on explicit conditions like attributes and location. The C-Adapter builds on that to handle functional descriptions and implicit reasoning. They mirror how humans move from catching keywords to deeper comprehension.

The team also designs four complementary distribution-alignment losses, aiming for the same object to be understood the same way, whether the instruction is a short description or a longer reasoning chain.
To support these, they build the HMPL-Instruct with 840k instructions, covering concept to reasoning, object to part, single-instance to multi-instance.

On simple instructions, SAM3-I outperforms the SAM3 Agent baseline by 31.3 absolute points in gIoU. On complex instructions, the margin is 22.6 points. It uses 1/8 of the parameters and requires only a single forward pass.

The work shows that segmentation can acquire complex language understanding through parameter-efficient adaptation, without giving up existing capabilities. With larger instruction data and more dialog-style interaction, general-purpose segmentation that follows real human instructions is starting to look practical.

04/16/2026

In an first-person (Ego) view, you can annotate an object in someoneโ€™s hand. Switch to a third-person (Exo) camera, and the same object shifts in position, scale, and appearance. It may be occluded by the hand, or confused with similar items nearby. Segmentation and correspondence quickly stop being reliable then.

This is the main challenge of cross-view . In real systems, it stalls critical workflows in multi-camera , video retrieval, and human-robot teaching. Even cannot handle this well. Its spatial prompting was never designed to transfer across views.

A recent Highlight paper, Vยฒ-SAM, caught our attention. It extends SAM2 to a unified cross-view object correspondence framework, without requiring camera poses, semantic labels, or explicit . The same object can be reliably re-identified and segmented across different viewpoints.

๐Ÿ“„ Paper: https://arxiv.org/abs/2511.20886
๐Ÿ  Project: https://jianchengpan.space/projects/V2-SAM/

The method splits the problem into two parts: where the object is, and what it looks like. Vยฒ-Anchor uses geometry-aware features from for cross-view matching, enabling SAM2's point-prompt capability in cross-view settings for the first time. Vยฒ-Visual introduces a Visual Prompt Matcher that aligns object appearance representations across views at both feature and structural levels.

On the Ego-Exo4D benchmark, Vยฒ-SAM sets a new record at 48.0 overall IoU, surpassing the previous best by 4.6 points, while using only 15M trainable parameters, less than 1% of the strong baseline ObjectRelator. On DAVIS-2017 video and the HANDAL-X robotic cross-view transfer task, Vยฒ-SAM leads by a wide margin. Zero-shot transfer to HANDAL-X reaches 77.2 IoU, showing strong generalization.

This work provides a practical, engineering-grounded answer to cross-view perception. It has clear potential as a general-purpose backbone for multi-camera understanding, embodied demonstration learning, and human-to-robot view transfer.

Want your business to be the top-listed Computer & Electronics Service in Irvine?
Click here to claim your Sponsored Listing.

Address


5319 University Drive , PMB 6368
Irvine, CA
92612

Alerts

Be the first to know and let us send you an email when Basic.AI posts news and promotions. Your email address will not be used for any other purpose, and you can unsubscribe at any time.

Subscribe

We will notify you when anything happens in Irvine.