Boston AI Builds Practical AI Systems for Voice, Vision and More
AI is changing quickly. Voice and vision are giving software a richer understanding of people and their surroundings, while advances in specialized models, local inference and multimodal AI are making new applications practical.
At Boston AI and ICS, we’re building systems that put these capabilities to work in real products and applications. Our work spans voice, vision, 3D imaging, specialized model development, local AI, robotics and more. Here are a few examples.
Building with Voice and Vision
When AI can see and listen, it can do a lot more than respond to a prompt. Boston AI combines expertise in computer vision, voice and AI engineering to build systems that can understand what’s happening around them and respond in useful ways.
These highlighted projects illustrate different ways AI can interpret the physical world, from reconstructing 3D environments to understanding where people are in a room or separating voices in a noisy space. They also highlight an important part of our approach: combining AI models with the engineering, processing, and application software needed to make them work in real-world conditions.
3D Photo Creation
What if a handful of photos could become an interactive 3D scene? This project turns just 3 - 20 photos from a phone or drone into a scene that can be explored in a web browser or exported to Blender. Using Meta’s VGGT-1B, PyTorch, CUDA, and a custom GPU renderer, the system estimates camera positions and depth and produces an initial view in about 20 seconds, while COLMAP and Ceres refine the scene in the background. It also supports 3D Gaussian splats, real-world measurements, and AI-based object detection and labeling with YOLO26 – all running on a single Windows PC with a 16 GB RTX 5060 Ti GPU.

HAL Audiomix
HAL Audiomix applies AI and signal processing to a very different problem: turning multiple microphones in a room into a single, clean voice stream for applications such as Zoom. The system processes 10 ms audio blocks in an average of just 1.9 ms, using Dugan-style gain sharing, noise gating, SNR analysis, auto-leveling, compression, and limiting to balance speakers while suppressing room noise. Built with Python, NumPy, SciPy, PyTorch, pyannote, CUDA and RNNoise, it can also map speaker locations, recognize enrolled voices, and calibrate microphone positions using acoustic timing measurements.

PhotoID
PhotoID reconstructs a room and the people inside it as a live 3D scene, updating every 1.2 seconds from three or more ordinary 1080p Wi-Fi cameras. Running on a single Windows PC with an RTX 5060 Ti 16 GB GPU, it combines VGGT-1B for depth estimation, YOLO11n-seg for real-time person detection, and OpenCV/ArUco calibration to maintain accurate geometry while separating moving people from the pre-measured room. A custom headless OpenGL renderer streams the reconstructed scene to a browser, with visual treatments ranging from sharp and soft imagery to VHS-style effects.
Training AI for Specific Applications
General-purpose models are powerful, but many products require AI that understands a particular type of data, environment or task. Boston AI builds and trains vision models for specific applications, including object recognition and specialized image analysis.
This work includes more than model training. We develop the surrounding data pipelines, evaluation processes, and software needed to integrate models into working systems. For instance, our Skin Cancer Detection System applies computer vision to specialized imagery and domain-specific problems. The AI-powered image recognition system is trained to identify potential signs of skin cancer with remarkable precision.
Running AI Where It Matters
Not every AI application belongs in the cloud. Products may need local inference for privacy, low latency, reliability, cost, or operation without a constant internet connection.
Waltham AI Appliance
Waltham is a hardened, private AI operating system and inference appliance built on the ASUS Ascent GX10, with an NVIDIA GB10 GPU and unified memory. It runs AI models and applications locally in isolated, hardware-aware arm64/CUDA containers, with strict memory controls and an OpenAI-compatible /v1 API for coding agents and other AI tools. A secure portal provides authentication, role-based access control, personal API tokens, isolated application networks, and detailed audit and model-usage logging.
Waltham is one example of how we approach local AI as a complete system rather than simply putting a model on a device. The hardware, model runtime, applications, security, user management, and APIs all have to work together.

Bringing AI Into the Product
Adding AI to a product is about more than connecting a model to an interface. The technology has to fit the way people actually interact with the product.
The Touch Table illustrates this challenge. Its vision system can identify objects placed on the table, while the original concept used touch to request more information. As the system developed, it became clear that putting too many functions into the touch interface made the interaction cumbersome. The team is now reworking the interaction model, bringing voice, vision, and touch together in a more natural experience.

AI in Systems That Do Something
Our AI work extends well beyond these examples. We’re applying AI and advanced computing to robotics, sensing, medical and specialized imaging, industrial equipment, and other real-world systems, working on projects that leverage different combinations of perception, intelligence, control and human interaction.
The common thread is practical: choosing the right models and technologies, integrating them with the rest of the system, and engineering them to work under real-world constraints. Whether AI is running in the cloud, on a dedicated appliance, or directly within a product, our goal is the same – to turn what AI can do into something useful.
If you’re in the Boston area and would like to see these and some of our other AI projects in person, join us on Wednesday, October 7 from 4 - 8 pm at our home office in Waltham for a special evening devoted to the practice and discussion of applied AI. The event is free. Registration is required.