Machine Vision Intelligent System Suite
RUIYI Visual CAST Platform
RUIYI Visual CAST sits between vision models — foundation models, detection models, classic algorithms and your own — and the systems that act on what they see. Connect, adapt, serve and track: one AP

Vision models have become very capable. Show one a picture and it can say what is in it, what looks wrong, and how that differs from the last one. For a lot of tasks, that is now good enough.
The hard part is the other half: getting what the model sees into the systems that do something about it.
A camera on a line catches a surface defect and the model returns a judgement. Then what. That judgement has to become a non-conformity record in the quality system, attached to a specific operation and batch; it has to decide whether the lot is released, reworked or held; it has to raise an alert to the shift leader; and it has to leave evidence that can be pulled up later. That stretch in the middle is where most vision projects actually stall: one adapter for the camera, another rewrite when the model changes, two deployment paths for edge and cloud, results that stop at the demo screen, and no one watching whether the model is still right three months later.
RUIYI Visual CAST covers that stretch. It does not build models and it does not replace business systems — it sits between them and organises models, devices, compute and results into one capability layer that business systems can call reliably. The system above integrates once. The models underneath can be swapped, added and recombined.
Connect. Adapt. Serve. Track. Bring models and devices in, adapt them to your data and your vocabulary, serve results through one API, and track versions, performance and cost — the models change, the integration above does not.
What it is
It is the platform layer between vision capability and the systems that run the business. Upward, business systems call vision capability through one interface. Downward, it registers models and algorithms, schedules compute and orchestrates inference.
The four letters are what it does:
C — Connect. Reaching down to models and algorithms: vision foundation models, purpose-built detection models, classic vision algorithms and OCR, and models the customer trained themselves — all onboarded the same way and managed in one place. Reaching sideways to devices and sources: IP cameras, industrial cameras, edge devices, recorded video and video platforms already in place.
A — Adapt. Translating difference into one language: camera protocols and stream formats from different vendors, model formats exported from different frameworks, and most importantly the translation from what a model outputs — something looks wrong here — into what a business system recognises: which station, which batch, which defect class, which non-conformity code, which action to trigger. It also covers describing a detection target in plain language and turning that into a working detector, rather than collecting and labelling a large sample set every single time.
S — Serve. Exposing capability through one API and SDK: single calls, batch jobs and live streams; several models chained into one pipeline — locate, then classify, then measure; compute allocated across cloud, edge and on-premise by policy; and results routed by rule to wherever they belong — the quality system, an alert channel, a dashboard, or the customer's own application.
T — Track. Models have versions that can be rolled out gradually and rolled back. Performance is watched continuously, with drift and degradation flagged for review before they spread. Who called what, when, with which model version, using how much compute, and producing which judgement — all recorded and searchable. Retention, redaction and access to image data are governed by policy.
Put plainly: the model is not the hard part. Getting what it sees into the systems that act is.
What gets in the way today
Teams rarely fail because the model is inaccurate. They fail on the plumbing around it, and on what happens after go-live.
Integration
Every model arrives with its own interface. A new scene or a new model supplier means integrating again: new authentication, new request shape, new response structure, new error codes.
The model becomes a lock-in. Once the business system is written against one supplier's interface, changing model means rebuilding — so the incumbent stays, even after something better exists.
Results never reach the systems that act. The model says there is a scratch in the image. What the business system needs is station 3, lot A2024-0713, defect class SCR-02, recommend hold. Nobody builds that mapping, so the result stays on a demo screen.
Devices
Cameras and protocols are a zoo. IP cameras, industrial cameras, edge devices and existing video platforms each have their own way in and their own stream format.
Edge and cloud are two separate builds. Low-latency work has to run at the edge while central management wants the cloud, and deployment, updates and monitoring tend to get written twice.
Operations
Nothing watches the model after release. Lighting changes, incoming material changes, a lens gets dirty — performance drifts quietly until complaints or a wave of missed defects force attention.
Versions are a mystery. Which site runs which version, what changed, and whether the previous one can be restored when something breaks — usually nobody can say.
Compute is allocated by guesswork. Several jobs contend for the same resources with no basis for priority, no trigger to scale, and no view of which job the spend belongs to.
Governance
Image data carries regulatory weight. How long it is kept, who may view it, whether faces leave the site, and where data sits in cross-border operations — requirements vary by region and need handling at the platform layer, not reinvented in every application.
Audit has nothing to hold on to. Was a judgement made by the model or corrected by a person, against which model version, and is the source image still there — when it matters, these cannot be produced.
Who it is for
Manufacturers. Want vision without building an algorithm team to get it: quality, safety and equipment checks stay with the business units and their integration partners, while CAST delivers the capability into the systems those teams already use.
System integrators. Need vision inside a project without standing up inference serving and interfaces from scratch each time: CAST gives one onboarding path and one delivery shape that carries between projects.
Solution vendors and software companies. Need vision inside their own product and delivered to their customers: CAST acts as the capability layer beneath the product, so model choice and iteration never disturb the interface they ship.
RUIYI's own product lines. Visual Quality Inspection, Visual Security Monitoring and the vision checks inside the catering suite run on the same capability layer — sharing model registry, compute scheduling, lifecycle and governance instead of each building its own.
What you get
Capabilities
C — Connect
Capability | What it does |
Model and algorithm registry | Vision foundation models, detection models, classic vision algorithms and OCR, and customer-trained models registered, versioned and discoverable in one place. |
Device and source onboarding | IP cameras, industrial cameras, edge devices, recorded video and existing video platforms brought in and managed together. |
Several models side by side | More than one model attached to the same scene, run in parallel or switched by rule, for comparison, gradual rollout and fallback. |
A — Adapt
Capability | What it does |
Protocol and format conversion | Camera protocols, stream formats and model formats from different vendors normalised, exposing one interface upward. |
Detection targets in plain language | A target and scene described in words becomes a working detector, reducing the dependence on large labelled sample sets for every new case. |
Results mapped to business objects | Model output mapped into the fields a business system recognises: station, batch, defect class, disposition and the action to trigger. |
Environment adaptation | One capability definition deployed to cloud, on-premise or edge runtimes according to latency and data-residency requirements. |
S — Serve
Capability | What it does |
Unified API and SDK | One interface covering single calls, batch jobs and live streams, independent of the model underneath. |
Pipeline orchestration | Models and rules chained into a flow — locate, classify, measure, judge — configurable and reusable across scenes. |
Compute scheduling | Cloud, edge and on-premise resources allocated and scaled by policy, with job priority and quotas. |
Result routing and closed loop | Results sent by rule to the quality system, an alert channel, a dashboard or the customer's own application, and triggering what follows. |
Human review channel | Low-confidence results routed to a review queue, with human decisions fed back and recorded. |
T — Track
Capability | What it does |
Version and release management | Model versions, gradual rollout and rollback, with per-site scope control. |
Performance monitoring and drift alerts | Behaviour tracked continuously, with drift and degradation identified and raised for review. |
Audit and traceability | Model version, input, judgement and operator recorded for every call and searchable afterwards. |
Data governance | Retention period, redaction policy and access rights for images and results configured by organisation and region. |
Cost and usage analysis | Compute and call volume reported by model, job and site to support capacity and cost decisions. |
How it works
Connect. Register models and algorithms — platform-provided, customer-owned or third-party — bring in cameras and data sources, and normalise protocols and formats.
Define the capability. Describe what to detect and how to judge it, in words or by configuration, and map the output into the fields and actions the business system already understands.
Compose the pipeline. Chain models and rules into a flow: capture, pre-process, infer, judge, route.
Choose where it runs. Cloud, on-premise or edge, with compute assigned by latency, bandwidth and data-residency requirements.
Integrate once. The business system calls the capability through the unified API. Later model changes stay invisible to it.
Run and review. Results land in the business system and trigger what follows; low-confidence items go to human review and the decision flows back.
Track and iterate. Versions, behaviour and cost are recorded continuously; degradation leads to retraining or replacement, released through gradual rollout with rollback available.
Integration & deployment
Business systems. MES, QMS, WMS, EHS and customer-owned applications, connected through the unified API and webhooks.
Device layer. IP cameras, industrial cameras, edge computing devices and existing video management systems.
Model layer. Platform-provided, customer-owned and third-party models, onboarded in a common format.
Industrial protocols. Connections to PLC, SCADA and shop-floor automation configured per project.
Deployment model — cloud, on-premise or hybrid — and the integration scope are agreed during scoping, usually starting with one scene or one line before extending to further sites.

