A machine vision inspection apparatus

The machine vision inspection apparatus addresses label inspection challenges with an AI-driven vision-language model and edge processing, ensuring fast, accurate, and adaptable inspection across varying product layouts and orientations.

WO2026114884A1PCT designated stage Publication Date: 2026-06-04VISKA AUTOMATION SYSTEMS LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
VISKA AUTOMATION SYSTEMS LTD
Filing Date
2025-11-25
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing machine vision systems face challenges in accurately inspecting products with varying label layouts and orientations, especially in low-volume, high-variety manufacturing environments, due to human error in manual checks and the complexity of configuring traditional systems for diverse languages, fonts, and layouts, which are exacerbated by dynamic printing and motion-induced text variations.

Method used

A machine vision inspection apparatus with a camera head mounted on a robot, utilizing an AI-powered vision-language model that processes tokenized text strings and captures images under varying conditions, dynamically adjusting settings and performing real-time inspections with a GPU-enabled edge processing unit for fast inference.

Benefits of technology

Enables accurate, real-time inspection with reduced setup time and human error, achieving high-speed defect detection and flexible configuration across diverse products, while minimizing the need for extensive data collection and neural network training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025084221_04062026_PF_FP_ABST
    Figure EP2025084221_04062026_PF_FP_ABST
Patent Text Reader

Abstract

A machine vision inspection system (1) has digital data processors (25) linked with an articulated arm robot (15) supporting a camera and illuminator head (10) for capturing images form a scene over a machine bed (3). The head also includes a system-on-module with a GPU as an inspection processor located adjacent the camera sensor (120), but in a different chamber (109) of the head, separated by a chamber wall (123) from a camera chamber (108). The object may be three dimensional components in a manufacturing line, or product labels. The processors move the robot to pre-set camera positions and control image acquisition and scene illumination for each of a plurality of inspection points. During setup, the processors capture a pre-set robot position, preset camera settings, preset illumination parameters, and a natural language prompt for each of the plurality of inspection points. During runtime, for each inspection point the processors cause the robot (15) to move to the pre-set position for the inspection point when an object is in the camera scene (3) and activate the camera illuminator head to illuminate the scene and capture an image of the object. The inspection processor (140) then executes artificial intelligence code to analyse the image and determine results for the scene, generate an output for the inspection point in a format determined by the prompt associated with the inspection point during setup.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] “A Machine Vision Inspection Apparatus”

[0002] Introduction

[0003] The present invention relates to machine vision inspection, especially for manufacturing lines.

[0004] In production lines there is often a requirement for real time inspection of products. Such inspection can be complex, for example to detect numbers and other data on labels which may be at different angles and having different levels of visibility on the line. Inspection requirements for the parts may include inspections for multiple sides and angles, and can include various visual quality aspects including dimensions and product identification. Accuracy of inspection is sometimes critical, especially for medical products, either devices or pharmaceuticals.

[0005] Some industries in particular suffer from problems of low product volumes, high product-to- product variation, and a large variety of products manufactured on the same manufacturing lines. In addition to this, these industries may have high product yields (i.e. defects do no occur often) thus leading to low volumes of defect sample data.

[0006] In the manufacture of pharmaceutical, medical, or food products there frequently exists the requirement to inspect product labelling. Especially within regulated industries such as medical device manufacturing, there is a need for accurate batch control, correlating material batch numbers with manufactured goods batch numbers. Thus, when starting a new batch of devices, it is necessary to verify the source of the materials, and at the end of each batch, to verify and validate the contents and tracking of each label applied to each product. In the medical device and pharmaceutical industries in particular, label and leaflet mix-up present a significant liability risk to the manufacturer (for example, a patient taking the incorrect dose of medication which causes them harm).

[0007] As a countermeasure to this risk, the labels on the products are manually inspected or are inspected by traditional machine vision systems. Manual checks are subject to human error and a lack of traceability, which presents obvious disadvantages. Traditional machine vision systems, utilising either OCR / Deep Learning / Edge Learning technologies are utilised to good effect, however, they must normally be configured with a recipe file unique to each label layout. Given the wide number of languages, fonts, marketing label styles and layouts, this presents a significant setup issue for the end user. Furthermore, in many cases the labels are printed dynamically by an inline inkjet printing system. The printed text is subject to much variation as the motion of the parts on the conveyor mean that the text can be presented as sparse, bold, straight, wavy, and in varying locations and orientations. This makes the reliable configuration and operation of such a system extremely challenging.

[0008] The present invention is directed towards providing an improved machine vision inspection apparatus.

[0009] Summary of the Invention

[0010] We describe a machine vision inspection system comprising: a camera head mounted on a robot, and comprising a camera and an illuminator; a robot controller with digital data processors, linked with the robot; an inspection processor comprising image capture circuits and an Al processor; wherein the robot controller is configured to move the robot to pre-set camera positions each for a scene and being associated with an inspection point, and the inspection processor is configured to control image acquisition and scene illumination for each inspection point, wherein the inspection processor is trained with artificial intelligence code according to a vision-language model.

[0011] Preferably: the vision-language model is based on an artificial intelligence model which accepts and processes a tokenised text string as an input, the model includes an image feature extraction tool, so that during an initial model training phase, the image feature extraction tool is trained for feature extraction under various scales, brightness, contrast, and colour configurations, and the model is configured to, during training, combine tokenised real language description strings derived from said prompts with corresponding extracted feature maps for given robot positions and camera settings, images;

[0012] Also, preferably, the robot controller and the inspection processor are configured to, for each of a plurality of inspection points: during setup, capture setup data for a pre-set robot position, preset camera settings, preset illumination parameters, and a natural language prompt for each of the plurality of inspection points, and storing said setup data; during runtime, for each inspection point: cause the robot to move to the pre-set position for the inspection point when an object is in the camera scene, activate the camera illuminator head to illuminate the scene and capture an image of the object, execute the artificial intelligence code to analyse the image and determine results for the scene, generate an output for the inspection point in a format determined by the prompt associated with the inspection point during setup.

[0013] In some preferred examples, the artificial intelligence software is configured to, during run-time, generate an output which is un-tokenised according to the prompt received during training, resulting in real language output.

[0014] In some preferred examples, the processors are configured to cause a plurality of images to be captured by the camera while the robot is in motion to the pre-set position for the current inspection point, and to process said images to provide temporal inspection for the inspection point.

[0015] In some preferred examples, the processors are configured to cause said images to be captured with different scene illumination conditions. In some preferred examples, said temporal inspection include (a) analysis of context of single images by correlating scene to real descriptive language, and (b) determination of defects on objects by utilising knowledge of temporal context of a series of images.

[0016] In some preferred examples, said data processors include a graphics processing unit, GPU, configured to perform convolutional calculations at a speed allowing the system to achieve an inference rate in less than 10 seconds.

[0017] In some preferred examples, the processors are configured to operate offline, at the edge of a manufacturing line inspection station.

[0018] In some preferred examples, the processors are configured to communicate with a cloud-based expert agent model during training for defining prompts to allows the user to automate the generation of the training prompts by entering a more general description of an object defect and to receive from the cloud-based model a recursively refined edge model prompt until desired results are achieved.

[0019] In some preferred examples, the inspection processor is located in a camera head housing together with a camera sensor and lens, and the robot controller comprises a separate robot processor linked with the inspection processor.

[0020] In some preferred examples, said camera head comprises a housing which has a camera chamber comprising the camera sensor and the lens mounted along an optical axis, and said camera chamber is separated by a wall from a processor chamber containing circuits including an Al processor.

[0021] In some preferred examples, the camera sensor is mounted to a substrate which is in turn mounted by support of a material including a ceramic.

[0022] In some preferred examples, the camera sensor support is mounted within the camera chamber in a configuration spaced apart from the chamber wall separating the camera chamber from the processor chamber.

[0023] In some preferred examples, the processor chamber comprises a carrier board supporting electronic components on a lower side facing a housing base and on an opposed side supporting spacers which support a system-on-module which includes a GPU for executing the inspection processing Al model software, and said system-on-module is linked by a series of thermally conducting components to a housing top wall which has external fins.

[0024] In some preferred examples, the series of heat transfer components include a heat transfer block which interfaces with system-on-module via a thermally conductive paste and interfaces with the housing via a thermally conductive foil. Preferably, the foil is of graphite material.

[0025] In some preferred examples, the system-on-module is pressed against the series of heat conducting components by a leaf spring acting between the carrier board and the system-on-module.

[0026] In some preferred examples, the camera head housing has an anodised coating on some or all of its external surfaces.

[0027] In some preferred examples, the coating has a thickness in the range of 0.02 mm to 0.05 mm. In some preferred examples, the coating has a reflectivity of less than 30% reflectivity in the visible light spectrum.

[0028] In some preferred examples, the camera substrate mount comprises Aluminium composited with Zirconium Tungstate, with a negative-to-low co-efficient of thermal expansion (CTE) between -2 pm / m / K and 2 pm / m / K.

[0029] In some preferred examples, the inspection processor and the robot controller are configured to act as agents which perform agentic communication with each other.

[0030] In some preferred examples, the inspection processor and the robot processor are configured to communicate with natural language prompts.

[0031] In some preferred examples, the robot controller is configured to cause camera head movement to intelligently bring a previously omitted part of an object into a visible scene, in response to a prompt from the robot controller to the inspection processor asking if a particular expected part of an object under inspection has been visible.

[0032] In some preferred examples, the inspection processor is configured to communicate by natural language prompts to at least one other inspection processor.

[0033] In some preferred examples, the inspection processor is configured to communicate to provide data concerning a particular view and to process corresponding data concerning a different view from the other inspection processor, and to generate an output which allows for omission of background objects, or which indicates a defect on a highly polished surface which is not discernible from all views of the object.

[0034] In some preferred examples, the inspection processor is configured to include, in its output data, data which is susceptible to human interpretation of the rationale of the robot processor.

[0035] In some preferred examples, the inspection processor is configured to dynamically adjust image input dimensions to optimise resolution of the image within available memory, by: downscaling an initial whole image according to a prompt to return a subregion position to provide a first inference as a JSON of bounding box co-ordinates, transforming the bounding box co-ordinates to map to the original image scale, and cropping the original full resolution image using the bounding box co-ordinates returned in the first inference, and passing to the model for a second inference, for a second inference, reading at the full original resolution to determine an output such as a Use By date on a label.

[0036] In some preferred examples, the inspection processor is configured to return data as a JSON bounding box co-ordinates and to transform the bounding box co-ordinates to map to the original image scale.

[0037] In some preferred examples, the inspection processor is configured to pass a cropped image to the model for a second inference at full resolution.

[0038] In some preferred examples, the inspection processor is configured to quantize a model to a lower bit size to reduce memory and processing requirements.

[0039] Additional Statements

[0040] According to the invention there is provided a machine vision inspection system comprising digital data processors linked with a robot supporting a camera and illuminator head, the processors being configured to move the robot to pre-set camera positions and to control image acquisition and scene illumination, and wherein the processors are configured to, for each of a plurality of inspection points: during training, capture a pre-set robot position, preset camera settings, preset illumination parameters, and a natural language prompt for each of the plurality of inspection points, during runtime, for each inspection point: cause the robot to move to the pre-set position for the inspection point when an object is in the camera scene, activate the camera illuminator head to illuminate the scene and capture an image of the object, execute artificial intelligence code to analyse the image and determine results for the scene, generate an output for the inspection point in a format determined by the prompt associated with the inspection point during setup. In some preferred examples, the vision-language model is based on an artificial intelligence model which accepts and processes a tokenised text string as an input.

[0041] In some preferred examples, the model includes an image feature extraction tool, so that during initial model training, an image feature extraction tool is trained on a varied dataset to ensure robust feature extraction under various scales, brightness, contrast, and colour configurations.

[0042] In some preferred examples, the artificial intelligence model is configured to, during training, combining tokenised real language description strings derived from said prompts with corresponding extracted feature maps for given robot positions and camera settings, images.

[0043] In some preferred examples, the artificial intelligence software is configured to, during run-time, generate and output which is un-tokenised according to the prompt received during training, resulting in real language output.

[0044] In some preferred examples, the processors are configured to cause a plurality of images to be captured by the camera while the robot is in motion to the pre-set position for the current inspection point, and to process said images to provide temporal inspection for the inspection point.

[0045] In some preferred examples, the processors are configured to cause said images to be captured with different scene illumination conditions.

[0046] In some preferred examples, said temporal inspection include (a) analysis of context of single images by correlating scene to real descriptive language, and (b) determination of defects on objects by utilising knowledge of temporal context of a series of images.

[0047] In some preferred examples, said data processors include a graphics processing unit, GPU, configured to perform convolutional calculations at a speed allowing the system to achieve an inference rate in less than 10 seconds.

[0048] In some preferred examples, the processors are configured to operate offline, such as entirely at the edge of a manufacturing line inspection station.

[0049] In some preferred examples, the processors are configured to communicate with a cloud-based expert agent model for defining prompts to allows the user to automate the generation of the training prompts by entering a more general description of an object defect and to receive from the cloud-based model a recursively refined edge model prompt until desired results are achieved.

[0050] Detailed Description of the Invention

[0051] The invention will be more clearly understood from the following description of some embodiments thereof, given by way of example only with reference to the accompanying drawings in which:

[0052] Fig. l is a front view of a machine vision inspection system of the invention;

[0053] Fig. 2 is a plan view of a camera head of the system, and Fig. 3 is a cross sectional vies of the camera head;

[0054] Figs. 4, 5, and 6 are sequences of displayed images and accompanying text indicating interfacing with the system; and

[0055] Fig. 7 is a flow diagram illustrating operation of the system in more detail.

[0056] Referring to Fig. 1 a system 1 of the invention comprises a main housing or cabinet 2 having a bed 3 for supporting an object under inspection, a camera head 100 mounted on an articulating arm robot 15, and a di splay / keyboard 20. The housing 2 contains a robot controller 25 with digital data processors including at least one GPU. The controller 25 and, the robot 15 and the head 100 may be mounted differently, such as alongside a production or packing line in a manufacturing environment.

[0057] In this example, the camera head 100 comprises not only a camera and image capture circuits, but also an inspection processor which receives the camera images in real time and processes them using Al models to generate an inspection output. In other embodiments the inspection controller is included in the cabinet 2, linked to the camera by wires. However, it is preferred that it be closer to the camera for the purposes of response time.

[0058] Referring to Figs. 2 and 3 the camera head 100 is shown in more detail. It comprises a rectangular housing 101 comprising a steel casing 102 on a base plate 103, these forming together a generally rectangular enclosure. On the outside there is an anodized coating 105 over all exposed surfaces of the housing 101. The anodized coating provides for very efficient heat transfer, as set out in more detail below. There are longitudinal heat transfer fins 106 and diagonal fins 107 at the corners on a top side of the head 100.

[0059] Importantly, the head 100 comprises inspection processor electronic hardware to execute the inspection software, including Al models. This is done locally within the head 100. The processors 25 in the cabinet 2 is primarily configured to control the robot 15 for optimum movement of the head 100, and so is referred to as a robot controller. In some embodiments the inspection circuits are solely in the cabinet 2, but in preferred embodiments they are within the head 100, and so are very much locally positioned to provide immediate inspection feedback. In the latter case, it is advantageous that the inspection processor is within the head 100 and the robot controller 25 is within the cabinet 2 or any housing which is linked with the articulating arm robot 15, which has a motor at each joint as is well known in the art. In this case, the inspection processor and the robot controller communicate with each other over a wired or wireless (preferably wired) interface in which the operations of image capture and analysis are performed within the head 100 and the robot control operations are within the processor 25, and the necessary data and commands are transmitted over this interface, as described in more detail below. They act as agents, which use natural language commands to interface.

[0060] Referring to Fig. 3, the camera head 100 comprises a camera chamber 108 and an inspection processor chamber 109. As described in more detail below, these chambers are provided so that there is minimal thermal expansion or contraction of supports for the camera sensor and lens, and that the inspection processor can be of a sufficiently powerful Al capacity for local inspection processing within the head 100, despite the fact that significant heat may be generated.

[0061] At a front side of the housing 101 and within the camera chamber 108 there is a camera sensor 120 facing outwardly on an optical axis A. The camera sensor 120 is on a substrate 121 which is supported on a lower end on the base 103 and at the upper end by a horizontal support 122. The latter acts as a spacer between the board 121 and a chamber wall 123 which is orthogonal to the optical axis. The wall 123 acts as a dividing wall to separate the inspection circuits (on the left) from the camera components on the right. The wall 123 forms the camera chamber 108 which is thermally isolated from the inspection circuits on the opposed side of the wall 123 (on the left as viewed in Fig. 3). The support / spacer 122 also provides a wired link to the inspection circuits. The wall 123 is of FR4 material. Within the camera chamber 108 the camera sensor 120 faces outwardly (distally) on the optical axis A, and the camera board 121 is supported on the distal side by an annular ceramic-based support 124. The ceramic support 124 extends around to encompass any or all of a full 360° circumference around the optical axis. An advantageous aspect of the support 124 is that it has low thermal expansion, and so have little or no expansion in response to heat generated within the inspection processor chamber 109. Also, around the optical axis is an optical tube 125 or collar which supports a lens 126. The fact that the camera sensor 120 is within the isolated chamber 108 and that it is supported on the distal side by the ceramic support 124 means that there is very little variation with temperature change in position of the sensor 120 along the optical axis, the distance along this axis between the sensor 120 and the lens 126 remaining constant.

[0062] Within the enclosure on the side opposed to the chamber wall 123 here is a carrier circuit board 130 linked to a USB interface 131. The board 130 supports inspection hardware components 132 facing towards the base 103. The components 132 consume very little power and do not generate any significant heat.

[0063] In one example the hardware components include the following:

[0064] Camera 120: 2MP - 12MP Colour CMOS Sensor

[0065] Lens 126: 12 mm Electrofocus Lens

[0066] Lighting (not shown, distally of the lens 126): White LEDs, diffuse ring lighting, continuous mode

[0067] Inspection Processor: ARM Processor with GPU System-on-Module.

[0068] Al Framework: Transformers™

[0069] Robotic controller 25: Universal Robots UR5e™

[0070] Control Software: Visible™

[0071] The GPU is integrated into a system module 140, which is supported above the carrier board 130 by support spacers 133 and 135 and a leaf spring 134. The supports 133 and 134 support the module 140 spaced apart from the board 130 by a space 136. There is a path from the module 140 to the housing fins 106, to ensure fast dissipation of heat from the module 140 without need for a fan. This path has the following interfaces:

[0072] 143. Heat transfer paste. This is non conductive electrically, and is easily applied to the to of the module 140 during assembly.

[0073] 144. Heat transfer metal plate, having a thickness of 4.8 mm and being of Aluminium. 145. Thin film or foil, of graphite and having a thickness of 50 pm, and having a thermal conductivity of 20 W / mK. This foil is easily applied during assembly, and provides for very effective heat spreading towards the casing 101 and hence the fins 106 and 107.

[0074] The head 100 is therefore constructed to perform the inspection processing locally as a complete unit, thereby allowing for a particularly fast response time to inspection events. In some examples the system has been tested for inspection of packages passing at a rate of 4 parts per second, and the response time for detection of a fault was of the order of 0.5 to 2.0 seconds. The processing is preformed by an Al-capable module within the head 100, and heat which is generated is conducted efficiently to the space above the head via the above path. This is away from the camera chamber, and so it does not affect the relative positions of the sensor 120 and the lens 128, and there is no need for a fan, which would cause unwanted vibration. The heat transfer arrangement allows efficient operation of the GPU without need for a link to a separate Al machine. All runtime processing can be done locally within the camera head 100, providing an excellent real time performance. This is very important, for example, for inspection points on a fast moving manufacturing production line. The heat transfer arrangement avoids need for a fan, thereby avoiding vibration, which might cause inaccuracy in image acquisition.

[0075] Moreover, as described in detail below, the processor 25 can be used solely for user interfacing and robot control. In its robot control operations, it can interface with the head 100, so that positioning of the head 100 is optimised for image capture for inspection, and the route to this position can be controlled to avoid obstacles.

[0076] The inspection processor module 140 executes algorithms for the visual aspect of inspection of various regions of the object (including data labels or physical components). Operation of the processors is driven by language prompts instead of or in addition to customised neural networks or specifically tuned algorithms.

[0077] Robotic inspection instructions are driven by voice using language explanations, which allows for a significant reduction in operator expertise required to configure the system.

[0078] The module 140 has an initial Al training phase, and then a startup phase in which it stores settings for each of a plurality of inspection points. It then has a runtime mode in which it performs real time inspection processing. The inspection processor is 140 programmed with Artificial Intelligence software for the analysis of images, and the format of what is outputted is governed by user’s prompts which are provided in text boxes or by audio. The use of language-driven algorithms eliminates the need for collection of large datasets for the specific tuning of the algorithm and the significant energy consumption required for training new neural networks. It is envisaged that the system allows increased speed of deployment of new inspections on the factory floor by a factor of up to 10 times. The Artificial Intelligence software is under the Transformers framework.

[0079] The inspection processor 140 is trained with artificial intelligence code according to a visionlanguage model which is based on an artificial intelligence model which accepts and processes a tokenised text string as an input. The model includes an image feature extraction tool, so that during an initial model training phase, the image feature extraction tool is trained for feature extraction under various scales, brightness, contrast, and colour configurations. Also, the model is configured to, during training, combine tokenised real language description strings derived from these prompts with corresponding extracted feature maps for given robot positions and camera settings.

[0080] In a setup phase the robot controller 25 is configured to move the robot 15 so that the head 100 is at pre-set camera positions each for a scene and being associated with an inspection point. The inspection processor 140 controls image acquisition and scene illumination for each inspection point, captures setup data for a pre-set head 100 position including presets camera settings, preset illumination parameters, and a natural language prompt, and stores the setup data;

[0081] During runtime, for each inspection point the inspection processor 140 and the robot controller 25: cause the head 100 to move to the pre-set position for the inspection point when an object is in the camera scene, activate the camera illuminator head to illuminate the scene and capture an image of the object, execute the artificial intelligence code to analyse the image and determine results for the scene, generate an output for the inspection point in a format determined by the prompt associated with the inspection point during setup.

[0082] In more detail, the processors 140 connect to the camera 120 and the robot 15 controller 25. A Vision Language model runs on its own service on the module 140 once the system starts. The user creates an Inspection Plan which comprises a number of Inspection Points, and each Inspection Point has a corresponding robotically controlled head 100 position, camera image acquisition parameters, and image processing parameters (in this case, a language prompt). Each Inspection Point is defined using an Inspection Point Wizard:

[0083] Move robot by hand to relevant position while reviewing the live camera image.

[0084] Optimise position and acquisition parameters.

[0085] Define Language Prompt and test.

[0086] The system then runs the Inspection Plan, and the images are collected and correlated to results returned by the processing algorithm - at which point the images and results can be stored to a database for later review by the user.

[0087] The vision-language model is based on a Transformer architecture, where the inputs to the model are:

[0088] - A tokenised text string

[0089] - An image

[0090] The architecture of the model includes an image feature extraction layer, similar to conventional feature extraction layers used in convolutional neural networks. During initial model training, the image feature extraction tool is trained on a varied dataset to ensure robust feature extraction under various scales, brightness, contrast, and colour configurations. During training of the vision language model, the model is trained by combining tokenised real language description strings with corresponding extracted feature maps for given images. During run-time, the output of the transformer architecture is then un-tokenised, resulting in real language output. This allows for operators to describe in real language what kinds of features / objects / defects they are looking for in an image, instead of collecting tens to hundreds of examples of defects and training a specific neural network for each complex inspection.

[0091] In another embodiment the system utilises a set of recorded images while the robotic positioning system is in motion. This allows for a variety of images to be collected in varying lighting conditions, which is useful for certain types of defects that may not be visualised in a single image. An example of this could be detection of a shallow scratch on a polished reflective convex surface, in which case the scratch may only be visible to the eye / camera as a small dynamic change in surface condition, as opposed to a single obvious defect. The temporal nature of the analysis allows for the detection of defects not possible even with multiple single images, as the image would also need to be compared to each other to evaluate the defect, which is complex to configure for each inspection type. The approach of using vision language models with a vision transformer type architecture, allows the model to a) analyse the context of single images by correlating scene to real descriptive language, and to b) determine defects on manufactured products by utilising knowledge of temporal context of a series of frames.

[0092] The processors 140 optimise Vision Language Models which are co-ordinated with the robotic positioning system 15, which can intelligently interpret the data in the image, reducing setup time from days / months to minutes. The system uses an optimised vision language model running on a robotic positioning system with integrated (Graphics Processing Unit) GPU processing capability. By utilising an Al model trained on a variety of image and text data, the system can perform robust OCR recognition and feature recognition on the acquired images of the labels and assemblies, producing a refined result from the language model as a result of optimising the prompt.

[0093] By utilising the on-board GPU, it can perform the convolutional calculations at high speed, allowing the system to achieve an inference rate in single digit seconds, allowing the operator to receive real time result feedback as to contents of the relevant information on the product present in the field of view of the camera. As the Al model is trained on a variety of scales, the model can detect text sizes commonly used in medical device, pharmaceutical, and food manufacturing, giving the system high flexibility for a computer vision-based system, and can also extract and correlate features in the images to trained text prompts, allowing the operator to describe in normal language what type of defect in the image they are looking to inspect for.

[0094] Fig. 4 illustrates an example use case for label reading. This shows examples of prompts that the user has designed and entered into the left-hand text box for each image. The system acquires the image and generates an output as shown in the right-hand text box. The user validates that the output works across variety of products and tweak prompt as needed.

[0095] Fig. 5 illustrates a use case for anomaly detection. In this case the Inspection Points are for images of a bolt in a slot or a through hole. These examples illustrate that the user can specify an output in a full sentence or as a simple Yes / No format. Fig. 6 is another set of examples, in this case for inspection of labels for which the Al software determines which of the dates is the expiry date. It then generates an answer according to the prompt provided by the user. Setup using method of the invention (for surface conditions, position and layouts - setup time <5 mins for each Inspection Point):

[0096] (a) Design prompt and enter prompt into text box. As shown in any of the left-hand text box columns of Figs. 4 to 6.

[0097] (b) The robot 15 is moved by the user to the best position, and the camera and illuminator parameters are adjusted for optimum image capture for each Inspection Point. The system acquires an image.

[0098] (c) The inspection processor 140 generates an output and populates the right-hand box with an answer according to the format requested. The user provide feedback according to the output to validate the output.

[0099] (d) Repeat (a) to (c) for each Inspection Point.

[0100] Referring to Fig. 7, in more detail in a method 200 implemented by the system the processors, perform the following steps:

[0101] 201, start-up a setup wizard to guide the user through providing the necessary inputs.

[0102] 202, positioning of the head 100 for optimum capture of images from an object on the bed 3. The bed 3 may for example be a temporary position of a conveyor belt which will be kept stationary for the purposes of inspection.

[0103] 203, optimization of the camera parameters, including all image acquisition parameters such as acquisition time and illumination intensity and wavelength / colour. Movement / adjustment of the robot 15 may be entirely manual or it may be controlled by a user interface.

[0104] 204, selection of the inspection type, of which two examples are shown: 210-213 for component inspection and 220-223 for label inspection.

[0105] Component Inspection

[0106] 210, the wizard prompts the user to check a sample component, to ensure that it is the correct type.

[0107] 211, the user enters the prompt, and examples of these are given in the left-hand columns of Figs. 4 to 6. In other examples the prompts are provided in audio format.

[0108] 212, the system, using the physical and optical parameters set in steps 102 and 103 captures an image and the executes the Al soft ware to analyse it.

[0109] 213, the system provides an output, and the user validates that it is correct.

[0110] 220-223, these steps are akin to the steps 210-213 except that in this case the object under inspection is a label, and the examples of Figs. 4 to 6 also apply. At this stage the system has been set up for either or both of a particular type of component and label, for at least one Inspection Point. For each Inspection Point the system has ben set to move the robot to the desired position to capture an image with the desired image capture parameters, to analyse the image and provide an output in the format instructed by the user prompt.

[0111] 250, this is the first step of a sequence for runtime operation, and the step is initialization for inspection for a component or label.

[0112] 251, start of inspection, retrieving the settings made in the steps preceding step 250.

[0113] 252, automatic movement of the robot to the desired position.

[0114] 253, capture of an image according to the Inspection Point optical parameters. Al analysis of this image.

[0115] 254, Logging of the results.

[0116] 255, repeating steps 252 to 254 for each inspection Point for the component or label.

[0117] 156, generating an inspection report with all desired answers, examples of which are set out in the right-hand columns of Figs. 4 to 6.

[0118] The processors 25 and 140 both include GPU processors which have the capacity to execute the software without need for cloud computing. By executing on an edge device at the location such as a factory it can easily interface with other systems such as factory systems. This allows a response time within 10 seconds, typically within 3 seconds. Because the setup includes physical movement of the camera head 100 to the optimum position and also optimisation of the camera and illumination parameters the processors are likely to correctly interpret the image data for each Inspection Point. Also, the output data will always be in the required format.

[0119] The general approach may be summarised as:

[0120] Acquire images of labels.

[0121] Enter a prompt for the required output.

[0122] Test the model and verify output.

[0123] Modify the prompt and verify it modifies the output in a desired manner.

[0124] Test across a variety of the acquired images.

[0125] The following provides additional detail about aspects of the system 1.

[0126] Anodised Surfaces for Improved Passive Heat Transfer

[0127] The outer surface of the head 100 is hard anodised with a layer 105 with a thickness in the range of 0.02 to 0.05 mm. The layer 105 is dyed with a dark colour giving an appearance of low reflectivity (< 30% reflectivity in the visible light spectrum). The function of the surface treatment is to provide a) a hard-wearing decorative surface for industrial use, and b) to increase the radiative heat emission rate of the enclosure to the surrounding air in all directions (the emissivity of the enclosure is improved by 30% compared to bare aluminium). The radiative heat transfer capacity of the enclosure allows for up to 30 W of continuous power consumption of the enclosure, not exceeding the general industrial safety limits of < 65 °C - which allows the system to operate with high computer performance with only passive cooling (assuming an ambient environmental temperature of < 25 °C). An advantage of passive cooling is that no moving parts are required for cooling, thus eliminating a source of vibration, improving image quality especially in high- magnification applications, and reducing the noise levels produced by the device which contribute to environmental noise and operator fatigue.

[0128] Camera sensor to Lens Spacing

[0129] The mounting of the camera sensor 120 and the lens 126 as described above provides a back focal length (BFL) of 17.526 mm. Camera-to-camera repeatability is important in industrial applications where cameras may need to be replaced over time, and it is important to ensure the exact same magnification and focus at the same lens configuration. This is important for high performance and sensitive Al detections, such as micro-particle detections where the defect size is of the order of < 5 pixels in diameter. The BFL (sometimes referred to as the Flange Focal Distance (FFD)) is precisely controlled by: i) the metal-ceramic composite material of the mounts 124, in one example Aluminium composited with Zirconium Tungstate, exhibiting a negative-to-low co-efficient of thermal expansion (CTE) between -2 pm / m / K and 2 pm / m / K, or Al-SiC between 6 pm / m / K and 8 pm / m / K, and / or ii) pre-compensating the component positioning to allow for thermal expansion at a nominal operating temperature of 50 °C, and / or iii) deliberately under sizing the mount during manufacturing such that its final BFL is determined via the addition of stainless steel shims as to achieve the correct BFL.

[0130] Camera to Robot Agentic Communication

[0131] In some embodiments, the module 140 and the processor 25 act as separate but interfacing agents to achieve dynamic navigation for inspection applications.

[0132] For the inspection it is necessary to control the robot pose, camera acquisition parameters, and image processing parameters (which determine if that inspection point is a Pass / Fail). However, a major challenge when inspecting large assemblies is that the camera head must navigate around objects protruding from assemblies to avoid a collision, which reduces throughput and risk of equipment / part damage. While it is possible to pre-plan robot trajectories from virtual environments (such as a robotic path planning system in simulation tools with CAD models of the assembly), this approach does not account for certain types of components on typical assemblies, such as wires, wiring harnesses, hoses, and other flexible components. As the camera head 100 runs Al models that can accept natural language prompts, it allows for a flexible and inter-operable communication concept between the robot and camera systems, which can still occur within existing traditional communication protocols (such as TCP / IP sockets or OPC / UA nodes).

[0133] An example use case: as part of an inspection plan the system must inspect a connector at the end of a hose - and in this scenario the hose is visible, but the end of the hose is outside of the current camera field of view. The robot agent prompts the camera agent (as programmed): “is the connector present on the hose?” - to which the camera agent would respond: “I can see the hose, but the end of the hose appears to be outside the field of view to the bottom right”. This information can then be interpreted by the robot agent 25, which moves the camera head 100 to the bottom right direction as to bring the end of the hose into view (and hence the inspection can be performed). In this way, the natural language prompt allows the two systems to interact in a highly flexible but intelligent way, without extensive integration (as would be required in existing industrial systems).

[0134] In another scenario, the camera head 100 can communicate with multiple other agentic systems via natural language, where multiple cameras can communicate together with multiple camera views to achieve a rational consensus. An example of this is in the inspection of highly reflective surfaces, which are often challenging in factory settings due to: a) background objects can be visible in the reflected image, and a system would need to determine what is a background reflection and what is a surface structure image, and b) minor defects on otherwise highly polished surfaces are often only visible over what is considered a “temporal visual inspection”.

[0135] In this type of temporal visual inspection, the camera heads are moved relative to an object, a variety of images captured while moving the part, and a rational consensus formed by analysing each image, analysing the difference between subsequent images in the sequence, and rational determination of which changes belong to the background / environment and which belong to the part surface, and then a final rational determination of if the surface flaw detected is of the correct significance for part Fail classification. This approach allows for improved sensitivity in detecting very minor surface defects (for example, a < 5 pm wide and deep scratch on a highly polished surface) and the natural language element of agentic communication allows for a flexibly and thorough “discussion” of the analysis in question. It also allows for simple human interpretation of the rationale of the robotic inspection system, which is a significant difficulty in the validation of inspection systems in regulated environments and applications - where traditional Al systems are treated as an unknown black box that can only be proven by extensive empirical experimentation.

[0136] A limitation of running LLM and VLM models on edge hardware is that limited GPU memory may be available to perform an analysis. To return a result in a timely manner it is highly important that the whole model is loaded in memory. However, during runtime execution, the memory requirements are highly dependent on the size and geometry of the input image. For example, passing an image with a 1920 x 1080 in 24-bit colour would require memory in excess of 20 GB. In the invention of some preferred embodiments, we have developed a method for dynamically adjusting the image input dimensions to optimise the resolution of the image within the available memory. To achieve this, the following approach was implemented.

[0137] Where the object in question is an image sub region, the sub-region is cropped and the image passed to the model with the prompt for inference. An example of this: in an image of a box on a conveyor, the user wants to read the Use by Date printed on a label on the box. The inference is performed in 3 steps: i) Initial whole image is downscaled to 300 x 200 pixels with a prompt to return the label position. The data returned by inference #1 is a JSON of bounding box co-ordinates. This input image is too low resolution after downscaling to read the Use by Date reliably, but is sufficient to determine the label position and runs inference quickly within the available memory. ii) The bounding box co-ordinates are transformed to map to the original image scale. ii) The original full resolution image is cropped using the bounding box co-ordinates returned in inference #1, and passed to the model for inference #2. As this image region is at the full original resolution, the Use by Date can be read reliably by Inference #2, and as the image was cropped, the inference was able to run in the available memory. This approach can be further optimised depending on the model size by: i) quantising the model from 32bit to 4-bit (depending on the model requirements), and ii) changing the image colour space from colour (24-bit) to monochrome (8-bit), depending on the application.

[0138] Another embodiment of the invention uses for the setup phase, to optimise the prompts sophisticated cloud-based model (with upwards of 70 B parameters) as an expert agent for defining the appropriate prompts to achieve the desired output result from Vision Language model running within the system. This allows the user to automate the generation of the precise prompts by entering a more general description of the product defect into the cloud LLM model. The cloudbased model can then recursively refine the edge model prompt until the desired results are achieved.

[0139] It will be appreciated that the invention greatly simplifies and speeds up training for machine vision inspection, as described above. This is much more effective than the prior well-known approach involving at least one month of data collection and analysis and ongoing maintenance. This prior approach typically involved acquiring an image, finding the location of the print area, and finding the location of the print in the print area. Assuming consistent text size and spacing, placing an OCR region tool over the middle line of text (expiry date). This prior approach typically then involved training the OCR tool on the printed text (at least 20 examples of each digit in each position). These steps must be repeated for each format, for example 25-MAR-2023 or 25 / 03 / 23. The use of language-driven algorithms eliminates the need for collection of large datasets for the specific tuning of the algorithm and the significant energy consumption required for training new neural networks. It is envisaged that the system allows increased speed of deployment of new inspections on the factory floor by a factor of up to 10 times.

[0140] The invention has the advantage of processing on the “edge” in the factory, reducing the need for the customer to send image data of their products to the cloud (which is an IP and cybersecurity concern).

[0141] The architecture of having the inspection processor in the camera head 100 provides major advantages. The immediate proximity to the camera sensor allows guaranteed latency, and avoidance of need to stream to a PC-type processor in the cabinet or elsewhere locally. Also, this architecture allows for modular scalability by providing multiple camera heads, and also interinspection processor communication, which may be beneficial for certain manufacturing situations. Via the addition of Wi-Fi capability through the addition of a USB dongle on the inspection processor, it is possible to avoid the need for signal cables routed through the robot arm or along the robot arm and the attendant risk that they twist excessively and become damaged. Also, power can easily be provided from the ends of a robot arm via common robot end-effector power supply ports.

[0142] Components of embodiments can be employed in other embodiments in a manner as would be understood by a person of ordinary skill in the art. The invention is not limited to the embodiments described but may be varied in construction and detail.

Claims

Claims1. A machine vision inspection system (1) comprising: a camera head (100) mounted on a robot (15), and comprising a camera and an illuminator; a robot controller with digital data processors (25), linked with the robot (15); an inspection processor comprising image capture circuits (132) and an Al processor (140); wherein the robot controller is configured to move the robot to pre-set camera positions each for a scene and being associated with an inspection point, and the inspection processor is configured to control image acquisition and scene illumination for each inspection point, wherein the inspection processor is trained with artificial intelligence code according to a vision-language model, and in which: the vision-language model is based on an artificial intelligence model which accepts and processes a tokenised text string as an input, the model includes an image feature extraction tool, so that during an initial model training phase, the image feature extraction tool is trained for feature extraction under various scales, brightness, contrast, and colour configurations, and the model is configured to, during training, combine tokenised real language description strings derived from said prompts with corresponding extracted feature maps for given robot positions and camera settings, images; wherein the robot controller and the inspection processor are configured to, for each of a plurality of inspection points: during setup, capture setup data for a pre-set robot position, preset camera settings, preset illumination parameters, and a natural language prompt for each of the plurality of inspection points, and storing said setup data; during runtime, for each inspection point: cause the robot (15) to move to the pre-set position for the inspection point when an object is in the camera scene (3), activate the camera illuminator head to illuminate the scene and capture an image of the object, execute the artificial intelligence code to analyse the image and determine results for the scene, generate an output for the inspection point in a format determined by the prompt associated with the inspection point during setup.

2. A machine vision inspection system as claimed in claim 1, wherein the artificial intelligence software is configured to, during run-time, generate an output which is un- tokenised according to the prompt received during training, resulting in real language output.

3. A machine vision inspection system as claimed in claim 1 or claim 2, wherein the processors are configured to cause a plurality of images to be captured by the camera while the robot is in motion to the pre-set position for the current inspection point, and to process said images to provide temporal inspection for the inspection point.

4. A machine vision inspection system as claimed in claim 3, wherein the processors are configured to cause said images to be captured with different scene illumination conditions.

5. A machine vision inspection system as claimed in any preceding claim, wherein said temporal inspection include (a) analysis of context of single images by correlating scene to real descriptive language, and (b) determination of defects on objects by utilising knowledge of temporal context of a series of images.

6. A machine vision inspection system as claimed in any preceding claim, wherein said data processors include a graphics processing unit, GPU, configured to perform convolutional calculations at a speed allowing the system to achieve an inference rate in less than 10 seconds.

7. A machine vision inspection system as claimed in any preceding claim, wherein the processors are configured to operate offline, at the edge of a manufacturing line inspection station.

8. A machine vision inspection system as claimed in any preceding claim, wherein the processors are configured to communicate with a cloud-based expert agent model during training for defining prompts to allows the user to automate the generation of the training prompts by entering a more general description of an object defect and to receive from the cloud-based model a recursively refined edge model prompt until desired results are achieved.

9. A machine vision inspection system as claimed in any preceding claim, wherein the inspection processor (132, 140) is located in a camera head housing (101) together with a camera sensor (120) and lens (126), and the robot controller comprises a separate robot processor (25) linked with the inspection processor.

10. A machine vision inspection system as claimed in claim 9, wherein said camera head (100) comprises a housing (101) which has a camera chamber (108) comprising the camera sensor (120) and the lens (126) mounted along an optical axis (A), and said camera chamber is separated by a wall (123) from a processor chamber (109) containing circuits (132, 140) including an Al processor.

11. A machine vision inspection system as claimed in claim 10, wherein the camera sensor (120) is mounted to a substrate (121) which is in turn mounted by support (124) of a material including a ceramic.

12. A machine vision inspection system as claimed in claim 10 or claim 11, wherein the camera sensor support (121) is mounted within the camera chamber (108) in a configuration spaced apart from the chamber wall (123) separating the camera chamber (108) from the processor chamber (109).

13. A machine vision inspection system as claimed in any of claims 10 to 12, wherein the processor chamber comprises a carrier board (130) supporting electronic components on a lower side facing a housing base (103) and on an opposed side supporting spacers (133, 134, 135) which support a system-on-module (140) which includes a GPU for executing the inspection processing Al model software, and said system-on-module is linked by a series of thermally conducting components to a housing top wall (102) which has external fins (106, 107).

14. A machine vision inspection system as claimed in claim 13, wherein the series of heat transfer components include a heat transfer block (144) which interfaces with system-on- module via a thermally conductive paste (143) and interfaces with the housing via a thermally conductive foil (145).

15. A machine vision inspection system as claimed in claim 14, wherein the foil is of graphite material.

16. A machine vision inspection system as claimed in any of claims 13 to 15, wherein the system-on-module (140) is pressed against the series of heat conducting components by a leaf spring (134) acting between the carrier board (130) and the system-on-module.

17. A machine vision inspection system as claimed in any of claims 10 to 16, wherein the camera head housing (101) has an anodised coating (105) on some or all of its external surfaces.

18. A machine vision inspection system as claimed in claim 17, wherein the coating (105) has a thickness in the range of 0.02 mm to 0.05 mm.

19. A machine vision inspection system as claimed in claim 18, wherein the coating (105) has a reflectivity of less than 30% reflectivity in the visible light spectrum.

20. A machine vision inspection system as claimed in any of claims 10 to 19, wherein the camera substrate mount comprises Aluminium composited with Zirconium Tungstate, with a negative-to-low co-efficient of thermal expansion (CTE) between -2 pm / m / K and 2 pm / m / K.

21. A machine vision inspection system as claimed in any of claims 10 to 20, wherein the inspection processor and the robot controller are configured to act as agents which perform agentic communication with each other.

22. A machine vision inspection system as claimed in claim 21, wherein the inspection processor and the robot processor are configured to communicate with natural language prompts.

23. A machine vision inspection system as claimed in claim 22, wherein the robot controller is configured to cause camera head movement to intelligently bring a previously omitted part of an object into a visible scene, in response to a prompt from the robot controller to the inspection processor asking if a particular expected part of an object under inspection has been visible.

24. A machine vision inspection system as claimed in claim any of claims 10 to 23, wherein the inspection processor is configured to communicate by natural language prompts to at least one other inspection processor.

25. A machine vision inspection system as claimed in claim 24, wherein the inspection processor is configured to communicate to provide data concerning a particular view and to process corresponding data concerning a different view from the other inspection processor, and to generate an output which allows for omission of background objects, or which indicates a defect on a highly polished surface which is not discernible from all views of the object.

26. A machine vision inspection system as claimed in claim 25, wherein the inspection processor is configured to include, in its output data, data which is susceptible to human interpretation of the rationale of the robot processor.

27. A machine vision inspection system as claimed in any of claims 10 to 26, wherein the inspection processor is configured to dynamically adjust image input dimensions to optimise resolution of the image within available memory, by: downscaling an initial whole image according to a prompt to return a subregion position to provide a first inference as a JSON of bounding box co-ordinates, transforming the bounding box co-ordinates to map to the original image scale, and cropping the original full resolution image using the bounding box co-ordinates returned in the first inference, and passing to the model for a second inference, for a second inference, reading at the full original resolution to determine an output such as a Use By date on a label.

28. A machine vision inspection system as claimed in claim 27, wherein the inspection processor is configured to return data as a JSON bounding box co-ordinates and to transform the bounding box co-ordinates to map to the original image scale.

29. A machine vision inspection system as claimed in claim 28, wherein the inspection processor is configured to pass a cropped image to the model for a second inference at full resolution.

30. A machine vision inspection system as claimed in any of claims 10 to 29, wherein the inspection processor is configured to quantize a model to a lower bit size to reduce memory and processing requirements.