Visual classification work system and electronic device

By using the automatic segmentation and dynamic scheduling model of the visual classification operation system, the problem of rapid new product launch and efficient identification in semi-automated factories has been solved, achieving high-accuracy product counting and classification, and reducing cost and time requirements.

CN121170548BActive Publication Date: 2026-05-05GUANGZHOU QIYIN INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU QIYIN INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2025-09-28
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In semi-automated factories, existing AI visual inspection technologies are struggling to adapt quickly to the demands of diverse product types and rapid iteration cycles. Traditional methods suffer from low accuracy, high cost, and insufficient stability, especially during new product launches and iterations, failing to meet the high precision requirements of industrial production.

Method used

A visual classification system is adopted, including a task subsystem, a model training subsystem, and a client operation interface. It utilizes the SAM2 segmentation model and a comprehensive similarity algorithm, combined with target tracking and Bayes' theorem, to achieve automatic product and background segmentation and rapid sample collection. The AI ​​model is dynamically scheduled through a scheduling module to enable rapid launch and efficient operation of new products.

Benefits of technology

It can count and classify new products within 15 minutes with an accuracy rate of over 99%, reducing the cost of manual operation and hardware modification, and improving the intelligent monitoring capability of the production line.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121170548B_ABST
    Figure CN121170548B_ABST
Patent Text Reader

Abstract

This invention provides a visual classification operation system and electronic device, including: an operation subsystem, a model training subsystem, and a client operation interface; the operation subsystem includes: a first operation module, a second operation module, and a scheduling module; the first operation module is used to encode and decode single-frame images of video and output product images; the second operation module is used to input single-frame images of video into a trained detection model and output product images; the scheduling module is used to schedule the use of the first and second operation modules according to the training maturity of the current production products; the model training subsystem is used to connect to a message server after startup and train the model of the visual classification operation system through message events. Without changing the hardware, it achieves the ability to count and classify products when switching to new product production within 15 minutes, with an accuracy rate of over 99%, laying a solid foundation for intelligent monitoring of the production process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence (AI) technology, and in particular to a visual classification system and electronic device. Background Technology

[0002] AI-powered industrial visual inspection (AIVI) is an innovative solution that utilizes artificial intelligence technologies, especially deep learning and computer vision, to perform high-precision, high-speed quality inspection of products on the production line. By mimicking the human visual system, it automatically identifies and analyzes key characteristics of products such as size, shape, color, and defects, significantly improving the efficiency of quality control in manufacturing.

[0003] On the technical front, continuous optimization of deep learning and machine learning algorithms has significantly improved AI's performance in areas such as image recognition and defect detection, saving substantial human resource costs. While the industry faces challenges such as data security issues and rapid technological updates, the integration of new technologies like 5G and the Internet of Things presents a very promising future for AI-powered industrial vision inspection services.

[0004] However, focusing on semi-automated factories, these factories remain the mainstay of low-cost production, with manual labor accounting for a large proportion. The randomness and variety of product placement pose significant challenges to online counting, sorting, and inspection. Traditional manual and sensor-based inspection methods have obvious drawbacks: manual inspection relies heavily on manpower, is costly, and suffers from low accuracy due to subjective judgment and fatigue; sensor-based inspection has strict requirements regarding product placement and angle, making it difficult to adapt to complex production scenarios. These problems severely hinder lean management and digital upgrades of production lines, leading to significant errors in factory labor cost accounting and inflated costs. For example, an electronics factory with multiple semi-automated production lines mainly producing various electronic components used manual reporting for output value calculation due to the wide variety and random placement of products. In the first quarter of 2024, the actual output value was 300,000 yuan less than the manually reported value, exposing the inaccuracies of manual inspection in counting and accounting, resulting in direct economic losses for the company.

[0005] Under the stringent requirement of 99% accuracy in industrial production, a significant portion of existing AI vision companies still rely on the outdated approach of manually collecting samples, manually labeling, and manually training and debugging. Some employ unsupervised or weakly supervised learning methods, but these suffer from insufficient stability due to uncontrollable output, failing to meet the demands of diverse product types and rapid iteration cycles. In these semi-automated factories, to survive and adapt to market demands, the products produced are varied and numerous. Some production lines even switch between several products daily, and dozens or even hundreds of products monthly. Some products are even produced only once, placing extremely high demands on iteration synchronization, which traditional vision training cycles can no longer satisfy. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide a visual classification system and electronic device to enable industrial production to quickly go online in the early stages of new product counting, classification and identification, and to operate efficiently in the later stages.

[0007] In a first aspect, embodiments of the present invention provide a visual classification operation system, comprising: an operation subsystem, a model training subsystem, and a client operation interface; the operation subsystem is used to count, classify, detect, and identify products on the production line; the operation subsystem includes: a first operation module, a second operation module, and a scheduling module; the first operation module is used to encode and decode single-frame images of video and output product images; the second operation module is used to input single-frame images of video into a trained detection model and output product images; the scheduling module is used to schedule the use of the first and second operation modules according to the training maturity of the current production products; the model training subsystem is used to connect to a message server after startup and perform model training for the visual classification operation system through message events; the client operation interface is used to establish real-time communication with the operation subsystem and display the output information of the operation subsystem; the client operation interface is also used to stop the video frame of the current frame and edit data based on the user's operation on the video frame of the current frame.

[0008] In optional embodiments of this application, the first task module is further configured to output background images, product information, and task data; wherein, the product information includes: segmentation rectangle, coordinates, directional bounding box information, segmentation pixel area, video frame number, and confidence level; if manual operation is involved, the first task module is further configured to output single-frame images of the video and product segmentation information; the second task module is further configured to output product information and task data.

[0009] In optional embodiments of this application, if the current product being manufactured is a new product, the scheduling module is used to use the first job module; if the current product being manufactured follows a preset job model, the scheduling module is used to use the second job module; if the second job module has an output anomaly during the use of the second job module, the first job module is used to process the video single frame image with the output anomaly, and the scheduling module is used to process the output results of the first job module and the second job module; wherein, the processing methods include: collecting sample datasets, anomaly reports, and transferring to manual verification.

[0010] In an optional embodiment of this application, the first operation is equipped with a segmentation model; the segmentation model is used to encode and decode single-frame images of real-time video from the production line inspection point at a fixed frame rate, and segment the single-frame images of the video according to the target size; the first operation module is also used to determine whether the segmentation result is a product image or a background image, and to rotate and perform secondary segmentation on the segmentation result in sequence; the first operation module is also used to process the result of the secondary segmentation display through a classification model; wherein, the processing method includes: error detection, classification and counting; the first operation module is also used to collect product images and product information of product images, and output the initial judgment result of the single-frame image of the video.

[0011] In an optional embodiment of this application, if the product currently being produced is a new product, the first operation module is used to notify the operator to perform a product segmentation and selection operation for the product currently being produced.

[0012] In an optional embodiment of this application, the segmentation model is used to segment multiple rectangles of a single frame image of the output video using a coordinate dot matrix, and the first task module is used to determine whether each rectangle is a product image or a background image by a similarity weighted judgment method; wherein, the similarity weighted judgment method includes: calculating multiple similarities through multiple similarity calculation methods and performing a weighted judgment on the multiple similarities.

[0013] In optional embodiments of this application, multiple similarity calculation methods include: target similarity calculation method and Hamming similarity calculation method; wherein, the target similarity calculation method is used to calculate the similarity between the first value and the second value using the following formula: s=|ab| / (|b|+ε), f=1-min(s,1); where s is the intermediate value of the target similarity calculation method, a is the first value, b is the second value, ε is the preset minimum representable value, f is the similarity of the target similarity calculation method, and min(s,1) is the smaller value between s and 1.

[0014] In an optional embodiment of this application, the model training subsystem is used to use the product image and background image output by the first task module as a training set, and to train the model of the visual classification task system based on the training set.

[0015] In an optional embodiment of this application, the model of the visual classification task system includes: a classification model of the first task module and a detection model of the second task module.

[0016] Secondly, embodiments of the present invention also provide an electronic device, which includes the aforementioned visual classification system.

[0017] The embodiments of the present invention bring the following beneficial effects:

[0018] This invention provides a visual classification system and electronic device that can be applied to pallet and belt production lines. Without changing the hardware, it enables the counting and classification of products for switching to new products within 15 minutes, with an accuracy rate of over 99%, laying a solid foundation for intelligent monitoring of the production process.

[0019] Other features and advantages of this disclosure will be set forth in the following description, or some features and advantages may be inferred from the description or determined without doubt, or may be learned by practicing the techniques described above.

[0020] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0021] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of the structure of a visual classification system provided in an embodiment of the present invention;

[0023] Figure 2 A schematic diagram illustrating the overall process of a visual classification system provided in an embodiment of the present invention;

[0024] Figure 3 This is a schematic diagram of the main process of SAM2 segmentation provided in an embodiment of the present invention;

[0025] Figure 4 This is a schematic diagram illustrating a product segmentation and selection operation process provided in an embodiment of the present invention;

[0026] Figure 5 This invention provides a schematic diagram of a process for determining and classifying outputs based on confidence levels and an AI model, as provided in an embodiment of the invention.

[0027] Figure 6 A schematic diagram of Python code provided for an embodiment of the present invention;

[0028] Figure 7 A schematic diagram of a rectangular marker provided in an embodiment of the present invention;

[0029] Figure 8 A schematic diagram of a secondary segmentation result provided in an embodiment of the present invention;

[0030] Figure 9 A schematic diagram of the first part of a secondary segmentation algorithm code provided in an embodiment of the present invention;

[0031] Figure 10 A schematic diagram of the second part of a secondary segmentation algorithm code provided in an embodiment of the present invention;

[0032] Figure 11 A schematic diagram of the third part of a secondary segmentation algorithm code provided in an embodiment of the present invention;

[0033] Figure 12 A schematic diagram of the first part of a comprehensive similarity algorithm code provided in an embodiment of the present invention;

[0034] Figure 13 A schematic diagram of the second part of a comprehensive similarity algorithm code provided in an embodiment of the present invention;

[0035] Figure 14 A schematic diagram of the third part of a comprehensive similarity algorithm code provided in an embodiment of the present invention;

[0036] Figure 15 This is a schematic diagram of the fourth part of a comprehensive similarity algorithm code provided in an embodiment of the present invention;

[0037] Figure 16 A flowchart illustrating an AI model hot-loading module for a task subsystem provided in an embodiment of the present invention;

[0038] Figure 17 A schematic diagram of the algorithm code for a posterior probability calculation function provided in an embodiment of the present invention;

[0039] Figure 18 This is a schematic diagram of an AI model training process provided in an embodiment of the present invention. Detailed Implementation

[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] Currently, under the stringent requirement of 99% accuracy in industrial production, a significant portion of AI vision companies still rely on the outdated approach of manually collecting samples, manually labeling, and manually training and debugging. Some employ unsupervised or weakly supervised learning methods, but these suffer from insufficient stability due to uncontrollable output, failing to meet the demands of diverse product types and rapid iteration cycles. In these semi-automated factories, to survive and adapt to market demands, the production of a wide variety of products is required. Some production lines may even switch between several products daily, and dozens or even hundreds of products monthly. Some products are produced only once, placing extremely high demands on iteration synchronization, which traditional vision training cycles can no longer satisfy.

[0042] Based on this, the present invention provides a visual classification system and electronic device, specifically an AI-powered high-speed visual classification system that enables industrial production to quickly launch in the early stages of new product counting, classification, and identification, and to operate efficiently in the later stages.

[0043] To facilitate understanding of this embodiment, a visual classification system disclosed in this embodiment of the invention will first be described in detail.

[0044] Example 1:

[0045] This invention provides a visual classification system that can be applied to software systems. See [link to relevant documentation]. Figure 1 The diagram shows the structure of a visual classification system, which includes a task subsystem, a model training subsystem, and a client interface.

[0046] The first subsystem is used to count, classify, inspect, and identify products on the production line.

[0047] The operation subsystem in this embodiment is the main part of the visual classification operation system. Its main function is to count, classify, detect and identify products on the production line, which will be referred to as operation below.

[0048] like Figure 1As shown, the operation subsystem includes: a first operation module, a second operation module, and a scheduling module; the first operation module is used to encode and decode single-frame images of video and output product images; the second operation module is used to input single-frame images of video into a trained detection model and output product images; the scheduling module is used to schedule the use of the first and second operation modules according to the training maturity of the current production product.

[0049] 1. First assignment module (hereinafter referred to as Module 1):

[0050] In some embodiments, the first task module is further configured to output a background image, product information, and task data; wherein, the product information includes: segmentation rectangle, coordinates, directional bounding box information, segmentation pixel area, video frame number, and confidence level; if manual operation is involved, the first task module is further configured to output a single-frame video image and product segmentation information.

[0051] Module 1: Integrated Segmentation, Collection, and Operation Module. Its function is to use SAM2 (SegmentAnything Model 2) to encode and decode single-frame images of video, automatically segment product images, collect and perform operations, and enable the ability to launch new products within 15 minutes. This module is the main contributor. The output of this module mainly includes: (1) Segmented product and background images and related information, including segmentation rectangle, coordinates, OBB (Oriented Bounding Box) information, segmented pixel area, video frame number, confidence level, etc. (2) If manual operation is required, the entire video image and product segmentation information are additionally saved. (3) Operation data.

[0052] 2. Second assignment module (hereinafter referred to as Module 2):

[0053] In some embodiments, the second operation module is also used to output product information and operation data.

[0054] Module 2: The main AI model module uses a pre-trained AI model as its core for operations. Compared to Module 1, it features high efficiency and strong generalization ability, and is comparable to existing technologies. It plays an important role in this system, but from a patent perspective, it is only an auxiliary module. The main outputs of this module are: (1) segmented product images and related information; (2) operation data.

[0055] 3. Scheduling Module:

[0056] In some embodiments, if the current product being manufactured is a new product, the scheduling module is used to use the first job module; if the current product being manufactured follows a preset job model, the scheduling module is used to use the second job module; if the second job module has an output anomaly during the use of the second job module, the first job module is used to perform a job on the video single frame image with the output anomaly, and the scheduling module is used to process the output results of the first job module and the second job module; wherein the processing methods include: collecting sample datasets, anomaly reports, and transferring to manual verification.

[0057] The scheduling module schedules the use of Module 1 and Module 2 based on the training maturity of the current production product. There are three main scheduling modes: Module 1 is used when a new product is launched; Module 2 is used when the job AI model evaluation is passed. If an output anomaly is received during the use of Module 2, such as a confidence level below a specified threshold, the single-frame video image needs to be processed using Module 1. The two outputs are compared, and appropriate processing is performed based on the comparison results, including collecting sample datasets, reporting anomalies, and forwarding to manual verification.

[0058] II. The model training subsystem is used to connect to the message server after startup and train the model of the visual classification system through message events.

[0059] The model training subsystem in this embodiment is an independent AI model training system. After startup, it connects to a message server. This system uses an MQTT (Message Queuing Telemetry Transport) message server, listens for topics, and can initiate AI model training by automatically reading the corresponding configuration through message events. Upon completion, it generates the AI ​​model used by this system, including model files such as *.pt or *.onnx, and automatically pushes a completion message. This AI model training system is an important component of this system, contributing to its automation and efficiency improvement.

[0060] 3. The client operation interface is used to establish real-time communication with the operation subsystem and display the output information of the operation subsystem; the client operation interface is also used to stop the video frame of the current frame and edit data based on the user's operation on the video frame of the current frame.

[0061] In this embodiment, the client interface is an interactive interface. For ease of deployment, HTML5 (Hypertext Markup Language) is used to implement the interactive functions, allowing it to be used independently or embedded. However, regardless of the implementation method, the UI (User Interface) must implement the following main functions:

[0062] (1) Establish real-time communication with the job subsystem, which can be done via WebSocket (a real-time communication protocol), MQTT or other methods.

[0063] (2) It can display the output information of the operation subsystem, including real-time video with product label boxes.

[0064] (3) It can stop the video frame of the current frame, perform coordinate click operations on the screen, and submit editing data through real-time communication.

[0065] The comparison of the product count and classification of new products with existing technologies is shown in Table 1.

[0066]

[0067] Table 1

[0068] This invention provides a visual classification system that can be applied to pallet and belt production lines. Without changing the hardware, it enables the counting and classification of products for switching to new products within 15 minutes, with an accuracy rate of over 99%, laying a solid foundation for intelligent monitoring of the production process.

[0069] Example 2:

[0070] This embodiment provides another visual classification system, implemented based on the above embodiment, focusing on the implementation of each subsystem of the visual classification system. The overall flow of the visual classification system can be found in [link to relevant documentation]. Figure 2 The diagram shows the overall process of a visual classification system.

[0071] In some embodiments, the first task includes a segmentation model; the segmentation model is used to encode and decode single-frame images of real-time video from production line inspection points at a fixed frame rate, and segment the single-frame images of video according to a target size; the first task module is also used to determine whether the segmentation result is a product image or a background image, and to rotate and perform secondary segmentation on the segmentation result in sequence; the first task module is also used to process the result of the secondary segmentation display through a classification model; wherein, the processing methods include: error detection, classification, and counting; the first task module is also used to collect product images and product information of product images, and output the initial judgment result of the single-frame image of video.

[0072] Module 1 in this embodiment is the core of the visual classification subsystem. It automatically collects samples, classifies them, and saves relevant parameters and data. It can perform SAM2 encoding and decoding on videos online or offline to segment product images and similarly shaped backgrounds. This data serves as sample data for AI model training and also provides basic production data for initial judgments through graphic similarity comparison.

[0073] Module 1 can encode and decode real-time video of production line inspection points at a fixed frame rate based on the SAM2 (Segment Anything Model 2) segmentation model. Since the video environment in industrial production is relatively fixed—for example, production line pallets or conveyor belts are not significantly different in appearance, and the appearance of the products is also uniform and fixed, with only differences on observable surfaces, and perhaps some variations in lighting, shadows, and angles—the main structures are basically similar. In terms of graphics technology, similarity can be used to distinguish products from the background. (See also...) Figure 3 The diagram shown illustrates the main process of SAM2 segmentation. The main technical steps involved in its implementation are described below:

[0074] (1) Automatically segment the product or background according to the target size. SAM2 segmentation itself does not have the ability to determine whether the segmented image is a product or a background, so manual selection of base samples is required when launching new products.

[0075] In some embodiments, if the product currently being produced is a new product, the first operation module is used to notify the operator to perform a product segmentation and selection operation for the product currently being produced.

[0076] Upon receiving the current product number, the system queries the database for existing product numbers. If it's a new product, it notifies a human to perform product segmentation and selection, generating multiple basic samples. The product segmentation and selection process is relatively simple. Clicking on the screen in the operation panel, module 1's SAM2 displays marked rectangles based on the input coordinates, similar to X-AnyLabeling (a data labeling tool). The operator simply selects the most suitable mark based on the product's shape and confirms submission. If the product has many faces, each with a different appearance, the process can be repeated multiple times to cover as many as possible. This process generally takes no more than 15 minutes; the simpler the shape, the fewer steps are required. After submission, the system records the product image, dimensions, and location information. Then, the system uses a coordinate array to fully decode the current image, marking all products and saving any accompanying backgrounds of similar sizes as negative samples. In addition, the system saves the current image for subsequent AI model training to verify model accuracy, providing a verification dataset for automation (hereinafter referred to as the verification dataset).

[0077] See Figure 4 The diagram illustrates a product segmentation and selection process. This process can produce: 1. A dataset of product and background samples, which can be used for similarity comparison and AI model training. 2. An image of the current video frame, along with product bounding box information and product ID, which can be used for validating the OBB object detection model after training.

[0078] In some embodiments, the segmentation model is used to segment multiple rectangles of a single frame image of the output video using a coordinate dot matrix, and the first task module is used to determine whether each rectangle is a product image or a background image by a similarity weighted judgment method; wherein, the similarity weighted judgment method includes: calculating multiple similarities through multiple similarity calculation methods and performing a weighted judgment on the multiple similarities.

[0079] Each coordinate group in SAM2 is divided into three rectangles of varying sizes, and the coordinate grid itself is divided into numerous rectangles. (See also...) Figure 5 The diagram shows a process for determining the output classification based on confidence level and AI model. It requires using a rectangular target similarity confidence threshold (here, confidence level refers to shape similarity) to filter out OBB rectangles that are similar in size to the product. Then, it uses the intersection-union ratio of OBB rectangles to remove overlapping rectangles with low confidence levels. Finally, it marks all products and the background and distinguishes them by color.

[0080] In some embodiments, multiple similarity calculation methods include: target similarity calculation method and Hamming similarity calculation method; wherein, the target similarity calculation method is used to calculate the similarity between the first value and the second value using the following formula: s=|ab| / (|b|+ε), f=1-min(s,1); where s is the intermediate value of the target similarity calculation method, a is the first value, b is the second value, ε is the preset minimum representable value, f is the similarity of the target similarity calculation method, and min(s,1) is the smaller value between s and 1.

[0081] The Python code for target similarity and rectangular target similarity can be found in [link to Python code]. Figure 6 The image shows a schematic diagram of Python code.

[0082] (2) Use comprehensive similarity to make a binary classification judgment, i.e., product or background.

[0083] The specific implementation involves rotating the OBB rectangle based on its angle before performing a second segmentation, removing redundant pixels, and then using a comprehensive similarity score. The calculation formula is: Gray target similarity × a + Hamming similarity × b (where a + b = 1). This means using color-sensitive gray target similarity and structure-sensitive Hamming similarity to determine whether the segmented image represents a product or background, and providing a comprehensive similarity score for each.

[0084] The rectangular marker for OBB in the original image can be found in [reference needed]. Figure 7 The diagram shows a rectangular marker. The result of rotating the OBB rectangle according to the angle and then performing a secondary segmentation can be found in [reference needed]. Figure 8 The diagram shows a result of a quadratic segmentation. The algorithm code for quadratic segmentation can be found in [link to algorithm]. Figure 9A schematic diagram of the first part of a quadratic segmentation algorithm code is shown below. Figure 10 A schematic diagram of the second part of a quadratic segmentation algorithm code is shown below. Figure 11 This diagram illustrates the third part of a secondary segmentation algorithm. The comprehensive similarity algorithm code can be found in [reference needed]. Figure 12 The diagram shown is a schematic of the first part of the code for a comprehensive similarity algorithm. Figure 13 The diagram shown is a schematic of the second part of a comprehensive similarity algorithm code. Figure 14 The diagram shown is a schematic of the third part of a comprehensive similarity algorithm code. Figure 15 The diagram shown is a schematic of the fourth part of a comprehensive similarity algorithm code, wherein... Figure 12-15 In the code, the function name is composite_similarity, which is used as the confidence level. The Hamming similarity is derived from the output value of the Hamming distance transformation.

[0085] (3) The segmented images are then judged again using the corresponding trained classification AI model (binary classification, product and background) (such as error detection, classification, counting, etc.) to obtain more accurate results.

[0086] (4) The task subsystem collects sample images according to specified rules and retains relevant image information (including timestamps, OBB information, product confidence scores, video frame numbers, etc.), and outputs the initial judgment results. When saving the output images, the filenames contain relevant information, including OBB size, pixel area (pixel area: each pixel segmented by SAM2 has a Mask confidence value; this value is valid only if it is greater than a threshold, defined as 1, and less than the threshold as 0; the sum of all 1s is the pixel area), angle, video frame number, etc., ensuring that other systems can reconstruct this information data according to the information encoding rules when reading the image data. The directory structure for saving the images and related information after output is as follows:

[0087] dataset / {product number} /

[0088] ---{0 Background} /

[0089] ---{1 product} /

[0090] ---{-1 Unknown} /

[0091] The {-1 Unknown} category requires manual verification, but the process is simple: just drag the image to either {0 Background} or {1 Product} according to the rules. If object detection or product error detection is needed, the product category can be further subdivided into smaller types; here, we'll only discuss the main categories. The resulting images (product and background) can be used as training datasets for AI models for object detection and classification of that product, eliminating the need for manual annotation. This step is a key technique for enabling new products to launch in a short timeframe.

[0092] This embodiment can also include an AI model hot-loading module in the task subsystem: after receiving an instruction, the module loads a new AI model, enabling more accurate visual judgments. (See also...) Figure 16 The diagram shows a flowchart of an AI model hot-loading module in a task subsystem.

[0093] To ensure an accuracy rate of over 99%, two key technologies are used. One is a target tracking algorithm, which determines whether the product is the same based on motion trajectory and features when the images are segmented from sequential video frames. The other is a algorithm that uses Bayes' theorem and the law of total probability to calculate the posterior probability. A single judgment cannot reach 99%, but the probability can be greater than 99% when multiple judgments are made.

[0094] 1. Target tracking algorithm: There are many target tracking algorithms available. In addition, the intersection-union ratio of OBB rectangles can be calculated, and the similarity of OBB rectangles, coordinates, angles, etc. can be comprehensively calculated using target similarity to track the target's trajectory. Furthermore, by comparing texture features, it is possible to track whether the target is the same one.

[0095] 2. Calculate the posterior probability using the product confidence score calculated for each frame. If the score is above 0.8, and this holds for 3 to 4 consecutive frames, the probability will be above 99%. The function for calculating the posterior probability can be found in [link to relevant documentation]. Figure 17 The diagram shows an algorithm for calculating the posterior probability. Running it in a Python editor will print: Posterior probability: 0.997999.

[0096] In some embodiments, the model training subsystem is used to use the product image and background image output by the first task module as a training set to train the model of the visual classification task system based on the training set.

[0097] The model training subsystem in this embodiment is an AI model training system that primarily uses automated training and secondarily uses manual training; it is a very important component of this system. See also... Figure 18The diagram shows a flowchart of an AI model training process. The core principle behind the high success rate of the AI ​​model training system is that it uses an image dataset generated by SAM2 segmentation. These images are all industrial products with relatively fixed graphic textures and some natural distinction from the background images. In addition, with accurate OBB labeling, the AI ​​classification model or AI object detection model trained on such images has a relatively high accuracy.

[0098] In some embodiments, the model of the visual classification task system includes: a classification model of the first task module and a detection model of the second task module.

[0099] In this embodiment, the model training subsystem is directly related to the system mainly in the automated training of two models: one is a single-product binary classification AI model, which distinguishes between products and background, and is used as an enhancement judgment for module 1; the other is a multi-product OBB object detection AI model, which is used as the main function of module 2, because this AI model can also segment product images, which is consistent with the output of module 1.

[0100] I. Training of AI Model for Single Product Classification: The main purpose of training the AI ​​model for single product classification is to enable Module 1 to achieve generalization ability and improve accuracy in a short period of time.

[0101] The steps for training a single-product classification AI model are as follows: 1. Listen to relevant topics on the message server. 2. Upon receiving a training topic message, obtain the product ID and AI model type (classification). 3. Read the configuration information based on the product ID. This includes the address of the sample image dataset. If the dataset is not on the local machine, it needs to be downloaded using a network file tool (such as SVN, Git, or FTP). 4. Split the dataset into a training dataset, a testing dataset, and a validation dataset according to a certain ratio. 5. Read the training parameters from the configuration and start training. 6. After training is complete, output the ONNX model and upload it to the location specified in the configuration. 7. Send a recommendation message to the message server.

[0102] At this stage, the training task is complete. Upon receiving the message, the assignment subsystem will load the latest AI classification model to improve generalization capabilities and accuracy. Module 1 will then re-assign the video recordings without AI model involvement, generating assignment data. If there are differing data results, a report will be generated for professionals to analyze and upgrade the similarity calculation or adjust the similarity parameters. This process continuously improves the overall system accuracy.

[0103] If 50 to 100 different images are collected as samples for a product, a single product model can be trained. According to actual tests, using a common graphics card such as an RTX 4060 to train a YOLO v8 (an object detection model) classification model with up to 300 images, it can be completed in less than 15 minutes at the fastest and about an hour at the slowest, thus ensuring that training can be completed in a short period of time.

[0104] II. Multi-Product Target Detection AI Model Training: Target detection AI model training aims to develop a more mature, efficient, and generalizable AI model suitable for a wider range of products. The training process is similar to single-product classification training, but includes an additional step: using a dedicated test dataset. Training can be initiated when: firstly, a sufficient amount of new product sample data has been accumulated; and secondly, the confidence level for older products is insufficient, meaning there is a new sample dataset requiring incremental training to improve accuracy.

[0105] The steps for training a multi-product object detection AI model are as follows: Listen to relevant topics on the message server. Upon receiving a training topic message, obtain the product number and AI model type (object detection). Read the configuration information based on the product number. This includes the address of the sample image dataset; if it's not on the local machine, it needs to be downloaded using a network file tool (such as SVN, Git, FTP, etc.). Split the dataset into a training dataset, a testing dataset, and a validation dataset according to a certain ratio. Read the training parameters from the configuration and start training. After training, validate the model using the dedicated validation dataset. If it fails, notify the administrator to investigate the cause, check the dataset, or adjust the training parameters before retraining. Output the ONNX model and upload it to the specified location in the configuration. Send a recommendation message to the message server.

[0106] The training cycle for the object detection AI model is longer than that for a single product, but this does not affect the operation of the system.

[0107] The visual task classification system provided in this invention allows for new product launches in just 15 minutes. Combined with the fixed background of the production line, as long as the product's color doesn't perfectly match the background, manual operation requires only a few screen taps. The SAM2 model segments the product image texture with OBB information, and then, with target tracking, posterior probability, and comprehensive similarity judgment, the probability of identifying the product is confirmed. This enables counting, classification, and subsequent detection, unaffected by background texture noise, achieving accuracy comparable to deep learning-based AI models. However, this stage lacks the generalization ability of neural networks, and errors may occur if the texture is too abnormal. Therefore, a single-product classification AI model can be automatically trained on a collected dataset, achieving generalization within an hour. This process eliminates the need for traditional manual annotation, significantly shortening the training time. Furthermore, due to confidence assessment, previously uncollected sample datasets can be continuously and automatically collected, allowing for retraining and generalization, thus accommodating a wider range of product textures.

[0108] Under normal circumstances, Module 1 can train a multi-product object detection model in one day. This model also has the ability to segment products, classify them, and output confidence scores, seamlessly replacing the work of Module 1. It can continue to run and iterate on the dataset with less memory and lower computing power. Throughout this process, since the system functions are already complete, a large amount of work does not require the intervention and maintenance of AI professionals. AI professionals are only needed in abnormal situations, resulting in a significant reduction in operating costs compared to existing technologies. This achieves four major advantages: short cycle iteration, high efficiency, high accuracy, and low cost.

[0109] The visual task classification system provided in this embodiment of the invention has the following main advantages:

[0110] 1. A technical approach was developed that combines SAM2 with rectangular similarity to achieve automatic segmentation of products and backgrounds of the same shape, as well as automatic collection of image samples.

[0111] 2. Combining segmentation with application: Outline segmentation with OBB rectangle information is used for training AI object detection models; rotating the OBB image texture according to its angle and then performing secondary segmentation to obtain a more compact texture has multiple uses: first, it can be used for feature and similarity comparison of pure algorithms without AI model training; second, it can be used for training classification models; and third, it can be used as input for classification models to obtain more accurate judgments.

[0112] 3. A comprehensive similarity comparison method using image texture was created, which takes into account both structural and color comparisons. It can also adjust the weights and integrated algorithms according to the characteristics of actual texture features, making the similarity comparison more in line with actual needs, thereby enabling image and background classification without training.

[0113] 4. A comparison method using the comprehensive similarity of image textures was created to achieve deduplication of image sample datasets, avoiding the infinite growth of datasets, improving the training efficiency of AI models, and achieving effective incremental training.

[0114] 5. The basic structure and process of this system were designed, an automated process was built, manual work was reduced, and accuracy was ensured.

[0115] 6. A mode was created to save all images and data of products segmented manually for automatic verification after the relevant AI model training is completed, which simplifies the operation steps and improves the success rate of AI model training.

[0116] In addition, the visual task classification system provided in this embodiment also has some alternative solutions:

[0117] 1. The code used in this system description is based on Python examples. To make the code easier to understand, tensor data is not used. The actual language used in the system's deployment is uncertain and depends on the client's needs, available hardware, and operating system, requiring conversion into corresponding programs. Some programs may be run on GPUs (Graphics Processing Units) to improve concurrency and efficiency. GPU languages ​​include CUDA (a parallel computing platform and programming model), GLSL (a high-level shading language based on C), and Vulkan (a cross-platform graphics API).

[0118] 2. The example only shows the target similarity + Hamming similarity algorithm for comprehensive similarity. In reality, more combined algorithms can be used to calculate the similarity based on the texture of the image. For example, if more consistency is emphasized, the SSIM (Structural Similarity Index) similarity algorithm should be added, or other similarity algorithms such as Gray-Level Co-occurrence Matrix (GLCM), Gray-Level Histogram, Local Binary Pattern (LBP), etc. This embodiment does not limit the specific similarity algorithm.

[0119] 3. In this embodiment, it is mentioned that after SAM2 segmentation, the shape is filtered first and then the binary classification judgment of product or background is performed. This is to save computing power (texture calculation consumes more computing power than shape calculation, and filtering can reduce texture calculation and thus reduce computing power). However, if we do not consider the waste of computing power, especially on GPU, it can be calculated in parallel, and the shape similarity, product texture similarity and background texture similarity are output at the same time.

[0120] 4. In this case study, the AI ​​model was mainly trained using YOLOv8, but the same capability can be achieved by using other YOLOvN models, building your own neural network, or using other training platforms.

[0121] 5. Although this system only emphasizes classification, this is only a foundation. If the product image dataset is further subdivided on this basis, it is possible to realize more application scenarios such as product counting, product classification, error detection, and even defect detection.

[0122] 6. This system does not explicitly mention unsupervised learning or weakly supervised learning because of considerations regarding the maturity of the model. After testing, the accuracy and cycle time of these two learning methods do not yet meet the requirements in industrial environments. If these two technologies prove to be mature and easy to use in the future, they can be incorporated into the AI ​​model training process.

[0123] Example 3:

[0124] This invention also provides an electronic device, including the visual classification system provided in the foregoing embodiments.

[0125] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and / or device described above can be referred to the corresponding process in the foregoing embodiments, and will not be repeated here.

[0126] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.

[0127] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0128] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0129] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A visual classification system, characterized in that, The visual classification system includes: a task subsystem, a model training subsystem, and a client operation interface; The operation subsystem is used to count, classify, detect and identify products on the production line; The operation subsystem includes: a first operation module, a second operation module, and a scheduling module; the first operation module is used to encode and decode single-frame images of video and output product images; the second operation module is used to input the single-frame images of video into a trained detection model and output the product images; the scheduling module is used to schedule the use of the first operation module and the second operation module according to the training maturity of the current production product. The model training subsystem is used to connect to the message server after startup and train the model of the visual classification task system through message events. The client operation interface is used to establish real-time communication with the operation subsystem and display the output information of the operation subsystem; the client operation interface is also used to stop the video frame of the current frame and edit data based on the user's operation on the video frame of the current frame; The first task module is also used to output background images, product information, and task data; wherein, the product information includes: segmentation rectangle, coordinates, directional bounding box information, segmentation pixel area, video frame number, and confidence level; if manual operation is involved, the first task module is also used to output the single-frame image of the video and product segmentation information; the second task module is also used to output the product information and the task data; If the currently produced product is a new product, the scheduling module is used to use the first job module; if the currently produced product follows a preset job model, the scheduling module is used to use the second job module; if the second job module has an output anomaly during the use of the second job module, the first job module is used to process the video single frame image with the output anomaly, and the scheduling module is used to process the output results of the first job module and the second job module; wherein the processing methods include: collecting sample datasets, anomaly reports, and transferring to manual verification; The first task module is equipped with a segmentation model; the segmentation model is used to encode and decode single-frame images of real-time video from production line inspection points at a fixed frame rate, and segment the single-frame images of the video according to a target size; the first task module is also used to determine whether the segmentation result is the product image or the background image, and to rotate and perform secondary segmentation on the segmentation result in sequence; the first task module is also used to process the results of the secondary segmentation display through a classification model; wherein, the processing methods include: error detection, classification, and counting; the first task module is also used to collect the product image and the product information of the product image, and output the initial judgment result of the single-frame image of the video.

2. The visual classification system according to claim 1, characterized in that, If the currently produced product is a new product, the first operation module is used to notify the operator to perform a product segmentation and selection operation for the currently produced product.

3. The visual classification system according to claim 1, characterized in that, The segmentation model is used to segment and output multiple rectangles of the single frame image of the video using a coordinate dot matrix. The first task module is used to determine whether each rectangle is a product image or a background image by weighted similarity judgment. The weighted similarity judgment method includes: calculating multiple similarities through multiple similarity calculation methods and weighting the multiple similarities.

4. The visual classification system according to claim 3, characterized in that, The various similarity calculation methods include: target similarity calculation method and Hamming similarity calculation method; The target similarity calculation method is used to calculate the similarity between the first value and the second value using the following formula: s=|ab| / (|b|+ε), f=1-min(s,1); where s is the intermediate value of the target similarity calculation method, a is the first value, b is the second value, ε is the preset minimum representable value, f is the similarity of the target similarity calculation method, and min(s,1) is the smaller value between s and 1.

5. The visual classification system according to claim 1, characterized in that, The model training subsystem is used to train the model of the visual classification task system based on the product image and background image output by the first task module as a training set.

6. The visual classification system according to claim 5, characterized in that, The model of the visual classification task system includes: the classification model of the first task module and the detection model of the second task module.

7. An electronic device, characterized in that, The electronic device includes: the visual classification system according to any one of claims 1-6.

Citation Information

Patent Citations

  • Visual processing and model training method and device, storage medium and program product

    CN114549904A

  • Food image classification method and system, storage medium and terminal

    CN114842266A