Visual classification operation system and electronic equipment

By using the SAM2 segmentation model and comprehensive similarity algorithm of the visual classification system, the problem of low accuracy in product counting and classification in semi-automated factories has been solved, enabling rapid launch and efficient identification of new products, achieving an accuracy rate of 99% and reducing operating costs.

CN121170548AActive Publication Date: 2025-12-19GUANGZHOU QIYIN INTELLIGENT TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511399434.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-12-19
Estimated Expiration
2045-09-28

AI Technical Summary

Technical Problem

In semi-automated factories, there are many types of products and their placement is highly random. Existing AI visual inspection technologies are difficult to achieve high-accuracy product counting, classification and inspection. Traditional methods are costly and lack stability, and cannot meet the needs of a wide variety of products and a fast iteration cycle.

Method used

A visual classification system is provided, including a task subsystem, a model training subsystem, and a client operation interface. It uses the SAM2 segmentation model and a comprehensive similarity algorithm to achieve automatic segmentation and recognition of product images. It combines target tracking algorithms and Bayes' theorem to improve accuracy, supporting the rapid launch of new products and efficient operation in the later stages.

Benefits of technology

Without changing the hardware, it achieves the ability to count and classify new products with an accuracy rate of over 99%, reducing operating costs and improving the level of intelligence in the production process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121170548A_ABST
    Figure CN121170548A_ABST
Patent Text Reader

Abstract

The invention provides a visual classification operation system and electronic equipment. The visual classification operation system comprises an operation subsystem, a model training subsystem and a client operation interface, the operation subsystem comprises a first operation module, a second operation module and a scheduling module; the first operation module is used for encoding and decoding a video single-frame image and outputting a product image; the second operation module is used for inputting a video single-frame image into the trained detection model and outputting a product image; the scheduling module is used for scheduling the application conditions of the first operation module and the second operation module according to the training maturity of the current production product; and the model training subsystem is used for accessing the message server after being started and carrying out model training of the visual classification operation system through a message event. Under the condition that hardware facilities are not changed, product counting, product classification and recognition and other capacities of switching new product production within 15 minutes are achieved, the accuracy rate reaches 99% or above, and a solid foundation is laid for intelligent monitoring of the production process.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence (AI), and particularly relates to a visual classification operation system and an electronic device. BACKGROUND

[0002] AI industrial visual inspection service (AIVI) is an innovative solution that utilizes artificial intelligence technology, particularly deep learning and computer vision, to perform high-precision and high-speed quality inspection on products on the production line. It simulates the human visual system to automatically identify and analyze key characteristics such as size, shape, color, and defects of products, significantly improving the efficiency of quality control in manufacturing.

[0003] At the technical level, the continuous optimization of deep learning and machine learning algorithms has significantly improved the performance of AI in image recognition and defect detection, saving a large amount of labor costs. Although the industry also faces some challenges, such as data security issues and rapid technology update iteration, with the integration of new technologies such as 5G and the Internet of Things, the development prospects of the AI industrial visual inspection service industry are very broad.

[0004] However, focusing on semi-automated factories, such factories are still the main force of low-cost production, and manual work accounts for a large proportion. The randomness of product placement and the variety of products make product online counting, classification, and detection face great challenges. Traditional manual and sensor inspection methods have obvious defects: manual inspection relies on a large number of manpower, with high cost, and is limited by human subjective judgment and fatigue, with low accuracy; sensor detection has strict requirements for product placement position, angle, etc., and is difficult to adapt to complex production scenes. These problems seriously hinder the lean management and digital upgrade of the production line, resulting in large errors in factory labor cost accounting and high virtual cost. For example, an electronics factory has multiple semi-automated production lines, mainly producing various electronic components. Due to the variety of products and random placement, manual reporting is used for output value accounting. In the first quarter of 2024, the actual output value accounting was 300,000 yuan less than manual reporting, exposing the inaccuracy of manual inspection in counting and accounting, causing direct economic losses to the enterprise.

[0005] Under the hard requirement of 99% accuracy rate in industrial production, a considerable part of AI vision enterprises in the prior art still follow the old route of manual sample collection, manual annotation, manual training and debugging, and a part of them use unsupervised learning or weakly supervised learning method to realize, but due to uncontrollable output, the stability is insufficient, and other factors, which cannot meet the requirements of the operation of many product types and fast iteration cycle. Because in these semi-automatic factories, in order to survive and adapt to market demand, the products to be produced are relatively miscellaneous and many, and individual production lines even change several products for production a day, and dozens or even hundreds of products a month, and the production of some products is one-time, which has very high requirements on iteration and synchronization. The traditional visual training cycle cannot meet the requirements. SUMMARY

[0006] Therefore, the purpose of the present application is to provide a visual classification operation system and an electronic device to enable industrial production to quickly go online in the early stage of new product counting and classification, and to efficiently run in the later stage.

[0007] In a first aspect, an embodiment of the present application provides a visual classification operation system, which comprises an operation subsystem, a model training subsystem and a client operation interface; the operation subsystem is used for counting, classifying, detecting and identifying products on a production line; the operation subsystem comprises a first operation module, a second operation module and a scheduling module; the first operation module is used for encoding and decoding a single frame of video image to output a product image; the second operation module is used for inputting the single frame of video image into a trained detection model to output the product image; the scheduling module is used for scheduling the use of the first operation module and the second operation module according to the training maturity of the current production product; the model training subsystem is used for starting to access a message server after the model training of the visual classification operation system is completed through a message event; and the client operation interface is used for establishing real-time communication with the operation subsystem to display the output information of the operation subsystem; and the client operation interface is also used for stopping the video picture of the current frame and editing data based on the operation of the user on the video picture of the current frame.

[0008] In an optional embodiment of the present application, the first operation module is also used for outputting a background image, product information and operation data; wherein the product information comprises a segmentation rectangle, coordinates, a direction bounding box information, a segmentation pixel area, a video frame number and a confidence; if there is manual operation, the first operation module is also used for outputting a single frame of video image and product segmentation information; and the second operation module is also used for outputting product information and operation data.

[0009] In an optional embodiment of the present application, if the current production product is a new product, the scheduling module is configured to use the first job module; if the current production product passes through the preset job model, the scheduling module is configured to use the second job module; in the process of using the second job module, if there is an output exception in the second job module, the first job module is used to perform a job on the single-frame image of the video of the output exception, and the scheduling module is configured to process according to the output results of the first job module and the second job module; wherein the processing mode includes collecting sample data sets, exception reports and transferring manual verification.

[0010] In an optional embodiment of the present application, the first job is provided with a segmentation model; the segmentation model is configured to encode and decode the single-frame image of the real-time video of the production line detection point according to a fixed frame rate, and segment the single-frame image according to a target size; the first job module is further configured to determine whether the segmentation result is a product image or a background image, and sequentially rotate and secondarily segment the segmentation result for display; the first job module is further configured to process the result of the secondarily segmented display through a classification model; wherein the processing mode includes error and omission detection, classification and counting; the first job module is further configured to collect product images and product information of the product images, and output an initial judgment result of the single-frame image.

[0011] In an optional embodiment of the present application, if the current production product is a new product, the first job module is configured to notify a manual to perform a product segmentation selection operation on the current production product.

[0012] In an optional embodiment of the present application, the segmentation model is configured to segment a plurality of rectangles of the single-frame image of the output video through a coordinate point array, and the first job module is configured to determine whether each rectangle is a product image or a background image through a similarity weighted judgment mode; wherein the similarity weighted judgment mode includes calculating a plurality of similarities through a plurality of similarity calculation modes and performing a weighted judgment on the plurality of similarities.

[0013] In an optional embodiment of the present application, the plurality of similarity calculation modes include a target similarity calculation mode and a Hamming similarity calculation mode; wherein the target similarity calculation mode is configured to calculate the similarity of a first value and a second value through the following formula: s=|a-b| / (|b|+ε), f=1-min(s,1); wherein s is an intermediate value of the target similarity calculation mode, a is the first value, b is the second value, ε is a preset minimum representable value, f is the similarity of the target similarity calculation mode, and min(s,1) is the smaller value of s and 1.

[0014] In an optional embodiment of the present application, the model training subsystem is configured to use the product image and the background image output by the first job module as a training set, and train the model of the visual classification job system based on the training set.

[0015] In an optional embodiment of the present application, the model of the visual classification job system comprises a classification model of the first job module and a detection model of the second job module.

[0016] In a second aspect, the embodiments of the present application further provide an electronic device, which comprises the visual classification job system described above.

[0017] The embodiments of the present application have the following beneficial effects: The embodiments of the present application provide a visual classification job system and an electronic device, which can be applied to pallet and belt production lines, realize product counting, product classification identification and other capabilities within 15 minutes of switching new product production without changing hardware facilities, and the accuracy rate reaches more than 99%, laying a solid foundation for intelligent monitoring of production processes.

[0018] Other features and advantages of the present disclosure will be described in the following description, or can be inferred or determined without doubt from the description, or can be known by implementing the above-mentioned technologies of the present disclosure.

[0019] In order to make the above-mentioned purposes, features and advantages of the present disclosure more obvious and easy to understand, the following preferred embodiments are described in detail below, and the accompanying drawings are described as follows. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or prior art description. Obviously, the drawings described below are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0021] Figure 1 A structural schematic diagram of a visual classification job system provided by the embodiments of the present application; Figure 2 A schematic diagram of the overall flow of a visual classification job system provided by the embodiments of the present application; Figure 3 A main flow schematic diagram of SAM2 segmentation provided by the embodiments of the present application; Figure 4 A product segmentation selection operation flow expansion schematic diagram provided by the embodiments of the present application; Figure 5 A classification flow schematic diagram according to confidence condition and AI model judgment output provided by the embodiments of the present application; Figure 6 A python code schematic diagram provided by the embodiments of the present application; Figure 7A schematic diagram of a rectangular mark provided for an embodiment of the present application; Figure 8 A schematic diagram of a secondary segmentation result provided for an embodiment of the present application; Figure 9 A schematic diagram of a first part of the algorithm code for secondary segmentation provided for an embodiment of the present application; Figure 10 A schematic diagram of a second part of the algorithm code for secondary segmentation provided for an embodiment of the present application; Figure 11 A schematic diagram of a third part of the algorithm code for secondary segmentation provided for an embodiment of the present application; Figure 12 A schematic diagram of a first part of the algorithm code for comprehensive similarity provided for an embodiment of the present application; Figure 13 A schematic diagram of a second part of the algorithm code for comprehensive similarity provided for an embodiment of the present application; Figure 14 A schematic diagram of a third part of the algorithm code for comprehensive similarity provided for an embodiment of the present application; Figure 15 A schematic diagram of a fourth part of the algorithm code for comprehensive similarity provided for an embodiment of the present application; Figure 16 A schematic diagram of the flow of the job subsystem AI model hot loading module provided for an embodiment of the present application; Figure 17 A schematic diagram of the algorithm code for the calculation function of the posterior probability provided for an embodiment of the present application; Figure 18 A schematic diagram of the flow of the AI model training provided for an embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be described below in connection with the drawings, which are apparently part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0023] At present, under the hard requirement of 99% accuracy rate in industrial production, a considerable part of AI vision enterprises in the prior art still follow the old route of manual sample collection, manual annotation, manual training and debugging, and a part uses unsupervised learning or weakly supervised learning method to realize, but due to uncontrollable output, the stability is insufficient and other factors, which cannot meet the operation demand of product variety and fast iteration cycle. Because in these semi-automatic factories, in order to survive and adapt to market demand, the products to be produced will be more miscellaneous, and individual production lines even change several products for production a day, and change dozens or even hundreds of products a month, and the production of some products is one-time, which has very high requirements on iteration period, and the traditional visual training period has been far unable to meet the requirements.

[0024] Based on this, the embodiment of the present application provides a visual classification operation system and electronic equipment, and specifically provides an AI vision ultra-fast classification operation system, which can make the industrial production quickly online in the early stage of new product counting and classification, and efficiently run in the later stage.

[0025] In order to facilitate the understanding of the present embodiment, first, a visual classification operation system disclosed by the present embodiment is introduced in detail.

[0026] Embodiment one: The present embodiment provides a visual classification operation system, which can be applied to a software system, referring to Figure 1 The structure schematic diagram of the visual classification operation system, the visual classification operation system comprises: an operation subsystem, a model training subsystem and a client operation interface.

[0027] I. The operation subsystem is used for counting, classifying, detecting and identifying products on a production line.

[0028] The operation subsystem in the present embodiment is the main part of the visual classification operation system, and the main function is to count, classify, detect and identify products on a production line, which is collectively referred to as operation hereinafter.

[0029] As shown in Figure 1 The operation subsystem comprises a first operation module, a second operation module and a scheduling module; the first operation module is used for encoding and decoding a single frame image of a video and outputting a product image; the second operation module is used for inputting the single frame image of the video into a trained detection model and outputting the product image; and the scheduling module is used for scheduling the use of the first operation module and the second operation module according to the training maturity of the current production product.

[0030] 1. The first operation module (hereinafter referred to as module 1) comprises: In some embodiments, the first job module is further configured to output the background image, product information, and job data; wherein the product information comprises: segmentation rectangle, coordinates, OBB (Oriented Bounding Box) information, segmentation pixel area, video frame number, and confidence; and if there is manual operation, the first job module is further configured to output the single-frame video image and product segmentation information.

[0031] Module 1: Segmentation, Collection, and Job Integration Module, which can use SAM2 (full name: SegmentAnything Model 2) to encode and decode single-frame video images, automatically segment product images, collect and perform jobs, and achieve the ability of 15-minute new product online. The main output of this module includes: (1) segmented product and background images and related information, including segmentation rectangle, coordinates, OBB (Oriented Bounding Box) information, segmentation pixel area (Segment Area), video frame number, confidence, etc. (2) If there is manual operation, the entire video image and product segmentation information are additionally saved. (3) Job data.

[0032] 2. Second job module (referred to as module 2, the same below): In some embodiments, the second job module is further configured to output product information and job data.

[0033] Module 2: AI Model Main Job Module, which uses a mature AI model as the core for job performance. Compared with module 1, it has the characteristics of high efficiency and strong generalization ability, and is similar to existing technologies. It plays an important role in the system, but is only an auxiliary module in the patent aspect. The main output of this module includes: (1) segmented product images and related information; (2) job data.

[0034] 3. Scheduling module: In some embodiments, if the current production product is a new product, the scheduling module is configured to use the first job module; if the current production product passes through a preset job model, the scheduling module is configured to use the second job module; in the process of using the second job module, if the second job module has an output exception, the first job module is used to perform jobs on the single-frame video image with the output exception, and the scheduling module is configured to process the output results of the first job module and the second job module; wherein the processing methods include: collecting sample data sets, exception reports, and transferring manual verification.

[0035] The scheduling module schedules the use of module 1 and module 2 according to the training maturity of the current production product, mainly in three scheduling modes: new product online, use module 1. Job AI model evaluation passes, use module 2. During the use of module 2, if the output is abnormal, such as the confidence is lower than the specified threshold, the single-frame picture of the video needs to be processed by module 1, the two output results are compared, and the comparison result is processed accordingly, including collecting sample data set, abnormal report, transferring to manual verification, etc.

[0036] II. The model training subsystem is used to start the message server after access, and the model training of the visual classification job system is performed through the message event.

[0037] The model training subsystem in this embodiment is an independent AI model training system, which accesses the message server after starting. The system uses the MQTT (Message Queue Telemetry Transport) message server to listen to the topic, and can start an AI model training through the message event and automatically read the corresponding configuration. After completion, the AI model used by the system will be generated, including *.pt or *.onnx model files, and the completion message will be automatically pushed. The AI model training system is an important part of the automation and efficiency improvement of the system.

[0038] III. The client operation interface is used to establish real-time communication with the job subsystem and display the output information of the job subsystem. The client operation interface is also used to stop the current frame of video picture and edit data based on the user's operation on the current frame of video picture.

[0039] The client operation interface in this embodiment is an operation interface that can interact with the user. Considering the convenience of deployment, H5 (Fifth Generation HyperText Markup Language) is used to realize the interaction function, that is, it can be used independently or embedded. However, no matter which way the UI (User Interface) is realized, the following main functions need to be realized: (1) Real-time communication with the job subsystem can be established through WebSocket (a real-time communication protocol), MQTT or other methods.

[0040] (2) The output information of the job subsystem can be displayed, including the real-time video with product annotation box.

[0041] (3) The current frame of video picture can be stopped, and the function of coordinate clicking operation on the picture and submitting editing data through real-time communication can be realized.

[0042] The comparison of product counting, classification and existing technology for new products can be shown in Table 1.

[0043]

[0044] Table 1 The embodiment of the present application provides a visual classification operation system, which can be applied to a tray type and a belt type production line, realizes product counting, product classification identification and other capabilities within 15 minutes of switching new product production without changing hardware facilities, and has an accuracy of more than 99%, thereby laying a solid foundation for intelligent monitoring of a production process.

[0045] Embodiment two: The embodiment provides another visual classification operation system, which is realized on the basis of the above embodiment, and mainly describes an implementation manner of each subsystem of the visual classification operation system. The overall flow of the visual classification operation system can be referred to the schematic diagram of the overall flow of a visual classification operation system shown in Figure 2

[0046] In some embodiments, the first operation is provided with a segmentation model; the segmentation model is used for encoding and decoding a single-frame image of a real-time video of a production line detection point according to a fixed frame rate, and segmenting the single-frame image according to a target size; the first operation module is further used for determining whether the segmentation result is a product image or a background image, and sequentially rotating and secondarily segmenting the segmentation result; the first operation module is further used for processing a result of the secondary segmentation by using a classification model; wherein the processing manner includes error and omission detection, classification and counting; the first operation module is further used for collecting the product image and product information of the product image, and outputting an initial judgment result of the single-frame image.

[0047] The module 1 in the embodiment is a core part of the visual classification operation subsystem, automatically collects samples, automatically classifies, and saves related parameters and data. The module 1 can encode and decode a video on-line or off-line, segment a product picture and a background with similar shapes, and provide basic production data by serving as sample data for AI model training and providing initial judgment by using graphic similarity comparison.

[0048] The module 1 can be based on a SAM2 (Segment Anything Model 2) segmentation model, and encode and decode a real-time video of a production line detection point according to a fixed frame rate. Since the video environment of industrial production is relatively fixed, for example, a production line tray or a conveying belt has little difference in appearance, the appearance of a product is uniform and fixed, and only the observable surface is different, at most, some light, shadow and angle changes, and the main structure is basically similar, and the product and the background can be distinguished by using similarity in graphic technology. Referring to a main flow schematic diagram of SAM2 segmentation shown in Figure 3 ​​(1) Automatically segment the product or background according to the target size. SAM2 segmentation itself does not have the ability to determine whether the segmented image is a product or a background, so manual selection of base samples is required when launching new products.

[0049] In some embodiments, if the product currently being produced is a new product, the first operation module is used to notify the operator to perform a product segmentation and selection operation for the product currently being produced.

[0050] Upon receiving the current product number, the system queries the database for existing product numbers. If it's a new product, it notifies a human to perform product segmentation and selection, generating multiple basic samples. The product segmentation and selection process is relatively simple. Clicking on the screen in the operation panel, module 1's SAM2 displays marked rectangles based on the input coordinates, similar to X-AnyLabeling (a data labeling tool). The operator simply selects the most suitable mark based on the product's shape and confirms submission. If the product has many faces, each with a different appearance, the process can be repeated multiple times to cover as many as possible. This process generally takes no more than 15 minutes; the simpler the shape, the fewer steps are required. After submission, the system records the product image, dimensions, and location information. Then, the system uses a coordinate array to fully decode the current image, marking all products and saving any accompanying backgrounds of similar sizes as negative samples. In addition, the system saves the current image for subsequent AI model training to verify model accuracy, providing a verification dataset for automation (hereinafter referred to as the verification dataset).

[0051] See Figure 4 The diagram illustrates a product segmentation and selection process. This process can produce: 1. A dataset of product and background samples, which can be used for similarity comparison and AI model training. 2. An image of the current video frame, along with product bounding box information and product ID, which can be used for validating the OBB object detection model after training.

[0052] In some embodiments, the segmentation model is used to segment multiple rectangles of a single frame image of the output video using a coordinate dot matrix, and the first task module is used to determine whether each rectangle is a product image or a background image by a similarity weighted judgment method; wherein, the similarity weighted judgment method includes: calculating multiple similarities through multiple similarity calculation methods and performing a weighted judgment on the multiple similarities.

[0053] Each coordinate group in SAM2 is divided into three rectangles of varying sizes, and the coordinate grid itself is divided into numerous rectangles. (See also...) Figure 5One of the shown according to the confidence case and AI model judgment output classification flow chart, need to use the rectangular target similarity confidence threshold (the confidence here is the shape similarity) Filter out the OBB rectangle similar to the product size, and remove the overlapping and low-confidence rectangle through the OBB rectangle intersection ratio calculation, then mark all products and background, and distinguish them with color.

[0054] In some embodiments, the plurality of similarity calculation methods include: a target similarity calculation method and a Hamming similarity calculation method; wherein the target similarity calculation method is used to calculate the similarity of the first value and the second value through the following formula: s=|a-b| / (|b|+ε), f=1-min(s,1); wherein s is the intermediate value of the target similarity calculation method, a is the first value, b is the second value, ε is the preset minimum representable value, f is the similarity of the target similarity calculation method, and min(s,1) is the smaller value of s and 1.

[0055] The python code of the target similarity and the rectangular target similarity can be seen from Figure 6 The shown is a schematic diagram of a python code.

[0056] (2) Use the comprehensive similarity to make a binary classification judgment, that is, whether it is a product or background.

[0057] The specific implementation is to rotate according to the OBB rectangular information according to the angle, then perform secondary segmentation, remove the redundant pixels, and then use the comprehensive similarity, and the calculation formula is: gray target similarity×a+Hamming similarity×b (wherein a+b=1). That is, the gray target similarity sensitive to color and the Hamming similarity sensitive to structure are used to judge whether the segmented picture is a product or background, and the respective comprehensive similarity is given.

[0058] The rectangular mark of the OBB on the original image can be seen from Figure 7 The shown is a schematic diagram of a rectangular mark. The result of the secondary segmentation according to the OBB rectangular information according to the angle can be seen from Figure 8 The shown is a schematic diagram of a secondary segmentation result. The algorithm code of the secondary segmentation can be seen from Figure 9 The shown is a schematic diagram of the first part of the algorithm code of the secondary segmentation, Figure 10 The shown is a schematic diagram of the second part of the algorithm code of the secondary segmentation, and Figure 11 The shown is a schematic diagram of the third part of the algorithm code of the secondary segmentation. The comprehensive similarity algorithm code can be seen from Figure 12 The shown is a schematic diagram of the first part of the comprehensive similarity algorithm code, Figure 13 The shown is a schematic diagram of the second part of the comprehensive similarity algorithm code, Figure 14Fig. 3 shows a schematic diagram of a third part of the code of the comprehensive similarity algorithm, Figure 15 Fig. 4 shows a schematic diagram of a fourth part of the code of the comprehensive similarity algorithm, Figures 12-15 The function name in the code is: composite_similarity, the confidence is converted from the Hamming distance to the output value.

[0059] (3) The segmented picture is used to judge again (such as error detection, classification, counting, etc.) by using the corresponding trained classification AI model (binary classification, product and background) to obtain more accurate results.

[0060] (4) The work subsystem collects sample pictures according to the specified rules, and retains the picture related information (including timestamp, OBB information, product confidence, video frame number, etc.), and outputs the initial judgment result. The picture output in the saving process has the related information in the file name, including OBB size, pixel area (pixel area: SAM2 divides each pixel into a mask confidence value, which is greater than the threshold value, and is defined as 1, and less than the threshold value, which is 0, and the sum of all 1s is the pixel area), angle, video frame number, etc., to ensure that other systems can restore these information data according to the information coding rules when reading the picture data. After the picture and related information are output, the directory structure of the saved directory is: dataset / {product number} / ---{0 background} / ---{1 product} / ---{-1 unknown} / The classification of {-1 unknown} needs to be verified by manual operation, which is simple, and only needs to drag the picture to {0 background} or {1 product} according to the rules to complete the classification. If target detection or product error detection is required, product classification can also be further divided into small types according to the requirements. Here, only the large classification is described. The pictures (products and backgrounds) produced here can be used as training data sets for product target detection, classification, etc. AI model, and no longer need manual annotation. This link is the key method for new products to be online in a short period.

[0061] In this embodiment, a work subsystem AI model hot loading module can also be set: after receiving the instruction, the module loads the new AI model, and then more accurate visual judgment can be performed. Please refer to Figure 16 Fig. 5 shows a flowchart of a work subsystem AI model hot loading module.

[0062] To ensure more than 99% accuracy, two key technologies are used, one is the target tracking algorithm, the product is divided into pictures in the order of video frames according to the motion trajectory combined with the feature judgment is the same; Two is to use Bayes theorem and total probability formula to write a posterior probability, a single judgment can not reach 99%, but the probability of multiple judgments can be greater than 99%.

[0063] 1, target tracking algorithm: there are many target tracking algorithms, in addition, the OBB rectangle can also be used to calculate the intersection ratio and the similarity of the target, and the similarity of OBB rectangle, coordinates and angle can also be tracked to the motion trajectory of the target, plus texture feature comparison, it can track whether the target is the same.

[0064] 2, calculate the posterior probability, use the product confidence calculated from each frame, as long as it reaches 0.8 or more, 3 to 4 frames in a row, it can reach more than 99%, the posterior probability calculation function can be seen in Figure 17 The algorithm code of a posterior probability calculation function is shown in the figure. Running in the python editor can print out: calculate posterior probability: 0.997999.

[0065] In some embodiments, the model training subsystem is used to train the model of the visual classification job system based on the product image and the background image output by the first job module as a training set.

[0066] The model training subsystem in this embodiment is an AI model training system mainly based on automation training and supplemented by manual training, which is a very important part of the system. Referring to Figure 18 An AI model training process diagram is shown. The core principle of the high success rate of the AI model training system is that it uses the picture data set generated by SAM2 segmentation. This picture is an industrial product, the graphic texture is relatively fixed, and it is naturally distinguished from the background picture. In addition, the precise OBB marking, such pictures are trained in AI vision, and the accuracy of the AI classification model or AI target detection model is relatively high.

[0067] In some embodiments, the model of the visual classification job system includes: the classification model of the first job module and the detection model of the second job module.

[0068] The model training subsystem in this embodiment is directly related to the system, mainly for automatic training of two models, one is a single product two-classification AI model, two-classification is to distinguish products and backgrounds, which is used as the enhanced judgment of module 1; The second is a multi-product OBB target detection AI model, which is mainly used as module 2, because this AI model can also segment product pictures, which is consistent with the output content of module 1.

[0069] I. Single product classification AI model training: The main purpose of single product classification training is to improve the accuracy of the generalization ability of module 1 in a short period of time.

[0070] The steps of single product classification AI model training are as follows: listen to the relevant topics of the message server. Receive training topic messages, get product number and AI model type (classification). Read configuration information according to product number. There is the address of the sample picture data set, if not on the local machine, it needs to be downloaded to the local machine through network file tools (such as SVN (centralized version control system), GIT (distributed version control system), FTP (file transfer protocol) and other tools). Split into training dataset (Training Dataset), testing dataset (Testing Dataset) and validation dataset (Validation Dataset) according to a certain proportion. Read training parameters from configuration, start training. After training, output onnx model, upload to configuration specified location. Recommend messages to message server.

[0071] At this step, the training task has been completed, and the job subsystem will load the latest AI classification model after receiving the message, and complete the generalization ability improvement accuracy. And module 1 will rework the video records that are not judged by AI model, and generate job data. If there are different data results, generate a report for professional personnel to analyze and upgrade similarity algorithm or adjust similarity parameters. In this way, the accuracy of the entire system can be continuously improved.

[0072] If a product collects 50 to 100 different pictures as samples, it can be trained with a single product model. After actual testing, using an ordinary graphics card such as RTX 4060 to train alone, it only takes less than 15 minutes to train a YOLO v8 (a target detection model) classification model within 300 pictures. A slow one can be completed within an hour or so, thereby ensuring that the training is completed in a short period of time.

[0073] II. Multi-product target detection AI model training: Target detection AI model training is to train a more mature, more efficient, and more powerful AI model with stronger generalization ability, which is suitable for more product production. The training process is similar to single product classification training, but with an additional step: using inspection special data set. The conditions for starting training are: first, the new product sample data set has accumulated to a certain number; second, the old product has found that the confidence is insufficient, that is, there is a new sample data set, which needs to be incrementally trained to improve the accuracy.

[0074] The steps of the multi-product target detection AI model training are as follows: listening to the message server related topic. Receive training topic message, get product number and AI model type (target detection). According to the product number, read the configuration information. There is the address of the sample picture data set. If not on the local machine, it needs to be downloaded to the local machine through network file tools (such as SVN, GIT, FTP, etc.). Split into training dataset, testing dataset and validation dataset according to a certain proportion. Read the training parameters from the configuration and start the training. After training, use the test special dataset for verification. If it is unqualified, the reason needs to be notified to the artificial to find out, check the dataset or adjust the training parameters and train again. Output onnx model, upload to the configuration specified bit position. Recommend the message to the message server.

[0075] Among them, the training period of the target detection AI model is longer than that of a single product, but it does not affect the operation of the system.

[0076] The above-mentioned visual work classification system provided by the embodiment of the application only needs 15 minutes of operation for new product online, combined with the fixed background of the production line. As long as the product and the background color are not completely the same, manual operation only needs to click a few times on the screen. After the SAM2 model is labeled and segmented, the product picture texture with OBB information is obtained. After target tracking and posterior probability are added, and the judgment of comprehensive similarity is made, the product probability can be confirmed, and counting, classification and subsequent other detection can be realized, which is not disturbed by background texture noise, and the accuracy rate is the same as that of the AI model after deep learning. This stage does not have the generalization ability of neural network, and if the texture is too abnormal, there is a probability of error. Therefore, after collecting a certain data set, a single product classification AI model can be automatically called for training, and the generalization ability can be realized within one hour. In this process, traditional manual labeling is not needed, so that the training process is greatly shortened. Moreover, because of the confidence judgment, sample data sets that have not been used before can be automatically collected, and the generalization ability can be realized after retraining, and more product textures can be compatible.

[0077] Under normal circumstances, module 1 can train a multi-product target detection model in one day. The target detection model also has the ability to segment products, classify and output confidence, and can seamlessly replace the work of module 1 to continue running and iterating the data set in a way of less memory and lower algorithm. In this whole process, since the system functions have been completed, a large amount of work does not need AI professionals to operate and maintain, and only in abnormal circumstances AI professionals need to participate, so that the operation cost is significantly reduced compared with the prior art, thereby realizing the four advantages of short cycle iteration, high efficiency, high accuracy and low cost.

[0078] The above-mentioned visual work classification system provided by the embodiment of the present application mainly has the following advantages: 1. A technical route is created to realize automatic segmentation of products and the same shape background and automatic collection of picture samples by using SAM2 combined with rectangular similarity.

[0079] 2. Segmentation is combined with application: the contour segmentation with OBB rectangular information is used for training of an AI target detection model; the OBB picture texture is rotated according to its angle and then subjected to secondary segmentation to obtain a more compact texture, which has multiple uses: 1. used for pure algorithm feature and similarity comparison without AI model training; 2. used for training of a classification model; 3. used as input of the classification model to obtain more accurate judgment.

[0080] 3. A comparison mode of comprehensive similarity of picture texture is created, which takes into account the comparison of structure and color, and the weight and integrated technical algorithm can be adjusted according to the characteristics of actual texture features, so that the similarity comparison is more in line with actual needs, thereby realizing classification and judgment of pictures and backgrounds in a no-training state.

[0081] 4. A comparison mode of comprehensive similarity of picture texture is created to realize the deduplication operation of picture sample data set, avoid infinite increase of the data set, improve the training efficiency of the AI model, and realize effective incremental training.

[0082] 5. The basic structure and process of the system are designed, an automatic process is constructed, manual work is reduced, and accuracy is ensured.

[0083] 6. A mode of saving all pictures and data of products segmented based on manual operation for automatic verification after completion of related AI model training is created, which simplifies the operation steps and improves the success rate of AI model training.

[0084] In addition, the above-mentioned visual work classification system provided by the embodiment also has some alternative solutions: 1. The codes used in the description of the system are all examples of python, in order to make the codes easier to understand, and tensor data is not used, the language used in the actual deployment of the system is uncertain, it also depends on the customer's needs, what kind of hardware, system running, etc., need to be converted into the corresponding program. Some programs are concurrent to improve efficiency, and will also run on GPU (graphics processing unit), GPU languages include Cuda (parallel computing platform and programming model), GLSL (high-level shading language based on C language), Vulkan (cross-platform graphics API), etc.

[0085] 2. The example only shows the target similarity + Hamming similarity algorithm for comprehensive similarity. In reality, more combined algorithms can be used to calculate the similarity based on the texture of the image. For example, if more consistency is emphasized, the SSIM (Structural Similarity Index) similarity algorithm should be added, or other similarity algorithms such as Gray-Level Co-occurrence Matrix (GLCM), Gray-Level Histogram, Local Binary Pattern (LBP), etc. This embodiment does not limit the specific similarity algorithm.

[0086] 3. In this embodiment, it is mentioned that after SAM2 segmentation, the shape is filtered first and then the binary classification judgment of product or background is performed. This is to save computing power (texture calculation consumes more computing power than shape calculation, and filtering can reduce texture calculation and thus reduce computing power). However, if we do not consider the waste of computing power, especially on GPU, it can be calculated in parallel, and the shape similarity, product texture similarity and background texture similarity are output at the same time.

[0087] 4. In this case study, the AI ​​model was mainly trained using YOLOv8, but the same capability can be achieved by using other YOLOvN models, building your own neural network, or using other training platforms.

[0088] 5. Although this system only emphasizes classification, this is only a foundation. If the product image dataset is further subdivided on this basis, it is possible to realize more application scenarios such as product counting, product classification, error detection, and even defect detection.

[0089] 6. This system does not explicitly mention unsupervised learning or weakly supervised learning because of considerations regarding the maturity of the model. After testing, the accuracy and cycle time of these two learning methods do not yet meet the requirements in industrial environments. If these two technologies prove to be mature and easy to use in the future, they can be incorporated into the AI ​​model training process.

[0090] Example 3: This invention also provides an electronic device, including the visual classification system provided in the foregoing embodiments.

[0091] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and / or device described above can be referred to the corresponding process in the foregoing embodiments, and will not be repeated here.

[0092] In addition, in the description of the embodiments of the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connection", "connecting" should be understood in a broad sense, for example, can be fixed connection, can also be detachable connection, or integrally connected; can be mechanical connection, can also be electrical connection; can be directly connected, can also be indirectly connected through an intermediate medium, can be the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0093] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the part of the prior art that essentially contributes to the present application or the part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, and various program code storage media.

[0094] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. In addition, the terms "first", "second", "third" are only for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0095] Finally, it should be noted that: the above-described embodiments are only specific embodiments of the present application, which are used to illustrate the technical solutions of the present application, and are not limiting, the protection scope of the present application is not limited thereto, although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art within the technical range disclosed by the present application can modify or easily think of changes to the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications, changes or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A visual classification system, characterized in that, The visual classification system includes: a task subsystem, a model training subsystem, and a client operation interface; The operation subsystem is used to count, classify, detect and identify products on the production line; The operation subsystem includes: a first operation module, a second operation module, and a scheduling module; the first operation module is used to encode and decode single-frame images of video and output product images; the second operation module is used to input the single-frame images of video into a trained detection model and output the product images; the scheduling module is used to schedule the use of the first operation module and the second operation module according to the training maturity of the current production product. The model training subsystem is used to connect to the message server after startup and train the model of the visual classification task system through message events. The client operation interface is used to establish real-time communication with the operation subsystem and display the output information of the operation subsystem; the client operation interface is also used to stop the video frame of the current frame and edit the data based on the user's operation on the video frame of the current frame.

2. The visual classification system according to claim 1, characterized in that, The first task module is also used to output background images, product information, and task data; wherein, the product information includes: segmented rectangle, coordinates, directional bounding box information, segmented pixel area, video frame number, and confidence level; If manual operation is involved, the first operation module is also used to output the single-frame image of the video and product segmentation information; The second operation module is also used to output the product information and the operation data.

3. The visual classification system according to claim 1, characterized in that, If the currently produced product is a new product, the scheduling module is used to utilize the first job module; If the currently produced product follows a preset work model, the scheduling module is used to utilize the second work module; If the second job module has an output anomaly during the use of the second job module, the first job module is used to perform a job on the video single frame image with the output anomaly. The scheduling module is used to process the output results of the first job module and the second job module. The processing methods include: collecting sample datasets, reporting anomalies, and transferring to manual verification.

4. The visual classification system according to claim 2, characterized in that, The first task is equipped with a segmentation model; the segmentation model is used to encode and decode single-frame images of real-time video from the production line inspection point at a fixed frame rate, and to segment the single-frame images of the video according to the target size. The first task module is also used to determine whether the segmentation result is the product image or the background image, and to rotate and perform secondary segmentation on the segmentation result in sequence for display. The first task module is also used to process the results of the secondary segmentation display using a classification model; the processing methods include: error detection, classification, and counting. The first task module is also used to collect the product image and the product information of the product image, and output the initial judgment result of the single frame image of the video.

5. The visual classification system according to claim 4, characterized in that, If the currently produced product is a new product, the first operation module is used to notify the operator to perform a product segmentation and selection operation for the currently produced product.

6. The visual classification system according to claim 4, characterized in that, The segmentation model is used to segment and output multiple rectangles of the single frame image of the video using a coordinate dot matrix. The first task module is used to determine whether each rectangle is a product image or a background image by weighted similarity judgment. The weighted similarity judgment method includes: calculating multiple similarities through multiple similarity calculation methods and weighting the multiple similarities.

7. The visual classification system according to claim 6, characterized in that, The various similarity calculation methods include: target similarity calculation method and Hamming similarity calculation method; The target similarity calculation method is used to calculate the similarity between the first value and the second value using the following formula: s=|ab| / (|b|+ε), f=1-min(s,1); where s is the intermediate value of the target similarity calculation method, a is the first value, b is the second value, ε is the preset minimum representable value, f is the similarity of the target similarity calculation method, and min(s,1) is the smaller value between s and 1.

8. The visual classification system according to claim 4, characterized in that, The model training subsystem is used to train the model of the visual classification task system based on the product image and background image output by the first task module as a training set.

9. The visual classification system according to claim 8, characterized in that, The model of the visual classification task system includes: the classification model of the first task module and the detection model of the second task module.

10. An electronic device, characterized in that, The electronic device includes: the visual classification system according to any one of claims 1-9.

Citation Information

Patent Citations

  • Visual processing and model training method and device, storage medium and program product

    CN114549904A

  • Food image classification method and system, storage medium and terminal

    CN114842266A

  • Image classification model training method, image classification method, device and equipment

    CN116863184A

  • Job recognition system, job recognition method, and program

    CN119337076A

  • Video classification method and device

    CN119851182A