Machine vision method based on intention instruction interaction driving

Through the machine vision method driven by interactively driven intent instructions, efficient switching and instruction execution between different visual models is achieved, and the problems of high operation difficulty and large video memory usage in the prior art are solved, which improves the inference efficiency of the algorithm and retains the recognition accuracy.

CN120163244APending Publication Date: 2025-06-17BEIJING SPACEFLIGHT TUOPUGAO SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510209686.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

When the task requirements of existing machine vision models change in different scenarios, they need to operate manually to switch the visual model, which increases the difficulty of operation. The trained model will call all parameters every time it is used, occupying a large video memory and making it difficult to apply on small devices.

Method used

Using a machine vision method based on interaction-driven intent instructions, the combination of intent analysis model and visual model can achieve efficient switching and instruction execution between different visual models, which reduces the threshold for operators and improves the inference efficiency of the algorithm.

Benefits of technology

It realizes efficient switching between different visual models, reduces operation difficulty, improves the inference efficiency of the algorithm, and retains the recognition accuracy of traditional visual algorithms, which is suitable for small devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163244A_ABST
    Figure CN120163244A_ABST
Patent Text Reader

Abstract

The invention discloses a machine vision method based on intention instruction interaction driving. According to the method, an intention text instruction data set is used for training a text encoder, so that the text encoder can encode text instructions into features which can be understood by a visual model; then, pre-training a visual model E by using the image data set Mg with the label and the cross entropy loss, so that the visual model E can output probability distribution Pg of an image sample; and finally, constructing an image-text combined data set O, training an image-text aggregation module by using region-text comparison loss and the data set O, receiving a text vector feature parameter C transmitted by the intention analysis model by using the module, and selecting a visual model Ea to execute a visual task. According to the method, an intention instruction interaction mode is introduced, the visual model is enabled to focus on a certain type of objects during task execution instead of detecting all training types at the same time, compared with a traditional visual algorithm, the reasoning speed is higher, fewer resources are occupied, and meanwhile, the detection precision of the visual algorithm can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of machine vision, and particularly relates to a machine vision method driven by intent instruction interaction. Background Art

[0002] As an important branch of artificial intelligence, machine vision algorithms simulate the human visual system to achieve efficient processing and analysis of images and videos. This technology covers multiple fields such as image processing, pattern recognition, and machine learning, aiming to extract useful features and information from images to achieve tasks such as automatic recognition, detection, and measurement.

[0003] Today's vision algorithms have various types of models to specifically solve practical problems in different application fields according to the characteristics of real scenes. For example, in security monitoring, object detection models are used to identify people or vehicles breaking into restricted areas; in the field of medical imaging, semantic segmentation can be used to effectively divide and classify different tissue regions in images, such as gray matter, white matter, and cerebrospinal fluid in brain magnetic resonance images, improving the diagnostic accuracy of doctors for diseases. However, existing models are independently deployed for specific scenarios. Whenever there are task requirements for different scenarios, manual operations are required to switch to other visual models to continue executing tasks, which increases the operation difficulty of staff and is time-consuming and laborious. In addition, previous vision algorithms perform visual tasks for pre-set object categories. In this way, all the parameters of the trained model are called for detection each time it is used, occupying a large amount of video memory, and there will be problems that it is difficult to apply to small devices. Summary of the Invention

[0004] In view of the above defects or deficiencies in the prior art, the present invention aims to provide a machine vision method driven by intent instruction interaction, including two parts: an intent parsing model and a vision model. The algorithm uses an intent instruction-driven method to switch between different vision models and execute relevant tasks of the objects contained in the instructions, achieving an efficient reasoning process and retaining the accuracy of the original vision model. The present invention enhances the ease of operation of the application end of vision algorithms, reduces the usage threshold of personnel, improves the reasoning efficiency of the algorithm, and has excellent recognition accuracy of traditional vision algorithms.

[0005] The purpose of the present invention is achieved as follows: A machine vision method driven by intent instruction interaction switches between different vision models through intent instructions, including the following steps:

[0006] Step S1, construct an intent parsing text dataset X driven by intent instructions t , the dataset X t includes intent instruction data, intent instruction labels, and intent instruction drivers;

[0007] Step S2, use the intent parsing text dataset X t and the intent parsing loss to train the intent parsing model I;

[0008] Step S3, use k-fold cross-validation to verify that the average test accuracy of the intent parsing model I reaches the set value;

[0009] Step S4, use the labeled image dataset M g and the cross-entropy loss to pre-train the visual model E so that it can output the probability distribution P of the image samples g ;

[0010] Step S5, construct the dataset O composed of image-text combinations, use the region-text alignment loss and the dataset O to train the image-text aggregation module, which is used to receive the text vector feature parameter C transmitted by the intent parsing model I, and select the corresponding model E in the visual model E according to this parameter a to perform visual tasks.

[0011] Furthermore, the calculation method of the intent parsing loss in step S2 is as follows:

[0012]

[0013] where x t is a piece of data in the intent parsing text dataset X t and y t is its corresponding label, n is the total number of dataset samples, and I(x t ) represents the vector feature encoded by the intent parsing model I for the text sample x t .

[0014] Furthermore, the calculation method using k-fold cross-validation in step S3 is as follows:

[0015]

[0016] where CV(K) represents the mean squared error of K-fold cross-validation, K represents the number of folds for splitting the dataset,

[0017] Y test represents the true value of the test, and Y pred represents the predicted value of the model.

[0018] Furthermore, the calculation method of the cross-entropy loss in step S4 is as follows:

[0019]

[0020] where Y g is the image dataset M gThe true distribution, P g is the probability distribution output by the visual model for the image dataset M g , is the cross-entropy loss.

[0021] Furthermore, the dataset composed of the image-text combinations constructed in step S5 is where Img i is an image sample, T i is the text statement manually labeled for it, and N is the total number of samples.

[0022] Furthermore, the calculation formula for the region-text alignment loss in step S5 is:

[0023]

[0024] where is the region-text alignment loss, (Img, T) i is an image-text pair sample in the image-text dataset O, (E(Img), Sim(C, T)) i is a feature pair composed of the image encoding feature and the text similarity feature. E(Img) is the feature encoded by the visual model for the image, and Sim(C, T) is the similarity calculation method between the text vector feature parameter C transmitted by the intent parsing model I and the text vector space T in the dataset O. And:

[0025]

[0026] Furthermore, the set value of the average test accuracy in step 3 is at least 95%. When the average test accuracy does not reach 95%, the proportion of each intent data in the dataset X t can be adjusted. If the accuracy rate driven by an intent instruction in the intent parsing model I is lower than 95%, then increase the amount of data for this intent instruction in the dataset X t so that it can be trained more sufficiently.

[0027] The beneficial effects of the present invention are as follows: A machine vision method based on intent instruction interaction driving according to the present invention realizes efficient switching between different visual models through intent instructions, has faster inference ability, and retains the accuracy of the original visual model. At the same time, the present invention enhances the operability of the visual algorithm application end, reduces the usage threshold of personnel, improves the computing efficiency of the algorithm, and retains the excellent recognition accuracy of the traditional visual algorithm.

[0028] The present invention will be further elaborated in detail below in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings

[0029] Figure 1It is a flowchart of a machine vision method based on intent instruction interaction drive according to the present invention;

[0030] Figure 2 It is a schematic diagram of the training of the intent parsing model I in step S2;

[0031] Figure 3 It is a schematic diagram of the k-fold cross-validation method in step S3;

[0032] Figure 4 It is a schematic diagram of training the image-text aggregation module using the region-text alignment loss and the image-text dataset O in step S5. Detailed implementation mode

[0033] Example 1:

[0034] As Figure 1 shown, a machine vision method based on intent instruction interaction drive, which realizes the switching between different vision models through intent instructions, includes the following steps:

[0035] Step S1, construct an intent parsing text dataset X based on intent instruction drive t , the dataset X t includes intent instruction data, intent instruction labels, and intent instruction drives;

[0036] For example, in the text dataset X t , the intent instruction data is: "Help me detect the wearing situation of work clothes", the intent instruction label is "detection-0, work clothes-1", and the intent instruction drive is "detect work clothes". Another example: "Device No. 1 is operating, which is very dangerous. Help me detect the people near Device No. 1", the intent instruction label is "detection-0, (Device No. 1, people)-1, distance estimation-2, (Device No. 1, people)-3", and the intent instruction drive is "detect Device No. 1 and people, and use the depth estimation model to judge the distance between the two".

[0037] Step S2, as Figure 2 shown, use the intent parsing text dataset X t and the intent parsing loss to train the intent parsing model I, and the calculation method of the intent parsing loss is:

[0038]

[0039] where x t is a piece of data in the intent parsing text dataset X t , y t is its corresponding label, n is the total number of dataset samples, and I(x t ) represents the text sample x tThe vector features encoded by the intent parsing model I.

[0040] The intent parsing loss is an important feedback signal for model training. During the training process, the optimization algorithm adjusts the model's parameters based on the loss value. If the loss value is large, it indicates that the model's performance is poor and significant adjustments to the parameters are needed; as training progresses, the loss value gradually decreases, and the model's ability to parse intents gradually improves. When the loss value converges to a small value, it means that the model has achieved good performance in the intent parsing task.

[0041] Step S3: Use k-fold cross-validation to verify the average test accuracy of the intent parsing model I and ensure it is above 95%. When the average test accuracy does not reach 95%, the proportion of each intent data in the dataset X can be adjusted. If the accuracy of an intent instruction driven by an intent instruction in the intent parsing model I is lower than 95%, then increase the t amount of data for that intent instruction in the dataset X t so that it can be more fully trained.

[0042] As Figure 3 shown, the light green and light gray parts in the figure represent all the samples of the dataset, the dark gray represents the part of the dataset divided for training, and the green part represents the part of the dataset divided for testing. In this example, taking the dataset divided into five folds as an example, each fold is used as the test data respectively, and the remaining four folds are used as the training data for experiments. The calculation method of k-fold cross-validation in step S3 is as follows:

[0043]

[0044] where CV(K) represents the mean squared error of k-fold cross-validation, K represents the number of folds into which the dataset is divided, and Y test represents the true value of the test, and Y pred represents the predicted value of the model.

[0045] The evaluation results using k-fold cross-validation are usually more accurate and stable. Because there are more data partitioning methods, it can better reflect the model's performance on different data subsets and reduce the accidental impact of data partitioning.

[0046] Step S4: Use the labeled image dataset M g and cross-entropy loss to pre-train the visual model E so that it can output the probability distribution P g of the image samples. The calculation method of the cross-entropy loss is as follows:

[0047]

[0048] where Y g is the image dataset Mg The true distribution, P g is the probability distribution output by the visual model for the image dataset M g , is the cross-entropy loss.

[0049] Step S5, constructing a dataset composed of image-text combinations. As Figure 4 shown, use the region-text alignment loss and the image-text dataset O to train the image-text aggregation module. After sufficient training, this module can be used to receive the text vector feature parameter C transmitted by the intent parsing model I, and select the corresponding model E in the visual model E according to this parameter a to perform visual tasks.

[0050] The dataset composed of image-text combinations constructed in step S5 is where Img i is an image sample, T i is the text statement manually marked for it, and N is the total number of samples. The calculation formula for the region-text alignment loss in step S5 is:

[0051]

[0052] where is the region-text alignment loss, (Img, T) i is an image-text pair sample in the image-text dataset O, (E(Img), Sim(C, T)) i is a feature pair composed of the image encoding feature and the text similarity feature. E(Img) is the feature encoded by the visual model for the image, and Sim(C, T) is the similarity calculation method between the text vector feature parameter C transmitted by the intent parsing model I and the text vector space T in the dataset O, and:

[0053]

[0054] The image-text dataset O composed of image-text combinations contains image-text pair samples, (Img, T) i is a pair of samples in the dataset O. Img is the image sample encoding feature, and T is the text feature description of the image sample. For example, (the image sample of work clothes, "work clothes"), the text "work clothes" is the text feature description of the work clothes image sample feature;

[0055] The intent instruction drive in the intent parsing model I is "detect work clothes", and the text vector feature parameter is C;

[0056] The trained image-text aggregation module can, through prior experience, select the "object detection" model E in the visual model E a, then compare the text vector feature parameter C with the text vector space T, search for the corresponding image samples of "work uniforms", and finally select and execute the relevant visual task of "detecting work uniforms".

[0057] A machine vision method driven by intent instruction interaction. First, use the intent text instruction dataset X t , train the text encoder so that it can encode the text instructions into features that the visual model can understand; then, use the labeled image dataset M g and the cross-entropy loss to pre-train the visual model E so that it can output the probability distribution P of the image samples g ; finally, construct the dataset O composed of images and texts, and use the region-text comparison loss and the dataset O to train the image-text aggregation module. This module is used to receive the text vector feature parameter C transmitted by the intent parsing model and select the corresponding model E in the visual model E according to this parameter a to execute the visual task. The present invention introduces the method of intent instruction interaction, which reduces the operation threshold of the operator, and enables the visual model to focus on a certain type of object when performing tasks, rather than detecting all training categories at the same time. Compared with the traditional visual algorithm, it has a faster inference speed, less resource occupancy, and can ensure the detection accuracy of the visual algorithm.

[0058] Finally, it should be noted that the above is only used to illustrate the technical solution of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred arrangement, those of ordinary skill in the art should understand that the technical solution of the present invention (such as the application of various formulas, the sequence of steps, etc.) can be modified or equivalently replaced without departing from the spirit and scope of the technical solution of the present invention.

Claims

1. A machine vision method based on intention instruction interaction driving, which realizes switching between different visual models through intention instructions, characterized in that: The steps include: Step S1: construct an intention parsing text dataset X driven by intention instructions t , dataset X t It includes intention instruction data, intention instruction label and intention instruction driver; Step S2: Use intent to parse text dataset X t and intent parsing loss Training intent parsing model I; Step S3, using a k-fold cross validation method to verify that the average test accuracy of the intent parsing model I reaches a set value; Step S4, using the labeled image dataset M g Pre-train the visual model E with cross entropy loss so that it can output the probability distribution P of image samples g ; Step S5, constructing a dataset O composed of images and texts, using the region-text comparison loss and the dataset O to train an image-text aggregation module, which is used to receive the text vector feature parameter C transmitted by the intent parsing model I, and select the corresponding model E in the visual model E according to the parameter a Perform visual tasks.

2. The machine vision method based on intention instruction interaction drive according to claim 1 is characterized in that: Intent parsing loss in step S2 The calculation method is: where x t Parsing text dataset X for intent t A piece of data, y t is its corresponding label, n is the total number of samples in the data set, I(x t ) represents the text sample x t Vector features encoded by intent parsing model I.

3. The machine vision method based on intention instruction interaction drive according to claim 1 is characterized in that: The calculation method using k-fold cross validation in step S3 is: Where CV(K) represents the mean square error of K-fold cross validation, K represents the number of folds into which the data set is divided, and Y test represents the true value of the test, Y pred Represents the predicted value of the model.

4. The machine vision method based on intention instruction interaction drive according to claim 1 is characterized in that: The cross entropy loss in step S4 is calculated as: where Y g The image dataset M g The true distribution of P g For the visual model to image dataset M g The probability distribution of the output, is the cross entropy loss.

5. The machine vision method based on intention instruction interaction drive according to claim 1 is characterized in that: The image-text combination dataset constructed in step S5 is Img i is the image sample, T i is a text sentence that is manually labeled, and N is the total number of samples.

6. The machine vision method based on intention instruction interaction drive according to claim 5 is characterized in that: The calculation formula of region-text comparison loss in step S5 is: in is the region-text comparison loss, (Img,T) i is an image-text pair sample in the image-text dataset O, (E(Img),Sim(C,T)) i is a feature pair composed of image encoding features and text similarity features, E(Img) is the feature of image encoding by the visual model, Sim(C,T) is the similarity calculation method between the text vector feature parameter C transmitted by the intention parsing model I and the text vector space T in the data set O, and:

7. The machine vision method based on intention instruction interaction drive according to claim 1 is characterized in that: The average test accuracy setting value in step 3 is at least 95%. When the average test accuracy does not reach 95%, the data set X can be adjusted. t If the accuracy of the intent instruction driven by a certain intent instruction in the intent parsing model I is lower than 95%, then increase the data set X t The amount of data for the intent command in the training data can be more fully trained.