Infrared point target identification method and system based on visual language model
Through the method based on visual language model, learning the language and visual features of infrared point target image datasets are solved, and the problem of low accuracy of infrared point target recognition is achieved, and higher recognition accuracy and lower manual labeling cost are achieved.
Patent Information
- Application Number
- CN202311607692.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the accuracy of infrared point target recognition is low, mainly due to the low infrared imaging clarity, strong background noise, few target pixels, and weak grayscale, which makes it difficult to extract the key target features.
Using a visual language model-based method, we learn the high-level semantic feature association between language and images by obtaining the language description, high-level semantic features and word vectors of target semantic categories of infrared point target image data sets, and combining the language semantic coding model and visual semantic coding model, and then learning the high-level semantic feature association between language and images, thereby improving the accuracy of infrared point target recognition.
It improves the accuracy of infrared point target recognition, reduces dependence on image or language labels, saves manual labeling costs, and improves the performance and generalization capabilities of the model in complex scenarios.
Smart Images

Figure CN120070840A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision image recognition, and in particular to an infrared point target recognition method and system based on a vision language model. Background Art
[0002] As a basic and core task in the field of computer vision, image recognition has developed rapidly in recent years. It requires extracting key features in an image and accurately judging the target categories in the image. Deep learning methods have made significant breakthroughs in image recognition tasks, with performance exceeding traditional methods. At the same time, the key features extracted by image recognition can also be used for subsequent tasks such as target detection and image segmentation. Its feature extraction ability directly affects the accuracy of subsequent tasks and plays a crucial role in promoting the development of other computer vision tasks.
[0003] At the level of target recognition methods. Traditional methods statistically analyze the apparent features of the target, extract key features such as mean, variance, contour, and area, and use manually designed thresholds or simple machine learning methods, such as AdaBoost, SVM, BP neural network, etc., as target recognition means. However, it requires certain target prior knowledge as a reference for the design of recognition methods and is limited by a large number of threshold tuning parameters, and cannot adapt to large dynamic scene changes. Existing deep learning methods adjust a large number of model parameters through multi-level convolutional operations and through true value labels and feedback optimization. By this learnable way, key features of the target are extracted to complete the target recognition task. However, it requires a large number of labeled data samples for training. In practice, a dataset with complete labels requires a large amount of human annotation cost to obtain, and the trained model only has good effects in scenarios with small differences from the test set data, and has insufficient adaptability to complex and diverse scenarios.
[0004] In particular, at the level of infrared point target recognition tasks. Discovering and effectively recognizing targets at a farther distance is a long-term goal pursued by infrared detection systems. However, due to factors such as long detection distance, strong background noise, infrared detector manufacturing process and materials, etc., the clarity of infrared imaging is low, the number of target pixels is small, and the gray level is weak, resulting in few types of extractable apparent features of infrared point targets, poor effectiveness, and lack of high-level semantic information, making it impossible to extract key features of the target, forming a difficult problem in infrared point target recognition, and the recognition accuracy is low. Summary of the Invention
[0005] The present invention provides an infrared point target recognition method and system based on a vision language model, which can solve the technical problem of low infrared point target recognition accuracy in the prior art.
[0006] According to one aspect of the present invention, an infrared point target recognition method based on a vision language model is provided, and the method includes:
[0007] Obtain word vectors containing the language descriptions, high-level semantic features, and target semantic categories of each image based on the infrared point target image dataset, prompt engineering, and prior information;
[0008] Input the word vectors containing the language descriptions, high-level semantic features, and target semantic categories of each image into the language semantic encoding model to obtain language high-level semantic feature encodings;
[0009] Input each image in the infrared point target image dataset into the visual semantic encoding model to obtain visual high-level semantic feature encodings;
[0010] Obtain the similarity between the language high-level semantic feature encoding and the visual high-level semantic feature encoding, and obtain the distance loss between the language high-level semantic feature encoding and the visual high-level semantic feature encoding based on the similarity;
[0011] Train the language semantic encoding model and the visual semantic encoding model based on the distance loss between the language high-level semantic feature encoding and the visual high-level semantic feature encoding to obtain an updated vision-language model;
[0012] Determine the category of the infrared point target to be detected according to the recognition requirement, and input the infrared point target to be detected and the category of the infrared point target to be detected into the updated vision-language model to realize the recognition of the infrared point target.
[0013] Preferably, obtaining word vectors containing the language descriptions, high-level semantic features, and target semantic categories of each image based on the infrared point target image dataset, prompt engineering, and prior information includes:
[0014] Based on prior information, make a preliminary judgment on each image in the infrared point target image dataset to obtain candidate keywords corresponding to each image, score the key confidence of each candidate keyword, and screen the candidate keywords based on the scoring results to obtain a unified expression of credible language description keywords for each image;
[0015] Use prompt engineering to enhance the language expression of the unified expression of credible language description keywords for each image to obtain statements containing the language descriptions, high-level semantic features, and target semantic categories of each image;
[0016] Convert the statements containing the language descriptions, high-level semantic features, and target semantic categories of each image into word vectors containing the language descriptions, high-level semantic features, and target semantic categories of each image by using word vector encoding.
[0017] Preferably, the unified expression of credible language description keywords for each image is obtained by the following formula:
[0018]
[0019] In the formula, Wkey For the unified expression of the keywords of the credible language descriptions of each image, W i is the candidate keyword with a confidence level higher than the screening threshold, and N is the number of keywords passed through the screening.
[0020] Preferably, the statements including the language descriptions, high-level semantic features, and target semantic categories of each image are obtained through the following formula:
[0021] P(W key ) = L description (W key ) + L semantics (W key ) + L category (W key ) ;
[0022] In the formula, P(W key ) is the statement including the language description, high-level semantic feature, and target semantic category of each image, L description (W key ) is the language description prompt statement, L semantics (W key ) is the high-level semantic feature prompt statement, L category (W key ) is the target semantic category prompt statement.
[0023] Preferably, the word vectors including the language descriptions, high-level semantic features, and target semantic categories of each image are obtained through the following formula:
[0024] W embeding (P(W key )) = D(P(W key ))
[0025] In the formula, W embeding (P(W key )) is the word vector including the language description, high-level semantic feature, and target semantic category of each image, P(W key ) is the statement including the language description, high-level semantic feature, and target semantic category of each image, and D(·) is the word vector dictionary encoding operation.
[0026] Preferably, the word vectors including the language descriptions, high-level semantic features, and target semantic categories of each image are input into the language semantic encoding model, and the obtained language high-level semantic feature encoding includes:
[0027] The input encoding layer in the language semantic encoding model processes the input word vector data through dimension conversion and position and category encoding to obtain the pre-encoded features of the language semantic encoding model;
[0028] The feature extraction layer in the language semantic encoding model extracts key features from the pre-encoded features of the language semantic encoding model through Transformer, obtaining the key semantic features of the language semantic encoding model;
[0029] The feature mapping layer in the language semantic encoding model refines and fuses the key semantic features of the language semantic encoding model through mapping transformation, obtaining the high-level semantic feature encoding of the language.
[0030] Preferably, inputting each image in the infrared point target image dataset into the visual semantic encoding model, obtaining the high-level visual semantic feature encoding includes:
[0031] The input encoding layer in the visual semantic encoding model processes each input image through dimension conversion and position and category encoding, obtaining the pre-encoded features of the visual semantic encoding model;
[0032] The feature extraction layer in the visual semantic encoding model extracts key features from the pre-encoded features of the visual semantic encoding model through Transformer, obtaining the key semantic features of the visual semantic encoding model;
[0033] The feature mapping layer in the visual semantic encoding model refines and fuses the key semantic features of the visual semantic encoding model through mapping transformation, obtaining the high-level visual semantic feature encoding.
[0034] Preferably, the distance loss between the high-level semantic feature encoding of the language and the high-level visual semantic feature encoding is obtained by the following formula:
[0035]
[0036] In the formula, L dis (I,T) is the distance loss between the high-level semantic feature encoding of the language and the high-level visual semantic feature encoding, S(I,T) is the similarity between the high-level semantic feature encoding of the target language of the same category and the high-level visual semantic feature encoding, I is the high-level visual semantic feature encoding, T is the high-level semantic feature encoding of the language that matches the high-level visual semantic feature encoding, S(I,T j ) is the similarity between the high-level semantic feature encoding of the target language of different categories and the high-level visual semantic feature encoding, T j is the high-level semantic feature encoding of the language that does not match the high-level visual semantic feature encoding, M is the number of high-level semantic feature encodings of the language that do not match, τ is the temperature parameter, exp(·) is the exponential function, and log(·) is the natural logarithm function.
[0037] Preferably, the infrared point target category to be detected is determined according to the recognition requirement, and the infrared point target to be detected and the infrared point target category to be detected are input into the updated vision-language model. The infrared point target recognition includes: determining the infrared point target category to be detected according to the recognition requirement, inputting the infrared point target category to be detected into the updated language semantic encoding model to update the word vectors, and recognizing the infrared point target to be detected based on the updated vision semantic encoding model; wherein, the updated vision-language model includes the updated language semantic encoding model and the updated vision semantic encoding model.
[0038] According to another aspect of the present invention, there is provided an infrared point target recognition system based on a vision-language model, the system comprising:
[0039] A language description and prompt module, configured to obtain word vectors including language descriptions, high-level semantic features, and target semantic categories of each image based on an infrared point target image dataset, prompt engineering, and prior information;
[0040] A training module, configured to input the word vectors including language descriptions, high-level semantic features, and target semantic categories of each image into a language semantic encoding model to obtain language high-level semantic feature encodings; configured to input each image in the infrared point target image dataset into a vision semantic encoding model to obtain vision high-level semantic feature encodings; configured to obtain the similarity between the language high-level semantic feature encodings and the vision high-level semantic feature encodings, and obtain the distance loss between the language high-level semantic feature encodings and the vision high-level semantic feature encodings based on the similarity; and further configured to train the language semantic encoding model and the vision semantic encoding model based on the distance loss between the language high-level semantic feature encodings and the vision high-level semantic feature encodings to obtain an updated vision-language model;
[0041] An inference module, configured to determine the infrared point target category to be detected according to the recognition requirement, and input the infrared point target to be detected and the infrared point target category to be detected into the updated vision-language model to implement infrared point target recognition.
[0042] According to still another aspect of the present invention, there is provided a computer device, including a memory, a processor, and an infrared point target recognition program based on a vision-language model stored in the memory and executable on the processor. When the processor executes the infrared point target recognition program based on the vision-language model, the above-mentioned method is implemented.
[0043] Applying the technical solution of the present invention, the infrared point target is transformed into high-level semantic features of language through prior information, as a supplement to the problems of few extractable feature types, poor effectiveness, and lack of high-level semantic information of the infrared point target appearance, solving the problem of difficult extraction of key features of the target by the visual model and improving the recognition accuracy of the infrared point target. In addition, since the two models are trained by learning the high-level semantic feature association between language and image, there is no need to label the images or language, which is an unsupervised training method, saving a large amount of manual annotation costs. At the same time, the low-cost training method enables the model to easily learn more data, and under the guidance of language features, improves the performance and generalization ability of the model in more complex and unknown scenarios. Therefore, the present invention uses prior information to learn high-level semantic features at the language level, which can better guide the visual model to extract and recognize the features of the infrared point target in detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings included are used to provide a further understanding of the embodiments of the present invention, which form a part of the specification, are used to illustrate the embodiments of the present invention, and together with the text description are used to explain the principles of the present invention. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0045] Figure 1 The flowchart of the infrared point target recognition method based on the vision-language model provided by an embodiment of the present invention is shown;
[0046] Figure 2 The flowchart of obtaining word vectors provided by an embodiment of the present invention is shown;
[0047] Figure 3 The flowchart of obtaining the high-level semantic feature encoding of language provided by an embodiment of the present invention is shown;
[0048] Figure 4 The flowchart of obtaining the high-level semantic feature encoding of vision provided by an embodiment of the present invention is shown;
[0049] Figure 5 The structural schematic diagram of the infrared point target recognition system based on the vision-language model provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and in no way limits the present invention and its application or use. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts belong to the scope of protection of the present invention.
[0051] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless otherwise clearly specified in the context, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0052] Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present invention. At the same time, it should be understood that, for the sake of convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship. Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and devices should be regarded as part of the authorization specification. In all the examples shown and discussed herein, any specific value should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0053] As Figure 1 shown, the present invention provides an infrared point target recognition method based on a vision-language model, and the method includes:
[0054] S10. Obtain word vectors including language descriptions, high-level semantic features, and target semantic categories of each image based on an infrared point target image dataset, prompt engineering, and prior information;
[0055] S20. Input the word vectors including language descriptions, high-level semantic features, and target semantic categories of each image into a language semantic encoding model to obtain a language high-level semantic feature encoding, that is, a text feature encoding;
[0056] S30. Input each image in the infrared point target image dataset into the visual semantic encoding model to obtain visual high-level semantic feature encodings, that is, image feature encodings;
[0057] S40. Obtain the similarity between the language high-level semantic feature encoding and the visual high-level semantic feature encoding, and obtain the distance loss between the language high-level semantic feature encoding and the visual high-level semantic feature encoding based on the similarity;
[0058] S50. Train the language semantic encoding model and the visual semantic encoding model based on the distance loss between the language high-level semantic feature encoding and the visual high-level semantic feature encoding to obtain an updated vision-language model;
[0059] S60. Determine the category of the infrared point target to be detected according to the recognition requirement, and input the infrared point target to be detected and the category of the infrared point target to be detected into the updated vision-language model to achieve infrared point target recognition.
[0060] In the present invention, the infrared point target is transformed into language high-level semantic features through prior information, as a supplement to the problems of few extractable feature types, poor effectiveness, and lack of high-level semantic information of the infrared point target appearance, so as to solve the problem that it is difficult for the visual model to extract key features of the target, and improve the recognition accuracy of the infrared point target. In addition, since the two models are trained by learning the high-level semantic feature association between language and image, there is no need to label the images or language, which is an unsupervised training method and saves a large amount of manual annotation costs. At the same time, the low-cost training method enables the model to easily learn more data, and under the guidance of language features, improves the performance and generalization ability of the model in more complex and unknown scenarios. Therefore, the present invention uses prior information to learn high-level semantic features at the language level, which can better guide the visual model to extract and recognize the features of point targets in infrared detection.
[0061] As Figure 2 shown, according to an embodiment of the present invention, in S10 of the present invention, obtaining the word vectors including the language description, high-level semantic features, and target semantic categories of each image based on the infrared point target image dataset, prompt engineering, and prior information includes:
[0062] S11. Make a preliminary judgment on each image in the infrared point target image dataset based on prior information to obtain candidate keywords corresponding to each image, score the key confidence of each candidate keyword, and screen the candidate keywords based on the scoring results to obtain a unified expression of credible language description keywords for each image;
[0063] S12. Use prompt engineering to enhance the language expression of the unified expression of credible language description keywords for each image to obtain statements including the language description, high-level semantic features, and target semantic categories of each image;
[0064] S13. Convert the statements containing the linguistic descriptions, high-level semantic features, and target semantic categories of each image into word vectors containing the linguistic descriptions, high-level semantic features, and target semantic categories of each image by using a word vector encoding method.
[0065] Specifically, in S11 of the present invention, the unified expression of the credible linguistic description keywords of each image is obtained through the following formula:
[0066]
[0067] In the formula, W key is the unified expression of the credible linguistic description keywords of each image, W i is the candidate keyword with a confidence level higher than the screening threshold, and N is the number of keywords passed through the screening.
[0068] Specifically, in S12 of the present invention, the statements containing the linguistic descriptions, high-level semantic features, and target semantic categories of each image are obtained through the following formula:
[0069] P(W key ) = L description (W key ) + L semantics (W key ) + L category (W key );
[0070] In the formula, P(W key ) is the statement containing the linguistic descriptions, high-level semantic features, and target semantic categories of each image, L description (W key ) is the linguistic description hint statement, L semantics (W key ) is the high-level semantic feature hint statement, and L category (W key ) is the target semantic category hint statement.
[0071] Specifically, in S13 of the present invention, the word vectors containing the linguistic descriptions, high-level semantic features, and target semantic categories of each image are obtained through the following formula:
[0072] W embeding (P(W key )) = D(P(W key ))
[0073] In the formula, W embeding (P(W key )) is the word vector containing the linguistic descriptions, high-level semantic features, and target semantic categories of each image, P(W keyIt is a statement containing the language description, high-level semantic features, and target semantic categories of each image, and D(·) is the word vector dictionary encoding operation.
[0074] As Figure 3 shown, according to an embodiment of the present invention, in S20 of the present invention, the word vectors containing the language description, high-level semantic features, and target semantic categories of each image are input into the language semantic encoding model, and the obtained language high-level semantic feature encoding includes:
[0075] S21. The input encoding layer in the language semantic encoding model processes the input word vector data through dimension conversion and position and category encoding to obtain the pre-encoded features of the language semantic encoding model;
[0076] S22. The feature extraction layer in the language semantic encoding model extracts the key features from the pre-encoded features of the language semantic encoding model through Transformer to obtain the key semantic features of the language semantic encoding model;
[0077] S23. The feature mapping layer in the language semantic encoding model refines and fuses the key semantic features of the language semantic encoding model through mapping transformation to obtain the language high-level semantic feature encoding.
[0078] As Figure 4 shown, according to an embodiment of the present invention, in S30 of the present invention, each image in the infrared point target image dataset is input into the visual semantic encoding model, and the obtained visual high-level semantic feature encoding includes:
[0079] S31. The input encoding layer in the visual semantic encoding model processes each input image through dimension conversion and position and category encoding to obtain the pre-encoded features of the visual semantic encoding model;
[0080] S32. The feature extraction layer in the visual semantic encoding model extracts the key features from the pre-encoded features of the visual semantic encoding model through Transformer to obtain the key semantic features of the visual semantic encoding model;
[0081] S33. The feature mapping layer in the visual semantic encoding model refines and fuses the key semantic features of the visual semantic encoding model through mapping transformation to obtain the visual high-level semantic feature encoding.
[0082] According to an embodiment of the present invention, in S40 of the present invention, the distance loss between the language high-level semantic feature encoding and the visual high-level semantic feature encoding is obtained through the following formula:
[0083]
[0084] In the formula, L dis(I, T) is the distance loss between the high-level semantic feature encoding of the language and the high-level semantic feature encoding of the vision. S(I, T) is the similarity between the high-level semantic feature encoding of the target language of the same category and the high-level semantic feature encoding of the vision. I is the high-level semantic feature encoding of the vision, T is the high-level semantic feature encoding of the language that matches the high-level semantic feature encoding of the vision, and S(I, T j ) is the similarity between the high-level semantic feature encoding of the target language of different categories and the high-level semantic feature encoding of the vision. T j is the high-level semantic feature encoding of the language that does not match the high-level semantic feature encoding of the vision. M is the number of high-level semantic feature encodings that do not match. τ is the temperature parameter, exp(·) is the exponential function, and log(·) is the natural logarithm function.
[0085] For example, the similarity between the high-level semantic feature encoding of the language and the high-level semantic feature encoding of the vision can adopt the cosine similarity. At this time, where ||·|| is the calculation of the second norm.
[0086] According to an embodiment of the present invention, in S50 of the present invention, the updated vision-language model includes an updated language semantic encoding model and an updated vision semantic encoding model.
[0087] According to an embodiment of the present invention, in S60 of the present invention, determining the category of the infrared point target to be detected according to the recognition requirement, and inputting the infrared point target to be detected and the category of the infrared point target to be detected into the updated vision-language model to realize infrared point target recognition includes: determining the category of the infrared point target to be detected according to the recognition requirement, inputting the category of the point target to be detected into the updated language semantic encoding model to update the word vector, and recognizing the infrared point target to be detected based on the updated vision semantic encoding model.
[0088] As Figure 5 shown, the present invention provides an infrared point target recognition system based on a vision-language model. The system includes:
[0089] A language description and prompt module, configured to obtain word vectors including language descriptions, high-level semantic features, and target semantic categories of each image based on an infrared point target image dataset, prompt engineering, and prior information;
[0090] A training module, configured to input the language description, high-level semantic features, and word vectors of target semantic categories that contain various images into a language semantic encoding model to obtain high-level language semantic feature encodings; configured to input each image in an infrared point target image dataset into a visual semantic encoding model to obtain high-level visual semantic feature encodings; configured to obtain the similarity between the high-level language semantic feature encodings and the high-level visual semantic feature encodings, and obtain a distance loss between the high-level language semantic feature encodings and the high-level visual semantic feature encodings based on the similarity; and further configured to train the language semantic encoding model and the visual semantic encoding model based on the distance loss between the high-level language semantic feature encodings and the high-level visual semantic feature encodings to obtain an updated vision-language model.
[0091] An inference module, configured to determine a category of an infrared point target to be detected according to a recognition requirement, and input the infrared point target to be detected and the category of the infrared point target to be detected into the updated vision-language model to implement infrared point target recognition.
[0092] The present invention also provides a computer device, including a memory, a processor, and an infrared point target recognition program based on a vision-language model stored in the memory and executable on the processor. When the processor executes the infrared point target recognition program based on the vision-language model, the method described in any one of the above is implemented.
[0093] In summary, the present invention provides a method and a system for infrared point target recognition based on a vision-language model. By using prior information to convert an infrared point target into high-level language semantic features, as a supplement to the few and ineffective extractable feature types of the infrared point target appearance and the lack of high-level semantic information, the problem that it is difficult for a visual model to extract key features of a target is solved, and the accuracy of infrared point target recognition is improved. In addition, since the two models are trained by learning the high-level semantic feature association between language and images, there is no need to label images or language, which is an unsupervised training method and saves a large amount of manual labeling costs. At the same time, the low-cost training method enables the model to easily learn more data, and under the guidance of language features, improves the performance and generalization ability of the model in more complex and unknown scenarios. Therefore, the present invention uses prior information to learn high-level semantic features at the language level, which can better guide the visual model to extract and recognize the features of point targets in infrared detection.
[0094] Parts not described in detail in the present invention are well-known technologies to those skilled in the art.
[0095] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by orientation words such as "front, rear, upper, lower, left, right", "lateral, vertical, perpendicular, horizontal" and "top, bottom", etc. is usually based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description. Without contrary description, these orientation words do not indicate and imply that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation on the protection scope of the present invention; the orientation words "inside, outside" refer to the inside and outside relative to the contour of each component itself.
[0096] For ease of description, spatial relative terms such as "above", "over", "on the upper surface", "above-mentioned", etc. can be used here to describe the spatial positional relationship of a device or feature shown in the drawings with other devices or features. It should be understood that the spatial relative terms are intended to cover different orientations in use or operation in addition to the orientation described in the drawings for the device. For example, if the device in the drawing is inverted, the device described as "above" or "over" other devices or structures will then be positioned "below" or "under" other devices or structures. Thus, the exemplary term "above" can include both the orientations of "above" and "below". The device can also be positioned in other different ways (rotated 90 degrees or in other orientations), and corresponding interpretations should be made for the spatial relative descriptions used here.
[0097] In addition, it should be noted that using words such as "first", "second", etc. to limit components is only for the convenience of distinguishing the corresponding components. Without additional statement, the above words have no special meaning. Therefore, it should not be construed as a limitation on the protection scope of the present invention.
[0098] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention should be included within the protection scope of the present invention.
Claims
1. An infrared point target recognition method based on a vision - language model, characterized in that, the method includes: Obtaining word vectors containing the language description, high - level semantic features, and target semantic categories of each image based on an infrared point target image dataset, prompt engineering, and prior information; Inputting the word vectors containing the language description, high - level semantic features, and target semantic categories of each image into a language semantic encoding model to obtain language high - level semantic feature encoding; Inputting each image in the infrared point target image dataset into a vision semantic encoding model to obtain vision high - level semantic feature encoding; Obtaining the similarity between the language high - level semantic feature encoding and the vision high - level semantic feature encoding, and obtaining the distance loss between the language high - level semantic feature encoding and the vision high - level semantic feature encoding based on the similarity; Training the language semantic encoding model and the vision semantic encoding model based on the distance loss between the language high - level semantic feature encoding and the vision high - level semantic feature encoding to obtain an updated vision - language model; Determining the category of the infrared point target to be detected according to the recognition requirement, and inputting the infrared point target to be detected and the category of the infrared point target to be detected into the updated vision - language model to achieve infrared point target recognition.
2. The method according to claim 1, characterized in that, Obtaining word vectors containing the language description, high - level semantic features, and target semantic categories of each image based on an infrared point target image dataset, prompt engineering, and prior information includes: Making a preliminary judgment on each image in the infrared point target image dataset based on prior information to obtain candidate keywords corresponding to each image, scoring the key confidence of each candidate keyword, and screening the candidate keywords based on the scoring results to obtain a unified expression of credible language description keywords for each image; Enhancing the language expression of the unified expression of credible language description keywords for each image by using prompt engineering to obtain statements containing the language description, high - level semantic features, and target semantic categories of each image; Converting the statements containing the language description, high - level semantic features, and target semantic categories of each image into word vectors containing the language description, high - level semantic features, and target semantic categories of each image by using word vector encoding.
3. The method according to claim 2, characterized in that, The unified expression of credible language description keywords for each image is obtained through the following formula: Where, W key is the unified expression of the language description keywords credible for each image, and W i is the candidate keyword with a confidence level higher than the screening threshold, and N is the number of keywords passed through the screening.
4. The method according to claim 2 or 3, characterized in that, The statements containing the language description, high - level semantic features, and target semantic categories of each image are obtained through the following formula: P(W key ) = L description (W key ) + L semantics (W key ) + L category (W key ); Wherein, P(W key ) is a statement containing the language description, high-level semantic features, and target semantic categories of each image, L description (W key ) is a language description prompt statement, L semantics (W key ) is a high-level semantic feature prompt statement, L category (W key ) is a target semantic category prompt statement.
5. The method according to any one of claims 2 - 4, characterized in that, The word vectors containing the language description, high - level semantic features, and target semantic categories of each image are obtained through the following formula: W embeding (P(W key )) = D(P(W key )) Where, W embeding (P(W key )) is a word vector containing the language description, high-level semantic features, and target semantic categories of each image, P(W key ) is a statement containing the language description, high-level semantic features, and target semantic categories of each image, and D(·) is a word vector dictionary encoding operation.
6. The method according to claim 1, characterized in that, Inputting the word vectors containing the language description, high - level semantic features, and target semantic categories of each image into a language semantic encoding model to obtain language high - level semantic feature encoding includes: The input encoding layer in the language semantic encoding model processes the input word vector data through dimension conversion and position and category encoding to obtain pre - encoded features of the language semantic encoding model; The feature extraction layer in the language semantic encoding model extracts key features from the pre-encoded features of the language semantic encoding model through Transformer to obtain the key semantic features of the language semantic encoding model; The feature mapping layer in the language semantic encoding model refines and fuses the key semantic features of the language semantic encoding model through mapping transformation to obtain the high-level semantic feature encoding of the language.
7. The method according to claim 1, characterized in that, inputting each image in the infrared point target image dataset into the visual semantic encoding model to obtain the high-level visual semantic feature encoding includes: The input encoding layer in the visual semantic encoding model processes each input image through dimension conversion and position and category encoding to obtain the pre-encoded features of the visual semantic encoding model; The feature extraction layer in the visual semantic encoding model extracts key features from the pre-encoded features of the visual semantic encoding model through Transformer to obtain the key semantic features of the visual semantic encoding model; The feature mapping layer in the visual semantic encoding model refines and fuses the key semantic features of the visual semantic encoding model through mapping transformation to obtain the high-level visual semantic feature encoding.
8. The method according to claim 1, characterized in that, obtaining the distance loss between the high-level language semantic feature encoding and the high-level visual semantic feature encoding through the following formula: Where, L dis (I, T) is the distance loss between the high-level semantic feature encoding of language and the high-level semantic feature encoding of vision, S(I, T) is the similarity between the high-level semantic feature encoding of the target language of the same category and the high-level semantic feature encoding of vision, I is the high-level semantic feature encoding of vision, T is the high-level semantic feature encoding of language that matches the high-level semantic feature encoding of vision, S(I, T j ) is the similarity between the high-level semantic feature encoding of the target language of different categories and the high-level semantic feature encoding of vision, T j is the high-level semantic feature encoding of language that does not match the high-level semantic feature encoding of vision, M is the number of high-level semantic feature encodings that do not match, τ is the temperature parameter, exp(·) is the exponential function, and log(·) is the natural logarithm function.
9. The method according to claim 1, characterized in that, determining the category of the infrared point target to be detected according to the recognition requirement, and inputting the infrared point target to be detected and the category of the infrared point target to be detected into the updated visual language model to realize the infrared point target recognition, including: determining the category of the infrared point target to be detected according to the recognition requirement, inputting the category of the point target to be detected into the updated language semantic encoding model to update the word vector, and recognizing the infrared point target to be detected based on the updated visual semantic encoding model; wherein, the updated visual language model includes the updated language semantic encoding model and the updated visual semantic encoding model.
10. An infrared point target recognition system based on a visual language model, characterized in that, the system includes: A language description and prompt module for obtaining word vectors including language descriptions, high-level semantic features and target semantic categories of each image based on the infrared point target image dataset, prompt engineering and prior information; A training module for inputting the word vectors including language descriptions, high-level semantic features and target semantic categories of each image into the language semantic encoding model to obtain the high-level language semantic feature encoding; for inputting each image in the infrared point target image dataset into the visual semantic encoding model to obtain the high-level visual semantic feature encoding; for obtaining the similarity between the high-level language semantic feature encoding and the high-level visual semantic feature encoding, and obtaining the distance loss between the high-level language semantic feature encoding and the high-level visual semantic feature encoding based on the similarity; and also for training the language semantic encoding model and the visual semantic encoding model based on the distance loss between the high-level language semantic feature encoding and the high-level visual semantic feature encoding to obtain the updated visual language model; A reasoning module, configured to determine the category of the infrared point target to be detected according to the recognition requirement, and input the infrared point target to be detected and the category of the infrared point target to be detected into the updated vision-language model, so as to realize the recognition of the infrared point target.
11. A computer device, characterized in that, it includes a memory, a processor, and an infrared point target recognition program based on a vision-language model stored on the memory and executable on the processor. When the processor executes the infrared point target recognition program based on the vision-language model, the method described in any one of claims 1 to 9 is realized.