Target person image recognition method based on multi-dimensional features

By adopting multi-dimensional feature image recognition methods in facial recognition technology, including video intelligent processing, feature optimization and intelligent detection modules, the recognition accuracy problems caused by low-quality images are solved, and efficient processing and accurate recognition of low-resolution images are achieved.

CN120047965APending Publication Date: 2025-05-27SHAANXI PUBLIC INFORMATION IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411944153.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is limited to low-quality images in facial recognition, resulting in a decrease in recognition accuracy, especially in surveillance videos, due to low resolution and blurred motion, the details of the face are unclear, affecting the recognition results.

Method used

The target personnel image recognition method based on multi-dimensional features is adopted, and the low-resolution image is super-resolution processed through the video intelligent processing module. The feature intelligent processing module optimizes the target person attributes for multiple rounds. The intelligent detection module uses the GroundingDINO algorithm to detect and recognize characters, and combines the Resnet algorithm to deduplicate and merge to ensure the accuracy of the recognition results.

Benefits of technology

It improves the processing ability of low-quality images, enhances the accuracy and efficiency of face recognition, and can effectively identify target personnel without the need for clear face images, improving the utilization rate of massive surveillance video streams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047965A_ABST
    Figure CN120047965A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, and discloses a multi-dimensional feature-based target person image recognition method, which comprises the following steps of: converting acquired fact information into attribute description of a target person, communicating feature attribute text description and image recognition capability of the target person, and identifying the target person. The classical problem of target person recognition in image recognition is solved based on cross-text and image modalities, the problem of target person retrieval under the condition that necessary conditions for providing recent photos of target persons are lacked in a traditional solution is solved, and the method has higher inclusiveness for monitoring camera shooting angles and the like. Therefore, the utilization rate of mass monitoring video streams is improved, fact information can be fully utilized, and the image recognition capability and the positioning efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly to a method for recognizing target personnel images based on multi-dimensional features. Background Art

[0002] In recent years, face recognition technology has been widely used due to its advantages such as non-contact, long-distance, and high accuracy. The application of big data and artificial intelligence has further improved the efficiency and accuracy of target personnel recognition. By analyzing massive data, AI algorithms can discover the behavior patterns and characteristics of criminal target personnel from it, and improve the speed of case detection. For example, the application of deep learning algorithms in face recognition enables the system to accurately extract and match face information from complex backgrounds.

[0003] When performing face recognition, it is necessary to provide a high-quality original image of the target personnel's face in order to accurately describe the face features of the target personnel, so as to perform face recognition comparison with the database. However, the quality and coverage of the database directly affect the recognition result. If the images in the database are not comprehensive or clear enough, the accuracy of recognition will be affected. Moreover, face extraction and recognition from surveillance recognition require a relatively high resolution of the surveillance video, and the face needs to be clearly distinguishable. However, the images obtained in actual cases often have quality problems for various reasons (such as the quality of surveillance equipment, environmental conditions, lighting conditions, camera installation location, shooting distance). The images in actual surveillance videos may have problems such as motion blur and low resolution, resulting in unclear face details, thus affecting the recognition result. These low-quality images may lead to recognition failure or misrecognition. Therefore, the present invention proposes a method for recognizing target personnel images based on multi-dimensional features. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for recognizing target personnel images based on multi-dimensional features to solve the above technical problems.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] A method for recognizing target personnel images based on multi-dimensional features, the method comprising the following steps:

[0007] Step S1, video acquisition: Real-time collect video data from a surveillance camera and store it in a specified database;

[0008] Step S2, page interaction terminal: According to the factual information in the video, mark the importance of the video content, divided into three levels: high, medium, and low;

[0009] Step S3, Video Intelligent Processing Module: The video processing adopts different frame extraction strategies according to the importance of the video. For high-importance video segments, high-frequency frame extraction is used; for medium-importance segments, medium-frequency frame extraction is used; and for low-importance segments, low-frequency frame extraction is used. The extracted images are denoised, and enhanced processing such as image super-resolution of low-resolution images is performed using the ESRGAN algorithm to improve the image quality;

[0010] Step S4, Feature Intelligent Processing Module: Two rounds of optimization are carried out. In the first round of optimization, the input target person attributes are connected into sentences to generate an initial prototype of the prompt; in the second round of optimization, the Qwen, TeleChat, and Llama models are used to comprehensively optimize the sentence semantics and voice, and finally three different prompts are output;

[0011] Step S5, Intelligent Detection Module: Based on the optimized prompts and the processed images, person detection and recognition are performed using the GroundingDINO algorithm, and the detection results are de-duplicated and merged based on the Resnet algorithm to ensure the accuracy of the final recognition results.

[0012] Step S6, Result Review and Output: The detection results are submitted to criminal investigation personnel for manual review and confirmation. According to the review results, through multiple iterative optimizations, the accuracy and efficiency of recognition are continuously improved. The final output recognition results include image detection frames, basic person attributes, and source video numbers.

[0013] As a further description of the solution of the present invention, the specific process of step S3 includes:

[0014] Step S31, Determine the range of the camera device based on factual information;

[0015] Step S32, Real-time collect video data from the surveillance camera and store it in a specified database;

[0016] Step S33, Determine the importance of the video stream in the interaction interface, that is, the likelihood of the target person appearing in the video;

[0017] Step S34, Adopt different frame extraction strategies according to different importance levels to extract key frames from the video stream to ensure that image segments with important information are captured;

[0018] Step S35, Use the Resnet algorithm to compare the similarity of the extracted frame images and remove images with a similarity greater than 98%;

[0019] Step S36, Preprocess the extracted image frames to improve the image quality.

[0020] As a further description of the solution of the present invention, the specific working process of step S4 is:

[0021] Step S41: On the interaction page, provide the target person's appearance attributes, clothing attributes, action attributes, and environmental attributes according to the factual information; Step S42: First-round prompt optimization: Use NLP technology to perform the first-round sentence combining and standardization on the attribute description text input by the user, including synonym replacement, keyword extraction, and semantic analysis, to ensure that the description information is accurate and consistent;

[0022] Step S43: Second-round prompt optimization: Use semantic large model technology to perform semantic expansion and machine interaction optimization on the prototype of the prompt sentence, and finally generate multiple prompt sentence organizational forms that can interact with GroundingDINO intelligently.

[0023] As a further description of the solution of the present invention, the specific working process of step S5 is as follows:

[0024] Step S51: Use the GroundingDino algorithm to implement the function of searching for pictures by text. According to the optimized attribute description text, retrieve similar images in the database;

[0025] Step S52: Use Resnet to compare the similarity of the retrieved images. Duplicate the images with a similarity of more than 95%, and output the target box images;

[0026] Step S53: Submit the candidate images to criminal investigation personnel for manual review, and confirm the identity of the target person according to the actual situation.

[0027] As a further description of the solution of the present invention, the specific process of determining the importance degree of the video stream includes:

[0028] Construct a calculation model for the video stream importance evaluation coefficient, and the expression is:

[0029]

[0030] In the formula, i represents the i-th video stream, P i is the person attribute index, E i is the environmental attribute index, α and β are the corresponding weight coefficients respectively, P i0 is the reference value of the set person attribute index, E i0 is the set environmental attribute reference value.

[0031] As a further description of the solution of the present invention, the acquisition process of the person attribute index includes:

[0032] Construct a calculation model for the person attribute index, and the expression is:

[0033]

[0034] Wherein, j is the j-th person in the i-th video stream, M is the total number of persons appearing in the i-th video stream, where j belongs to M, S is the number of person attribute items set by the system, k is the k-th person attribute, where k belongs to S, Q j is the person attribute of the j-th person, H k is the parameter value of the k-th person attribute, H k0 is the parameter reference value of the k-th person attribute, ρ k is the weight coefficient of the k-th person attribute.

[0035] As a further description of the solution of the present invention, the process of obtaining the environmental attribute index includes:

[0036] Construct an environmental attribute index calculation model, and the expression is:

[0037]

[0038] In the formula, t 1 is the start time of the i-th video stream, t 2 is the end time of the i-th video stream, f i (t) is the change function of the environmental parameter value of the i-th video stream, f 0 (t) is the standard environmental parameter value change function preset by the system.

[0039] As a further description of the solution of the present invention, the specific process of determining the importance of the video stream further includes:

[0040] Compare the video stream importance evaluation coefficient σ i of the i-th video stream with the set target values σ 1 and σ 2 ;

[0041] If σ i is less than or equal to σ 1 , then the importance of the i-th video stream is low;

[0042] If σ 1 is less than σ i and less than σ 2 , then the importance of the i-th video stream is medium;

[0043] If σ 2 is greater than or equal to σ i , then the importance of the i-th video stream is high.

[0044] Advantages of the present invention:

[0045] 1. End-to-end method flow of target person retrieval from video stream to detection box

[0046] Convert the obtained factual information into the attribute description of the target person, connect the text description of the target person's characteristic attributes and the image recognition ability, and realize the classic problem of target person recognition in image recognition based on cross-text and image modalities. Solve the problem of target person retrieval in the case of the lack of the necessary condition of providing a recent photo of the target person in the traditional solution, and have higher inclusiveness for monitoring camera angles, etc., rather than having to capture a clear face image of the target person, so as to improve the utilization rate of the massive monitoring video stream, be able to make full use of the factual information, and improve the image recognition ability and positioning efficiency. Based on the feature text attributes (appearance attributes, clothing attributes, action attributes, environmental attributes), use NLP technology to optimize the attribute words in multiple rounds of statements to ensure that the prompt words are more accurate and rich semantically, and convert them into prompt words that can better interact with the algorithm model. Use multi-modal GroundingDINO for the alignment of image features and text features, and output the retrieval box of the target person.

[0047] 2. Guiding and Optimizing Method Based on the Attribute Description of the Target Person

[0048] The interaction page can input the appearance attributes, clothing attributes, action attributes and environmental attributes of the target person. The four aspects of attribute prompts provide description guidance for the target person to the image recognition personnel. Based on the four-part attribute description, use NLP technology to perform the first-round statement optimization on the attribute words, and use semantic large models such as Qwen, TeleChat, and Llama to perform the second-round optimization on the statement prompt words to ensure that the prompt words can better interact with the multi-modal model semantically, provide multiple efficient model guiding prompt words, so as to comprehensively describe the characteristics of the target person, and align and fuse the text feature vectors in the high-dimensional space and the image feature vectors.

[0049] 3. Solve the Problem of Fuzzy Retrieval of Target Persons Based on the Open World GroundingDINO Model

[0050] Based on the factual information, conduct multi-perspective descriptions of the target person, use the multi-modal GroundingDINO model to align the image and text features, and output the retrieval box of the target person. This helps to solve the problems of such portrait missing and fuzzy feature retrieval, makes the person tracking have a wider retrieval path, and helps to greatly improve the working efficiency of image recognition, especially for time-critical immediate events such as missing persons.

[0051] 4. Segmenting and Extracting Frames Strategy Based on Video Importance

[0052] The videos collected from the surveillance video stream can classify the importance of the video lines according to the factual information of the events. Based on different importance levels, different frame sampling strategies are adopted, and the Resnet network is used to remove images with similarity above 98% from the sampled images, so as to effectively balance the detection accuracy and detection speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The present invention will be further described below in conjunction with the accompanying drawings.

[0054] Figure 1 is a partial flowchart of the target person image recognition method based on multi-dimensional features provided by the present invention;

[0055] Figure 2 is a partial flowchart of the video intelligent processing method of the present invention;

[0056] Figure 3 is a partial flowchart of the feature intelligent processing method of the present invention;

[0057] Figure 4 is a partial flowchart of the intelligent detection method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0059] Please refer to Figures 1-4 , the present invention is a target person image recognition method based on multi-dimensional features, which is characterized in that the method includes the following steps:

[0060] Step S1, video acquisition: Real-time video data is collected from the surveillance camera and stored in the specified database;

[0061] Step S2, page interaction end: According to the factual information in the video, the importance of the video content is marked and divided into three levels: high, medium, and low; the user can input relevant attributes of the target person according to the video picture and factual information, such as gender, age, height, clothing, actions, description of the surrounding environment, etc. The algorithm image recognition result will be output at the page interaction end, and an artificial review interaction function will be provided.

[0062] Step S3, Video Intelligent Processing Module: The video processing adopts different frame extraction strategies according to the importance of the video. For high-importance video segments, high-frequency frame extraction is used; for medium-importance segments, medium-frequency frame extraction is used; and for low-importance segments, low-frequency frame extraction is used. The extracted images are denoised, and enhanced processing such as image super-resolution of low-resolution images is performed using the ESRGAN algorithm to improve the image quality;

[0063] Step S4, Feature Intelligent Processing Module: Two rounds of optimization are carried out. In the first round of optimization, the input target person attributes are connected into sentences to generate an initial prototype of the prompt; in the second round of optimization, the Qwen, TeleChat, and Llama models are used to comprehensively optimize the sentence semantics and voices, and finally three different prompts are output;

[0064] Step S5, Intelligent Detection Module: Based on the optimized prompts and the processed images, person detection and recognition are performed using the GroundingDINO algorithm, and the detection results are de-duplicated and merged based on the Resnet algorithm to ensure the accuracy of the final recognition results.

[0065] Step S6, Result Review and Output: The detection results are submitted to criminal investigation personnel for manual review and confirmation. According to the review results, through multiple iterations of optimization, the accuracy and efficiency of recognition are continuously improved. The final output recognition results include image detection frames, basic person attributes, and source video numbers.

[0066] Through the above technical solutions, the obtained factual information is converted into an attribute description of the target person, connecting the text description of the target person's characteristic attributes and the image recognition ability, realizing the classic problem of target person recognition in image recognition based on cross-text and image modalities, solving the problem of target person retrieval in the traditional solution when the necessary condition of providing a recent photo of the target person is missing, having higher inclusiveness for monitoring camera angles, etc., rather than having to capture a clear face image of the target person, thereby improving the utilization rate of the massive monitoring video stream, being able to make full use of factual information, and improving the image recognition ability and positioning efficiency. Based on the feature text attributes (appearance attributes, clothing attributes, action attributes, environmental attributes), the NLP technology is used to optimize the attribute words in multiple rounds of sentences to ensure that the prompts are more accurate and rich semantically, and are converted into prompts that can better interact with the algorithm model. The multi-modal GroundingDINO is used as the alignment of image features and text features to output the retrieval frame of the target person.

[0067] As a further description of the solution of the present invention, the specific process of step S3 includes:

[0068] Step S31, Determine the range of the camera device based on the factual information;

[0069] Step S32: Real-time collect video data from the surveillance camera and store it in the specified database;

[0070] Step S33: Determine the importance degree of the video stream in the interaction interface, that is, the probability of the target person appearing in the video;

[0071] Step S34: Adopt different frame extraction strategies according to different importance degrees, and extract key frames from the video stream to ensure capturing image segments with important information;

[0072] Step S35: Use the Resnet algorithm to compare the similarity of the extracted frame images, and remove the images with a similarity greater than 98%;

[0073] Step S36: Preprocess the extracted image frames, including denoising, deblurring, image enhancement, etc. For example, use the ESRGAN algorithm to perform image super-resolution enhancement processing on low-resolution images to improve the image quality.

[0074] As a further description of the solution of the present invention, the specific working process of step S4 is as follows:

[0075] Step S41: On the interaction page, provide the appearance attributes, clothing attributes, action attributes and environmental attributes of the target person according to the factual information; such as gender, age, height, body type, clothing, etc.

[0076] Step S42: The first round of prompt optimization: Use NLP technology to perform the first round of sentence optimization and standardization of the attribute description text input by the user, including synonym replacement, keyword extraction, semantic analysis, etc., to ensure that the description information is accurate and consistent;

[0077] Step S43: The second round of prompt optimization: Use semantic large model technology to perform semantic expansion and machine interaction optimization on the prototype of the prompt sentence, such as semantic large models such as Qwen, TeleChat, Llama, etc., and finally generate multiple prompt sentence organizational forms that can interact with GroundingDINO intelligently.

[0078] As a further description of the solution of the present invention, the specific working process of step S5 is as follows:

[0079] Step S51: Use the GroundingDino algorithm to implement the function of searching for pictures by text. According to the optimized attribute description text, retrieve similar images in the database; The GroundingDino algorithm can effectively align the text description with the image features to achieve efficient text-to-image retrieval.

[0080] Step S52: Use Resnet to compare the similarity of the retrieved images, remove the duplicates of the images with a similarity above 95%, and output the target box images;

[0081] Step S53: Submit the candidate images to criminal investigation personnel for manual review, and confirm the identity of the target person according to the actual situation.

[0082] Through the above technical solution, the obtained factual information is converted into an attribute description of the target person, connecting the text description of the target person's characteristic attributes and the image recognition ability, realizing the solution of the classic problem of target person recognition in image recognition based on cross-text and image modalities, solving the problem of target person retrieval in the case of the absence of the necessary condition of providing recent photos of the target person in the traditional solution, having higher inclusiveness for monitoring camera angles, etc., rather than having to capture a clear face image of the target person, thereby improving the utilization rate of the massive monitoring video stream, being able to make full use of factual information, and improving the image recognition ability and positioning efficiency. Based on the feature text attributes (appearance attributes, dressing attributes, action attributes, environmental attributes), using NLP technology to optimize the attribute words in multiple rounds of statements to ensure that the prompt words are more accurate and rich semantically, and converting them into prompt words that can better interact with the algorithm model, using multi-modal GroundingDINO as the alignment of image features and text features, and outputting the retrieval box of the target person.

[0083] As a further description of the solution of the present invention, the specific process of determining the importance degree of the video stream includes:

[0084] Construct a calculation model for the importance evaluation coefficient of the video stream, and the expression is:

[0085]

[0086] In the formula, i represents the i-th video stream, P i is the character attribute index, E i is the environmental attribute index, α and β are the corresponding weight coefficients respectively, P i0 is the reference value of the set character attribute index, E i0 is the set environmental attribute reference value.

[0087] As a further description of the solution of the present invention, the acquisition process of the character attribute index includes:

[0088] Construct a calculation model for the character attribute index, and the expression is:

[0089]

[0090] In the formula, j is the j-th person in the i-th video stream, M is the total number of people appearing in the i-th video stream, where j belongs to M, S is the number of character attribute items set by the system, k is the k-th character attribute, where k belongs to S, Q j is the character attribute of the j-th person, H k is the parameter value of the k-th character attribute, Hk0 is the parameter reference value for the k-th character attribute, ρ k is the weight coefficient for the k-th character attribute.

[0091] It should be noted that the character attributes are the gender, age, height, body type, clothing, etc. set by the system. The parameter values of the character attributes are input into the trained neural network model according to the described factual information and the factual information in the video, and the parameter values of each attribute are output.

[0092] As a further description of the solution of the present invention, the process of obtaining the environmental attribute index includes:

[0093] Construct an environmental attribute index calculation model, and the expression is:

[0094]

[0095] In the formula, t 1 is the start time of the i-th video stream, t 2 is the end time of the i-th video stream, f i (t is the environmental parameter value change function of the i-th video stream, f 0 (t) is the standard environmental parameter value change function preset by the system.

[0096] It should be noted that the method for obtaining the environmental parameter value change function is: taking Δt as the time period, obtaining the data of the number of key environmental features described by the factual information appearing in each period changing with time.

[0097] As a further description of the solution of the present invention, the specific process of determining the importance degree of the video stream further includes:

[0098] Compare the video stream importance evaluation coefficient σ i of the i-th video stream with the set target values σ 1 and σ 2 ;

[0099] If σ i is less than or equal to σ 1 , then the importance degree of the i-th video stream is low;

[0100] If σ 1 is less than σ i and less than σ 2 , then the importance degree of the i-th video stream is medium;

[0101] If σ 2 is greater than or equal to σ i , then the importance degree of the i-th video stream is high.

[0102] Through the above technical solution, this embodiment provides a method for determining the importance degree of a video stream. Based on the described factual information and the factual information in the video, the value of the character attribute index is obtained. According to the data of the change over time of the number of key environmental features described by the factual information appearing in each period of time, the value of the environmental attribute index is obtained. Then, through the formula Calculate the evaluation coefficient of the importance degree of the video stream, and compare the evaluation coefficient σ i of the i-th video stream with the set target values σ 1 and σ 2 ; if σ i is less than or equal to σ 1 , then the importance degree of the i-th video stream is low; if σ 1 is less than σ i and less than σ 2 , then the importance degree of the i-th video stream is medium; if σ 2 is greater than or equal to σ i , then the importance degree of the i-th video stream is high.

[0103] The above has described an embodiment of the present invention in detail, but the content described is only a preferred embodiment of the present invention and cannot be considered as limiting the implementation scope of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.

Claims

1. A method for target person image recognition based on multi-dimensional features, characterized in that: The method comprises the following steps: Step S1, video acquisition: real-time acquisition of video data from surveillance cameras and storage in a designated database; Step S2, page interaction end: according to the factual information in the video, the importance of the video content is marked and divided into three levels: high, medium and low; Step S3, video intelligent processing module: video processing adopts different frame extraction strategies according to the importance of the video. High-frequency frame extraction is used for high-importance video segments, medium-frequency frame extraction is used for medium-importance segments, and low-frequency frame extraction is used for low-importance segments. The extracted images are denoised, and the ESRGAN algorithm is used to perform image super-resolution and other enhancement processing on low-resolution images to improve image quality. Step S4, feature intelligent processing module: perform two rounds of optimization. The first round of optimization connects words into sentences according to the input target person attributes to generate the initial prompt prototype; the second round of optimization uses Qwen, TeleChat, LLama models to perform all-round optimization on sentence semantics and voice, and finally outputs three different prompts; Step S5, intelligent detection module: perform person detection and recognition based on the GroundingDINO algorithm according to the optimized prompt words and processed images, and deduplicate and merge the detection results based on the Resnet algorithm to ensure the accuracy of the final recognition results. Step S6, result review and output: The detection results are submitted to criminal investigation personnel for manual review and confirmation. According to the review results, multiple iterations of optimization are performed to continuously improve the accuracy and efficiency of recognition. The final output recognition results include the image detection frame, basic attributes of the person, and the source video number.

2. The method for target person image recognition based on multi-dimensional features according to claim 1, characterized in that: The specific process of step S3 includes: Step S31, determining the range of the camera equipment according to the factual information; Step S32: Collect video data from the surveillance camera in real time and store it in a specified database; Step S33, determining the importance of the video stream in the interactive interface, that is, the probability of the target person appearing in the video; Step S34: adopt different frame extraction strategies according to different importance levels to extract key frames from the video stream to ensure that image segments with important information are captured; Step S35, using the Resnet algorithm to perform similarity comparison on the extracted frame images, and removing images with a similarity greater than 98%; Step S36: pre-process the extracted image frames to improve image quality.

3. The method for target person image recognition based on multi-dimensional features according to claim 1, characterized in that: The specific working process of step S4 is as follows: Step S41: On the interactive page, provide the target person's appearance attributes, clothing attributes, action attributes, and environment attributes according to the factual information; Step S42, first round of prompt optimization: using NLP technology to perform the first round of sentence optimization and standardization on the attribute description text input by the user, including synonym replacement, keyword extraction, and semantic analysis, to ensure that the description information is accurate and consistent; Step S43, the second round of prompt optimization: use the semantic big model technology to perform semantic expansion and machine interaction optimization on the prompt sentence prototype, and finally generate multiple prompt organization forms that can interact with GroundingDINO intelligently.

4. The method for target person image recognition based on multi-dimensional features according to claim 1, characterized in that: The specific working process of step S5 is as follows: Step S51, using the GroundingDino algorithm to implement the text-based image search function, and searching similar images in the database according to the optimized attribute description text; Step S52, using Resnet to compare the similarity of the retrieved images, deduplicating images with a similarity of more than 95%, and outputting the target frame image; Step S53: Submit the candidate image to criminal investigation personnel for manual review to confirm the identity of the target person based on actual conditions.

5. The method for target person image recognition based on multi-dimensional features according to claim 2, characterized in that: The specific process of determining the importance of the video stream includes: Construct a calculation model for the evaluation coefficient of video stream importance, the expression is: Where i represents the i-th video stream, P i is the character attribute index, E i is the environmental attribute index, α and β are the corresponding weight coefficients, P i0 is the reference value of the character attribute index, E i0 It is the reference value of the set environment property.

6. The method for target person image recognition based on multi-dimensional features according to claim 5, characterized in that: The process of obtaining the character attribute index includes: Construct a character attribute index calculation model, the expression is: In the formula, j is the jth character in the i-th video stream, M is the total number of characters appearing in the i-th video stream, where j belongs to M, S is the number of character attributes set by the system, k is the kth character attribute, where k belongs to S, Q j is the character attribute of the jth character, H k is the parameter value of the kth character attribute, H k0 is the parameter reference value of the k-th character attribute, ρ k is the weight coefficient of the kth character attribute.

7. The method for target person image recognition based on multi-dimensional features according to claim 5, characterized in that: The process of obtaining the environmental attribute indicators includes: Construct the calculation model of environmental attribute indicators, the expression is: Where t1 is the start time of the i-th video stream, t2 is the end time of the i-th video stream, and f i (t) is the change function of the environment parameter value of the ith video stream, and f0(t) is the change function of the standard environment parameter value preset by the system.

8. The method for target person image recognition based on multi-dimensional features according to claim 5, characterized in that: The specific process of determining the importance of the video stream also includes: The video stream importance evaluation coefficient σ of the i-th video stream i Compare with the set target values ​​σ1 and σ2; If σ i If it is less than or equal to σ1, the importance of the i-th video stream is low; If σ1 is less than σ i If it is less than σ2, the importance of the i-th video stream is medium; If σ2 is greater than or equal to σ i , then the importance of the i-th video stream is high.