Virtual target recognition method, device, electronic device and storage medium

Through feature extraction and pre-trained models combined with style difference, the problem of low accuracy in virtual target recognition is solved, and the accuracy and reliability of recognition is improved.

CN119672444BActive Publication Date: 2025-06-03SHANDONG INSPUR INNOVATION & ENTREPRENEURSHIP TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510185928.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-03
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify virtual targets, resulting in low recognition accuracy.

Method used

By obtaining the sample images to be identified, performing feature extraction processing, constructing the data to be identified, and using the pre-trained target recognition model and style recognition model, combined with the degree of style difference, comprehensive judgment is made.

Benefits of technology

Improve the accuracy and reliability of virtual target recognition, and reduce misjudgment and misjudgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672444B_ABST
    Figure CN119672444B_ABST
Patent Text Reader

Abstract

The present application provides a virtual target recognition method, apparatus, electronic device, and storage medium. The virtual target recognition method includes: obtaining a sample image to be recognized; performing feature extraction processing on the sample image to obtain image features; constructing data to be recognized based on the image features; processing the data to be recognized based on a pre-trained target recognition model to obtain a target recognition result of the target in the sample image; processing the image information of the target in the sample image and the image information of the background region in the sample image based on a pre-trained style recognition model to obtain a first style recognition result of the sample image and a second style recognition result of the target; determining a style difference degree between the target and the sample image according to the first style recognition result and the second style recognition result; and determining the recognition results of the sample image and the target according to the style difference degree and the target recognition result. The present application relates to the technical field of computer vision and can improve the accuracy of recognizing virtual targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a method, device, electronic device, and storage medium for virtual target recognition. Background Art

[0002] Currently, with the development of deep learning technology, through training with large-scale data sets and continuously optimizing the model generation ability, the model can capture more detailed features of real human images, such as facial contours, expression changes, skin color differences, etc. This enables the model to better understand human language and visual information, and thus generate more realistic and aesthetically pleasing virtual images. The virtual images generated by generative large models are getting closer and closer to real human images in appearance. However, since the facial features or postures generated based on artificial intelligence can finely simulate real targets, it is difficult to visually distinguish virtual images from real human images. As a result, users cannot distinguish real targets (such as real human bodies, real animals, etc.) and virtual targets (such as virtual human bodies, virtual animals, etc.) in the image, which may lead to a decrease in the accuracy of virtual target recognition. Summary of the Invention

[0003] In view of the above, it is necessary to propose a method, device, electronic device, and storage medium for virtual target recognition to solve the technical problem of low accuracy in virtual target recognition.

[0004] This application provides a method for virtual target recognition, which is applied to an electronic device. The method includes: obtaining a sample image to be recognized; performing feature extraction processing on the sample image to obtain image features; constructing data to be recognized according to the image features; processing the data to be recognized based on a pre-trained target recognition model to obtain a target recognition result of the target in the sample image; processing the image information of the target and the image information of the background area in the sample image based on a pre-trained style recognition model to obtain a first style recognition result of the sample image and a second style recognition result of the target; determining the style difference degree between the target and the sample image according to the first style recognition result and the second style recognition result; and determining the recognition results of the sample image and the target according to the style difference degree and the target recognition result.

[0005] In some embodiments, the method further includes training the target recognition model, and the training of the target recognition model includes: constructing first training data according to pre-stored training images; wherein, the training images correspond to first label data; determining a first confidence level of the corresponding first training data according to the first label data; the first confidence level is used to characterize the credibility of the target in the training image; inputting the first training data into a pre-constructed first initial recognition model to obtain a first prediction result corresponding to the first training data; determining a first loss value of the first initial recognition model according to the first confidence level, the first prediction result and the first label data; updating the initial recognition model based on the backpropagation algorithm until the first loss value meets a preset condition, then stopping updating the first initial recognition model to obtain a target recognition model trained to a convergent state.

[0006] In some embodiments, the determining the first confidence level of the corresponding first training data according to the first label data includes:

[0007] ;

[0008] wherein, R represents the first confidence level of the corresponding first training data; i is the index of the target in the training image indicated by the first label data, and m is the number of targets in the training image indicated by the first label data; j is the index of the limb part of the target indicated by the first label data, and n is the number of limb parts of the target indicated by the first label data; is the number of pixel points of the limb part with index j in the target with index i in the training image indicated by the first label data; is the number of pixel points of the trunk part in the target with index i in the training image; is the score of the limb part with index j in the target with index i in the training image indicated by the first label data; is the score of the facial feature of the target with index i in the training image indicated by the first label data.

[0009] In some embodiments, the determining the first loss value of the first initial recognition model according to the confidence level, the first prediction result and the first label data includes:

[0010] ;

[0011] Among them, Loss1 represents the first loss value of the initial recognition model; R represents the first confidence; e represents the natural constant; i is the index of the target in the training image indicated by the first label data, and m is the number of targets in the training image indicated by the first label data; j is the index of the limb part of the target indicated by the first label data, and n is the number of limb parts of the target indicated by the first label data; is the number of pixel points of the limb part with index j in the target with index i in the training image, is the number of pixel points of the limb part with index j in the target with index i in the training image indicated by the first prediction result; is the number of pixel points of the torso part in the target with index i in the training image; is the score of the limb part with index j in the target with index i in the training image indicated by the first label data, is the score of the limb part with index j in the target with index i in the training image indicated by the first prediction result; is the score of the facial feature of the target with index i in the training image indicated by the first label data, is the score of the facial feature of the target with index i in the training image indicated by the first prediction result.

[0012] In some embodiments, the method further includes training the style recognition model, and the training of the style recognition model includes: constructing second training data according to the pre-stored training images; wherein, the training images correspond to second label data; the second label data includes a first style vector and a second style vector; determining the second confidence of the corresponding second training data according to the second label data; inputting the second training data into a pre-constructed second initial recognition model to obtain a second prediction result corresponding to the second training data; the second prediction result includes a first predicted style vector and a second predicted style vector; determining the second loss value of the second initial recognition model according to the second confidence, the second prediction result and the second label data; updating the second initial recognition model based on the backpropagation algorithm until the second loss value meets a preset condition, and then stopping updating the initial recognition model to obtain a style recognition model trained to a convergent state.

[0013] In some embodiments, the determining the second confidence of the corresponding second training data according to the second label data includes:

[0014] ;

[0015] Wherein, S represents the second confidence corresponding to the second training data; z is the index of the target in the training image indicated by the second label data, and v is the number of targets in the training image indicated by the second label data; x is the index of the dimension in the first style vector indicated by the second label data, and y is the number of dimensions in the first style vector indicated by the second label data; is the value of the dimension with index x in the first style vector corresponding to the target with index z in the training image indicated by the second label data; is the value of the dimension with index x in the second style vector indicated by the second prediction result.

[0016] In some embodiments, determining the second loss value of the second initial recognition model according to the second confidence, the second prediction result and the second label data includes:

[0017] ;

[0018] Wherein, Loss2 represents the second loss value of the second initial recognition model; S represents the second confidence corresponding to the second training data; e represents the natural constant; z is the index of the target in the training image indicated by the second label data, and v is the number of targets in the training image indicated by the second label data; x is the index of the dimension in the first style vector indicated by the second label data, and y is the number of dimensions in the first style vector indicated by the second label data; is the value of the dimension with index x in the first style vector corresponding to the target with index z in the training image indicated by the second label data, is the value of the dimension with index x in the first predicted style vector corresponding to the target with index z in the training image indicated by the second prediction result; is the value of the dimension with index x in the second style vector indicated by the second label data, is the value of the dimension with index x in the second predicted style vector indicated by the second prediction result.

[0019] An embodiment of the present application further provides a virtual target recognition device, which includes: an acquisition module for acquiring a sample image to be recognized; a feature extraction module for performing feature extraction processing on the sample image to obtain image features; a construction module for constructing data to be recognized according to the image features; a first recognition module for processing the data to be recognized based on a pre-trained target recognition model to obtain a target recognition result of the target in the sample image; a second recognition module for processing the image information of the target and the image information of the background area in the sample image based on a pre-trained style recognition model to obtain a first style recognition result of the sample image and a second style recognition result of the target; a determination module for determining the style difference degree between the target and the sample image according to the first style recognition result and the second style recognition result; the determination module is further configured to determine the recognition results of the sample image and the target according to the style difference degree and the target recognition result.

[0020] An embodiment of the present application further provides an electronic device, which includes a processor and a memory, and the processor is used to implement the virtual target recognition method when executing a computer program stored in the memory.

[0021] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the virtual target recognition method are implemented.

[0022] It can be seen from the above technical solutions that in the embodiment of the present application, the key features of the sample image are accurately obtained through feature extraction, providing data support for subsequent target recognition and style recognition. The pre-trained target recognition model can accurately judge whether the target human body features in the sample image are reasonable and can capture the effective information of the target from the deep image features. In the process of analyzing the style of the sample image, not only the overall style of the sample image is concerned, but also the style recognition of the target part in the image is particularly carried out, ensuring a comprehensive analysis of the image style. By comparing the difference degree between the target style and the overall style of the sample image, the consistency and authenticity of the image can be evaluated more carefully. Combining the target recognition result and the style difference degree can comprehensively judge whether the target in the sample image is a virtual target. This multi-dimensional judgment method helps to reduce misjudgment and missed judgment, and improves the accuracy and reliability of the judgment. Description of the Drawings

[0023] Figure 1 FIG. is an application scenario diagram of a virtual target recognition method provided by an embodiment of the present application.

[0024] Figure 2It is a flowchart of a virtual target recognition method provided by an embodiment of the present application.

[0025] Figure 3 It is a schematic diagram of a sample image provided by an embodiment of the present application.

[0026] Figure 4 It is a schematic diagram of a target provided by an embodiment of the present application.

[0027] Figure 5 It is a functional module diagram of a virtual target recognition device provided by an embodiment of the present application.

[0028] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0029] In order to more clearly understand the purpose, features and advantages of the present application, the present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other. Many specific details are set forth in the following description in order to provide a thorough understanding of the present application. The described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0030] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present application, "a plurality of" means two or more, unless otherwise specifically defined.

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments, and are not intended to limit this application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0032] An embodiment of the present application provides a virtual target recognition method, which can be applied to one or more electronic devices. An electronic device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.

[0033] The electronic device can be any electronic product that can perform human-computer interaction with customers. For example, a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet Protocol Television (IPTV), a smart wearable device, etc.

[0034] The electronic device may further include a network device and / or a client device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.

[0035] The network where the electronic device is located includes, but is not limited to, the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), etc.

[0036] As Figure 1 As shown, a virtual target recognition method provided by the present application can be applied to an electronic device 100, and the electronic device 100 is communicatively connected to a server 200. The electronic device 100 is configured to obtain a sample image to be recognized from the server 200. The electronic device 100 performs feature extraction processing on the sample image to obtain image features, and constructs data to be recognized according to the image features. The electronic device 100 further processes the data to be recognized based on a pre-trained target recognition model to obtain a target recognition result of the target in the sample image, and processes the image information of the target in the sample image and the image information of the background area in the sample image based on a pre-trained style recognition model to obtain a first style recognition result of the sample image and a second style recognition result of the target. The electronic device 100 further determines the style difference degree between the target and the sample image according to the first style recognition result and the second style recognition result, and determines the recognition results of the sample image and the target according to the style difference degree and the target recognition result.

[0037] AsFigure 2 As shown, it is a flowchart of a virtual target recognition method provided by an embodiment of the present application. According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted. A virtual target recognition method provided by an embodiment of the present application includes the following steps.

[0038] S20, obtain a sample image to be recognized.

[0039] In an embodiment of the present application, the sample image to be recognized can be an image with a virtual target collected in real time by an electronic device. For example, the electronic device can read the sample image from a local file system, network, camera, or other image acquisition devices. After obtaining the sample image, preprocessing (such as scaling, cropping, denoising, etc.) can also be performed on the sample image to ensure that the image quality meets the requirements of subsequent processing. And the preprocessed sample image is cached in a specified directory or database, and corresponding indexes or metadata are established to facilitate subsequent retrieval and processing.

[0040] As Figure 3 shown is a schematic diagram of a sample image 300 provided by an embodiment of the present application. The sample image 300 includes multiple targets. For example, targets 310, 320, 330, 340, and 350. In order to determine whether the multiple targets in the sample image are virtual targets generated based on artificial intelligence technology, feature extraction can be performed on the sample image to determine the deep features in the sample image, and based on the deep features, it can be recognized whether the targets in the sample image are virtual targets generated based on artificial intelligence technology.

[0041] S21, perform feature extraction processing on the sample image to obtain image features.

[0042] In an embodiment of the present application, a pre-trained neural network model can be used as a feature extractor. The sample image is input into the feature extractor to obtain the deep image features in the sample image. Exemplarily, the image features include edge features, texture features, shape features, semantic information features, etc. Exemplarily, the pre-trained neural network model can be a VGG network or a ResNet network, and the present application does not make any limitations in this regard.

[0043] S22, construct data to be recognized according to the image features.

[0044] In an embodiment of the present application, the image features can be used to represent the deep information in the sample image. For example, the image features include the main content and objects (such as people, landscapes, animals, etc.) in the sample image; for another example, the image features also include the style information in the sample image, such as color, texture, lighting, etc., which can provide data support for subsequent style recognition of the sample image.

[0045] In an embodiment of the present application, information integration can be performed based on different types of information included in the image features, integrating the content and style information into a comprehensive description of the sample image for subsequent processing and analysis. Exemplarily, the information characterizing the main content and objects of the sample image in the image features can be a vector data, and the relevant data characterizing the style information of the sample image can also be a vector data. The data to be recognized corresponding to the sample image can be obtained by splicing the two vector data.

[0046] S23. Process the data to be recognized based on a pre-trained target recognition model to obtain the target recognition result of the target in the sample image.

[0047] In an embodiment of the present application, the data to be recognized includes a feature vector or an image patch characterizing the deep information of the sample image. The input of the target recognition model is the data to be recognized corresponding to the sample image. The target recognition model is used to perform inference calculations based on the data to be recognized to obtain information such as the position, category, and attributes of the target in the sample image, as well as information such as the limb parts, trunk parts, and facial features of the target. For example, when the target in the sample image is a human body, the target recognition model is used to determine information such as the position, category, and attributes of the human body in the sample image, as well as the pixel points corresponding to the limb parts of the human body, the pixel points corresponding to the trunk parts, and the facial features. The output of the target recognition model is the target recognition result corresponding to the target in the sample image, where the target recognition result is used to characterize the rationality of the human body features in the sample image. For example, the target recognition result can be used to determine the rationality of the human body proportion according to the number of pixel points of the limb parts and the number of pixel points of the trunk parts of the human body, and the target recognition result can also be used to determine the rationality of the structure of the limb parts according to the number of pixel points of the limb parts of the human body.

[0048] As Figure 4 shown is a schematic diagram of a target 400 in a sample image provided by an embodiment of the present application. Among them, the target 400 at least includes a facial area 410, a trunk part 420, a limb part 430, and a limb part 440. Among them, the limb part 430 can be used to characterize the arm part of the target 400, and the limb part 440 can be used to characterize the palm part of the target 400. The target recognition result can be used to characterize the rationality of the facial features of the target 400 according to the pixel points in the facial area 410, and the target recognition result can also be used to determine the rationality of the human body proportion according to the number of pixel points of the limb part 430 and the number of pixel points of the trunk part 420 of the target 400, and the target recognition result can also be used to determine the rationality of the structure of the limb part according to the number of pixel points of the limb part 440 of the human body.

[0049] In an embodiment of the present application, the target recognition model may be a pre-trained convolutional neural network model, or a residual neural network model, or a fully connected neural network model. The present application does not limit this.

[0050] In an embodiment of the present application, the method for training a target recognition model includes: constructing first training data according to pre-stored training images; wherein, the training images correspond to first label data; determining a first confidence level of the corresponding first training data according to the first label data; the first confidence level is used to characterize the credibility of the target in the training image; inputting the first training data into a pre-constructed first initial recognition model to obtain a first prediction result corresponding to the first training data; determining a first loss value of the first initial recognition model according to the first confidence level, the first prediction result, and the first label data; updating the initial recognition model based on the backpropagation algorithm until the first loss value meets a preset condition, then stopping updating the first initial recognition model to obtain a target recognition model trained to a convergent state. Among them, the preset condition may be that the first loss value is less than or equal to a preset first termination threshold. For example, when the preset first termination threshold is 0.1, in the case where the first loss value is less than or equal to 0.1, it can be determined that the difference between the first prediction result output by the first initial recognition model and the first label data is small, and the accuracy of the first prediction result output by the first initial recognition model is high, then the update of the first initial recognition model can be stopped to obtain a target recognition model trained to a convergent state.

[0051] In an embodiment of the present application, the pre-stored training images may be multiple images containing human targets, and the first training data may be deep image features obtained after feature extraction of the training images. The first label data may be pre-annotated data, which is used to indicate the pixel points of the trunk part corresponding to the target in the training image, and is also used to indicate the pixel points of the limb parts corresponding to the target in the training image, and is also used to indicate the score of the limb area corresponding to the target in the training image, and is also used to indicate the score of the facial area corresponding to the target in the training image. Among them, the score of the limb area is used to characterize the rationality of the structure of the limb area. For example, when the limb area is the hand area, if the hand area contains 4 fingers, the score of this limb area is relatively low; the score of the facial area is used to characterize the rationality of the structure of the facial area. For example, when the eye distance between the two eyes in the facial area is large, the score of the facial area is relatively low.

[0052] In an embodiment of the present application, in order to improve the accuracy of training the target recognition model, the first confidence level of the first training data is also determined according to the first label data in the training image. Specifically, the calculation method of the first confidence level satisfies the following relational expression: ;

[0053] Wherein, R represents the first confidence level corresponding to the first training data; i is the index of the target in the training image indicated by the first label data, m is the number of targets in the training image indicated by the first label data; j is the index of the limb part of the target indicated by the first label data, and n is the number of limb parts of the target indicated by the first label data; is the number of pixel points of the limb part with index j in the target with index i in the training image indicated by the first label data; is the number of pixel points of the trunk part in the target with index i in the training image; is the score of the limb part with index j in the target with index i in the training image indicated by the first label data; is the score of the facial feature of the target with index i in the training image indicated by the first label data.

[0054] Wherein, the lower the ratio between the number of pixel points of the target limb part and the number of pixel points of the trunk part, the greater the morphological difference between the limb part and the trunk part of the target, and the lower the first confidence level; the lower the score of the limb part of the target, the lower the rationality of the structure of the limb part, and the lower the first confidence level; the lower the score of the facial feature of the target, the lower the rationality of the facial feature of the target, and the lower the first confidence level. When the first confidence level is lower, it indicates that the credibility of the target in the training image is lower. Then, when training the target recognition model according to the training data corresponding to the training image, the credibility of the obtained result is lower, and the loss value of the target recognition model is also greater.

[0055] In an embodiment of the present application, determining the first loss value of the first initial recognition model according to the confidence level, the first prediction result, and the first label data includes:

[0056] ;

[0057] Wherein, Loss1 represents the first loss value of the initial recognition model; R represents the first confidence level; e represents the natural constant; i is the index of the target in the training image indicated by the first label data, m is the number of targets in the training image indicated by the first label data; j is the index of the limb part of the target indicated by the first label data, and n is the number of limb parts of the target indicated by the first label data; is the number of pixel points of the limb part with index j in the target with index i in the training image, is the number of pixel points of the limb part with index j in the target with index i in the training image indicated by the first prediction result; is the number of pixel points of the torso part in the target with index i in the training image; is the score of the limb part with index j in the target with index i in the training image indicated by the first label data, is the score of the limb part with index j in the target with index i in the training image indicated by the first prediction result; is the score of the facial features of the target with index i in the training image indicated by the first label data, is the score of the facial features of the target with index i in the training image indicated by the first prediction result. Among them, the higher the first loss value, the higher the degree of difference between the first prediction result recognized by the first initial recognition model and the first label data, indicating that the accuracy of the first initial recognition model is lower; the lower the first loss value, the lower the degree of difference between the first prediction result recognized by the first initial recognition model and the first label data, indicating that the accuracy of the first initial recognition model is higher.

[0058] S24. Process the image information of the target in the sample image and the image information of the background area in the sample image based on a pre-trained style recognition model to obtain a first style recognition result of the sample image and a second style recognition result of the target.

[0059] In an embodiment of the present application, a pre-trained style recognition model is used to process a sample image, and the sample image is segmented into a target part and a background part for style recognition respectively. The image information of the target part and the image information related to the background part are input into the style recognition model to obtain a first style recognition result of the sample image and a second style recognition result of the target. Among them, the first style recognition result of the sample image is used to characterize the style information of the target corresponding to the sample image, and the second style recognition result is used to characterize the style information of the background part corresponding to the sample image.

[0060] In an embodiment of the present application, the style recognition model can be a pre-trained convolutional neural network model, a residual neural network model, or a fully connected neural network model. The present application does not make any limitations in this regard.

[0061] In an embodiment of the present application, the method for training a style recognition model includes: constructing second training data according to pre-stored training images; wherein, the training images correspond to second label data; the second label data includes a first style vector and a second style vector; determining a second confidence level of the corresponding second training data according to the second label data; inputting the second training data into a pre-constructed second initial recognition model to obtain a second prediction result corresponding to the second training data; the second prediction result includes a first predicted style vector and a second predicted style vector; determining a second loss value of the second initial recognition model according to the second confidence level, the second prediction result and the second label data; updating the second initial recognition model based on the backpropagation algorithm until the second loss value meets a preset condition, then stopping updating the initial recognition model to obtain a style recognition model trained to a convergent state. Among them, the higher the second loss value, the higher the degree of difference between the second prediction result recognized by the second initial recognition model and the second label data, which indicates that the accuracy of the second initial recognition model is lower; the lower the second loss value, the lower the degree of difference between the second prediction result recognized by the second initial recognition model and the second label data, which indicates that the accuracy of the second initial recognition model is higher. Among them, the preset condition may be that the second loss value is less than or equal to a preset second termination threshold. For example, when the preset second termination threshold is 0.01, in the case where the second loss value is less than or equal to 0.01, it can be determined that the difference between the second prediction result output by the second initial recognition model and the second label data is small, and the accuracy of the second prediction result output by the second initial recognition model is high, then the update of the second initial recognition model can be stopped to obtain a style recognition model trained to a convergent state.

[0062] In an embodiment of the present application, in order to improve the accuracy of training the style recognition model, a second confidence level of the corresponding second training data is also determined according to the second label data. Specifically, the calculation method of the second confidence level satisfies the following relational expression:

[0063] ;

[0064] Among them, S represents the second confidence level corresponding to the second training data; z is the index of the target in the training image indicated by the second label data, v is the number of targets in the training image indicated by the second label data; x is the index of the dimension in the first style vector indicated by the second label data, y is the number of dimensions in the first style vector indicated by the second label data; is the value of the dimension with index x in the first style vector corresponding to the target with index z in the training image indicated by the second label data; is the value of the dimension with index x in the second style vector indicated by the second prediction result.

[0065] Among them, in the second tag data, the lower the similarity between the first style vector and the second style vector, the greater the difference between the style of the target and the style of the background part in the training image, and the lower the second confidence level. Therefore, when training the style recognition model according to the training data corresponding to the training image, the lower the credibility of the obtained result, the greater the loss value of the style recognition model.

[0066] In an embodiment of the present application, determining the second loss value of the second initial recognition model according to the second confidence level, the second prediction result, and the second tag data includes:

[0067] ;

[0068] Among them, Loss2 represents the second loss value of the second initial recognition model; S represents the second confidence level corresponding to the second training data; e represents the natural constant; z is the index of the target in the training image indicated by the second tag data, and v is the number of targets in the training image indicated by the second tag data; x is the index of the dimension in the first style vector indicated by the second tag data, and y is the number of dimensions in the first style vector indicated by the second tag data; is the value of the dimension with index x in the first style vector corresponding to the target with index z in the training image indicated by the second tag data, is the value of the dimension with index x in the first predicted style vector corresponding to the target with index z in the training image indicated by the second prediction result; is the value of the dimension with index x in the second style vector indicated by the second tag data, is the value of the dimension with index x in the second predicted style vector indicated by the second prediction result.

[0069] S25. Determine the style difference degree between the target and the sample image according to the first style recognition result and the second style recognition result.

[0070] In an embodiment of the present application, the first style recognition result and the second style recognition result can be compared to determine the differences in aspects such as color features, texture features, and lighting features in the sample image. Specifically, the style difference degree between the first style recognition result and the second style recognition result can be determined based on the cosine similarity, and the style difference degree can also be determined based on the KL divergence between the first style recognition result and the second style recognition result. The present application does not limit the specific calculation method of the style difference degree. Specifically, the difference between the style of the target and the overall style of the sample image can be evaluated according to the style difference degree. When the style difference degree is higher, it indicates that the difference between the style of the target and the overall style of the sample image is higher, and the possibility that the target is a virtual target generated by artificial intelligence technology is greater; when the style difference degree is lower, it indicates that the difference between the style of the target and the overall style of the sample image is lower, and the possibility that the target is a virtual target generated by artificial intelligence technology is lower.

[0071] S26. Determine the recognition results of the sample image and the target according to the style difference degree and the target recognition result.

[0072] In an embodiment of the present application, it can be determined whether the target in the sample image to be recognized is a virtual target generated based on artificial intelligence according to the style difference degree and the target recognition result. Specifically, the style difference degree and the target recognition result can be combined and compared with a preset empirical threshold. If the style difference degree and the target recognition result exceed the preset empirical threshold, it indicates that the style difference degree is large and the human body features are unreasonable, which means that the target in the sample image is a virtual target generated based on artificial intelligence.

[0073] As can be seen from the above technical solutions, the embodiment of the present application accurately obtains the key features of the sample image through feature extraction, providing data support for subsequent target recognition and style recognition. The pre-trained target recognition model can accurately judge whether the target human body features in the sample image are reasonable and can capture the effective information of the target from the deep image features. In the process of analyzing the style of the sample image, not only the overall style of the sample image is concerned, but also the style recognition of the target part in the image is particularly carried out to ensure a comprehensive analysis of the image style. By comparing the difference degree between the target style and the overall style of the sample image, the consistency and authenticity of the image can be evaluated more carefully. Combining the target recognition result and the style difference degree can comprehensively judge whether the target in the sample image is a virtual target. This multi-dimensional judgment method helps to reduce misjudgment and missed judgment and improves the accuracy and reliability of the judgment.

[0074] Please refer to Figure 5 , Figure 5It is a functional block diagram of a virtual target recognition device provided by an embodiment of the present application. A virtual target recognition device 51 includes an acquisition module 510, a feature extraction module 520, a construction module 530, a first recognition module 540, a second recognition module 550, and a determination module 560. The module / unit referred to in the present application means a series of computer-readable instruction segments that can be executed by a processor 13 and can complete a fixed function, and is stored in a memory 12. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0075] The acquisition module 510 is configured to acquire a sample image to be recognized.

[0076] The feature extraction module 520 is configured to perform feature extraction processing on the sample image to obtain image features.

[0077] The construction module 530 is configured to construct data to be recognized according to the image features.

[0078] The first recognition module 540 is configured to process the data to be recognized based on a pre-trained target recognition model to obtain a target recognition result of the target in the sample image.

[0079] The second recognition module 550 is configured to process the image information of the target in the sample image and the image information of the background area in the sample image based on a pre-trained style recognition model to obtain a first style recognition result of the sample image and a second style recognition result of the target.

[0080] The determination module 560 is configured to determine a style difference degree between the target and the sample image according to the first style recognition result and the second style recognition result.

[0081] The determination module 560 is further configured to determine a recognition result of the sample image and the target according to the style difference degree and the target recognition result.

[0082] In some embodiments, the first recognition module 540 is further configured to construct first training data according to pre-stored training images; wherein, the training images correspond to first label data; determine a first confidence level of the corresponding first training data according to the first label data; the first confidence level is used to characterize the credibility of the target in the training image; input the first training data into a pre-constructed first initial recognition model to obtain a first prediction result corresponding to the first training data; determine a first loss value of the first initial recognition model according to the first confidence level, the first prediction result and the first label data; update the initial recognition model based on the backpropagation algorithm until the first loss value meets a preset condition, and then stop updating the first initial recognition model to obtain a target recognition model trained to a convergent state.

[0083] In some embodiments, the first recognition module 540 is further configured to determine the first confidence level of the corresponding first training data according to the first label data, including:

[0084] ;

[0085] wherein, R represents the first confidence level of the corresponding first training data; i is the index of the target in the training image indicated by the first label data, and m is the number of targets in the training image indicated by the first label data; j is the index of the limb part of the target indicated by the first label data, and n is the number of limb parts of the target indicated by the first label data; is the number of pixel points of the limb part with index j in the target with index i in the training image indicated by the first label data; is the number of pixel points of the torso part in the target with index i in the training image; is the score of the limb part with index j in the target with index i in the training image indicated by the first label data; is the score of the facial feature of the target with index i in the training image indicated by the first label data.

[0086] In some embodiments, the first recognition module 540 is further configured to determine the first loss value of the first initial recognition model according to the confidence level, the first prediction result and the first label data, including: ;

[0087] Among them, Loss1 represents the first loss value of the initial recognition model; R represents the first confidence level; e represents the natural constant; i is the index of the target in the training image indicated by the first label data, and m is the number of targets in the training image indicated by the first label data; j is the index of the limb part of the target indicated by the first label data, and n is the number of limb parts of the target indicated by the first label data; is the number of pixel points of the limb part with index j in the target with index i in the training image, is the number of pixel points of the limb part with index j in the target with index i in the training image indicated by the first prediction result; is the number of pixel points of the torso part in the target with index i in the training image; is the score of the limb part with index j in the target with index i in the training image indicated by the first label data, is the score of the limb part with index j in the target with index i in the training image indicated by the first prediction result; is the score of the facial feature of the target with index i in the training image indicated by the first label data, is the score of the facial feature of the target with index i in the training image indicated by the first prediction result.

[0088] In some embodiments, the second recognition module 550 is further configured to construct second training data according to the pre-stored training images; wherein, the training images correspond to second label data; the second label data includes a first style vector and a second style vector; determine the second confidence level of the corresponding second training data according to the second label data; input the second training data into a pre-constructed second initial recognition model to obtain a second prediction result corresponding to the second training data; the second prediction result includes a first predicted style vector and a second predicted style vector; determine the second loss value of the second initial recognition model according to the second confidence level, the second prediction result and the second label data; update the second initial recognition model based on the backpropagation algorithm until the second loss value meets a preset condition, and then stop updating the initial recognition model to obtain a style recognition model trained to a convergent state.

[0089] In some embodiments, the second recognition module 550 is further configured to determine the second confidence level of the corresponding second training data according to the second label data, including:

[0090] ; wherein, S represents the second confidence corresponding to the second training data; z is the index of the target in the training image indicated by the second label data, and v is the number of targets in the training image indicated by the second label data; x is the index of the dimension in the first style vector indicated by the second label data, and y is the number of dimensions in the first style vector indicated by the second label data; is the value of the dimension with index x in the first style vector corresponding to the target with index z in the training image indicated by the second label data; is the value of the dimension with index x in the second style vector indicated by the second prediction result.

[0091] In some embodiments, the second recognition module 550 is further configured to determine the second loss value of the second initial recognition model according to the second confidence, the second prediction result, and the second label data, including:

[0092] ;

[0093] wherein, Loss2 represents the second loss value of the second initial recognition model; S represents the second confidence corresponding to the second training data; e represents the natural constant; z is the index of the target in the training image indicated by the second label data, and v is the number of targets in the training image indicated by the second label data; x is the index of the dimension in the first style vector indicated by the second label data, and y is the number of dimensions in the first style vector indicated by the second label data; is the value of the dimension with index x in the first style vector corresponding to the target with index z in the training image indicated by the second label data, is the value of the dimension with index x in the first predicted style vector corresponding to the target with index z in the training image indicated by the second prediction result; is the value of the dimension with index x in the second style vector indicated by the second label data, is the value of the dimension with index x in the second predicted style vector indicated by the second prediction result.

[0094] Please refer to Figure 6 , which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 100 includes a memory 12 and a processor 13. The memory 12 is used to store computer-readable instructions, and the processor 13 is configured to execute the computer-readable instructions stored in the memory to implement the virtual target recognition method according to any one of the above embodiments.

[0095] In an embodiment of the present application, the electronic device 100 further includes a bus and a computer program stored in the memory 12 and executable on the processor 13, such as a virtual target recognition program.

[0096] Figure 6 Only the electronic device 100 having the memory 12 and the processor 13 is shown. Those skilled in the art can understand that Figure 6 the shown structure does not constitute a limitation on the electronic device 100, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0097] Combined with Figure 2 , the memory 12 in the electronic device 100 stores a plurality of computer-readable instructions to implement a virtual target recognition method. The processor 13 can execute the plurality of instructions to achieve: obtaining a sample image to be recognized; performing feature extraction processing on the sample image to obtain image features; constructing data to be recognized according to the image features; processing the data to be recognized based on a pre-trained target recognition model to obtain a target recognition result of the target in the sample image; processing the image information of the target and the image information of the background area in the sample image based on a pre-trained style recognition model to obtain a first style recognition result of the sample image and a second style recognition result of the target; determining a style difference degree between the target and the sample image according to the first style recognition result and the second style recognition result; and determining the recognition results of the sample image and the target according to the style difference degree and the target recognition result.

[0098] Specifically, for the specific implementation method of the above instructions by the processor 13, reference may be made to Figure 2 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.

[0099] Those skilled in the art can understand that the schematic diagram is only an example of the electronic device 100 and does not constitute a limitation on the electronic device 100. The electronic device 100 may be of a bus structure or a star structure. The electronic device 100 may also include more or fewer other hardware or software than shown, or different component arrangements. For example, the electronic device 100 may also include input / output devices, network access devices, etc.

[0100] It should be noted that the electronic device 100 is only an example. Other existing or future electronic products that can be adapted to the present application should also be included in the protection scope of the present application and are incorporated herein by reference.

[0101] Among them, the memory 12 includes at least one type of readable storage medium, which can be non-volatile or volatile. The readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. The memory 12 can be an internal storage unit of the electronic device 100 in some embodiments, such as the mobile hard disk of the electronic device 100. The memory 12 can also be an external storage device of the electronic device 100 in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a FlashCard, etc. equipped on the electronic device 100. The memory 12 can not only be used to store application software installed in the electronic device 100 and various types of data, such as the code of a virtual target recognition program, etc., but also be used to temporarily store data that has been output or will be output.

[0102] In some embodiments, the processor 13 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions, including a combination of one or more central processing units (CPU), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 13 is the control core (Control Unit) of the electronic device 100, connecting various components of the entire electronic device 100 through various interfaces and lines. By running or executing programs or modules stored in the memory 12 (such as executing a virtual target recognition program), and calling data stored in the memory 12, it performs various functions of the electronic device 100 and processes data.

[0103] The processor 13 executes the operating system of the electronic device 100 and various installed application programs. The processor 13 executes the application program to implement the steps in each of the above embodiments of the virtual target recognition method, such as Figure 2 the steps shown.

[0104] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 100. For example, the computer program may be divided into an acquisition module 510, a feature extraction module 520, a construction module 530, a first recognition module 540, a second recognition module 550, and a determination module 560.

[0105] The integrated units implemented in the form of software function modules as described above may be stored in a computer-readable storage medium. The above-mentioned software function modules stored in a storage medium include several instructions for causing a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor (Processor) to execute a part of the virtual target recognition method described in each embodiment of the present application.

[0106] If the modules / units integrated in the electronic device 100 are implemented in the form of software function units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it may also be completed by a computer program instructing relevant hardware devices. The computer program may be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned various method embodiments may be implemented.

[0107] Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory, and other memories, etc.

[0108] Further, the computer-readable storage medium mainly includes a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.

[0109] The bus can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, in Figure 6 only one arrow is used to represent it, but it does not mean that there is only one bus or one type of bus. The bus is arranged to implement the connection and communication between the memory 12 and at least one processor 13, etc.

[0110] The embodiment of the present application also provides a computer-readable storage medium (not shown in the figure). Computer-readable instructions are stored in the computer-readable storage medium, and the computer-readable instructions are executed by a processor in an electronic device to implement the virtual target recognition method described in any one of the above embodiments.

[0111] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.

[0112] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0113] In addition, in each embodiment of the present application, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a hardware plus software functional module.

[0114] In addition, obviously, the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices described in the specification can also be implemented by one unit or device through software or hardware. Words such as first and second are used to represent names and do not represent any specific order.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A virtual target recognition method, applied to electronic equipment, characterized in that: The method comprises: Obtaining a sample image to be identified; Performing feature extraction processing on the sample image to obtain image features; Constructing data to be identified according to the image features; Processing the data to be identified based on a pre-trained target recognition model to obtain a target recognition result of the target in the sample image; the method also includes training the target recognition model, and the training of the target recognition model includes: constructing first training data based on a pre-stored training image; wherein the training image corresponds to first label data; determining a first confidence level of the corresponding first training data based on the first label data; the first confidence level is used to characterize the credibility of the target in the training image; inputting the first training data into a pre-constructed first initial recognition model to obtain a first prediction result corresponding to the first training data; determining a first loss value of the first initial recognition model based on the first confidence level, the first prediction result and the first label data; updating the initial recognition model based on a back-propagation algorithm until the first loss value meets a preset condition, stopping updating the first initial recognition model, and obtaining a target recognition model trained to a convergence state; wherein determining the first loss value includes: ; Wherein, Loss1 represents the first loss value of the initial recognition model; R represents the first confidence; e represents a natural constant; i is the index of the target in the training image indicated by the first label data, and m is the number of targets in the training image indicated by the first label data; j is the index of the limb part of the target indicated by the first label data, and n is the number of limb parts of the target indicated by the first label data; is the number of pixels of the limb part with index j in the target with index i in the training image, is the number of pixels of the limb part with index j in the target with index i in the training image indicated by the first prediction result; is the number of pixels in the torso of the target with index i in the training image; is the score of the limb part with index j in the target with index i in the training image indicated by the first label data, is the score of the limb part with index j in the target with index i in the training image indicated by the first prediction result; is the score of the facial feature of the target with index i in the training image indicated by the first label data, is the score of the facial feature of the target with index i in the training image indicated by the first prediction result; Processing image information of the target in the sample image and image information of a background area in the sample image based on a pre-trained style recognition model to obtain a first style recognition result of the sample image and a second style recognition result of the target; Determining a style difference between the target and the sample image according to the first style recognition result and the second style recognition result; The recognition result of the sample image and the target is determined according to the style difference and the target recognition result.

2. The virtual target recognition method according to claim 1, characterized in that: Determining the first confidence level of the first training data corresponding to the first label data includes: ; Wherein, R represents the first confidence corresponding to the first training data; i is the index of the target in the training image indicated by the first label data, and m is the number of targets in the training image indicated by the first label data; j is the index of the limb part of the target indicated by the first label data, and n is the number of limb parts of the target indicated by the first label data; is the number of pixels of the limb part with index j in the target with index i in the training image indicated by the first label data; is the number of pixels in the torso of the target with index i in the training image; is the score of the limb part with index j in the target with index i in the training image indicated by the first label data; is the score of the facial features of the target with index i in the training image indicated by the first label data.

3. The virtual target recognition method according to claim 1, characterized in that: The method further includes training the style recognition model, wherein the training the style recognition model includes: Constructing second training data according to a pre-stored training image; wherein the training image corresponds to second label data; and the second label data includes a first style vector and a second style vector; Determine a second confidence level of the corresponding second training data according to the second label data; Inputting the second training data into a pre-built second initial recognition model to obtain a second prediction result corresponding to the second training data; the second prediction result includes a first prediction style vector and a second prediction style vector; Determine a second loss value of the second initial recognition model according to the second confidence, the second prediction result and the second label data; The second initial recognition model is updated based on the back propagation algorithm until the second loss value meets a preset condition, and then the updating of the initial recognition model is stopped to obtain a style recognition model trained to a convergent state.

4. The virtual target recognition method according to claim 3, characterized in that: Determining the second confidence level of the corresponding second training data according to the second label data includes: ; Wherein, S represents the second confidence corresponding to the second training data; z is the index of the target in the training image indicated by the second label data, and v is the number of targets in the training image indicated by the second label data; x is the index of the dimension in the first style vector indicated by the second label data, and y is the number of dimensions in the first style vector indicated by the second label data; is the value of the dimension with index x in the first style vector corresponding to the target with index z in the training image indicated by the second label data; is the numerical value of the dimension indexed by x in the second style vector indicated by the second prediction result.

5. The virtual target recognition method according to claim 3, characterized in that: The determining the second loss value of the second initial recognition model according to the second confidence, the second prediction result and the second label data comprises: ; Wherein, Loss2 represents the second loss value of the second initial recognition model; S represents the second confidence corresponding to the second training data; e represents a natural constant; z is the index of the target in the training image indicated by the second label data, and v is the number of targets in the training image indicated by the second label data; x is the index of the dimension in the first style vector indicated by the second label data, and y is the number of dimensions in the first style vector indicated by the second label data; is the value of the dimension with index x in the first style vector corresponding to the target with index z in the training image indicated by the second label data, is the value of the dimension with index x in the first predicted style vector corresponding to the target with index z in the training image indicated by the second prediction result; is the value of the dimension indexed by x in the second style vector indicated by the second label data, is the numerical value of the dimension with index x in the second prediction style vector indicated by the second prediction result.

6. A virtual target recognition device, characterized in that: The device comprises a module for implementing the virtual target recognition method according to any one of claims 1 to 5, and the device comprises: An acquisition module, used for acquiring a sample image to be identified; A feature extraction module, used to perform feature extraction processing on the sample image to obtain image features; A construction module, used to construct the data to be identified according to the image features; A first recognition module, used for processing the to-be-recognized data based on a pre-trained target recognition model to obtain a target recognition result of the target in the sample image; A second recognition module, used for processing the image information of the target in the sample image and the image information of the background area in the sample image based on a pre-trained style recognition model to obtain a first style recognition result of the sample image and a second style recognition result of the target; A determination module, configured to determine a style difference between the target and the sample image according to the first style recognition result and the second style recognition result; The determination module is further configured to determine the sample image and the recognition result of the target according to the style difference and the target recognition result.

7. An electronic device, characterized in that: The electronic device comprises a processor and a memory, and the processor is used to implement the virtual target recognition method as claimed in any one of claims 1 to 5 when executing a computer program stored in the memory.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the virtual target recognition method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Portrait style recognition method and device and computer readable storage medium

    CN110555481A

  • Synthetic image recognition method and device and image recognition model training method and device

    CN114820467A