Identity attribute recognition method, computer readable storage medium and intelligent device
By combining image information and voice information to identify the identity attributes of the target, the problem of the inability to accurately identify the identity attributes of children in the prior art is solved, and the accuracy and scope of application of identification are improved.
Patent Information
- Application Number
- PCT/CN2024/136012
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-26
AI Technical Summary
Existing identity attribute recognition methods cannot accurately and effectively identify the age-group attributes of different users, such as whether they are children, especially when facial features are not distinguished or there is occlusion or poor light.
By acquiring the image information and voice information of the target, the first predicted identity category and first position information of the target are identified based on the image information, the second position information of the target is determined based on the voice information, and the identity attribute of the target is determined based on at least based on the information. This method combines multimodal information to improve recognition accuracy.
It effectively avoids the shortcomings of identity attribute recognition based on single-modal information, improves the accuracy of identity attribute recognition, and expands applicable scenarios.
Smart Images

Figure CN2024136012_26062025_PF_FP_ABST
Abstract
Description
Identity attribute recognition method, computer-readable storage medium, and intelligent device
[0001] This application claims priority to Chinese patent application No. 202311777729.5 filed on December 21, 2023, entitled “A method for identifying identity attributes, a computer-readable storage medium and a smart device”. The entire contents of the above Chinese patent application are incorporated into this application by reference. Technical Field
[0002] The present application relates to the field of data processing technology, and specifically provides an identity attribute recognition method, a computer-readable storage medium, and an intelligent device. Background Art
[0003] Currently, identity recognition primarily relies on determining gender, using voice or facial features. Because voice and facial features vary significantly between genders, single-modal recognition methods based on voice or facial features can be used to identify identity with high accuracy.
[0004] However, existing identity attribute recognition cannot accurately and effectively identify different user age groups, such as whether they are children. When identifying children's identity attributes, single-modal identification methods, such as facial feature-based identification, are not conducive to accurate identification due to the lack of distinction between children's and adults' facial features, facial occlusion, and poor lighting. In addition, single-modal identification based on voice also has the problem that adult women and children have similar fundamental frequencies, making it difficult to identify children's identity attributes. Summary of the Invention
[0005] This application aims to solve the above technical problems, that is, to solve the problem of low accuracy of existing identity attribute recognition methods.
[0006] In a first aspect, the present application provides an identity attribute recognition method, comprising:
[0007] Obtain target image and voice information;
[0008] identifying a first predicted identity category and first location information of the target based on the image information, and determining second location information of the target based on the voice information;
[0009] An identity attribute of the target is determined based on at least the first predicted identity category, the first location information, and the second location information.
[0010] In some embodiments, the method further comprises: obtaining a first confidence level of the first predicted identity category;
[0011] Determining the identity attribute of the target based on at least the first predicted identity category, the first location information, and the second location information includes:
[0012] An identity attribute of the target is determined based on the first predicted identity category, the first confidence level, the first location information, and the second location information.
[0013] In some embodiments, acquiring image information of the target includes acquiring multiple frames of image information of the target;
[0014] The determining the identity attribute of the target based on the first predicted identity category, the first confidence level, the first location information, and the second location information includes:
[0015] Comparing the first position information and the second position information of each frame of image information in the multiple frames of image information to obtain a first position comparison result of the multiple frames of image information;
[0016] Comparing a first confidence level of each frame of image information in the plurality of frames of image information with a first preset threshold range;
[0017] Obtaining a first confidence level based on a comparison result of the multiple frames of image information and the first preset threshold range;
[0018] An identity attribute of the target is determined based on the first predicted identity category, the first position comparison result, and the first confidence level.
[0019] In some embodiments, the method further comprises: identifying a second predicted identity category of the target based on the speech information;
[0020] The determining the identity attribute of the target based at least on the first predicted identity category, the first location information, and the second location information includes:
[0021] An identity attribute of the target is determined based on the first predicted identity category, the first location information, the second predicted identity category, and the second location information.
[0022] In some embodiments, the method further comprises: obtaining a second confidence level of the first predicted identity category; and identifying the second predicted identity category of the target based on the voice information and obtaining a third confidence level of the second predicted identity category;
[0023] The determining the identity attribute of the target based at least on the first predicted identity category, the first location information, and the second location information includes:
[0024] An identity attribute of the target is determined based on the first predicted identity category, the first location information, the second confidence level, the second predicted identity category, the third confidence level, and the second location information.
[0025] In some embodiments, acquiring image information of the target includes acquiring multiple frames of image information of the target;
[0026] The determining the identity attribute of the target based on the first predicted identity category, the first location information, the second confidence level, the second predicted identity category, the third confidence level, and the second location information includes:
[0027] Comparing the first position information and the second position information of each frame of image information in the multiple frames of image information to obtain a second position comparison result of the multiple frames of image information;
[0028] Comparing the second confidence level of each frame of image information in the multiple frames of image information with a second preset threshold range; and obtaining a second confidence level based on the comparison result of the multiple frames of image information and the second preset threshold range;
[0029] Comparing the third confidence level with a third preset threshold range; obtaining a third confidence level according to the comparison result;
[0030] An identity attribute of the target is determined based on the first predicted identity category, the second predicted identity category, the second position comparison result, the second confidence level, and the third confidence level.
[0031] In some embodiments, identifying the second predicted identity category of the target based on the voice information includes:
[0032] The speech information is recognized using a pre-trained neural network model to obtain a second predicted identity category of the target; the second predicted identity category includes a child and / or an adult.
[0033] In some embodiments, identifying the first predicted identity category of the target based on the image information includes:
[0034] The image information is identified using a pre-trained target identity attribute recognition model to obtain a first predicted identity category of the target, where the first predicted identity category includes a child and / or an adult.
[0035] In a second aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the identity attribute recognition method described in any one of the above items is implemented.
[0036] In a third aspect, the present application provides a smart device comprising:
[0037] at least one processor;
[0038] and, a memory communicatively coupled to the at least one processor;
[0039] The memory stores a computer program, and when the computer program is executed by the at least one processor, it implements any of the above-mentioned identity attribute recognition methods.
[0040] By employing the above technical solution, the present application can obtain image and voice information of a target; identify the target's first predicted identity category and first location information based on the image information; and determine the target's second location information based on the voice information; and determine the target's identity attributes based on at least the first predicted identity category, first location information, and second location information. This method effectively avoids the shortcomings of identity attribute recognition based on unimodal target information, improves the accuracy of identity attribute recognition, and has a wider range of applicable scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The preferred embodiments of the present application are described below with reference to the accompanying drawings, in which:
[0042] FIG1 is a flow chart showing the main steps of an identity attribute recognition method provided in an embodiment of the present application;
[0043] FIG2 is a schematic diagram of a process for determining the identity attributes of a target according to an embodiment of the present application;
[0044] FIG3 is a flow chart of an identity attribute recognition method provided by a preferred embodiment of the present application;
[0045] FIG4 is a flow chart of an identity attribute recognition method provided by another preferred embodiment of the present application;
[0046] FIG5 is a schematic flow chart of a method for implementing step S46 provided in an embodiment of the present application;
[0047] FIG6 is a flowchart of a method for identifying a second confidence level of a child's attributes based on image results and sound source localization results provided by the present application;
[0048] FIG7 is a flowchart of a method for identifying a third confidence level of a child's attribute based on a speech attribute result provided by a specific example of the present application;
[0049] FIG8 is a schematic diagram of the structure of a smart device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0050] Some embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the scope of protection of the present application.
[0051] Referring to FIG. 1 , FIG. 1 is a flow chart showing the main steps of an identity attribute recognition method provided in an embodiment of the present application, which may include:
[0052] Step S11: Acquire the target's image information and voice information;
[0053] Step S12: identifying a first predicted identity category and first location information of the target based on the image information, and determining second location information of the target based on the voice information;
[0054] Step S13: Determine the identity attribute of the target based on at least the first predicted identity category, the first location information, and the second location information.
[0055] In some embodiments, step S11 may specifically be obtaining the target's image information and voice information based on pre-stored data; in other embodiments, step S11 may also be obtaining the target's image information and voice information in real time based on an image acquisition device and a voice acquisition device.
[0056] In some embodiments, the first predicted identity category and first location information of the target identified based on the image information in step S12 may be specifically:
[0057] Using a pre-trained target identity attribute recognition model to recognize the image information, and obtaining a first predicted identity category of the target, the first predicted identity category including a child and / or an adult;
[0058] Based on the image information, the target detection algorithm is used to obtain the first position information of the target.
[0059] As an example, the target detection algorithm may include a YOLO (You Only Look Once) algorithm. In other embodiments, other algorithms in the art that can achieve the same function may also be used to obtain the first position information of the target.
[0060] In some embodiments, determining the second position information of the target based on the voice information in step S12 may specifically be determining the second position information of the target based on the voice information using a sound source localization technology.
[0061] In some embodiments, step S13 may specifically be determining the identity attribute of the target based on the first predicted identity category, the first location information, and the second location information.
[0062] In some embodiments, determining the target's identity attribute based on the first predicted identity category, the first location information, and the second location information can specifically include determining whether the first location information and the second location information are consistent; and if the first location information and the second location information are consistent, determining the first predicted identity category as the target's identity attribute. Combining the determination of whether the first location information and the second location information are consistent facilitates determining the validity of the first predicted identity category.
[0063] As a specific example, in a driving device scenario, when the first predicted identity category is a child, the first location information is the driver's seat, and the second location information is the rear seat of the driving device, and the first location information and the second location information are inconsistent, it can be determined that the reliability of the current first predicted identity category is low. When the first predicted identity category is a child, the first location information and the second location information are both the rear seat of the driving device, and the first location information and the second location information are consistent, the first predicted identity category can be determined as the identity attribute of the target.
[0064] In other embodiments, in order to obtain the target identity attributes more accurately, it can also include: obtaining a first confidence level of the first predicted identity category; the first confidence level can simultaneously obtain the first predicted identity category and the first confidence level of the target when using a pre-trained target identity attribute recognition model to identify image information.
[0065] Step S13 may also be specifically to determine the identity attribute of the target based on the first predicted identity category, the first confidence level, the first location information and the second location information.
[0066] In some embodiments, based on the first predicted identity category, the first confidence level, the first location information, and the second location information, determining the identity attribute of the target can be specifically as follows: when the first location information and the second location information are consistent, judging whether the first confidence level is greater than a preset first confidence threshold; if so, determining the first predicted identity category as the identity attribute of the target.
[0067] In other embodiments, obtaining the image information of the target in step S11 may specifically be obtaining multiple frames of image information of the target. Referring to FIG2 , FIG2 is a schematic diagram of a process for determining the identity attributes of a target according to an embodiment of the present application. Based on the first predicted identity category, the first confidence level, the first location information, and the second location information, determining the identity attributes of the target may specifically include:
[0068] Step S131: comparing the first position information and the second position information of each frame of image information in the multiple frames of image information to obtain a first position comparison result of the multiple frames of image information;
[0069] Step S132: comparing the first confidence level of each frame of image information in the multiple frames of image information with a first preset threshold range;
[0070] Step S133: obtaining a first confidence level based on a comparison result of the multiple frames of image information and a first preset threshold range;
[0071] Step S134: determining the identity attribute of the target according to the first predicted identity category, the first position comparison result, and the first confidence level.
[0072] In some embodiments, step S133 can be specifically as follows: when the first confidence of at least one frame of image information in multiple frames of image information is greater than the high threshold value of the first preset threshold range, determining the first confidence level to be high; when the first confidence of multiple frames of image information are all less than the high threshold value and at least one frame is greater than the low threshold value of the first preset threshold range, determining the first confidence level to be medium; when the first confidence of multiple frames of image information are all less than the low threshold value, determining the first confidence level to be low.
[0073] Correspondingly, in step S134 , when the first position comparison result is consistent and the first confidence level is high, the first predicted identity category may be determined as the identity attribute of the target.
[0074] In the above embodiment, determining the identity attributes of the target based on multi-frame image information is helpful to avoid accidental errors and improve the reliability of identity attribute recognition.
[0075] The above is a method for identifying an identity attribute provided in an embodiment of the present application. This method obtains image and voice information of a target; identifies the target's first predicted identity category and first location information based on the image information; and determines the target's second location information based on the voice information; and determines the target's identity attributes based on at least the first predicted identity category, first location information, and second location information. This method effectively avoids the shortcomings of identifying an identity attribute based on unimodal target information, improves the accuracy of identity attribute identification, and has a wider range of applicable scenarios.
[0076] Referring to FIG. 3 , FIG. 3 is a flow chart of an identity attribute recognition method provided by a preferred embodiment of the present application, which may include:
[0077] Step S31: Acquire the target's image information and voice information;
[0078] Step S32: identifying a first predicted identity category and first location information of the target based on the image information;
[0079] Step S33: determining second location information of the target based on the voice information; and identifying a second predicted identity category of the target based on the voice information;
[0080] Step S34: Determine the identity attribute of the target based on the first predicted identity category, the first location information, the second predicted identity category, and the second location information.
[0081] Among them, determining the second location information of the target based on the voice information in steps S31, S32 and step S33 can be implemented in the same manner as steps S11 and S12. For the sake of brevity, they are not repeated here. For details, please refer to the description above.
[0082] In some embodiments, the second predicted identity category of the target identified based on voice information in step S33 can be specifically: using a pre-trained neural network model to identify the voice information to obtain the second predicted identity category of the target; the second predicted identity category includes children and / or adults.
[0083] In some embodiments, step S34 can be specifically to determine whether the first location information and the second location information are consistent, and to determine whether the first predicted identity category and the second predicted identity category are consistent; when the first location information and the second location information are consistent and the first predicted identity category and the second predicted identity category are consistent, the first predicted identity category or the second predicted identity category can be determined as the identity attribute of the target.
[0084] The above is an identity attribute recognition method provided by a preferred embodiment of the present application, which can achieve the same beneficial effects as the corresponding embodiment of Figure 1, and by further combining the voice information to obtain a second predicted identity category, the identity attribute is identified based on the second predicted identity category, which is conducive to further improving the accuracy of recognition and avoiding misidentification caused by occlusion and other situations, and has a wider scope of application.
[0085] In some preferred embodiments, in order to improve the accuracy of identity attribute recognition, a second confidence level of the first predicted identity category and a third confidence level of the second predicted identity category may also be obtained.
[0086] Referring to FIG4 , FIG4 is a flow chart of an identity attribute recognition method provided by another preferred embodiment of the present application, which may include:
[0087] Step S41: Acquire the target's image information and voice information;
[0088] Step S42: identifying a first predicted identity category of the target and a second confidence level of the first predicted identity category based on the image information;
[0089] Step S43: identifying first position information of the target based on the image information;
[0090] Step S44: determining second location information of the target based on the voice information;
[0091] Step S45: identifying a second predicted identity category of the target based on the voice information and a third confidence level of the second predicted identity category;
[0092] Step S46: Determine the identity attribute of the target based on the first predicted identity category, the first location information, the second confidence level, the second predicted identity category, the third confidence level, and the second location information.
[0093] Among them, steps S41, S43, and S44 can be implemented in the same manner as the corresponding parts in steps S31-S33. For the sake of brevity, they are not repeated here. For details, please refer to the description above.
[0094] It should be noted that, in some embodiments, steps S43 and S42 may be performed simultaneously.
[0095] In some embodiments, step S42 may specifically be to use a pre-trained target identity attribute recognition model to recognize the image information and simultaneously obtain a first predicted identity category and a second confidence level of the target. The first predicted identity category includes children and / or adults.
[0096] In some embodiments, step S45 may specifically include: using a pre-trained neural network model to recognize the voice information, and simultaneously obtaining a second predicted identity category and a third confidence level of the target. The second predicted identity category includes a child and / or an adult.
[0097] In some embodiments, step S46 may specifically include:
[0098] comparing the first predicted identity category to the second predicted identity category;
[0099] comparing the first location information with the second location information;
[0100] Comparing the second confidence level with a second preset threshold range, and obtaining a second confidence level according to the comparison result;
[0101] Comparing the third confidence level with a third preset threshold range, and obtaining a third confidence level according to the comparison result;
[0102] When the first predicted identity category and the second predicted identity category are consistent, the first location information and the second location information are consistent, and at least one of the second confidence level and the third confidence level is high, the first predicted identity category or the second predicted identity category is determined as the identity attribute of the target.
[0103] In some embodiments, acquiring the image information of the target in step S41 may include acquiring multiple frames of image information of the target. Accordingly, referring to FIG5 , FIG5 is a flow chart of a method for implementing step S46 provided in an embodiment of the present application, which may include:
[0104] Step S461: Comparing the first position information and the second position information of each frame of image information in the multiple frames of image information to obtain a second position comparison result of the multiple frames of image information;
[0105] Step S462: comparing the second confidence level of each frame of image information in the multiple frames of image information with the second preset threshold range; obtaining a second confidence level based on the comparison result between the multiple frames of image information and the second preset threshold range;
[0106] Step S463: Compare the third confidence level with the third preset threshold range; obtain a third confidence level based on the comparison result;
[0107] Step S464: Determine the identity attribute of the target according to the first predicted identity category, the second predicted identity category, the second position comparison result, the second confidence level, and the third confidence level.
[0108] In some embodiments, the second confidence level obtained in step S462 based on the comparison results of multiple frames of image information and the second preset threshold range can be specifically determined as follows: when the second confidence of at least one frame of image information in the multiple frames of image information is greater than the high threshold value of the second preset threshold range, the second confidence level is determined to be high; when the second confidence of multiple frames of image information are all less than the high threshold value and at least one frame is greater than the low threshold value of the second preset threshold range, the second confidence level is determined to be medium; when the second confidence of multiple frames of image information are all less than the low threshold value, the second confidence level is determined to be low.
[0109] In some embodiments, step S463 can be specifically as follows: when the third confidence is greater than the high threshold value of the third preset threshold range, determining that the third confidence level is high; when the third confidence is less than the high threshold value of the third preset threshold range and greater than the low threshold value of the third preset threshold range, determining that the third confidence level is medium; when the third confidence is less than the low threshold value of the third preset threshold range, determining that the third confidence level is low.
[0110] In some embodiments, step S464 may be specifically as follows: when the first predicted identity category and the second predicted identity category are consistent, the second position comparison result is consistent, and at least one of the second confidence level and the third confidence level is high, the first predicted identity category or the second predicted identity category is determined as the identity attribute of the target.
[0111] As a specific example, the identity attribute recognition method provided in this application can be used to identify children. Referring to Figure 6, Figure 6 is a flow chart of a method for identifying a second confidence level of a child's attributes based on image results and sound source localization results provided in a specific example of this application, wherein the child's identity attribute is identified based on the image results (such as the first predicted identity attribute, the first location information, and the second confidence level) obtained for each frame of image information in the multi-frame image information and the sound source localization result (i.e., the second location information) based on the audio information. When the multi-frame image information and the sound source localization results from the 1st to the nth frame all indicate the presence of a child and the image matches the sound source position, the child score corresponding to each frame of image information (which can be represented by the second confidence level) is compared with the high threshold value and the low threshold value of the second preset threshold range to obtain a second confidence level. The second confidence level may include high confidence, medium confidence, and low confidence.
[0112] Figure 7 is a flow chart of a method for identifying the third confidence level of child attributes based on voice attribute results provided in a specific example of the present application, wherein the voice attribute result is the second predicted identity category and the third confidence level obtained based on the voice information. When the result is a child and the child score (which can be represented by the third confidence level) is greater than the high threshold value of the third preset threshold range, the third confidence level can be high confidence; when the result is a child and the child score is less than the high threshold value of the third preset threshold range and greater than the low threshold value, the third confidence level can be medium confidence; when the result is a child and the child score is less than the low threshold value of the third preset threshold range, the third confidence level can be low confidence.
[0113] The above is an identity attribute recognition method provided by another preferred embodiment of the present application. By combining the predicted identity category, location information and confidence level of multiple modalities of image information and semantic information, it is helpful to overcome the shortcomings of identity attribute recognition under a single modality, improve the accuracy of identity attribute recognition, and has a wider range of applicable scenarios.
[0114] It will be understood by those skilled in the art that all or part of the processes in the method for implementing the above-mentioned embodiments of the present application can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium, etc. that can carry the computer program code.
[0115] In another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program. When executed by a processor, the computer program implements the identity attribute recognition method described in any of the above embodiments. The computer-readable storage medium may be a storage device formed by various electronic devices. Optionally, in the embodiments of the present application, the computer-readable storage medium is a non-transitory computer-readable storage medium.
[0116] Another aspect of the present application provides an intelligent device, which may include at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program, and when the computer program is executed by the at least one processor, it implements the identity attribute recognition method described in any of the above embodiments. The intelligent device described in the present application may include a driving device, a smart car, a robot, and other devices. Referring to Figure 8, Figure 8 exemplarily shows that the memory 81 and the processor 82 are communicatively connected via a bus.
[0117] In some embodiments of the present application, the smart device further includes at least one sensor configured to sense information. The sensor is communicatively coupled to any of the processors described herein. Optionally, the smart device further includes an autonomous driving system configured to guide the smart device to autonomously drive or provide assisted driving. The processor communicates with the sensor and / or autonomous driving system to perform the identity attribute recognition method described in any of the above embodiments.
[0118] In some embodiments, the memory and the processor are connected via a bus, and only one memory and only one processor are provided.
[0119] In other embodiments, the smart device may include multiple memories and multiple processors. The program that executes the identity attribute recognition method of any of the above-described embodiments may be divided into multiple subroutines, each of which may be loaded and executed by a processor to perform different steps of the above-described method embodiments. Specifically, each subroutine may be stored in a different memory, and each processor may be configured to execute the programs in one or more memories to collectively implement the identity attribute recognition method of the above-described method embodiments.
[0120] The relevant user personal information that may be involved in the various embodiments of this application is strictly in accordance with the requirements of laws and regulations, following the principles of legality, legitimacy and necessity, and based on the reasonable purposes of business scenarios, to process the personal information that users actively provide during the use of products / services or generated due to the use of products / services, as well as the personal information obtained with the user's authorization.
[0121] The user personal information processed by this application will vary depending on the specific product / service scenario and must be based on the specific scenario in which the user uses the product / service. It may involve the user's voice information, image information, account information, device information, driving information, vehicle information, or other related information. This application will treat the user's personal information and its processing with a high degree of diligence.
[0122] This application attaches great importance to the security of user personal information and has taken reasonable and feasible security protection measures that comply with industry standards to protect user information and prevent personal information from being accessed, disclosed, used, modified, damaged or lost without authorization.
[0123] Thus far, the technical solutions of the present application have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of the present application is obviously not limited to these specific embodiments. Without departing from the principles of the present application, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present application.
Claims
1. A method for identifying identity attributes, characterized in that: include: Obtain the target's image and voice information; identifying a first predicted identity category and first location information of the target based on the image information, and determining second location information of the target based on the voice information; An identity attribute of the target is determined based on at least the first predicted identity category, the first location information, and the second location information.
2. The method according to claim 1, characterized in that: The method further includes: obtaining a first confidence level of the first predicted identity category; Determining the identity attribute of the target based at least on the first predicted identity category, the first location information, and the second location information includes: An identity attribute of the target is determined based on the first predicted identity category, the first confidence level, the first location information, and the second location information.
3. The method according to claim 2, characterized in that Acquiring image information of the target includes acquiring multiple frames of image information of the target; The determining the identity attribute of the target based on the first predicted identity category, the first confidence level, the first location information, and the second location information includes: Comparing the first position information and the second position information of each frame of image information in the multiple frames of image information to obtain a first position comparison result of the multiple frames of image information; Comparing a first confidence level of each frame of image information in the multiple frames of image information with a first preset threshold range; Obtaining a first confidence level according to a comparison result between the multiple frames of image information and the first preset threshold range; An identity attribute of the target is determined based on the first predicted identity category, the first position comparison result, and the first confidence level.
4. The method according to claim 1, characterized in that The method further includes: identifying a second predicted identity category of the target based on the speech information; The determining the identity attribute of the target based at least on the first predicted identity category, the first location information, and the second location information includes: An identity attribute of the target is determined based on the first predicted identity category, the first location information, the second predicted identity category, and the second location information.
5. The method according to claim 1, characterized in that The method further includes: obtaining a second confidence level of the first predicted identity category; and identifying a second predicted identity category of the target based on the voice information and obtaining a third confidence level of the second predicted identity category; The determining the identity attribute of the target based at least on the first predicted identity category, the first location information, and the second location information includes: An identity attribute of the target is determined based on the first predicted identity category, the first location information, the second confidence level, the second predicted identity category, the third confidence level, and the second location information.
6. The method according to claim 5, characterized in that Acquiring image information of the target includes acquiring multiple frames of image information of the target; The determining the identity attribute of the target based on the first predicted identity category, the first location information, the second confidence level, the second predicted identity category, the third confidence level, and the second location information includes: Comparing the first position information and the second position information of each frame of image information in the multiple frames of image information to obtain a second position comparison result of the multiple frames of image information; Comparing the second confidence level of each frame of image information in the multiple frames of image information with the second preset threshold range; obtaining a second confidence level according to the comparison result of the multiple frames of image information and the second preset threshold range; Comparing the third confidence level with a third preset threshold range; and obtaining a third confidence level according to the comparison result; An identity attribute of the target is determined based on the first predicted identity category, the second predicted identity category, the second position comparison result, the second confidence level, and the third confidence level.
7. The method according to claim 4, characterized in that The step of identifying the second predicted identity category of the target based on the voice information comprises: The speech information is recognized using a pre-trained neural network model to obtain a second predicted identity category of the target; the second predicted identity category includes a child and / or an adult.
8. The method according to any one of claims 1 to 7, characterized in that The identifying a first predicted identity category of the target based on the image information comprises: The image information is identified using a pre-trained target identity attribute recognition model to obtain a first predicted identity category of the target, where the first predicted identity category includes a child and / or an adult.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the identity attribute recognition method according to any one of claims 1 to 8 is implemented.
10. A smart device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores a computer program, and when the computer program is executed by the at least one processor, the identity attribute recognition method described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Intelligent gender and emotion recognition detection system and method based on vision and voice
CN105700682A
Face recognition enhancement method and device, electronic equipment and storage medium
CN113963396A
Driver identity verification method and device, computer equipment and storage medium
CN114241459A
Character type determination method and device, electronic equipment and storage medium
CN115862597A
Identity attribute recognition method, computer readable storage medium and intelligent equipment
CN117765623A
Cited By
Pet behavior training system based on AI vision
CN120982435A