Face recognition method and device, computer device, storage medium and program product

By acquiring facial images and collecting features of people inside the vehicle, and combining confidence scores and cabin sensor data, weighted processing and nonlinear mapping techniques were used to solve the problem of decreased facial recognition accuracy in complex in-vehicle environments, achieving higher recognition accuracy.

CN122435658APending Publication Date: 2026-07-21CHONGQING LANDIAN AUTOMOBILE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING LANDIAN AUTOMOBILE TECHNOLOGY CO LTD
Filing Date
2026-06-16
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, the complex and ever-changing cabin environment inside a vehicle leads to a decrease in the accuracy of facial recognition, resulting in misjudgments and recognition errors.

Method used

By acquiring facial images and collecting features of people inside the vehicle, the predicted probabilities of different candidate attributes under each preset attribute type are determined. Combined with confidence scores and cabin sensor data, weighted processing and nonlinear mapping techniques are used to improve the accuracy of facial recognition.

Benefits of technology

In complex environments, accurately identifying personnel attributes improves the accuracy of facial recognition and reduces misjudgments and recognition errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122435658A_ABST
    Figure CN122435658A_ABST
Patent Text Reader

Abstract

The application relates to a face recognition method and device, computer equipment, a storage medium and a program product, and relates to the technical field of image data processing. The method comprises the following steps: acquiring a face image of a person in a vehicle and a collection feature corresponding to a face region in the face image; determining a personnel attribute of the person as a prediction probability of different candidate attributes under each preset attribute type according to the face image and the collection feature; for each preset attribute type, determining a confidence degree of the face image under the preset attribute type according to the prediction probability of the different candidate attributes under the preset attribute type; and determining a target attribute of the person under each preset attribute type according to the confidence degree under each preset attribute type and the prediction probability of the different candidate attributes. The method can accurately determine the target attribute of the person under each preset attribute type, and improves the accuracy of face recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image data processing technology, and in particular to a face recognition method, apparatus, computer equipment, storage medium, and program product. Background Technology

[0002] With the development of vehicles, recognizing occupants based on their facial images has become an important technology for ensuring vehicle safety and personalized interaction.

[0003] In related technologies, computer vision technology is commonly used to recognize facial images of people inside vehicles. However, the complex and ever-changing cabin environment inside a vehicle can easily lead to a decrease in facial recognition accuracy, such as misidentifying children as adults, identifying elderly passengers with excessively inaccurate ages, and incorrectly identifying gender attributes.

[0004] Therefore, improving the accuracy of facial recognition is a problem that deserves attention. Summary of the Invention

[0005] Based on this, this application addresses the aforementioned technical problems by providing a face recognition method, apparatus, computer equipment, storage medium, and program product that can improve the accuracy of face recognition.

[0006] Firstly, this application provides a face recognition method, including:

[0007] Acquire facial images of people inside the vehicle and the features corresponding to the facial regions in the facial images;

[0008] Based on facial images and collected features, the personnel attributes of individuals are determined as the predicted probabilities of different candidate attributes under each preset attribute type;

[0009] For each preset attribute type, the confidence level of the face image under the preset attribute type is determined based on the predicted probability of different candidate attributes under the preset attribute type.

[0010] Based on the confidence level of each preset attribute type and the predicted probability of different candidate attributes, the target attribute of the person under each preset attribute type is determined.

[0011] By employing the aforementioned face recognition method, when determining the predicted probabilities of different candidate attributes under various preset attribute types for a person's attributes, the influence of the collected features (such as face occlusion rate and lighting suitability) on face recognition can be analyzed by introducing the collected features corresponding to the face region. This allows for more accurate face recognition in the complex and ever-changing environment inside a vehicle cabin, and more accurate prediction probabilities of the person's attributes. Furthermore, the confidence level of the face image under each preset attribute type reflects the reliability of the predicted probabilities of different candidate attributes under each preset attribute type. By combining the confidence level under each preset attribute type and the predicted probabilities of different candidate attributes, the target attribute of the person under each preset attribute type can be determined more accurately, thereby improving the accuracy of face recognition.

[0012] In one embodiment, there are multiple preset attribute types. The target attribute of a person under each preset attribute type is determined based on the confidence level of that preset attribute type and the predicted probability of different candidate attributes. This includes: when the confidence levels of different preset attribute types meet preset conditions, determining the target attribute of a face image under each preset attribute type based on the confidence level of that preset attribute type and the predicted probability of different candidate attributes; when the confidence levels of different preset attribute types do not meet preset conditions, acquiring cockpit sensor data corresponding to the person's location; and determining the target attribute of the person under each preset attribute type based on the cockpit sensor data and the predicted probability of different candidate attributes under each preset attribute type.

[0013] In this embodiment, when the confidence levels of different preset attribute types meet the preset conditions, the reliability of the predicted probabilities of different candidate attributes under each preset attribute type for personnel attributes is relatively high. Therefore, the target attribute of the face image under each preset attribute type can be determined directly based on the visual analysis results of the face image (confidence level and predicted probabilities of different candidate attributes under each preset attribute type). When the confidence levels of different preset attribute types do not meet the preset conditions, the reliability of the predicted probabilities of different candidate attributes under each preset attribute type for personnel attributes is relatively low. Therefore, the target attribute of the face image under each preset attribute type is determined by combining cockpit sensing data and visual analysis results of the face image. The above-mentioned scheme of using different target attribute determination methods based on the matching of confidence levels and preset conditions of different preset attribute types can achieve relatively accurate target attributes even with different confidence levels, thereby improving the accuracy of face recognition.

[0014] In one embodiment, determining that the confidence levels of different preset attribute types meet preset conditions includes: determining that the confidence levels of different preset attribute types meet preset conditions when the confidence levels of all preset attribute types are not less than the corresponding confidence thresholds.

[0015] In this embodiment, the confidence level of different preset attribute types is not less than the corresponding confidence threshold as a preset condition for the confidence level of different preset attribute types to meet the preset condition. This can realize the classification of confidence level, thereby providing a basis for the subsequent adoption of the corresponding target attribute determination method.

[0016] In one embodiment, the target attribute of the person under each preset attribute type is determined based on cockpit sensing data and the predicted probability of different candidate attributes under each preset attribute type, including: determining the prior probability of different candidate attributes under each preset attribute type based on cockpit sensing data; and determining the target attribute of the person under each preset attribute type based on the predicted probability and prior probability of different candidate attributes under each preset attribute type.

[0017] In this embodiment, by determining the prior probabilities of different candidate attributes under each preset attribute type based on cockpit sensing data, the predicted probabilities of different candidate attributes under the corresponding preset attribute type can be corrected by combining the prior probabilities of different candidate attributes under each preset attribute type, thereby obtaining a more accurate target attribute.

[0018] In one embodiment, the target attribute of a person under each preset attribute type is determined based on the predicted probability and prior probability of different candidate attributes under each preset attribute type, including: for each candidate attribute, weighting the predicted probability and prior probability of the candidate attribute to obtain the target probability of the candidate attribute; and for each preset attribute type, determining the target attribute of a person under each preset attribute type based on the target probabilities of different candidate attributes under the preset attribute type.

[0019] In this embodiment, by weighting the predicted probability and prior probability of the candidate attribute, the target probability of the candidate attribute can be obtained more accurately, thus obtaining a more accurate target attribute.

[0020] In one embodiment, based on the face image and the acquired features, determining the personnel attributes of a person as the predicted probabilities of different candidate attributes under each preset attribute type includes: extracting basic face features of the face region in the face image; and performing nonlinear mapping processing on the acquired features to obtain acquisition weight features; wherein, the acquisition weight features are used to characterize the weight of different acquisition features; performing weighted processing on the basic face features and the acquisition weight features to obtain the target face features of the face region in the face image; and determining the personnel attributes of a person as the predicted probabilities of different candidate attributes under each preset attribute type based on the target face features.

[0021] In this embodiment, by weighting the basic facial features and the collected weight features, the target facial features are obtained. The impact of the collected features (such as the face occlusion rate, lighting suitability, etc.) on facial recognition can be analyzed. Thus, in the complex and ever-changing environment inside the vehicle cabin, the face can be recognized more accurately, and the predicted probability of the personnel attributes can be obtained more accurately.

[0022] In one embodiment, each preset attribute type includes a first preset attribute type and a second preset attribute type, wherein the number of different candidate attributes under the first preset attribute type is greater than the number of different candidate attributes under the second preset attribute type; for each preset attribute type, the confidence level of the face image under the preset attribute type is determined based on the predicted probabilities of different candidate attributes under the preset attribute type, including: for each first preset attribute type, determining the maximum probability value and the second maximum probability value among the predicted probabilities of different candidate attributes under the first preset attribute type, and determining the confidence level of the face image under the preset attribute type based on the difference between the determined maximum probability value and the second maximum probability value; for each second preset attribute type, determining the maximum probability value among the predicted probabilities of different candidate attributes under the second preset attribute type, and determining the confidence level of the face image under the preset attribute type based on the determined maximum probability value.

[0023] In this embodiment, for preset attribute types with different numbers of candidate attributes, the confidence level of a face image under the preset attribute type is determined in different ways, which can obtain a more accurate confidence level, thereby laying the foundation for subsequent determination of the target attributes of a person under each preset attribute type.

[0024] In one embodiment, determining the predicted probabilities of a person's attributes as different candidate attributes under each preset attribute type based on the face image and the collected features includes: inputting the face image and the collected features into a face recognition model to obtain the predicted probabilities of a person's attributes as different candidate attributes under each preset attribute type.

[0025] In this embodiment, by inputting the face image and collected features into the face recognition model, the face recognition model can be used to analyze and process the face image and collected features, thereby quickly and accurately obtaining the predicted probability of different candidate attributes under each preset attribute type.

[0026] In one embodiment, the face recognition model is trained through the following steps: acquiring sample face images, sample collection features of the face region in the sample face images, and the real attributes of the sample person corresponding to the sample face image under each preset attribute type; inputting the sample face images and sample collection features into an initial model to obtain the sample prediction probabilities of the person attributes of the sample person as different candidate attributes under each preset attribute type; and adjusting the parameters of the initial model according to the real attributes corresponding to each preset attribute type and the sample prediction probabilities of different candidate attributes under each preset attribute type to obtain the face recognition model.

[0027] In this embodiment, during the model training phase, the initial model is tuned based on the true attributes corresponding to each preset attribute type and the sample prediction probabilities of different candidate attributes under each preset attribute type. This allows the trained face recognition model to output more accurate prediction probabilities, laying the foundation for more accurate determination of the target attributes of a person under each preset attribute type in the future.

[0028] In one embodiment, the initial model is tuned based on the true attribute corresponding to each preset attribute type and the sample prediction probabilities of different candidate attributes under the preset attribute type. This includes: for each preset attribute type, determining a first loss value corresponding to the preset attribute type based on the true attribute corresponding to the preset attribute type and the sample prediction probabilities of different candidate attributes; determining a second loss value based on the sample prediction probabilities of different candidate attributes under each preset attribute type; determining a target loss value based on the first loss value and the second loss value; and tuning the initial model if the target loss value does not meet the preset convergence condition.

[0029] In this embodiment, the initial model is tuned based on the first loss corresponding to each preset attribute type and the second loss value determined based on the sample prediction probability of different candidate attributes under each preset attribute type. This enables the trained face recognition model to predict the probability of different candidate attributes under different preset attribute types, thus realizing the multi-task prediction function and laying the foundation for more accurate determination of the target attributes of a person under each preset attribute type.

[0030] Secondly, this application also provides a face recognition device, comprising:

[0031] The acquisition module is used to acquire facial images of people inside the vehicle and the acquisition features corresponding to the facial regions in the facial images;

[0032] The first determining module is used to determine the predicted probability of different candidate attributes under each preset attribute type based on the face image and the collected features.

[0033] The second determining module is used to determine the confidence level of a face image under a preset attribute type based on the predicted probability of different candidate attributes under the preset attribute type for each preset attribute type.

[0034] The third determination module is used to determine the target attribute of a person under each preset attribute type based on the confidence level under each preset attribute type and the predicted probability of different candidate attributes.

[0035] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described above.

[0036] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.

[0037] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in any of the above aspects.

[0038] Regarding the beneficial effects of any of the technical solutions in the second to fifth aspects mentioned above, refer to the beneficial effects of the corresponding technical solutions in the first aspect; repeated examples will not be listed here. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 This is a schematic diagram illustrating an optional application environment for a model training method in one embodiment.

[0041] Figure 2 This is a schematic diagram of an optional process for a face recognition method in one embodiment;

[0042] Figure 3 This is a schematic diagram of an optional process for determining the predicted probability in one embodiment;

[0043] Figure 4 This is a schematic diagram of an optional training process for a face recognition model in one embodiment;

[0044] Figure 5This is a schematic diagram of an optional flow of a face recognition method in another embodiment;

[0045] Figure 6 This is a schematic diagram of an optional structure of a face recognition device in one embodiment;

[0046] Figure 7 This is a schematic diagram of an optional internal structure of a computer device in one embodiment. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application.

[0048] The terms "first," "second," etc., used in this application may be used to describe various elements, but these elements are not limited by these terms. These terms are used only to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0049] The model training method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. Terminal 102 can collect facial images of people inside the vehicle and send them to server 104. Server 104 can receive the facial images of people inside the vehicle, obtain the collected features corresponding to the facial regions in the facial images; based on the facial images and collected features, determine the predicted probabilities of different candidate attributes under each preset attribute type for the person's attributes; for each preset attribute type, determine the confidence level of the facial image under the preset attribute type based on the predicted probabilities of different candidate attributes under the preset attribute type; and determine the target attribute of the person under each preset attribute type based on the confidence level and the predicted probabilities of different candidate attributes. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and vehicle terminals. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0050] The user information (including but not limited to user device information, user personal information (facial images, age, gender) and data (including but not limited to data used for analysis, stored data, and displayed data) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0051] In one exemplary embodiment, such as Figure 2 As shown, a face recognition method is provided, including the following steps:

[0052] S210, acquire the facial images of the people inside the vehicle and the features corresponding to the facial regions in the facial images.

[0053] The persons inside the vehicle may include at least one person inside the vehicle.

[0054] The face region can be understood as the area that includes the facial features.

[0055] Among them, the collection features can be understood as the features of the collection conditions corresponding to the face region.

[0056] Optionally, the features collected may include at least one of the following: face scale, face occlusion rate, lighting suitability, seating position, and head posture features.

[0057] Face scale can be understood as the size of the face region in a face image. The greater the imaging distance of the face region, the smaller the face size. For example, when the imaging distance of the face region is 0.5 meters to 1 meter, the face size is approximately 128×128 pixels; when the imaging distance of the face region is 1.5 meters to 2.5 meters, the face size is approximately 64×64 pixels.

[0058] Face occlusion rate can be understood as the ratio of the occluded area to the total face area. The occluded area can include at least one area such as the area covered by a seatbelt across the face, the area covered by a mask, the area covered by sunglasses, the area covered by a child seat, or the area covered by a hand.

[0059] Lighting suitability can be understood as the light quality of the face area, used to characterize the uniformity of light in the facial features area, whether the brightness is appropriate, whether there is backlight from the sunroof, whether there is strong windshield light, whether there is dim light at night, whether there are sudden changes in light and dark in the tunnel, whether there is coloration of the cabin ambient light, and whether there is infrared supplementary light, etc.

[0060] Seat location can be understood as the position of the seat where the person corresponding to the facial region is located. For example, seat location can be represented based on seat location encoding.

[0061] Head posture features can be understood as the posture features of a person's head corresponding to the face region. For example, posture features may include at least one of head pitch angle and head yaw angle.

[0062] In some embodiments, facial images of occupants inside the vehicle can be captured by an image acquisition device within the vehicle cabin. After capturing the facial images, the image acquisition device can send them to a controller on the vehicle side. For example, the image acquisition device may include an Occupant Monitoring System (OMS) camera. The OMS camera can be a red-green-blue (RGB) camera or an infrared binocular camera. Optionally, the OMS camera can be mounted above the rearview mirror, covering the driver's seat, front passenger seat, and the entire rear cabin area.

[0063] In some embodiments, a facial region image can be cropped based on facial key points in a facial image captured by an image acquisition device, and the facial region image can be scaled to a preset size.

[0064] In some embodiments, the face scale can be obtained based on the pixel data of the face region in the face region image.

[0065] In some embodiments, the area of ​​the occluded region can be determined based on the face region image, and the face occlusion rate can be determined based on the area of ​​the occluded region and the total area of ​​the face region.

[0066] In some embodiments, the face region image can be converted into a grayscale image, and the lighting suitability can be obtained based on the pixel data of the grayscale image.

[0067] In some embodiments, the coordinate positions of facial key points in the facial region image and the corresponding depth information of the facial region can be determined based on the facial region image. The seat position can be determined based on the coordinate positions of facial key points in the facial region image, the corresponding depth information of the facial region, and the pressure sensing data corresponding to the seat.

[0068] In some embodiments, head pose features can be determined based on key point coordinate data of a face region image and standard 3D face key point coordinate data.

[0069] S220, Based on the face image and collected features, determine the personnel attributes of the person as the predicted probabilities of different candidate attributes under each preset attribute type.

[0070] Optionally, the preset attribute type may include at least one of the age attribute type and the gender attribute type. Of course, the preset attribute type may also include other attribute types besides the age attribute type and the gender attribute type; there are no specific limitations on this.

[0071] Among them, different candidate attributes under the age attribute type can include different candidate age ranges, such as 0-12 years old, 13-17 years old, 18-40 years old, 41-60 years old, and 60 years old and above. Different candidate attributes under the gender attribute type can include different candidate gender attributes, specifically including female and male.

[0072] Among them, personnel attributes can be understood as the attributes under each preset attribute type corresponding to a person, such as gender and age range.

[0073] For example, when the preset attribute types include gender attribute type and age attribute type, and the different candidate attributes under the gender attribute type include female and male, and the different candidate attributes under the age attribute type include 0-12 years old, 13-17 years old, 18-40 years old, 41-60 years old and over 60 years old, the predicted probability of a person's gender can be determined as female and male respectively, and the predicted probability of a person's age range can be determined as 0-12 years old, 13-17 years old, 18-40 years old, 41-60 years old and over 60 years old respectively, based on the facial image and collected features.

[0074] S230, for each preset attribute type, determine the confidence level of the face image under the preset attribute type based on the predicted probability of different candidate attributes under the preset attribute type.

[0075] For example, for the age attribute type, the confidence level of a face image under the age attribute type can be determined based on the predicted probability of different age intervals under the age attribute type.

[0076] For gender attribute categories, the confidence level of a face image under a gender attribute category can be determined based on the predicted probability of different genders under that category.

[0077] In some embodiments, the confidence level of a face image under a gender attribute category can be assigned based on the difference between the predicted probabilities of different genders under the gender attribute category.

[0078] If the difference between the predicted probabilities of different genders under the gender attribute category is greater than the difference threshold, the preset first confidence level can be used as the confidence level of the face image under the gender attribute category.

[0079] For example, if the predicted probability of a person being female is 80% and the predicted probability of a person being male is 20%, and the difference (60%) between the predicted probability of a person being female and the predicted probability of a person being male is greater than 50%, then the confidence level of the face image under the gender attribute category can be determined to be 90%.

[0080] If the difference between the predicted probabilities of different genders under the gender attribute category is no greater than the difference threshold, a preset second confidence level can be used as the confidence level of the face image under the gender attribute category. The second confidence level is lower than the first confidence level.

[0081] For example, if the predicted probability of a person being female is 60% and the predicted probability of a person being male is 40%, and the difference (20%) between the predicted probability of a person being female and the predicted probability of a person being male is less than 50%, then the confidence level of the face image under the gender attribute category can be determined to be 50%.

[0082] S240, based on the confidence level under each preset attribute type and the predicted probability of different candidate attributes, determine the target attribute of the person under each preset attribute type.

[0083] Based on the confidence level of the age attribute type and the predicted probability of different age intervals, the target age range for the individual can be determined. Based on the confidence level of the gender attribute type and the predicted probability of each gender, the target gender for the individual can be determined.

[0084] In some embodiments, for each preset attribute type, if the confidence level under the preset attribute type is greater than the confidence threshold, the candidate attribute with the highest predicted probability can be used as the target attribute of the person under the preset attribute type; if the confidence level under the preset attribute type is not greater than the confidence threshold, the face image of the person in the same vehicle and the collected features corresponding to the face region in the face image can be reacquired.

[0085] By employing the aforementioned face recognition method, when determining the predicted probabilities of different candidate attributes under various preset attribute types for a person's attributes, the influence of the collected features (such as face occlusion rate and lighting suitability) on face recognition can be analyzed by introducing the collected features corresponding to the face region. This allows for more accurate face recognition in the complex and ever-changing environment inside a vehicle cabin, and more accurate prediction probabilities of the person's attributes. Furthermore, the confidence level of the face image under each preset attribute type reflects the reliability of the predicted probabilities of different candidate attributes under each preset attribute type. By combining the confidence level under each preset attribute type and the predicted probabilities of different candidate attributes, the target attribute of the person under each preset attribute type can be determined more accurately, thereby improving the accuracy of face recognition.

[0086] Based on the technical solutions of the above embodiments, this application also provides another optional embodiment. In this optional embodiment, the number of preset attribute types is refined into multiple types, and the step of determining the target attribute in S240 is refined. Specifically, the step of determining the target attribute includes:

[0087] 1) When the confidence level of different preset attribute types meets the preset conditions, the target attribute of the face image under each preset attribute type is determined based on the confidence level under each preset attribute type and the prediction probability of different candidate attributes.

[0088] In some embodiments, determining that the confidence levels of different preset attribute types meet preset conditions includes: determining that the confidence levels of different preset attribute types meet preset conditions when the confidence levels of all preset attribute types are not less than the corresponding confidence thresholds.

[0089] For example, if the confidence level of the age attribute type is not less than the age confidence level threshold and the confidence level of the gender attribute type is not less than the gender confidence level threshold, the target attribute of the face image under each preset attribute type can be determined based on the confidence level under each preset attribute type and the prediction probability of different candidate attributes.

[0090] In other embodiments, the confidence level of different preset attribute types meeting the preset conditions may include: determining that the confidence level of different preset attribute types meets the preset conditions when the confidence level of at least one preset attribute type is not less than the corresponding confidence level threshold.

[0091] For example, if the confidence level of the age attribute type is not less than the age confidence level threshold, and / or the confidence level of the gender attribute type is not less than the gender confidence level threshold, the target attribute of the face image under each preset attribute type can be determined based on the confidence level under each preset attribute type and the prediction probability of different candidate attributes.

[0092] 2) When the confidence level of different preset attribute types does not meet the preset conditions, obtain the cockpit sensor data corresponding to the location of the person, and determine the target attribute of the person under each preset attribute type based on the cockpit sensor data and the predicted probability of different candidate attributes under each preset attribute type.

[0093] In some embodiments, the confidence level of different preset attribute types does not meet the preset conditions, including: the confidence level of at least one preset attribute type is less than the corresponding confidence threshold.

[0094] For example, if the confidence level of the age attribute type is less than the age confidence level threshold, and / or the confidence level of the gender attribute type is less than the gender confidence level threshold, cockpit sensor data corresponding to the location of the person can be obtained.

[0095] The cabin sensing data can be understood as the sensing data output by sensors in the cabin. For example, the cabin sensing data may include at least one of the following: seat pressure data, seat position data, seat belt status, and body domain sensing data. Seat pressure data can be understood as the pressure data output by the pressure sensor corresponding to the seat. Seat position data can be understood as at least one of the following: fore-and-aft displacement data output by the seat rail displacement sensor, seat height data output by the height sensor, and seat tilt angle data output by the tilt sensor. Seat belt status may include either the seat belt fastened status or the unfastened status obtained based on the Controller Area Network (CAN) bus. Body domain sensing data can be understood as the sensor data controlled by the body domain controller.

[0096] In some embodiments, cockpit sensing data output by sensors in the cockpit can be acquired when the confidence levels of different preset attribute types do not meet preset conditions.

[0097] Generally, cockpit sensor data differs depending on the individual's attributes. For example, the seat pressure data for a 10-year-old is less than that for a 30-year-old. In this embodiment, by using cockpit sensor data and the predicted probabilities of different candidate attributes under each preset attribute type, the target attribute of the individual under each preset attribute type can be obtained more accurately.

[0098] In some embodiments, the prior probabilities of different candidate attributes under each preset attribute type can be determined based on cockpit sensor data; and the target attribute of the person under each preset attribute type can be determined based on the predicted probability and prior probability of different candidate attributes under each preset attribute type.

[0099] In some embodiments, determining the prior probabilities of different candidate attributes under each preset attribute type based on cockpit sensing data may include: determining the prior probabilities of different candidate attributes under each preset attribute type based on the prior probabilities of different candidate attributes corresponding to different cockpit sensing data, and the cockpit sensing data corresponding to the current location of the vehicle occupants.

[0100] Among them, the prior probabilities of different candidate attributes under each preset attribute type corresponding to different cockpit sensor data can be obtained by establishing a correspondence between the cockpit sensor data of the reference personnel and the prior probabilities of different candidate attributes under each preset attribute type after obtaining the cockpit sensor data of the reference personnel and the prior probabilities of different candidate attributes of the reference personnel under each preset attribute type.

[0101] In some embodiments, based on the cockpit sensor data corresponding to the current location of the vehicle occupants, different candidate attributes under each preset attribute type that match the cockpit sensor data corresponding to the current location of the vehicle occupants can be queried from the prior probabilities of different candidate attributes under each preset attribute type corresponding to different cockpit sensor data.

[0102] In some embodiments, determining the target attribute of a person under each preset attribute type based on the predicted probability and prior probability of different candidate attributes under each preset attribute type may include: for each candidate attribute, obtaining the target probability of the candidate attribute based on the predicted probability and prior probability of the candidate attribute; and for each preset attribute type, determining the target attribute of a person under each preset attribute type based on the target probabilities of different candidate attributes under the preset attribute type.

[0103] In some embodiments, for each candidate attribute, the maximum value between the predicted probability and the prior probability of the candidate attribute can be used as the target probability of the candidate attribute. For example, if the predicted probability of the person in the car being male is 80% and the prior probability of the person in the car being male is 60%, the target probability of the person in the car being male can be determined to be 80%.

[0104] In some embodiments, for each candidate attribute, the predicted probability and prior probability of the candidate attribute can be weighted to obtain the target probability of the candidate attribute. For example, for each candidate gender attribute, the predicted probability and prior probability of the candidate gender attribute can be weighted and summed to obtain the target probability of the candidate gender attribute. For example, the weight coefficient corresponding to the predicted probability can be 0.5, and the weight coefficient corresponding to the prior probability can be 0.5. For instance, the predicted probability of the occupant being male is 80%, and the predicted probability of the occupant being female is 20%. The prior probability of the occupant being male is 60%, and the prior probability of the occupant being female is 40%. The target probability of the occupant being male can be obtained using the following formula: 80%×0.5+60%×0.5=70%, and the target probability of the occupant being female can be obtained using the following formula: 20%×0.5+40%×0.5=30%.

[0105] In some embodiments, for each preset attribute type, the candidate attribute with the highest target probability can be used as the target attribute for the person under each preset attribute type. In the example above, the target probability of the person in the car being male is 70%, and the target probability of the person in the car being female is 30%, thus it can be determined that the person in the car is male.

[0106] In this embodiment, when the confidence levels of different preset attribute types meet the preset conditions, the reliability of the predicted probabilities of different candidate attributes under each preset attribute type for personnel attributes is relatively high. Therefore, the target attribute of the face image under each preset attribute type can be determined directly based on the visual analysis results of the face image (confidence level and predicted probabilities of different candidate attributes under each preset attribute type). When the confidence levels of different preset attribute types do not meet the preset conditions, the reliability of the predicted probabilities of different candidate attributes under each preset attribute type for personnel attributes is relatively low. Therefore, the target attribute of the face image under each preset attribute type is determined by combining cockpit sensing data and visual analysis results of the face image. The above-mentioned scheme of using different target attribute determination methods based on the matching of confidence levels and preset conditions of different preset attribute types can achieve relatively accurate target attributes even with different confidence levels, thereby improving the accuracy of face recognition.

[0107] Based on the technical solutions of the above embodiments, this application also provides another optional embodiment, in which the step of determining the predicted probability in S220 is refined.

[0108] See Figure 3 The steps for determining the predicted probability, as shown, include:

[0109] S310, extract the basic facial features of the face region in the face image.

[0110] Among them, basic facial features can be understood as the basic feature vectors obtained by extracting features from the facial regions of a facial image.

[0111] S320 performs nonlinear mapping processing on the collected features to obtain the collected weight features.

[0112] Among them, the acquisition weight feature is used to characterize the weight of different acquisition features.

[0113] In some embodiments, the acquired features can be normalized features. The original acquired features of the obtained face region image can be normalized to obtain normalized features. For example, face scale, face occlusion rate, and lighting suitability can be normalized separately. The normalized face scale ranges from 0 to 1, with smaller face scales corresponding to rear-seat passengers in the cabin; the normalized face occlusion rate ranges from 0 to 1, with a higher occlusion rate indicating more severe occlusion; the normalized lighting suitability ranges from 0 to 1, with a higher lighting suitability indicating more suitable lighting.

[0114] In some embodiments, each seat position may be encoded, and the seat position may include at least one of the following: driver's seat, front passenger seat, rear left, rear middle, and rear right.

[0115] In some embodiments, two fully connected layers can be used to perform a linear weighted transformation on the collected features to map the collected features to features of a new dimension. Then, the Sigmoid activation function is used to perform a non-linear mapping on the features of the new dimension to obtain the probability of each collected feature, which is represented by 0-1.

[0116] In some embodiments, when the collected features include normalized face scale, normalized face occlusion rate, normalized illumination suitability, and seat position encoding, the collected features can be nonlinearly mapped according to the following formula:

[0117]

[0118] in, Indicates the collected weight features; denoted as Sigmoid activation function; MLP represents two fully connected layers; s represents the normalized face scale; o represents the normalized face occlusion rate; l represents the normalized lighting suitability; p represents the seat position encoding.

[0119] S330 performs weighted processing on the basic facial features and the collected weight features to obtain the target facial features of the facial region in the facial image.

[0120] The target facial features can be understood as the facial feature vector obtained after weighting the collected features.

[0121] In some embodiments, the basic facial features can be multiplied element-wise using the collected weight features to obtain the target facial features of the facial region in the facial image.

[0122] In some embodiments, the target facial features can be obtained according to the following formula:

[0123]

[0124] in, Indicates the target facial features; Based on basic facial features, This indicates an element-wise multiplication operation. This indicates the weighted features collected.

[0125] S340, based on the target facial features, determine the personnel attributes of the person as the predicted probabilities of different candidate attributes under each preset attribute type.

[0126] In some embodiments, the probabilities of different candidate attributes under each preset attribute type corresponding to different preset facial features can be obtained. Based on the target facial features, the predicted probabilities of different candidate attributes under each preset attribute type corresponding to the target facial features can be queried from the probabilities of different candidate attributes under each preset attribute type corresponding to different preset facial features.

[0127] In the above embodiments, the target face features are integrated with the collection weight features corresponding to the basic face features and the collection features. Based on the target face features, the predicted probability of different candidate attributes under each preset attribute type can be determined more accurately.

[0128] Based on the technical solutions of the above embodiments, this application also provides another optional embodiment. In this optional embodiment, each preset attribute type is refined into a first preset attribute type and a second preset attribute type. The number of different candidate attributes under the first preset attribute type is greater than the number of different candidate attributes under the second preset attribute type. Correspondingly, the confidence determination step in S230 is refined. Specifically, the confidence determination step includes:

[0129] 1) For each first preset attribute type, determine the maximum probability value and the second maximum probability value among the predicted probabilities of different candidate attributes under the first preset attribute type, and determine the confidence level of the face image under the preset attribute type based on the difference between the determined maximum probability value and the second maximum probability value.

[0130] In some embodiments, the first preset attribute type can be understood as the age attribute type, and the second preset attribute type can be understood as the gender attribute type.

[0131] For example, for the age attribute type, the maximum and second-highest probability values ​​among the predicted probabilities of different candidate age ranges can be determined. The confidence level of the face image under the age attribute type is then determined based on the difference between the determined maximum and second-highest probability values. For instance, the predicted probability for a person in a car aged 0-12 is 0%, for 13-17 is 10%, for 18-40 is 70%, for 41-60 is 20%, and for those over 60 is 0%. The difference between the predicted probability for 18-40 and 41-60, i.e., 50%, can be used as the confidence level of the face image under the age attribute type.

[0132] In some embodiments, the confidence level of a face image under the age attribute type can be determined according to the following formula:

[0133] Conf age =P top1 -P top2

[0134] Among them, Conf age P represents the confidence level of a face image under the age attribute type; top1 P represents the maximum predicted probability among different candidate age ranges. top2 This represents the second-highest probability value among the predicted probabilities of different candidate age ranges.

[0135] 2) For each second preset attribute type, determine the maximum probability value among the predicted probabilities of different candidate attributes under the second preset attribute type, and determine the confidence level of the face image under the preset attribute type based on the determined maximum probability value.

[0136] For example, for the gender attribute type, the maximum probability value among the predicted probabilities of different candidate gender attributes under the gender attribute type can be used as the confidence level of the face image under the gender attribute type.

[0137] In some embodiments, the confidence level of a face image under a preset attribute type can be determined according to the following formula:

[0138] Conf gender =max(P female ,P male )

[0139] Among them, Conf gender P represents the confidence level of a face image under the gender attribute type; female P represents the predicted probability of a female; male This represents the predicted probability of males; max(·) indicates taking the maximum value.

[0140] In this embodiment, for preset attribute types with different numbers of candidate attributes, the confidence level of a face image under the preset attribute type is determined in different ways, which can obtain a more accurate confidence level, thereby laying the foundation for subsequent determination of the target attributes of a person under each preset attribute type.

[0141] Based on the technical solutions of the above embodiments, this application also provides another optional embodiment, in which the step of determining the prediction probability in S220 is refined. Specifically, the step of determining the prediction probability includes: inputting the face image and collected features into the face recognition model to obtain the prediction probability of the person's attributes as different candidate attributes under each preset attribute type.

[0142] In some embodiments, the face recognition model can be deployed on the vehicle's controller. After the face image and collected features are input into the face recognition model, the face recognition model can output the predicted probability of different candidate attributes under each preset attribute type through one forward inference.

[0143] For example, a face recognition model can output the predicted probability corresponding to different gender attributes and the predicted probability corresponding to different age ranges.

[0144] In this embodiment, by inputting the face image and collected features into the face recognition model, the face recognition model can be used to analyze and process the face image and collected features, thereby quickly and accurately obtaining the predicted probability of different candidate attributes under each preset attribute type.

[0145] In some embodiments, such as Figure 4 As shown, a face recognition model can be trained through the following steps:

[0146] S410, acquire sample face images, sample collection features of face regions in sample face images, and the real attributes of sample personnel corresponding to sample face images under each preset attribute type.

[0147] In some embodiments, sample facial images of sample occupants inside the sample vehicle can be acquired using image acquisition devices within the vehicle cabin. The process for acquiring sample facial images is similar to the aforementioned process for acquiring facial images, and will not be repeated here.

[0148] In some embodiments, sample face images can be obtained from a database. These sample face images in the database are collected and stored in accordance with relevant regulations.

[0149] In some embodiments, the process of acquiring sample collection features is similar to the process of acquiring collection features described above, and will not be repeated here.

[0150] In some embodiments, the actual attributes of the sample face image corresponding to the sample person under each preset attribute type may be pre-annotated.

[0151] S420: Input the sample face image and sample collection features into the initial model to obtain the sample prediction probability of different candidate attributes under each preset attribute type for the personnel attributes of the sample personnel.

[0152] The initial model can understand the face recognition model to be trained.

[0153] In some embodiments, the initial model can be deployed on a server. After the sample face image and sample collection features are input into the initial model, the initial model can output the sample prediction probability of the person's attributes as different candidate attributes under each preset attribute type through one forward inference.

[0154] S430: Based on the true attributes corresponding to each preset attribute type and the sample prediction probabilities of different candidate attributes under the preset attribute type, the parameters of the initial model are adjusted to obtain the face recognition model.

[0155] In some embodiments, for each preset attribute type, a first loss value corresponding to the preset attribute type is determined based on the true attribute corresponding to the preset attribute type and the sample prediction probabilities of different candidate attributes; a second loss value is determined based on the sample prediction probabilities of different candidate attributes under each preset attribute type; a target loss value is determined based on the first loss value and the second loss value; and if the target loss value does not meet the preset convergence condition, the initial model is tuned.

[0156] The preset convergence conditions may include at least one of the following: the change in the target loss value between two adjacent iterations is less than a threshold, the gradient approaches 0, and the preset number of iterations is reached.

[0157] In some embodiments, for each sample face image, the first sub-loss value corresponding to the sample face image can be determined based on the probability label value of the true gender attribute of the sample face image, the sample prediction probability of different candidate gender attributes, and the sample collection weight features corresponding to the sample face image. Then, the first sub-loss values ​​corresponding to each sample face image are weighted and summed to obtain the weighted summation result. Finally, the first loss value corresponding to the gender attribute type is determined based on the average value of the weighted summation result.

[0158] In some embodiments, for a gender attribute type, the probability label value of the true gender attribute among different candidate gender attributes can be set to 1, and the probability label value of other candidate gender attributes can be set to 0. Based on the probability label values ​​of different candidate gender attributes and the sample prediction probabilities of different candidate gender attributes, the first loss value corresponding to the gender attribute type can be determined.

[0159] In some embodiments, the first loss value corresponding to the gender attribute type can be determined according to the following formula:

[0160]

[0161] in, The first loss value corresponds to the gender attribute type; N represents the number of sample face images. This represents the sample acquisition weight feature corresponding to the sample acquisition features of the i-th sample face image; This represents the probability label value of the true gender attribute of the i-th sample face image (e.g., the probability label value for male is 1, and the probability label value for female is 0). The above formula represents the predicted probability of samples with different candidate gender attributes; the parameters in the formula are dimensionless data.

[0162] In some embodiments, the probability fusion result corresponding to the sample face image is determined based on the probability label value of the kth candidate age interval corresponding to each sample face image and the sample prediction probability of the kth candidate age interval. The sample collection weight features corresponding to each sample face image are weighted and summed to obtain the weighted sum result. Then, the product of the probability fusion result and the weighted sum result is averaged to obtain the first loss value corresponding to the age attribute type.

[0163] In some embodiments, for the age attribute type, the probability marker for candidate age intervals greater than the actual age can be 0, and the probability marker for candidate age intervals not greater than the actual age can be 1. Based on the probability markers of different candidate age attributes and the sample prediction probabilities of different candidate age attributes, the first loss value corresponding to the age attribute type can be determined.

[0164] In some embodiments, the first loss value corresponding to the age attribute type can be determined according to the following formula:

[0165]

[0166] in, The first loss value corresponds to the age attribute type; N represents the number of sample face images; This represents the sample acquisition weight feature corresponding to the sample acquisition features of the i-th sample face image; This represents the probability label value of the k-th candidate age interval corresponding to the i-th sample face image (for example, when the real age is greater than or equal to each age value in the k-th candidate age interval, the probability label value of the k-th candidate age interval is 1, and when the real age is less than each age value in the k-th candidate age interval, the probability label value of the k-th candidate age interval is 0). This represents the predicted probability of the k-th candidate age range corresponding to the i-th sample face image; the parameters in the above formula are dimensionless data.

[0167] In some embodiments, the sample prediction probabilities of different gender attributes under the gender attribute type can be mapped to a linearly independent target feature space to obtain the first feature projection data corresponding to the gender attribute type; and the sample prediction probabilities of different age intervals under the age attribute type can be mapped to a linearly independent target feature space to obtain the second feature projection data corresponding to the age attribute type; and a second loss value is determined based on the first feature projection data and the second feature projection data.

[0168] In some embodiments, the product of the transpose of the first feature projection data corresponding to the gender attribute type and the second feature projection data corresponding to the age attribute type can be determined, and then the Frobenius norm can be taken on the product result to obtain the second loss value.

[0169] In some embodiments, the second loss value can be determined according to the following formula:

[0170]

[0171] in, Indicates the second loss value; This represents the first feature projection data corresponding to the gender attribute type; This represents the second feature projection data corresponding to the age attribute type; denoted as Frobenius norm; the parameters in the above formula are dimensionless data.

[0172] In some embodiments, the first loss value corresponding to the gender attribute type, the first loss value corresponding to the age attribute type, and the second loss value can be weighted and summed to obtain the target loss value. The first weight coefficient of the first loss value corresponding to the gender attribute type, the second weight coefficient of the first loss value corresponding to the age attribute type, and the third weight coefficient of the second loss value can be preset based on experience.

[0173] In some embodiments, a third loss value can be determined based on the first weight coefficient of the first loss value corresponding to the gender attribute type, the second weight coefficient of the first loss value corresponding to the age attribute type, and the third weight coefficient of the second loss value. For example, the third loss value, i.e., the L2 regularization term, can be determined based on the sum of the squares of the first, second, and third weight coefficients. Then, the first, second, and third loss values ​​corresponding to the gender attribute type, the age attribute type, and the first and third loss values ​​are weighted and summed to obtain the target loss value.

[0174] In some embodiments, the target loss value can be expressed as:

[0175]

[0176] in, Indicates the target loss value. This represents the weighting coefficient of the first loss value corresponding to the gender attribute type. This represents the weighting coefficient of the first loss value corresponding to the age attribute type. This represents the weighting coefficient of the second loss value. This represents the third loss value. The weighting coefficients represent the third loss value; the parameters in the above formula are dimensionless data. For example, You can use 1.0. You can take 1.2. You can take 0.5. Can be taken .

[0177] In some embodiments, the quantization error loss can be determined by weighting and summing the first loss value corresponding to the gender attribute type, the first loss value, the second loss value, the third loss value corresponding to the age attribute type, and the quantization error constraint loss to obtain the target loss value.

[0178] Among them, quantization error loss can be understood as the difference loss caused by numerical approximation when compressing model weights and activation values ​​from floating-point data to low-precision integers during the forward and backward propagation processes of the face recognition model training process.

[0179] In some embodiments, a quantization-aware training mechanism can be introduced during the training of a face recognition model, which introduces quantization noise into the forward and backward propagation processes, while also incorporating quantization error loss.

[0180] The quantization formula can be expressed as:

[0181] ,

[0182] The inverse quantization formula can be expressed as:

[0183] ,

[0184] Where W represents the 32-bit floating-point weight parameters generated during the training of the face recognition model. The value of W is determined based on the structure of the face recognition model, the training stage, and the data distribution, and has no fixed value; A represents the 32-bit floating-point activation value generated during the forward propagation of the network; W q W represents the quantized value of the 8-bit signed integer data type (INT8) corresponding to W. q The value range of A is [-127, 127]; q This represents the INT8 quantized value of A. q The value range of S is [0, 255]; w S represents the scaling factor corresponding to W. w It is determined based on the maximum absolute value of the current layer W, that is... S a S represents the scaling factor corresponding to A. a It is based on the maximum value of the current layer A. and minimum value Certain, that is Z w Z represents the zero-point offset corresponding to W, where W is quantized using symmetric quantization. w Take 0; Z a This represents the zero-point offset corresponding to A, where, , This indicates rounding a floating-point number to the nearest integer. The above S... w and Z w During the training of the face recognition model, S is determined based on the distribution of the weight parameter W. a and Z a During the training of the face recognition model, the distribution of activation value A is used to determine and map the data between 32-bit floating-point data and 8-bit signed integer data, thereby reducing the accuracy loss caused by the quantization process.

[0185] The quantization error loss can be expressed as: Accordingly, based on the target loss value and the quantization error loss, the total loss value can be obtained:

[0186]

[0187] in, The weighting coefficients representing the quantization error loss, for example, A value of 0.1 can be used. The parameters in the above formula are dimensionless data.

[0188] In some embodiments, the training process of a face recognition model may include an initialization phase and a cockpit scene training phase.

[0189] During the initialization phase, the initial model can be initialized to enable it to extract general facial features. This means that when a face image is input into the initial model, the basic facial features of the face region can be output.

[0190] During the cockpit scene training phase, sample facial images and sample collection features of sample personnel inside the sample vehicle can be input into the pre-trained initial model to perform end-to-end joint training of the gender branch network and age branch network in the initial model.

[0191] The gender branch network can include two fully connected layers, which can output the predicted probabilities for different gender attributes. The age branch network can include two fully connected layers, which can output the predicted probabilities for different age attributes.

[0192] In some embodiments, in order to improve the face recognition capability of the face recognition model, the face features and personnel attributes corresponding to face images with confidence scores below the confidence threshold under each preset attribute type can be saved to the local encrypted storage area after user authorization, for use in retraining the face recognition model.

[0193] Based on the technical solutions of the above embodiments, this application also provides another optional embodiment, in which the face recognition method is described in detail.

[0194] See Figure 5 The facial recognition method shown includes:

[0195] S501, acquire the facial images of people inside the vehicle and the features corresponding to the facial regions in the facial images.

[0196] S502, input the face image and collected features into the face recognition model to extract the basic face features of the face region in the face image.

[0197] S503, perform nonlinear mapping processing on the acquired features to obtain the acquired weight features; wherein, the acquired weight features are used to characterize the weights of different acquired features.

[0198] S504 performs weighted processing on the basic facial features and the collected weight features to obtain the target facial features of the facial region in the facial image.

[0199] S505, based on the target facial features, determine the personnel attributes of the person as the predicted probability of different candidate gender attributes under the gender attribute type and the predicted probability of different candidate age ranges under the age attribute type.

[0200] S506. Based on the maximum probability value among the predicted probabilities of different candidate gender attributes, determine the confidence level of the face image under the gender attribute type, and based on the maximum and second-highest probability values ​​among the predicted probabilities of different candidate age ranges, determine the confidence level of the face image under the age attribute type.

[0201] S507, if the confidence level under the gender attribute type is not less than the gender confidence threshold and the confidence level under the age attribute type is not less than the age confidence threshold, the target attribute of the face image under each preset attribute type is determined based on the confidence level under each preset attribute type and the prediction probability of different candidate attributes.

[0202] S508, if the confidence level under the gender attribute type is less than the gender confidence level threshold, and / or if the confidence level under the age attribute type is less than the age confidence level threshold, acquire the cockpit sensor data corresponding to the location of the person.

[0203] S509 determines the prior probability of different candidate attributes under each preset attribute type based on cockpit sensor data.

[0204] S510: For each candidate attribute, the predicted probability and prior probability of the candidate attribute are weighted to obtain the target probability of the candidate attribute.

[0205] S511, for each preset attribute type, determine the target attribute of the person under each preset attribute type based on the target probability of different candidate attributes under the preset attribute type.

[0206] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0207] Based on the same inventive concept, this application also provides a face recognition device for implementing the face recognition method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more face recognition device embodiments provided below can be found in the limitations of the face recognition method described above, and will not be repeated here.

[0208] In one exemplary embodiment, such as Figure 6 As shown, a face recognition device is provided, including: an acquisition module 610, a first determination module 620, a second determination module 630, and a third determination module 640, wherein:

[0209] The acquisition module 610 is used to acquire the facial images of people inside the vehicle and the acquisition features corresponding to the facial regions in the facial images;

[0210] The first determining module 620 is used to determine the predicted probability of different candidate attributes under each preset attribute type based on the face image and the collected features.

[0211] The second determining module 630 is used to determine the confidence level of a face image under a preset attribute type based on the predicted probability of different candidate attributes under the preset attribute type for each preset attribute type.

[0212] The third determining module 640 is used to determine the target attribute of the person under each preset attribute type based on the confidence level under each preset attribute type and the prediction probability of different candidate attributes.

[0213] In one embodiment, there are multiple preset attribute types; the third determining module 640 includes: a first determining unit, configured to determine the target attribute of the face image under each preset attribute type based on the confidence level of each preset attribute type and the prediction probability of different candidate attributes when the confidence level of different preset attribute types meets preset conditions; an acquisition unit, configured to acquire cockpit sensor data corresponding to the location of the person when the confidence level of different preset attribute types does not meet preset conditions; and a second determining unit, configured to determine the target attribute of the person under each preset attribute type based on the cockpit sensor data and the prediction probability of different candidate attributes under each preset attribute type.

[0214] In one embodiment, determining that the confidence levels of different preset attribute types meet preset conditions includes: determining that the confidence levels of different preset attribute types meet preset conditions when the confidence levels of all preset attribute types are not less than the corresponding confidence thresholds.

[0215] In one embodiment, the second determining unit is specifically used to: determine the prior probability of different candidate attributes under each preset attribute type based on cockpit sensing data; and determine the target attribute of the person under each preset attribute type based on the predicted probability and prior probability of different candidate attributes under each preset attribute type.

[0216] In one embodiment, the second determining unit is specifically used to: for each candidate attribute, weight the predicted probability and prior probability of the candidate attribute to obtain the target probability of the candidate attribute; and for each preset attribute type, determine the target attribute of the person under each preset attribute type based on the target probabilities of different candidate attributes under the preset attribute type.

[0217] In one embodiment, the first determining module 620 includes: an extraction unit for extracting basic facial features of a face region in a face image; a mapping unit for performing nonlinear mapping processing on the acquired features to obtain acquisition weight features; wherein the acquisition weight features are used to characterize the weights of different acquired features; a weighting unit for performing weighted processing on the basic facial features and the acquisition weight features to obtain target facial features of a face region in the face image; and a third determining unit for determining the predicted probabilities of different candidate attributes under each preset attribute type based on the target facial features.

[0218] In one embodiment, each preset attribute type includes a first preset attribute type and a second preset attribute type, wherein the number of different candidate attributes under the first preset attribute type is greater than the number of different candidate attributes under the second preset attribute type; the second determining module 630 includes: a fourth determining unit, configured to, for each first preset attribute type, determine the maximum probability value and the second maximum probability value among the predicted probabilities of different candidate attributes under the first preset attribute type, and determine the confidence level of the face image under the preset attribute type based on the difference between the determined maximum probability value and the second maximum probability value; and a fifth determining unit, configured to, for each second preset attribute type, determine the maximum probability value among the predicted probabilities of different candidate attributes under the second preset attribute type, and determine the confidence level of the face image under the preset attribute type based on the determined maximum probability value.

[0219] In one embodiment, the first determining module 620 includes: an input unit, used to input a face image and collected features into a face recognition model to obtain the predicted probabilities of different candidate attributes under each preset attribute type for the person's personnel attributes.

[0220] In one embodiment, the face recognition model is trained through the following steps: acquiring sample face images, sample collection features of the face region in the sample face images, and the real attributes of the sample person corresponding to the sample face image under each preset attribute type; inputting the sample face images and sample collection features into an initial model to obtain the sample prediction probabilities of the person attributes of the sample person as different candidate attributes under each preset attribute type; and adjusting the parameters of the initial model according to the real attributes corresponding to each preset attribute type and the sample prediction probabilities of different candidate attributes under each preset attribute type to obtain the face recognition model.

[0221] In one embodiment, the initial model is tuned based on the true attribute corresponding to each preset attribute type and the sample prediction probabilities of different candidate attributes under the preset attribute type. This includes: for each preset attribute type, determining a first loss value corresponding to the preset attribute type based on the true attribute corresponding to the preset attribute type and the sample prediction probabilities of different candidate attributes; determining a second loss value based on the sample prediction probabilities of different candidate attributes under each preset attribute type; determining a target loss value based on the first loss value and the second loss value; and tuning the initial model if the target loss value does not meet the preset convergence condition.

[0222] The various modules in the aforementioned face recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0223] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and databases. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database of this computer device stores facial images, acquired features, predicted probabilities of different candidate attributes under various preset attribute types for personnel attributes, confidence scores of facial images under the preset attribute types, target attributes of personnel under each preset attribute type, cockpit sensor data corresponding to the personnel's location, prior probabilities of different candidate attributes under each preset attribute type, target probabilities of each candidate attribute, acquired weight features, sample facial images, sample acquired features, sample predicted probabilities of different candidate attributes under each preset attribute type for personnel attributes, and first loss values, second loss values, and target loss values ​​corresponding to each preset attribute type. The input / output interface of this computer device is used for information exchange between the processor and external devices. The communication interface of this computer device is used for communication with external terminals via network connection. When the computer program is executed by the processor, it implements a facial recognition method.

[0224] Those skilled in the art will understand that Figure 7The structure shown is a block diagram of a partial structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0225] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0226] In one exemplary embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method embodiments.

[0227] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0228] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program mentioned can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0229] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0230] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A face recognition method, characterized in that, The method includes: Acquire facial images of people inside the vehicle and the collected features corresponding to the facial regions in the facial images; Based on the facial image and the collected features, the personnel attributes of the person are determined as the predicted probabilities of different candidate attributes under each preset attribute type; the number of preset attribute types is multiple. For each preset attribute type, the confidence level of the face image under the preset attribute type is determined based on the predicted probability of different candidate attributes under the preset attribute type. When the confidence level of different preset attribute types meets the preset conditions, the target attribute of the face image under each preset attribute type is determined based on the confidence level under each preset attribute type and the prediction probability of different candidate attributes. If the confidence level of different preset attribute types does not meet the preset conditions, obtain the cockpit sensor data corresponding to the location of the person. Based on the cockpit sensor data and the predicted probabilities of different candidate attributes under each preset attribute type, the target attribute of the person under each preset attribute type is determined.

2. The method according to claim 1, characterized in that, The confidence levels of the different preset attribute types satisfy the preset conditions, including: If the confidence level of different preset attribute types is not less than the corresponding confidence level threshold, then the confidence level of different preset attribute types is determined to meet the preset conditions.

3. The method according to claim 1, characterized in that, Based on the cockpit sensor data and the predicted probabilities of different candidate attributes under each preset attribute type, the target attribute of the person under each preset attribute type is determined, including: Based on the cockpit sensor data, determine the prior probability of different candidate attributes under each preset attribute type; Based on the predicted probability and prior probability of different candidate attributes under each preset attribute type, the target attribute of the person under each preset attribute type is determined.

4. The method according to claim 3, characterized in that, Based on the predicted probability and prior probability of different candidate attributes under each preset attribute type, the target attribute of the person under each preset attribute type is determined, including: For each candidate attribute, the predicted probability and prior probability of the candidate attribute are weighted to obtain the target probability of the candidate attribute. For each preset attribute type, the target attribute of the person under each preset attribute type is determined based on the target probability of different candidate attributes under the preset attribute type.

5. The method according to any one of claims 1-4, characterized in that, Based on the facial image and the acquired features, the personnel attributes of the person are determined as the predicted probabilities of different candidate attributes under each preset attribute type, including: Extract basic facial features from the facial region of the face image; and, The acquisition features are subjected to nonlinear mapping processing to obtain acquisition weight features; wherein, the acquisition weight features are used to characterize the weights of different acquisition features; The basic facial features and the collected weight features are weighted to obtain the target facial features of the facial region in the facial image; Based on the target facial features, the personnel attributes of the person are determined as the predicted probabilities of different candidate attributes under each preset attribute type.

6. The method according to any one of claims 1-4, characterized in that, Each preset attribute type includes a first preset attribute type and a second preset attribute type, wherein the number of different candidate attributes under the first preset attribute type is greater than the number of different candidate attributes under the second preset attribute type; For each preset attribute type, the confidence level of the face image under the preset attribute type is determined based on the predicted probabilities of different candidate attributes under the preset attribute type, including: For each first preset attribute type, determine the maximum probability value and the second maximum probability value among the predicted probabilities of different candidate attributes under the first preset attribute type, and determine the confidence level of the face image under the preset attribute type based on the difference between the determined maximum probability value and the second maximum probability value. For each second preset attribute type, determine the maximum probability value among the predicted probabilities of different candidate attributes under the second preset attribute type, and determine the confidence level of the face image under the preset attribute type based on the determined maximum probability value.

7. The method according to any one of claims 1-4, characterized in that, Based on the facial image and the acquired features, the personnel attributes of the person are determined as the predicted probabilities of different candidate attributes under each preset attribute type, including: The face image and the collected features are input into the face recognition model to obtain the predicted probabilities of different candidate attributes under each preset attribute type for the person's attributes.

8. The method according to claim 7, characterized in that, The face recognition model is trained through the following steps: Acquire sample face images, sample collection features of face regions in the sample face images, and the real attributes of the sample person corresponding to the sample face images under each preset attribute type; The sample face image and the sample collection features are input into the initial model to obtain the sample prediction probability of different candidate attributes under each preset attribute type for the personnel attributes of the sample personnel. Based on the true attributes corresponding to each preset attribute type and the sample prediction probabilities of different candidate attributes under the preset attribute type, the initial model is adjusted to obtain a face recognition model.

9. The method according to claim 8, characterized in that, Based on the true attribute corresponding to each preset attribute type and the sample prediction probabilities of different candidate attributes under the preset attribute type, the initial model is tuned, including: For each preset attribute type, a first loss value corresponding to the preset attribute type is determined based on the true attribute corresponding to the preset attribute type and the sample prediction probability of different candidate attributes. The second loss value is determined based on the sample prediction probability of different candidate attributes under each preset attribute type; Based on the first loss value and the second loss value, a target loss value is determined, and if the target loss value does not meet the preset convergence condition, the initial model is tuned.

10. A face recognition device, characterized in that, The device includes: The acquisition module is used to acquire facial images of people inside the vehicle and the acquisition features corresponding to the facial regions in the facial images; The first determining module is used to determine the predicted probability of different candidate attributes under each preset attribute type based on the face image and the collected features; the number of preset attribute types is multiple. The second determining module is used to determine the confidence level of the face image under the preset attribute type based on the predicted probability of different candidate attributes under the preset attribute type for each preset attribute type. The third determining module is used to determine the target attribute of the face image under each preset attribute type based on the confidence level of each preset attribute type and the prediction probability of different candidate attributes when the confidence level of different preset attribute types meets the preset conditions; and to obtain the cockpit sensor data corresponding to the location of the person when the confidence level of different preset attribute types does not meet the preset conditions; and to determine the target attribute of the person under each preset attribute type based on the cockpit sensor data and the prediction probability of different candidate attributes under each preset attribute type.

11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.