In-vehicle user emotion recognition method and device, vehicle and medium
By performing perspective correction and preprocessing on in-vehicle user facial images, a frontal view image is generated, which solves the problem of inaccurate facial information from the side view and improves the accuracy of emotion recognition and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2026-03-03
AI Technical Summary
In the in-vehicle environment, due to factors such as the camera installation angle and user posture, the captured user facial images are usually from a side view, resulting in inaccurate or even missing facial information, poor accuracy in emotion type recognition, and a poor user experience.
By correcting the perspective of the user's face image, using facial key points to determine three-dimensional rotation and translation information, the image is mapped to 3D space to generate a frontal view of the user's face image. Noise reduction, brightness correction, and illumination separation are then performed to improve image quality, which is then used for emotion recognition.
It achieves a complete and accurate reflection of the user's facial information, improves the accuracy of emotion type recognition, and enhances the user experience.
Smart Images

Figure CN121600569A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of in-vehicle intelligent interaction technology, and in particular to an in-vehicle user emotion recognition method, device, vehicle and medium. Background Technology
[0002] With the development of automotive intelligence, in-vehicle intelligent interaction technology has received widespread attention. During vehicle operation, facial images of users inside the vehicle can be collected, and the emotional types of users can be identified based on these images. This allows for more tailored interactive actions to meet the user's needs, thus enhancing the user experience.
[0003] In related technologies, emotion recognition often relies on images taken from a frontal view of the user. However, in the in-vehicle environment, due to factors such as the camera's installation angle and the user's posture, the captured facial images are usually from a side view, resulting in inaccurate or even missing facial information. This leads to poor accuracy in emotion type recognition and a poor user experience.
[0004] It should be noted that the information disclosed in the background section of this application is intended only to enhance the understanding of the general background of this application, and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention
[0005] This application provides a method, device, vehicle, and medium for in-vehicle user emotion recognition, which helps to solve the problem that the accuracy of user emotion type recognition is poor and the user experience is not good because the user's facial images are usually from a side view and the user's facial information is inaccurate or even missing.
[0006] In a first aspect, embodiments of this application provide an in-vehicle user emotion recognition method, including: Based on the facial key points in the user's face image, the viewpoint of the user's face image is corrected to determine the frontal view of the user's face image; Emotion recognition is performed on the user's facial image from the frontal view to determine the user's emotion type.
[0007] In one possible implementation, the step of performing viewpoint correction on the user's face image based on facial key points in the user's face image to determine a frontal view of the user's face image includes: Based on the facial key points in the user's face image, determine the user's three-dimensional rotation and translation information; Based on the user's three-dimensional rotation and translation information, the user's face image is corrected to determine a frontal view of the user's face image.
[0008] In one possible implementation, determining the user's three-dimensional rotation and translation information based on facial key points in the user's face image includes: Based on the facial key points in the user's face image and the corresponding key points in the standard 3D face model, the user's three-dimensional rotation and translation information are determined.
[0009] In one possible implementation, the step of performing viewpoint correction on the user's facial image based on the user's three-dimensional rotation and translation information to determine a frontal view of the user's facial image includes: The user's facial image is mapped to 3D space to determine the user's 3D facial model; Based on the user's 3D rotation and translation information, determine the viewpoint correction projection parameters of the user's 3D face model; Based on the user's 3D face model and the viewpoint correction projection parameters, a frontal view user face image is determined.
[0010] In one possible implementation, before performing viewpoint correction on the user's face image based on facial key points in the user's face image to determine the frontal view user's face image, the method further includes: The acquired face image is subjected to noise reduction processing, brightness correction processing for overexposed areas, and / or illumination separation and reflection enhancement processing for shadow areas to determine the user's face image.
[0011] In one possible implementation, the overexposed area brightness correction process includes: According to the formula: Determine the correction index; Where γ is the correction index, M0 is the brightness reference, and M is the median brightness within a preset range around the current pixel; According to the formula: Brightness correction processing is performed on overexposed areas of the acquired image; Among them, I out To output pixel brightness, M max For maximum brightness, I in This is the input pixel brightness.
[0012] In one possible implementation, the shadow area illumination separation and reflection enhancement processing includes: According to the formula: Determine the single-scale reflection components. Among them, R k (x,y) represents a single-scale component, I(x,y) represents the input value of the acquired image, and F... k (x,y) is a Gaussian wrapping function; According to the formula: Perform shadow area illumination separation and reflection enhancement processing on the acquired image data; Among them, R MSR (x,y) represents the multi-scale fusion component, R k (x,y) represents the single-scale reflection component, N is the number of scales, and ω k It is a single-scale weight.
[0013] Secondly, embodiments of this application provide an in-vehicle user emotion recognition device, comprising: A frontal view image determination device is used to correct the view of a user's face image based on facial key points in the user's face image, and determine a frontal view user face image. The user emotion type determination device is used to perform emotion recognition based on the user's facial image from the frontal view to determine the user's emotion type.
[0014] Thirdly, embodiments of this application provide a vehicle, including: A controller, wherein the controller is configured to be used in any of the first aspects of the method.
[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a controller, implements the method described in any one of the first aspects.
[0016] In this embodiment of the application, the collected user face image is subjected to perspective correction. The corrected frontal perspective can completely and accurately reflect the user's facial information. Emotion recognition is performed on the user's face image from the corrected frontal perspective, which improves the accuracy of the user's emotion type recognition and enhances the user experience. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a structural diagram illustrating an application scenario provided in an embodiment of this application. Figure 2 A flowchart illustrating an in-vehicle user emotion recognition method provided in an embodiment of this application; Figure 3 A flowchart illustrating another in-vehicle user emotion recognition method provided in an embodiment of this application; Figure 4A flowchart illustrating another in-vehicle user emotion recognition method provided in an embodiment of this application; Figure 5 A flowchart illustrating another in-vehicle user emotion recognition method provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an in-vehicle user emotion recognition device provided in an embodiment of this application; Figure 7 This is a structural schematic diagram of a vehicle provided in an embodiment of this application. Detailed Implementation
[0019] To better understand the technical solution of this application, the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0020] It should be understood that the described embodiments are merely some, not all, of the embodiments in this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0021] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0022] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0023] With the development of automotive intelligence, in-vehicle intelligent interaction technology has received widespread attention.
[0024] See Figure 1 This is a structural diagram illustrating an application scenario provided in an embodiment of this application. Figure 1 As shown, the vehicle 100 includes a controller 101 and a user image acquisition device 102. Specifically, the controller 101 can acquire the user's facial image through the user image acquisition device 102. Furthermore, the controller 101 identifies the emotional type of the user in the vehicle based on the facial image, and then provides more user-friendly interactive actions based on the user's emotional type, which helps to improve the user experience.
[0025] It should be pointed out that, as Figure 1In the application scenarios shown, vehicle 100 includes, but is not limited to, sedans, SUVs, and MPVs; controller 101 includes, but is not limited to, microcontroller units (MCUs) and system-on-chips (SoCs); user image acquisition devices 102 include, but are not limited to, visible light cameras and infrared cameras. This application embodiment does not impose specific limitations on these.
[0026] In related technologies, emotion recognition often relies on images taken from a frontal view of the user. However, in the in-vehicle environment, due to factors such as the camera's installation angle and the user's posture, the captured facial images are usually from a side view, resulting in inaccurate or even missing facial information. This leads to poor accuracy in emotion type recognition and a poor user experience.
[0027] To address the aforementioned issues, this application provides an in-vehicle user emotion recognition method. The method corrects the perspective of the collected user face image, and the corrected frontal perspective can fully and accurately reflect the user's facial information. By performing emotion recognition on the user's face image from the corrected frontal perspective, the accuracy of the user's emotion type recognition is improved, thereby enhancing the user experience.
[0028] See Figure 2 This is a flowchart illustrating an in-vehicle user emotion recognition method provided in an embodiment of this application. It can be applied to... Figure 1 The application scenarios shown are as follows: Figure 2 As shown, it mainly includes the following steps.
[0029] S201: Based on the facial key points in the user's face image, perform perspective correction on the user's face image to determine the frontal view of the user's face image.
[0030] As we can understand, facial key points refer to specific points, or sets of points, that locate facial features in a user's face image. Facial key points are typically located in areas that reflect the geometric shape and structural features of the user's face, such as facial features and contours. In this way, the location of these points allows us to obtain facial information about the user, including overall contours and local details, which is a crucial foundation for in-vehicle user emotion recognition.
[0031] Specifically, representative locations and regions can be selected based on the geometric shape and structural features of the user's face. In one possible implementation, facial key points detected and identified from the user's face image include multiple predefined facial key points in regions such as the eye region, nose region, lip region, and contour edge region. Of course, those skilled in the art can also use other facial key point combination schemes, such as sets of key points with different numbers and locations, depending on the specific needs of the application. This application does not impose specific limitations on these combinations.
[0032] Viewpoint correction can be performed based on the location information of facial key points in the current user's face image. Specifically, the spatial distribution of these key points under the current viewpoint can be analyzed, and combined with the standard distribution of facial key points under a frontal viewpoint, the difference between the current viewpoint and the frontal viewpoint can be obtained. Based on this difference, the user's face image is adjusted so that the distribution of facial key points conforms to the characteristics of a frontal viewpoint, thereby completing the viewpoint correction.
[0033] S202: Perform emotion recognition based on the user's face image from the frontal view to determine the user's emotion type.
[0034] In practical applications, a user's current facial information can be obtained from their facial image, and emotion recognition can be performed based on this information to determine the user's emotion type. Furthermore, the accuracy of emotion type recognition is improved when the user's facial image is a complete and accurate frontal view that reflects their facial information.
[0035] Specifically, facial features related to emotions can be obtained from a frontal view of the user's face image. These features are then matched and associated with preset emotion type feature information. Combined with preset judgment rules, the user's emotion type is finally determined.
[0036] Of course, those skilled in the art can also choose other methods to perform emotion recognition based on frontal view user facial images to determine the user's emotion type, depending on the actual situation. For example, a frontal view user facial image, or feature information extracted from the image, can be used as input. A neural network model trained on a dataset of facial images labeled with different emotion types can be used to analyze the input information and finally output the user's emotion type. This application does not impose specific limitations on this.
[0037] Understandably, once a user's emotional type is determined, corresponding feedback can be provided based on that type. For example, based on the user's emotional type, the feedback module can generate appropriate prompts (such as driver fatigue warnings, emotional reassurance prompts, etc.) to help drivers and passengers maintain a good emotional state and safety awareness inside the vehicle. It can also integrate with other in-vehicle intelligent systems (such as navigation and entertainment systems) to automatically adjust system settings based on the user's emotional state.
[0038] See Figure 3 This is a flowchart illustrating another in-vehicle user emotion recognition method provided in an embodiment of this application. Figure 3 As shown, in Figure 2 Based on the method embodiment shown, step S201 specifically includes the following steps.
[0039] S301: Determine the user's three-dimensional rotation and translation information based on the facial key points in the user's face image.
[0040] As mentioned above, the spatial distribution of facial key points in the current viewpoint can be analyzed. By combining this with the standard distribution of facial key points in a frontal viewpoint, the differences between the current viewpoint and the frontal viewpoint can be obtained. These differences can be described using 3D rotation and translation information.
[0041] Specifically, a spatial coordinate system can be established based on the image acquisition device. The three-dimensional rotation information reflects the rotational parameters of the user's face in three-dimensional space relative to a preset frontal viewpoint, typically including the rotation angles of the face around the three coordinate axes of the spatial coordinate system. For example, the pitch angle θ around the x-axis. pitch This typically reflects the vertical tilt of the face; the yaw angle θ around the y-axis. yaw This typically reflects the left and right turning of the face; the roll angle θ around the z-axis. roll These angles typically reflect the left and right tilting of the face. These angles directly reflect the difference between the current viewing angle and the frontal viewing angle of the face.
[0042] Of course, those skilled in the art can also choose other methods to reflect three-dimensional rotation information according to the actual situation. For example, by setting the user image acquisition device in certain specific positions inside the vehicle, the information can be reflected solely by the pitch angle θ around the x-axis. pitch yaw angle θ around the y-axis yaw and the roll angle θ about the z-axis roll Any two or only one of them reflects three-dimensional rotation information, and the embodiments of this application do not impose specific limitations on this.
[0043] Similarly, translation information refers to the positional offset parameters of the user's face in three-dimensional space relative to a preset reference position. It typically includes pixel offset within the imaging plane and depth offset along the imaging optical axis. The offset within the imaging plane reflects the positional deviation of the face in the image, while the depth offset along the imaging optical axis reflects the change in face size due to distance differences.
[0044] Of course, those skilled in the art can also choose other ways to reflect translation information according to the actual situation. For example, by setting the user image acquisition device in certain specific positions inside the vehicle, translation information can be reflected only by pixel offset in the plane or depth offset along the imaging optical axis. This application embodiment does not impose specific limitations on this.
[0045] In one possible implementation, the user's three-dimensional rotation and translation information are determined based on facial key points in the user's face image and corresponding key points in a standard 3D face model.
[0046] In practical applications, a standard 3D face model usually refers to a general model that is generated based on a large amount of human facial 3D scan data and includes the standard 3D coordinates of key facial points from a frontal view.
[0047] In this way, the user's 3D rotation and translation information can be determined based on facial key points in the user's face image and corresponding key points in the standard 3D face model. Specifically, the facial key points in the user's face image are matched with the corresponding key points in the standard 3D face model to obtain the coordinates of the facial key points in the user's face image and the coordinates of the corresponding key points in the standard 3D face model. A spatial geometric solution algorithm suitable for matching 2D image points with 3D spatial points is adopted. The coordinates of the facial key points in the user's face image and the coordinates of the corresponding key points in the standard 3D face model are input to solve for the 3D rotation and translation information.
[0048] For example, the Perspective-n-Point (PnP) algorithm is used to solve for the rotation and translation information of facial key points in a user's face image and their corresponding key points in a standard 3D face model. Specifically, Euler angles are used to represent the 3D rotation information, which is determined by the formula shown below.
[0049] Where, θ yaw Let θ be the yaw angle. pitch Let θ be the pitch angle. roll For roll angle, 3D model These are the coordinates of key points in a standard 3D face model, and in 2D... keypointThese are the coordinates of facial key points in the user's face image.
[0050] Of course, those skilled in the art can also choose other methods to determine the user's three-dimensional rotation and translation information based on facial key points in the user's face image, depending on the actual situation. For example, the facial key points in the user's face image can be matched with the coordinates of standard 2D key points, the three-dimensional rotation angle can be determined by solving the perspective transformation matrix, and the translation information can be obtained through a depth-sensing sensor such as an image acquisition device and the translation component in the perspective transformation matrix. This application does not impose specific limitations on this.
[0051] S302: Based on the user's three-dimensional rotation and translation information, perform perspective correction on the user's face image to determine a frontal view of the user's face image.
[0052] It is understood that the three-dimensional rotation information and translation information are information describing the difference between the current viewpoint and the frontal viewpoint. Based on the user's three-dimensional rotation information and translation information, the viewpoint of the user's face image can be corrected to determine the frontal viewpoint of the user's face image.
[0053] Specifically, the adjustment direction and magnitude of the user's facial image can be determined based on three-dimensional rotation and translation information. For example, the yaw angle θ in the three-dimensional rotation information can be used... yaw Positive and negative values and numerical magnitudes determine the adjustment direction and magnitude of the left and right rotation of the user's face image. Translation information determines the translation of the user's face image, and scaling control parameters determine the translation direction, distance, and scaling degree of the user's face image.
[0054] Of course, those skilled in the art can also choose other methods to perform perspective correction on the user's face image based on the facial key points in the user's face image, and determine the frontal view of the user's face image, such as 2D affine transformation based on facial key points, or correction based on the symmetry characteristics of facial key points, etc. The embodiments of this application do not impose specific limitations on this.
[0055] See Figure 4 This is a flowchart illustrating another in-vehicle user emotion recognition method provided in an embodiment of this application. Figure 4 As shown, in Figure 3 Based on the method embodiment shown, step S302 specifically includes the following steps.
[0056] S401: Map the user's face image to 3D space to determine the user's 3D face model.
[0057] In practical applications, a user's 3D face model usually refers to a 3D model that can accurately reflect the user's current facial information (including the three-dimensional geometric shape and structural features of the face), covering the three-dimensional spatial information of the user's facial features (such as eyes, nose, and lips) and contours.
[0058] It is understandable that, based on the user's facial information in the user's face image, combined with the statistical features or geometric transformation rules of 3D faces, the facial information carried by the two-dimensional image can be extended into a three-dimensional spatial representation, thereby constructing a 3D face model that matches the current user's facial features.
[0059] For example, a 3D facial statistical deformation model based on Principal Component Analysis (PCA) can be used to map a user's facial image into 3D space. Specifically, this is achieved through the formula... The user's facial image is mapped into a 3D space, where S is the user's 3D face. For an average 3D face, k is the threshold for the amount of facial morphological variation information, and α i U is the PCA coefficient. i These are the deformation basis vectors.
[0060] Understandably, an average 3D face representing the baseline shape of a typical face can be obtained through statistical analysis of 3D facial scan data. PCA is used to extract a series of deformation basis vectors U that can describe facial variations from these facial data. i Considering model accuracy and computational cost, the top k deformation basis vectors U that best reflect the main facial differences are selected. i And assign PCA coefficient α to each deformation basis vector. i To reflect the degree of influence of the corresponding variations on the user's face, these weighted deformation basis vectors are superimposed onto the average 3D face. The above determines the user's 3D face model.
[0061] Of course, those skilled in the art can also choose other methods to map user face images to 3D space to determine the user's 3D face model, depending on the actual situation. For example, a general 3D face model with predefined parameters (such as the BlendShape model) can be used, and the model parameters (such as facial feature positions, contour curvature, etc.) can be adjusted through optimization algorithms to match the model with the facial key points of the user's face image, thereby determining the user's 3D face model.
[0062] S402: Determine the viewpoint correction projection parameters of the user's 3D face model based on the user's 3D rotation and translation information.
[0063] In practical applications, the viewpoint correction projection parameters are a set of key parameters used to control the projection process of a user's 3D face model onto a frontal view 2D image. Their function is to ensure that the adjusted projection result of the user's 3D face model accurately presents facial features from a frontal view.
[0064] Specifically, the viewpoint correction projection parameters of the user's 3D face model can be determined based on 3D rotation and translation information. For example, the reverse rotation parameters can be determined based on 3D rotation information. For instance, if the current 3D face model is affected by a yaw angle θ... yaw If there is a left or right steering deviation, then determine the counter-rotation parameter (such as -θ) that can counteract the deviation. yaw (The amount of rotation). This determines the angle of the 3D face model's frontal view in 3D space, eliminating viewpoint shift caused by the angle. Position and scale adjustment parameters are determined based on translation information. For example, based on the offset within the imaging plane, the displacement parameters required to translate the 3D face model to a preset area for the frontal view are calculated; based on the depth offset along the optical axis, the scaling parameters required to adjust the model to a preset model distance are calculated. Thus, the viewpoint correction projection parameters of the user's 3D face model can be determined based on the reverse rotation parameters, scaling parameters, and displacement parameters.
[0065] S403: Determine the frontal view user face image based on the user's 3D face model and the viewpoint correction projection parameters.
[0066] Understandably, the viewing angle of a user's 3D face model can be adjusted based on the viewing angle correction projection parameters. For example, the inverse rotation parameter included in the viewing angle correction projection parameters can guide the user's 3D face model to complete posture correction in three-dimensional space, reducing the angular deviation between the current viewing angle and the frontal viewing angle, so that the model's facial orientation and the relative positions of facial features conform to the typical characteristics of a frontal view, such as horizontally symmetrical eyes, a centered nose bridge, and no lateral deviation of the facial contours; furthermore, the model is adjusted to a preset spatial position and scale based on the position and scale adjustment parameters in the viewing angle correction projection parameters.
[0067] Secondly, the user's 3D face model, adjusted for different viewpoints, is projected into a 2D image. A perspective projection algorithm can be used to map the spatial coordinates of the 3D model to a 2D imaging plane, converting it into pixel coordinate distribution to generate an initial frontal 2D image. This 2D image includes the user's facial features from a frontal viewpoint, i.e., a frontal view user face image.
[0068] Of course, those skilled in the art can also choose other methods according to the actual situation to determine the frontal view user face image based on the user's 3D face model and the viewpoint correction projection parameters. For example, before projection, based on the user's 3D face model, a neural renderer and texture reconstruction are used to map the texture information in the two-dimensional image onto the surface of the user's 3D face model. Through viewpoint correction projection parameters and frontal view projection technology, the three-dimensional model is projected onto the frontal view to generate the frontal view user face image. This application does not impose specific limitations on this.
[0069] Meanwhile, those skilled in the art can also choose other methods to perform perspective correction on the user's face image based on the user's three-dimensional rotation and translation information, and determine the frontal view user face image, depending on the actual situation. For example, the three-dimensional rotation information can be converted into rotation and shearing coefficients of a 2D image, and the translation information can be converted into translation and scaling parameters of a 2D image, integrated to form an affine or perspective transformation matrix, and applied to the user's face image to determine the frontal view user face image. This application does not impose specific limitations in this regard.
[0070] See Figure 5 This is a flowchart illustrating another in-vehicle user emotion recognition method provided in an embodiment of this application. Figure 5 As shown, in Figure 2 Based on the method embodiment shown, before step S201, the following steps are also included.
[0071] S501: Perform noise reduction processing, brightness correction processing for overexposed areas, and / or illumination separation and reflection enhancement processing on the acquired face image to determine the user's face image.
[0072] In practical applications, capturing facial images typically refers to acquiring images through user image capture devices. Understandably, in a vehicle interior environment, directly capturing a user's face may be affected by the complex environment, resulting in noise interference, localized overexposure, and shadow coverage. This makes the image unsuitable for subsequent facial landmark detection and emotion recognition. Therefore, preprocessing techniques such as noise reduction, brightness correction for overexposed areas, and / or shadow area correction can be used to optimize image quality.
[0073] Specifically, invalid noise signals in the acquired face images can be removed using filtering algorithms suitable for face images, preventing noise from interfering with subsequent facial feature extraction. For example, guided filtering can be used for noise reduction; for instance, the scaling factor and compensation factor can be determined using the formula shown below. and Determine the proportionality coefficient a k and compensation coefficient b k , Among them, u kσ is the average pixel value within the window. k 2 Let cov(I) be the pixel variance within the window. noisy I noisy ) k Let ε be the covariance between the input image and itself, and ε be the regularization parameter.
[0074] Through formula Denoising is performed on the acquired face images, where I denoised For the denoised image, I noisy The input image contains noise.
[0075] Understandable, cov(I) noisy I noisy ) k It can reflect pixel correlation, σ k 2 The variance of pixels within the output window is used to define regions where large variance corresponds to edge regions and regions where small variance corresponds to smooth regions. ε is a regularization parameter to prevent computational instability. k The average pixel value within the window is used to compensate for and prevent distortion. Finally, through linear combination, a denoised image is obtained that removes invalid noise while preserving facial details.
[0076] In one possible implementation, according to the formula: Determine the correction index; Where γ is the correction index, M0 is the brightness reference, and M is the median brightness within a preset range around the current pixel; According to the formula: Brightness correction processing is performed on overexposed areas of the acquired image; Among them, I out To output pixel brightness, M max For maximum brightness, I in This is the input pixel brightness.
[0077] Understandably, a reference standard for normal brightness is defined based on the brightness benchmark M0; M is the median brightness within a preset range around the current pixel, reflecting the actual brightness level of the local area. The correction index γ is determined based on the reference standard for normal brightness and the median brightness within the preset range around the current pixel. The correction index γ can be used to quantify the degree of deviation between the current local brightness and the benchmark brightness. The more severe the overexposure (the greater the deviation between M and M0), the more effectively γ can compress the dynamic range of brightness in the overexposed area. Subsequent power function calculations are performed on the input pixel brightness I... in By scaling proportionally, overexposed high brightness values are mapped to a reasonable range while preserving details of brightness changes, ultimately restoring the brightness of overexposed areas to normal and avoiding the loss of facial details due to brightness saturation.
[0078] In one possible implementation, according to the formula: Determine the single-scale reflection components; Among them, R k (x,y) represents a single-scale component, I(x,y) represents the input value of the acquired image, and F... k (x,y) is a Gaussian wrapping function; According to the formula: Perform shadow area illumination separation and reflection enhancement processing on the acquired image data; Among them, R MSR (x,y) represents the multi-scale fusion component, R k (x,y) represents the single-scale reflection component, N is the number of scales, and ω k It is a single-scale weight.
[0079] In practical applications, the Gaussian wrapping function F can be used. k (x,y) simulates light distribution at different scales to separate light components.
[0080] Understandably, the single-scale reflection component formula separates facial reflection information independent of illumination through logarithmic field subtraction (the illumination component is surrounded by a Gaussian function F). k (x,y) modeling, where the difference represents the facial reflection features; multi-scale fusion then weights and superimposes the reflection components extracted from different scales, preserving fine details at small scales while fusing the overall contour at large scales, ultimately enhancing the reflection information in shadowed areas. This achieves the preservation of facial details covered by shadows while separating and removing lighting interference.
[0081] Corresponding to the above embodiments, this application also provides an in-vehicle user emotion recognition device.
[0082] See Figure 6 This is a schematic diagram of the structure of an in-vehicle user emotion recognition device provided in an embodiment of this application. Figure 6 As shown, the in-vehicle user emotion recognition device 600 includes a frontal view image determination device 601 and a user emotion type determination device 602.
[0083] The frontal view image determination device 601 is used to perform view correction on the user's face image based on the facial key points in the user's face image, and determine the frontal view user face image. User emotion type determination device 602 is used to perform emotion recognition based on the user's face image from the frontal view and determine the user's emotion type.
[0084] In one possible implementation, the frontal view image determining device 601 is further configured to determine the user's three-dimensional rotation and translation information based on facial key points in the user's face image; and to perform viewpoint correction on the user's face image based on the user's three-dimensional rotation and translation information to determine the frontal view user face image.
[0085] In one possible implementation, the frontal view image determination device 601 is further configured to determine the user's three-dimensional rotation and translation information based on the facial key points in the user's face image and the corresponding key points in the standard 3D face model.
[0086] In one possible implementation, the frontal view image determining device 601 is also used for The user's facial image is mapped to 3D space to determine the user's 3D facial model; Based on the user's 3D rotation and translation information, determine the viewpoint correction projection parameters of the user's 3D face model; Based on the user's 3D face model and the viewpoint correction projection parameters, a frontal view user face image is determined.
[0087] For details regarding the embodiments of this application, please refer to the description of the above method embodiments. For the sake of brevity, these details will not be repeated here.
[0088] Corresponding to the above embodiments, this application also provides a vehicle.
[0089] See Figure 7 This is a structural schematic diagram of a vehicle provided in an embodiment of this application. Figure 7 As shown, vehicle 700 includes controller 701. Controller 701 is configured to perform some or all of the steps in the method embodiments.
[0090] For details regarding the embodiments of this application, please refer to the description of the above method embodiments. For the sake of brevity, these details will not be repeated here.
[0091] Corresponding to the above embodiments, this application also provides a computer-readable storage medium, wherein the computer-readable storage medium may store a program, and when the program runs, it can control the device where the computer-readable storage medium is located to execute some or all of the steps in the above method embodiments. In specific implementation, the computer-readable storage medium may be a magnetic disk, an optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0092] For details regarding the embodiments of this application, please refer to the description of the above method embodiments. For the sake of brevity, these details will not be repeated here.
[0093] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, the simultaneous existence of A and B, or the existence of B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects have an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0094] Those skilled in the art will recognize that the units and algorithm steps described in the embodiments disclosed herein can be implemented using electronic hardware, computer software, or a combination of electronic hardware and software. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0095] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the above-described apparatus, controller, and computer storage medium can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0096] In the several embodiments provided in this application, any function, if implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0097] The above description is merely a specific embodiment of this application. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application. The protection scope of this application should be determined by the protection scope of the claims.
Claims
1. A method for recognizing user emotions inside a vehicle, characterized in that, include: Based on the facial key points in the user's face image, the viewpoint of the user's face image is corrected to determine the frontal view of the user's face image; Emotion recognition is performed on the user's facial image from the frontal view to determine the user's emotion type.
2. The method according to claim 1, characterized in that, The step of correcting the viewpoint of the user's face image based on facial key points in the user's face image to determine a frontal view of the user's face image includes: Based on the facial key points in the user's face image, determine the user's three-dimensional rotation and translation information; Based on the user's three-dimensional rotation and translation information, the user's face image is corrected to determine a frontal view of the user's face image.
3. The method according to claim 2, characterized in that, The step of determining the user's three-dimensional rotation and translation information based on facial key points in the user's face image includes: Based on the facial key points in the user's face image and the corresponding key points in the standard 3D face model, the user's three-dimensional rotation and translation information are determined.
4. The method according to claim 2, characterized in that, The step of correcting the user's facial image based on the user's three-dimensional rotation and translation information to determine a frontal view of the user's facial image includes: The user's facial image is mapped to 3D space to determine the user's 3D facial model; Based on the user's 3D rotation and translation information, determine the viewpoint correction projection parameters of the user's 3D face model; Based on the user's 3D face model and the viewpoint correction projection parameters, a frontal view user face image is determined.
5. The method according to claim 1, characterized in that, Before determining the frontal view user face image by performing perspective correction on the user's face image based on facial key points in the user's face image, the process further includes: The acquired face image is subjected to noise reduction processing, brightness correction processing for overexposed areas, and / or illumination separation and reflection enhancement processing for shadow areas to determine the user's face image.
6. The method according to claim 5, characterized in that, The brightness correction process for the overexposed areas includes: According to the formula: Determine the correction index; Where γ is the correction index, M0 is the brightness reference, and M is the median brightness within a preset range around the current pixel; According to the formula: Brightness correction processing is performed on overexposed areas of the acquired image; Among them, I out To output pixel brightness, M max For maximum brightness, I in This is the input pixel brightness.
7. The method according to claim 5, characterized in that, The shadow area illumination separation and reflection enhancement processing includes: According to the formula: Determine the single-scale reflection components. Among them, R k (x,y) represents a single-scale component, I(x,y) represents the input value of the acquired image, and F... k (x,y) is a Gaussian wrapping function; According to the formula: Perform shadow area illumination separation and reflection enhancement processing on the acquired image data; Among them, R MSR (x,y) represents the multi-scale fusion component, R k (x,y) represents the single-scale reflection component, N is the number of scales, and ω k It is a single-scale weight.
8. An in-vehicle user emotion recognition device, characterized in that, include: A frontal view image determination device is used to perform perspective correction on a user's face image based on facial key points in the user's face image, and determine a frontal view user face image. The user emotion type determination device is used to perform emotion recognition based on the user's facial image from the frontal view to determine the user's emotion type.
9. A vehicle, characterized in that, include: A controller configured to perform the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1-7.