Voice image facial expression acquisition method based on visual large model

By using an edge-cloud collaborative architecture and a large visual model, combined with a multi-member selection logic based on gender and age weighted sorting and an adaptive transmission mechanism, the problem of uniform appearance and lagging expression-driven behavior in in-vehicle voice interaction systems has been solved. This has enabled high-fidelity, low-latency personalized voice image generation, improving the accuracy and robustness of in-vehicle human-computer interaction.

CN121686539APending Publication Date: 2026-03-17CHINA FAW CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-30
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing in-vehicle voice interaction systems suffer from monotonous appearances, delayed expression-driven interaction, and difficulty in accurately identifying the interaction target in complex scenarios with multiple passengers. Furthermore, fluctuations in network bandwidth lead to data transmission delays and packet loss, failing to meet the requirements for low-latency synchronization.

Method used

Adopting an edge-cloud collaborative architecture, it utilizes a large visual model for high-performance cognitive analysis in the cloud and in-vehicle terminal for real-time rendering. Combined with a multi-member selection logic based on gender and age weighting and an adaptive transmission mechanism, it ensures high-fidelity and low-latency voice image generation.

Benefits of technology

It achieves high-fidelity, low-latency voice image generation for specific interactive objects in an in-vehicle environment, improving personalization and interaction accuracy, and ensuring system robustness and smooth user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121686539A_ABST
    Figure CN121686539A_ABST
Patent Text Reader

Abstract

The invention discloses a voice image facial expression acquisition method based on a visual large model, and relates to the field of computer vision and intelligent cabin man-machine interaction. Uploading the collected video stream to a cloud server deployed with an artificial intelligence cognitive model based on a visual large model; receiving structured member information which is fed back by the cloud and contains member genders, ages and expression probabilities; executing multi-member optimization logic based on the information, and locking a target interaction object and extracting facial features by calculating a sorting score; and performing deformation on the parameterized model by utilizing facial features to construct a personalized grid, and converting real-time expression data into a weight to drive the grid to displace so as to generate a synchronous voice image. According to the method, accurate locking and high-fidelity expression driving of interaction objects in a complex scene are realized through an end-cloud collaboration and optimization strategy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer vision and intelligent cockpit human-computer interaction, and in particular to a voice avatar facial expression collection method based on a visual large model. BACKGROUND

[0002] Vehicle-mounted voice interaction systems gradually integrate visual avatars for enhancing the feedback perception of human-computer interaction. Existing vehicle-mounted avatar generation schemes mainly rely on local computing resources of vehicle-mounted terminals for processing. Limited by the allocation of computing resources and power consumption of vehicle-mounted hardware, the visual analysis algorithm deployed locally usually performs dimensionality reduction processing on facial feature points or adopts a simplified model. This processing method reduces the accuracy of facial feature extraction, making the generated voice avatar face grid deformation details missing, and the expression state difficult to map the subtle facial changes of users in real time. Moreover, the avatar model constructed has too high a degree of generalization of parameters, making it difficult to generate personalized avatars with differences for different users' facial features.

[0003] In addition, there are often multi-member co-riding scenarios in vehicle-mounted spaces. Existing visual capture systems often only filter according to distance or position information when multiple facial targets are detected within the field of view, lacking a multi-dimensional optimization strategy based on semantic attributes such as gender and age characteristics. This mechanism is prone to cause interaction object locking errors or frequent focus switching between different members in complex scenarios, leading to confusion in interaction logic. At the same time, if cloud computing power is directly introduced to improve analysis accuracy, network bandwidth fluctuates dynamically in a vehicle-mounted mobile network environment.

[0004] Conventional video stream transmission schemes lack adaptive adjustment mechanisms for bandwidth utilization, and are prone to cause data transmission congestion or packet loss when network conditions are limited, resulting in delay in the rendering and driving data of voice avatars, which cannot meet the real-time requirements of low-latency synchronization for vehicle-mounted human-computer interaction. SUMMARY

[0005] The present application provides a voice avatar facial expression collection method based on a visual large model, aiming to solve the technical problems of uniform avatars, lagging expression driving, and difficulty in accurately locking interaction objects in complex scenarios with multiple passengers in existing vehicle-mounted voice interaction systems.

[0006] The technical solution adopted by the present application is as follows:

[0007] The application discloses a voice image face expression collection method based on a visual large model, which comprises the following steps: first, a vehicle terminal controls an in-vehicle monitoring system camera to collect video stream data in response to a cloud cognitive analysis mode determined based on hardware permissions and network status, and uploads the video stream data to a cloud server deploying an artificial intelligence cognitive model based on a visual large model through a wireless communication unit; then, the vehicle terminal receives a structured member information list fed back by the cloud server, wherein the structured member information list is generated after the artificial intelligence cognitive model performs multi-task parallel processing on the video stream data, and contains the number of in-vehicle members, gender classification labels of each member, predicted age values and real-time expression category probabilities.

[0008] On this basis, the vehicle terminal performs multi-member optimization logic based on the structured member information list, locks a target interactive object from multiple in-vehicle members by calculating a sorting score combined with the gender classification labels and the predicted age values, and extracts a face feature vector and real-time expression data of the target interactive object. Finally, the vehicle terminal performs geometric deformation on a parameterized standard face model by using the face feature vector to construct a personalized static grid, converts the real-time expression data into expression basis weights to drive the personalized static grid to perform real-time vertex displacement, and generates and displays a voice image synchronized with the target interactive object on a display unit.

[0009] Further, in order to ensure the stability of the system, the process of determining the cloud cognitive analysis mode in the above method comprises the following steps: the vehicle terminal queries a permission state code corresponding to a hardware identifier of the in-vehicle monitoring system camera, and simultaneously detects a network registration state and a signal quality parameter of the wireless communication unit. When the permission state code indicates that the calling permission is turned on, the network registration state is connected, and the signal quality parameter is higher than a preset data transmission threshold, it is determined that the system enters the cloud cognitive analysis mode; otherwise, if the permission is not turned on or the network is disconnected, the system enters a local default mode and displays a default voice image data package stored locally.

[0010] Further, in order to ensure the stability of the system, the process of determining the cloud cognitive analysis mode in the above method comprises the following steps: the vehicle terminal queries a permission state code corresponding to a hardware identifier of the in-vehicle monitoring system camera, and simultaneously detects a network registration state and a signal quality parameter of the wireless communication unit. When the permission state code indicates that the calling permission is turned on, the network registration state is connected, and the signal quality parameter is higher than a preset data transmission threshold, it is determined that the system enters the cloud cognitive analysis mode; otherwise, if the permission is not turned on or the network is disconnected, the system enters a local default mode and displays a default voice image data package stored locally.

[0011] Further, regarding the generation of the structured member information list, the artificial intelligence cognitive model performs face target detection on the single-frame images of the video stream data and counts the number of valid detection boxes to determine the number of members. At the same time, the images in each valid detection box are feature-encoded to obtain a feature vector. The feature vector is respectively input into an age regression analysis branch network to output a predicted age value, input into an expression classification branch network to output a real-time expression category probability containing multiple basic expression probability values, and identify the corresponding gender characteristics to generate a gender classification label.

[0012] Further, the present application proposes a specific multi-member preference logic. The vehicle-mounted terminal calculates a ranking score for each in-vehicle member in the structured member information list. The ranking score is the weighted sum of the gender item score and the age item score. Specifically, the numerical mapping rule of the gender classification label is set such that the value corresponding to a female member is less than the value corresponding to a male member, and the gender weight coefficient used to calculate the gender item score is set to be greater than the product of the age weight coefficient used to calculate the age item score and a preset upper limit value of human age. The vehicle-mounted terminal performs a minimization search operation and selects the in-vehicle member with the smallest ranking score as the target interactive object.

[0013] Further, in order to ensure the coherence and robustness of expression-driven, after extracting the target interactive object data, the vehicle-mounted terminal checks the expression validity status identifier in the real-time expression data. If the identifier indicates invalid or no detection, a default expression control instruction pointing to a preset silent state parameter set is generated; if the identifier indicates valid, a real-time expression driving instruction is generated according to the real-time expression category probability.

[0014] Further, in constructing the personalized speech avatar geometry grid, the vehicle-mounted terminal analyzes the facial feature vector to obtain shape weight coefficients corresponding to a plurality of pre-set shape orthogonal bases. Then, the cumulative sum of the initial geometry shape vector of the parameterized standard face model and each shape orthogonal base multiplied by the corresponding shape weight coefficient is calculated, and the vertex coordinates of the parameterized standard face model are updated according to the calculation result.

[0015] Further, in driving the grid to change expressions, the vehicle-mounted terminal maps the real-time expression data into a target weight vector corresponding to a plurality of basic expression bases, and obtains the actual output weight vector of the last rendering frame. The target weight vector and the actual output weight vector are linearly interpolated using a preset smoothing coefficient to obtain a final application weight vector of the current rendering frame. Then, the displacement superposition value of each basic expression base multiplied by the corresponding final application weight vector is calculated, and the displacement superposition value is applied to the reconstructed personalized speech avatar geometry grid.

[0016] Further, in order to improve the visual authenticity, before display, the vehicle terminal extracts the average skin color value of the target interactive object from the video stream data and analyzes the red, green and blue color components, assigns the color components to the diffuse reflection color attribute of the personalized voice image geometric grid, and uses a programmable shading pipeline for rasterization processing to calculate the final pixel color value in combination with the lighting influence of the virtual light source.

[0017] The application also provides a computer device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above methods when executing the computer program.

[0018] The application realizes high-fidelity, low-delay voice image generation and expression driving of a specific interactive object in a vehicle environment by using the powerful feature extraction capability of a visual large model to obtain high-dimensional face information and combining specific optimization strategies and adaptive transmission mechanisms through an end-cloud collaborative architecture.

[0019] Through the above scheme, the following beneficial technical effects are obtained:

[0020] The application adopts an end-cloud collaborative processing architecture, deploys high-computing-power cognitive analysis tasks based on a visual large model in the cloud, and retains rendering driving tasks with high real-time requirements in the vehicle terminal. In this way, on the one hand, the powerful deep feature extraction capability of the visual large model is fully utilized, which can capture subtle facial expression changes and multi-dimensional attributes of in-vehicle members, improving the authenticity and personalization of the voice image; on the other hand, the computing burden of the vehicle terminal is effectively reduced, avoiding the decline of recognition accuracy or system lag caused by local hardware computing power bottlenecks, and achieving a balance between high-precision visual analysis and low-delay real-time rendering.

[0021] The application innovatively designs a multi-member optimization logic based on gender and age weighted sorting, which can quickly and accurately lock the target interactive object in a complex scene with multiple people in the vehicle by quantitatively scoring structured member information. Through specific numerical mapping and weight setting, this mechanism effectively solves the problems of target loss, frequent focus switching or misidentification of non-interactive personnel in the traditional vehicle vision system in a multi-face environment, ensuring that the generated voice image always keeps pace with the current core user, and greatly improving the accuracy and concentration of intelligent cockpit human-computer interaction.

[0022] This application constructs an adaptive link transmission guarantee mechanism. It monitors the uplink bandwidth utilization of the network in real time through the vehicle terminal and dynamically adjusts the quantization parameters of the video encoder based on the monitoring results to control the upload bitrate. This feature enables the system to proactively adapt to fluctuations in the vehicle mobile network environment, automatically reducing the data stream size to prevent congestion and packet loss when bandwidth is limited. This ensures the continuity and real-time nature of the video data required for cloud analysis, thereby guaranteeing the timeliness of expression-driven command issuance and enhancing the overall system robustness and user experience smoothness. Attached Figure Description

[0023] Figure 1 This is a flowchart illustrating a method for capturing voice-based facial expressions based on a large visual model, according to the present invention.

[0024] Figure 2 This is a schematic diagram illustrating the architectural principle of the interaction between the vehicle terminal and the cloud server and the multi-member selection logic of the present invention.

[0025] Figure 3 This is a schematic diagram of the personalized mesh construction and real-time expression-driven rendering process of the voice image of the present invention. Detailed Implementation

[0026] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] See attached document Figure 1 -Appendix Figure 3 This invention provides a method for capturing voice-based facial expressions based on a large visual model. The method is applied to a capture device, which includes an in-vehicle terminal, an in-vehicle monitoring system camera, a wireless communication unit, a cloud server, and a display unit. The in-vehicle terminal establishes electrical or data communication connections with the in-vehicle monitoring system camera, the wireless communication unit, and the display unit. The wireless communication unit establishes a remote communication connection with the cloud server.

[0028] The in-vehicle monitoring system cameras are installed in specific locations inside the vehicle, and their field of view covers the driver and passenger area. These cameras are configured to capture real-time image data and video stream data of the vehicle's interior environment, including facial images of the occupants. The raw video stream data is then transmitted to an onboard terminal for processing.

[0029] The vehicle-mounted terminal, acting as the local processing core, is configured to perform system initialization checks and data distribution. It internally stores default voice and image data. The terminal is configured to detect the access permission status of the in-vehicle monitoring system cameras and the vehicle's current network connection status. Based on these conditions, the terminal determines whether to invoke the default voice and image data or initiate a cloud-based analysis process. When access permission is enabled and the network connection is active, the terminal transmits the video stream data captured by the in-vehicle monitoring system cameras to the cloud server via its wireless communication unit.

[0030] An artificial intelligence (AI) cognitive model is deployed on the cloud server. This AI cognitive model is a large-scale visual model based on deep learning. It is configured to receive video stream data and perform frame-level or segment-level analysis. Specifically, the AI ​​cognitive model performs feature extraction operations, extracting information such as the number of occupants in the vehicle, facial expressions of occupants at each location, and age information of occupants at each location from the video stream data.

[0031] The AI ​​cognitive model or in-vehicle terminal is equipped with a member selection strategy logic. When the AI ​​cognitive model identifies two or more occupants, the member selection strategy logic filters based on the age and gender characteristics of the occupants at each location. The filtering rule is set to prioritize female occupants with the lowest age among the multiple occupants as the target occupants. The cloud server then transmits the extracted facial feature data and real-time expression data of the target occupants back to the in-vehicle terminal.

[0032] The in-vehicle terminal includes a rendering engine. The rendering engine is configured to receive facial feature data and real-time expression data of the target member. Using the facial feature data, the rendering engine generates a voice-image facial model that resembles the target member's face. The rendering engine further uses real-time expression data to drive the facial mesh or skeleton of the voice-image facial model, enabling the voice-image facial model to present a visual effect consistent with the target member's current expression. The display unit is configured to display the rendered voice-image facial model in real-time on the human-computer interaction interface.

[0033] The present invention provides a method for capturing voice image facial expressions based on a large visual model, which performs a strict environmental self-check procedure during the system startup phase.

[0034] During vehicle ignition and startup or the power-on initialization of the in-vehicle infotainment system, the vehicle terminal first performs a hardware access permission verification step for the in-vehicle monitoring system camera. The vehicle terminal initiates a query request to the vehicle operating system's permission management service, which includes the hardware identifier of the in-vehicle monitoring system camera. The vehicle terminal receives the permission status code returned by the permission management service. The vehicle terminal compares the received permission status code with a preset authorization status code.

[0035] If the permission status code does not match the preset authorization status code, it indicates that the access permission for the in-vehicle monitoring system camera is not enabled or has been disabled by the user. In this state, the vehicle terminal executes the default image loading instruction. The vehicle terminal reads the pre-stored default voice image data package from local non-volatile memory. The default voice image data package contains standardized 3D face model data and basic texture map data. The vehicle terminal loads the default voice image data package into the graphics rendering pipeline and presents the default voice image with fixed appearance characteristics on the display unit.

[0036] If the permission status code matches the preset authorization status code, it indicates that access permission for the in-vehicle monitoring system camera has been granted. The vehicle terminal then triggers the network connection status detection step. The vehicle terminal sends a link status query command to the wireless communication unit. The wireless communication unit detects the current connection status of the cellular mobile network or wireless local area network and detects the physical link connectivity of the data transmission channel.

[0037] The vehicle-mounted terminal determines whether the conditions for cloud interaction are met based on the signal quality parameters and network registration status fed back by the wireless communication unit. Only when the network registration status shows "connected" and the signal quality parameters are higher than the preset data transmission threshold, the vehicle-mounted terminal determines that the current network environment meets the requirements for video stream upload. If the network registration status shows "disconnected," or the signal quality parameters are lower than the preset data transmission threshold, the vehicle-mounted terminal determines that the network connection is unavailable. In the state of unavailable network connection, the vehicle-mounted terminal blocks the video stream data upload channel and instead executes the default image loading command, calling the local default voice image data packet for display.

[0038] To accurately describe the above initialization detection logic, the system's operating mode is set as follows: Set the permission status variable for the in-vehicle monitoring system camera to... ,in This indicates that permission is enabled. This indicates that permissions are disabled. The network connection status variable is set to... ,in 1 indicates that the network is connected and stable. This indicates a network disconnection or instability. System operating mode. The decision logic follows the Boolean operation formula as follows: The specific judgment rules are as follows:

[0039] ;

[0040] in, This represents the logical AND operation. Represents a logical OR operation; This indicates the cloud-based cognitive analysis mode, in which the system activates the in-vehicle monitoring system camera to capture video and upload it to the cloud server. This indicates the local default mode. In this mode, the system directly calls the default voice image resources stored locally on the vehicle terminal. Through the above logic, the system ensures that the subsequent visual large model analysis process will only be started when both hardware permissions and network environment are ready, so as to avoid invalid calculation requests or program errors in the absence of network or permissions.

[0041] After the vehicle terminal confirms that the operating mode is cloud-based cognitive analysis mode, it sends a video capture activation command to the in-vehicle monitoring system camera. The in-vehicle monitoring system camera responds to the video capture activation command, activating its internal optical sensor components and image signal processing circuitry.

[0042] The in-vehicle monitoring system camera performs continuous optical imaging of the vehicle's interior space according to pre-set image acquisition parameters. These parameters include, but are not limited to, the number of horizontal pixels, the number of vertical pixels, the frame sampling frequency, and pixel depth. In a preferred embodiment, the in-vehicle monitoring system camera sets the number of horizontal pixels to 1920, the number of vertical pixels to 1080, and the frame sampling frequency to 30 frames per second. The raw video signal generated by the in-vehicle monitoring system camera contains continuous time-series images, with each frame recording the facial texture features, facial geometric features, and ambient lighting information of the vehicle occupants.

[0043] To reduce wireless transmission bandwidth consumption and improve transmission efficiency, in-vehicle monitoring system cameras or vehicle terminals are equipped with video encoders. The video encoders use efficient video coding standards (such as H.264 or H.265) to perform inter-frame and intra-frame compression on the original video signal, generating a compressed video bitstream. The vehicle terminal divides the compressed video bitstream into multiple data payload units and adds a header information to each data payload unit. The header information includes the vehicle's unique identification code, timestamp synchronization information, and data packet sequence number.

[0044] The vehicle-mounted terminal controls the wireless communication unit to establish an encrypted communication tunnel with the data access gateway of the cloud server. The vehicle-mounted terminal then uploads the encapsulated video stream data packets to the cloud server through this encrypted communication tunnel. During the upload process, the vehicle-mounted terminal implements a dynamic bitrate adjustment strategy to adapt to fluctuations in network bandwidth.

[0045] Set the available uplink bandwidth of the network at the current moment to (Unit: Mbps), the real-time bitrate of the video stream is (Unit: Mbps). The vehicle-mounted terminal periodically calculates bandwidth utilization. : The vehicle-mounted terminal has a preset bandwidth utilization safety threshold. .

[0046] Vehicle-mounted terminals based on bandwidth utilization With bandwidth utilization safety threshold The comparison results are used to adjust the quantization parameters of the video encoder. Quantization parameters Real-time bitrate of the video stream There is a negative correlation. The adjustment logic is as follows:

[0047] ;

[0048] ;

[0049] in, It is a proportional adjustment coefficient; Adjust the step size for quantization parameters; The current quantization parameter; These are the adjusted quantization parameters. When bandwidth utilization is detected... Exceeding the bandwidth utilization safety threshold At that time, the vehicle terminal increases the quantization parameters. This reduces the real-time bitrate of the video stream. This prevents video stream data packets from becoming congested or lost during transmission. Through this dynamic transmission mechanism, the system ensures that the cloud server receives continuous and complete video stream data, providing a stable data source for subsequent artificial intelligence cognitive analysis.

[0050] After receiving the video stream data uploaded by the vehicle terminal, the cloud server decodes the video stream data into a sequence of single-frame images and inputs the sequence of single-frame images into an artificial intelligence cognitive model deployed on the cloud server. The artificial intelligence cognitive model uses a deep convolutional neural network or a large visual model structure based on the Transformer architecture to perform multi-task parallel processing on the input single-frame image sequence.

[0051] The AI ​​cognitive model first performs a face detection task. It extracts features from a single frame of image, generating a high-dimensional feature map. The model then generates multiple candidate regions on the feature map and calculates a confidence score for each region indicating that it contains a face. Using a non-maximum suppression algorithm, the model removes redundant candidate regions with an overlap rate exceeding a preset threshold, retaining regions with confidence scores higher than a valid detection threshold as the final face detection boxes. Finally, the model counts the number of retained face detection boxes to output information about the number of occupants in the vehicle.

[0052] To accurately describe the validity determination of member information, we define that for the first... The set of candidate targets output by the AI ​​cognitive model for the frame image is as follows: For each candidate target in the set The artificial intelligence cognitive model calculates its confidence score. Set the effective detection threshold to... Number of vehicle occupants effectively detected. The following counting formula is used to derive:

[0053] ;

[0054] in, For indicator functions, when the condition The function value is 1 if it is true, and 0 otherwise; if the calculated number of occupants in the vehicle... A value of 0 indicates that the AI ​​cognitive model has not extracted any valid facial information from the current video frame. In this case, the cloud server returns an empty result command to the in-vehicle terminal, triggering the in-vehicle terminal to maintain or switch to the default voice avatar.

[0055] When the number of occupants in the car When the value is greater than or equal to 1, the AI ​​cognitive model performs key point localization and attribute analysis on the image region within each valid face detection box. The AI ​​cognitive model locates the key points of the facial features and maps the face region image into a fixed-dimensional facial feature vector.

[0056] For extracting age information for members at each location, the AI ​​cognitive model inputs facial feature vectors into an age regression analysis branch network. This branch network outputs the predicted age for that member. For extracting facial expression information for members at each location, the AI ​​cognitive model inputs facial feature vectors into an expression classification branch network. This branch network analyzes the micro-motion units of facial muscles and outputs a probability distribution vector of the member's current expression state.

[0057] Setting the first The facial feature vectors of the effective members are The member's predicted age and facial expression state vector Represented as: ;

[0058] ;

[0059] in, This represents the age regression mapping function. This represents the facial expression classification mapping function. It is a vector containing probability values ​​for various basic emotional expressions (such as joy, sadness, anger, surprise, calmness, etc.). The artificial intelligence cognitive model selects... The expression category with the highest probability value is selected as the member's current expression. Simultaneously, the AI ​​cognitive model extracts the member's gender characteristics for subsequent member selection strategies. The cloud server sends the extracted structured data packet, containing member ID, predicted age, gender identifier, and expression status label, back to the in-vehicle terminal.

[0060] After the AI ​​cognitive model completes the detection and feature extraction of the occupants in the vehicle, the system enters the decision-making stage of determining the target interaction object. This stage is executed by the logic processing unit inside the cloud server or the in-vehicle terminal, aiming to solve the problem of voice image attribution in the presence of multiple people and achieve personalized interaction logic for each individual.

[0061] The logic processing unit first reads the structured member list output by the artificial intelligence cognitive model. This list contains a unique identity index, gender classification label, and predicted age for each detected occupant. The logic processing unit then determines the validity of each occupant within the structured member list. .

[0062] When the number of valid members When the value is equal to 1, the logic processing unit directly marks the unique member as the target interaction object and locks the member's facial feature data for subsequent image rendering.

[0063] When the number of valid members When the number is greater than or equal to 2, the logic processing unit activates the multi-member selection algorithm. The multi-member selection algorithm is configured to execute a hierarchical selection strategy based on gender priority and age values. The technical objective of this strategy is to preferentially select members with younger female characteristics as the mapping source for the speech image.

[0064] To quantify the screening process, we define the first member in the structured member list... The attribute vector of each member is ;

[0065] in, Indicates gender category labels, Represents the predicted age value. Defines gender category labels. The numerical mapping rule is as follows:

[0066] When a member is identified as female ;

[0067] When a member is identified as male .

[0068] The logic processing unit calculates the optimal sorting score for each member in the list. Optimal sorting score The calculation formula is as follows:

[0069] ;

[0070] in, This is the gender weighting coefficient. This is the age weighting coefficient. To ensure that gender characteristics have absolute screening priority, that is, to ensure that all female members have a higher priority than all male members, it is necessary to set... The value is far greater than the maximum contribution that the age factor can easily generate. Let's assume the theoretical maximum age of the vehicle occupants is... (For example, 100 years old), then the weighting coefficients must meet the following constraints:

[0071] ;

[0072] In one specific embodiment, set , The logic processing unit calculates the optimal ranking score for all members. Then, a minimum search operation is performed to select the member with the lowest preferred ranking score as the target interaction object. The index of the target interaction object is then defined. The following is confirmed:

[0073] ;

[0074] in, The number of valid members; Calculate the optimal ranking score for all members.

[0075] Based on the above calculation logic, when both female and male occupants are present in the vehicle, due to the female occupants'... Their base score was much lower than that of the male members. The system will prioritize selection from the female member set; within the same-sex member set, because... A positive number, the age value The smaller the value, the higher the total score. The lower the age, the more likely the system will target the youngest member.

[0076] Once the target interaction object is identified, the logic processing unit isolates and outputs the expression state label and facial feature vector of the target interaction object, discarding data of other non-target members to reduce the computational load on the subsequent rendering engine. If the expression state label of the target interaction object indicates that no valid expression was detected or the expression confidence is below the threshold, the logic processing unit will generate a default expression control instruction; if the expression state label of the target interaction object indicates a valid expression category, the logic processing unit will generate the corresponding real-time expression driving instruction.

[0077] After the AI ​​cognitive model completes the detection and feature extraction of the occupants in the vehicle and generates a structured dataset containing information on all valid occupants, the cloud server or in-vehicle terminal executes the decision logic to determine the target interaction object. This decision logic aims to solve the problem of which occupant's facial features and expression should be mapped to when there are multiple occupants in the vehicle.

[0078] The cloud server or vehicle terminal first reads the number of valid members parameter from the structured data set. When the number of valid members parameter is equal to 1, the cloud server or vehicle terminal directly marks this unique member as the target interaction object and extracts the facial feature data and expression information of the target interaction object for subsequent processing.

[0079] When the number of valid members is greater than or equal to 2, the cloud server or vehicle terminal activates the multi-member selection algorithm. This algorithm is configured to execute a hierarchical selection strategy based on gender priority and age. The technical logic of this strategy is to prioritize female members with the lowest age among multiple candidate members as the target interaction object.

[0080] To implement the digital calculation of this screening strategy, the structured data set is set to contain... A set has members. For any member in the set Define its gender characteristic value as Define its age characteristic value as .

[0081] Gender characteristics The numerical mapping rules are set as follows: when the artificial intelligence cognitive model identifies the member as female, The value is assigned to 0; when the artificial intelligence cognitive model identifies the member as male, The value was assigned to 1. Age characteristic value. The predicted age value is directly obtained from the output of the artificial intelligence cognitive model.

[0082] Cloud servers or in-vehicle terminals use a weighted sorting formula to calculate each member. Optimal ranking score The weighted sorting formula is defined as follows:

[0083] ;

[0084] in, This is the gender weighting coefficient. Age-weighted coefficient This is a preset theoretical upper limit for human age (e.g., 100). To ensure that the gender selection criteria have absolute priority, that is, to ensure that the preferred ranking score of all female members is necessarily higher than (lower than) the preferred ranking score of any male member, the weighting coefficients must satisfy the inequality constraint: .

[0085] In a specific computational implementation, set , Under this parameter setting, female members ( The score for male members is determined solely by the age item, which ranges from [0, 1]; while the score for female members is determined solely by the age item. The score range is between [100, 101]. The cloud server or vehicle terminal iterates through the computation set. The optimal ranking score of all members And perform a minimized lookup operation to determine the index of the target interactive object. : .

[0086] Through the above calculations, the system can accurately filter out target interaction objects that meet the rule of prioritizing younger women. In extreme cases where there are identical scores (e.g., two women of the same age), the system can perform secondary sorting based on the horizontal coordinate position of their detection boxes in the image or their detection confidence.

[0087] Determine the index of the target interaction object Subsequently, the cloud server or vehicle terminal isolates the first data from the structured dataset. Facial feature data and expression status information of each member. For the set The middle index is not equal to Other member data is discarded or archived by the system and not transmitted to the rendering engine, thereby reducing data transmission load and subsequent computational overhead for graphics rendering. Finally, the system outputs a uniquely identified target member data packet to the voice-image generation module.

[0088] The voice image generation scheme provided by this invention adopts parametric three-dimensional face reconstruction technology to construct a virtual image in real time based on facial feature data sent by the cloud server.

[0089] The rendering engine integrated within the vehicle terminal first receives the facial feature vector of the target member, extracted and transmitted by an artificial intelligence cognitive model. This facial feature vector is a high-dimensional numerical vector that encodes the facial geometric topology, proportional distribution of facial features, and basic facial texture features of the target member. The rendering engine has a pre-built standard benchmark face model. The standard benchmark face model consists of a quantitative set of three-dimensional vertices and triangular facet indices connecting these vertices, representing the general average geometric shape of a face.

[0090] The rendering engine performs shape-based mesh deformation operations to generate a personalized speech avatar geometry mesh. The engine stores multiple pre-trained shape orthogonal basis vectors, each representing a specific dimension of facial geometry (e.g., face width, nose length, eye spacing). The rendering engine resolves the target member's facial feature vectors into shape weight coefficients corresponding to each shape orthogonal basis vector.

[0091] To accurately describe the mesh deformation process, the geometric shape vector of the standard baseline face model is set as follows: It contains The three-dimensional coordinates of each vertex, i.e. The rendering engine settings store... An orthogonal basis of shape, denoted as Each of them Let the shape weight coefficient vector obtained from the analysis be... .

[0092] Geometric grid shape for personalized voice avatars The following can be calculated using the linear combination formula: .

[0093] By adjusting the shape weight coefficient The value of the value causes the rendering engine to drive the vertices of the standard baseline face model to shift, thereby making By approximating the real facial features of the target members in terms of geometric contours, a structured model that is personalized for each individual can be achieved.

[0094] After constructing the geometric mesh, the rendering engine performs texture mapping and material rendering operations. Based on the age and gender characteristics extracted from the AI ​​cognitive model, the rendering engine selects the corresponding skin material shader parameters. The rendering engine extracts the red, green, and blue color components of the target member's average facial skin tone from the raw video stream data. This red, green, and blue color component is then assigned to the diffuse color attribute.

[0095] The rendering engine utilizes a programmable shading pipeline to shape the geometric mesh. Rasterization is then performed. During rasterization, the rendering engine calculates the lighting effect of virtual light sources on the mesh surface, and combines diffuse color attributes, specular reflection intensity, and ambient occlusion parameters to calculate the final color value of each screen pixel. Finally, the display unit outputs the rasterized rendered 2D image frame to the in-vehicle human-machine interface, presenting a digital voice image with appearance features highly similar to the target occupants inside the vehicle.

[0096] In the voice-based facial expression acquisition system based on a large visual model, the expression-driven and synchronization logic is executed by the rendering engine inside the vehicle terminal. The aim is to map the real member expressions recognized by the cloud server to the virtual voice-based facial model in real time, so as to achieve emotional interaction and resonance in the visual dimension.

[0097] The rendering engine periodically parses real-time facial expression information of target members from received data packets. This real-time facial expression information includes an expression validity status flag and an expression category probability vector. The rendering engine first checks the expression validity status flag.

[0098] When the expression validity status indicator is invalid or not detected, or when the system experiences data packet loss due to network latency, the rendering engine executes the default expression reset logic. The rendering engine smoothly transitions the expression control parameters of the voice-image facial model to a preset set of silent state parameters. This preset set of silent state parameters corresponds to a neutral expression with relaxed facial muscles or a standby expression with subtle breathing animation. This logic ensures that the voice-image does not exhibit facial distortion or a frozen state during data interruptions or recognition blind spots.

[0099] When the expression validity status indicator is valid, the rendering engine executes real-time expression mapping logic. The rendering engine employs a linear drive technique based on hybrid deformation. The voice-image face model is predefined. There are 1 basic expression base, and each basic expression base represents an independent local deformation of facial muscles (e.g., the left corner of the mouth is raised, the eyebrows are raised, the jaw is opened, etc.).

[0100] The rendering engine converts the emoji category probability vectors sent from the cloud server into corresponding... Real-time facial weight vectors based on a base of basic facial expressions. (Settings) The basic expression base is Each of them All are 3D translation vectors with the same number of vertices as the voice-image face model. The real-time expression weight vector is set to... ,in .

[0101] To prevent visual abruptness or flickering caused by excessive facial expression changes between adjacent video frames, the rendering engine introduces a temporal linear interpolation smoothing algorithm. The target weight vector for the current frame is set as... The actual output weight vector of the previous rendered frame is The final applied weight vector for the current rendered frame. The calculation is as follows:

[0102] ;

[0103] in, This is the smoothing coefficient, and its value ranges from (0, 1). The smaller the value, the smoother the facial expression change; The higher the value, the more sensitive the facial expression response. In a preferred embodiment, Set to 0.6.

[0104] The final applied weight vector is calculated. Then, the rendering engine calculates and overlays the final face mesh shape after the surface changes. Combining the static geometric shapes of the aforementioned image modeling section... The final vertex position is calculated using the following formula: .

[0105] The rendering engine uses the calculated The vertex buffer data in the video memory is updated and submitted to the graphics processor for rendering. Through the above process, if the occupants of the vehicle exhibit a laughing expression, the artificial intelligence cognitive model identifies this category, and the rendering engine immediately increases the weight of the corresponding open mouth and upturned corners of the mouth as the base values. This allows the voice-image facial model to simultaneously display a laughing visual effect on the screen, thus achieving precise synchronization between text-to-image conversion and dynamic expressions.

[0106] The present invention provides a voice image facial expression acquisition system based on a large visual model, which includes an initialization detection module, a data acquisition and transmission module, a visual cognition analysis module, a strategy processing module, and a rendering driver module.

[0107] The initialization detection module establishes a communication interface with the underlying hardware abstraction layer of the vehicle operating system. It is configured to periodically poll the hardware permission status register of the in-vehicle monitoring system camera and read the network link status parameters of the wireless communication unit. Based on the read hardware permission status register values ​​and network link status parameters, the initialization detection module generates a system operating mode control signal. When the hardware permission status register value indicates permission is disabled, or the network link status parameters indicate network unavailability, the initialization detection module outputs a local default mode signal. When the hardware permission status register value indicates permission is enabled, and the network link status parameters indicate network connectivity, the initialization detection module outputs a cloud analysis mode signal.

[0108] The data acquisition and transmission module connects to the in-vehicle monitoring system's camera and wireless communication unit. Responding to the cloud-based analysis mode signal, the data acquisition and transmission module activates the in-vehicle monitoring system's camera to record video streams. The data acquisition and transmission module integrates a video compression encoding unit and a traffic control unit. The traffic control unit monitors the network's uplink bandwidth utilization in real time and dynamically adjusts the quantization parameters of the video compression encoding unit based on the bandwidth utilization. The data acquisition and transmission module then sends the compressed video stream data packets to the wireless communication unit.

[0109] The visual cognition analysis module is configured to acquire deep learning analysis results for video stream data packets. In one implementation, the visual cognition analysis module receives a structured list of occupants returned by a cloud server. This structured list includes statistics on the number of occupants in the vehicle, a gender classification label for each occupant, predicted age values, and real-time probability distributions of facial expression categories.

[0110] The strategy processing module is connected to the visual cognition analysis module. The strategy processing module is configured to execute target object selection logic in multi-member scenarios. It reads a structured list of member information. When the number of members in the vehicle is greater than or equal to two, the strategy processing module iterates through and calculates the selection ranking score for each member. The strategy processing module applies a preset weighted algorithm to map gender classification labels and predicted age values ​​into scalar scores. Based on the minimum score criterion, the strategy processing module identifies the target interaction object and outputs the facial feature data and real-time expression status data of the target interaction object.

[0111] The rendering driver module is connected to the display unit. Internally, it stores a parametric standard face model and a set of variable basis vectors. The rendering driver module receives facial feature data and real-time expression data of the target interactive object. It uses the facial feature data to calculate geometric weight coefficients and transforms the parametric standard face model into a personalized static mesh similar to the target interactive object's face. Finally, it uses the real-time expression data to calculate expression blending weight coefficients and performs vertex displacement superposition on the personalized static mesh.

[0112] The final output of the rendering driver module shows the vertex positions of the dynamic 3D model. Determined by the following superposition logic:

[0113] ;

[0114] in, To determine the initial vertex positions for a parameterized standard face model, The first determined by facial feature data The weight of a shape base, For the first Vertex displacement vectors corresponding to each shape base; The first determined by real-time facial expression state data The weight of each facial expression base, For the first Each expression base corresponds to a vertex displacement vector. The rendering driver module submits the calculated dynamic 3D model to the graphics processing unit for rendering, and presents the final voice image on the display unit.

[0115] The method and device for capturing voice-based facial expressions based on a large visual model provided by this invention can be specifically implemented as an in-vehicle electronic device. The in-vehicle electronic device, at the hardware physical level, includes a central processing unit, system memory, communication bus, camera interface circuit, wireless communication module, and display driver circuit.

[0116] The central processing unit, system memory, camera interface circuit, wireless communication module, and display driver circuit are interconnected via a communication bus. The communication bus is configured to transmit control signals, address signals, and data signals between the aforementioned hardware components.

[0117] The system memory includes volatile random access memory and non-volatile storage media. The non-volatile storage media stores computer-executable instructions, operating system software, and default voice image data resources. The computer-executable instructions are configured to, when loaded and executed by the central processing unit, cause the central processing unit to execute the logical steps of the voice image facial expression acquisition method based on the visual large model described in the foregoing embodiments, including initialization detection, video stream data processing, member selection strategy calculation, and expression-driven control.

[0118] The central processing unit (CPU) is the core of computing and control for in-vehicle electronic devices. The CPU integrates a general-purpose processor core and a graphics processing core. The general-purpose processor core is responsible for performing logical judgments, data parsing, and system scheduling tasks. The graphics processing core is configured to execute 3D graphics rendering pipeline operations, using received facial feature data and expression weight coefficients to perform vertex transformations, lighting calculations, and texture mapping on the voice-based facial model.

[0119] The camera interface circuit provides a physical connection port (such as a MIPICSI interface or a USB interface) for the in-vehicle monitoring system camera. The camera interface circuit is configured to receive the raw video signal captured by the in-vehicle monitoring system camera, perform analog-to-digital conversion or format decapsulation on the raw video signal, and store the processed digital video frames in the system memory's video buffer for reading by the central processing unit or transmission via the wireless communication module.

[0120] The wireless communication module includes a radio frequency transceiver and a baseband processor, supporting cellular mobile communication protocols (such as LTE, 5G NR) or wireless local area network protocols (such as Wi-Fi 6). Under the control of the central processing unit, the wireless communication module is configured to establish a remote data transmission link between the vehicle's electronic devices and the cloud server, enabling uplink transmission of video stream data and downlink reception of cloud analysis results.

[0121] The display driver circuit establishes a physical connection with the human-machine interface display screen inside the vehicle. The display driver circuit is configured to receive frame buffer data output from the graphics processing core, convert the frame buffer data into scanning signals and data signals to drive the pixel array of the human-machine interface display screen, thereby displaying a voice image that dynamically adjusts to changes in the facial features and expressions of the occupants in real time on the human-machine interface display screen.

[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for collecting facial expressions of a voice avatar based on a visual large model, characterized in that, The method comprises the following steps: S1, the vehicle terminal determines the cloud cognitive analysis mode in response to the hardware permission and the network state, controls the in-vehicle monitoring system camera to collect video stream data, and uploads the video stream data to a cloud server deploying an artificial intelligence cognitive model based on a visual large model through a wireless communication unit; S2, the vehicle terminal receives a structured member information list fed back by the cloud server, the structured member information list is generated after the artificial intelligence cognitive model performs multi-task parallel processing on the video stream data, and contains the number of in-vehicle members, gender classification labels of each member, predicted age values, and real-time expression category probabilities; S3, the vehicle terminal executes a multi-member optimization logic based on the structured member information list, locks a target interactive object from multiple in-vehicle members by calculating a sorting score combined with the gender classification labels and the predicted age values, and extracts a face feature vector and real-time expression data of the target interactive object; S4, the vehicle terminal performs geometric deformation on a parameterized standard face model using the face feature vector to construct a personalized static grid, and converts the real-time expression data into expression base weights to drive the personalized static grid to perform real-time vertex displacement, thereby generating and displaying a voice image synchronized with the target interactive object.

2. The voice image face expression collection method based on a visual large model according to claim 1, characterized in that, The cloud cognitive analysis mode in the S1 step comprises: The vehicle terminal queries a permission status code corresponding to a hardware identifier of the in-vehicle monitoring system camera, and detects a network registration state and a signal quality parameter of the wireless communication unit; If the permission status code indicates that the calling permission is turned on, and the network registration state is connected and the signal quality parameter is higher than a preset data transmission threshold, the vehicle terminal determines that the system enters the cloud cognitive analysis mode; If the permission status code indicates that the calling permission is not turned on, or the network registration state is disconnected, the vehicle terminal determines that the system enters a local default mode, and calls a default voice image data package stored locally for display.

3. The voice image face expression collection method based on a visual large model according to claim 1, characterized in that, The S1 step comprises: The vehicle terminal monitors the network uplink bandwidth utilization rate in real time, and the network uplink bandwidth utilization rate is a ratio of a real-time code rate of the video stream data to a network available uplink bandwidth; When the network uplink bandwidth utilization rate exceeds a preset bandwidth utilization rate safety threshold, the vehicle terminal calculates a quantization parameter adjustment step according to a difference between the network uplink bandwidth utilization rate and the bandwidth utilization rate safety threshold; The vehicle terminal increases the quantization parameter of the video encoder using the quantization parameter adjustment step to reduce the real-time code rate of the video stream data, and sends the compressed video stream data to the cloud server.

4. The voice image face expression collection method based on a visual large model according to claim 1, characterized in that, The generation process of the structured member information list in the S2 step comprises: The artificial intelligence cognitive model performs face target detection on a single frame image of the video stream data, and counts the number of valid detection boxes to generate the number of in-vehicle members; The artificial intelligence cognitive model encodes the features of the image in each valid detection box to obtain a feature vector; The artificial intelligence cognitive model inputs the feature vector into an age regression analysis branch network to output the predicted age value, and inputs the feature vector into an expression classification branch network to output the real-time expression category probability including multiple basic expression probability values; The artificial intelligence cognitive model identifies the gender feature corresponding to the feature vector to generate the gender classification label.

5. The voice image face expression collection method based on a visual large model according to claim 1, characterized in that, The S3 step includes: The vehicle terminal calculates the ranking score for each in-vehicle member in the structured member information list, which is a weighted sum of the gender item score and the age item score; The numerical value mapping rule of the gender classification label is set to be smaller for female members than for male members, and the gender weight coefficient used to calculate the gender item score is set to be greater than the product of the age weight coefficient used to calculate the age item score and a preset upper limit of human age in theory; The vehicle terminal performs a minimization search operation to select the in-vehicle member with the smallest ranking score as the target interactive object.

6. The voice image face expression collection method based on a visual large model according to claim 1, characterized in that, The S3 step also includes: The vehicle terminal isolates the data of the target interactive object from the structured member information list and discards the member data of non-target interactive objects; The vehicle terminal checks the expression validity status identifier in the real-time expression data; If the expression validity status identifier indicates invalidity or no detection, the vehicle terminal generates a default expression control instruction pointing to a preset silent state parameter set; If the expression validity status identifier indicates validity, the vehicle terminal generates a real-time expression driving instruction according to the real-time expression category probability.

7. The voice image face expression collection method based on a visual large model according to claim 1, characterized in that, The S4 step includes: The vehicle terminal analyzes the face feature vector to obtain the shape weight coefficient corresponding to a plurality of pre-set shape orthogonal bases; The vehicle terminal calculates the cumulative sum of the initial geometry shape vector of the parameterized standard face model and each shape orthogonal base multiplied by the corresponding shape weight coefficient; The vehicle terminal updates the vertex coordinates of the parameterized standard face model according to the calculation result to generate the personalized voice avatar geometry grid.

8. The voice image face expression collection method based on a visual large model according to claim 1, characterized in that, The S4 step includes: The vehicle terminal maps the real-time expression data into a target weight vector corresponding to a plurality of basic expression bases; The vehicle terminal obtains the actual output weight vector of the previous rendering frame and performs linear interpolation calculation on the target weight vector and the actual output weight vector using a preset smoothing coefficient to obtain the final application weight vector of the current rendering frame; The vehicle terminal calculates the displacement superposition value of each basic expression base multiplied by the corresponding final application weight vector and applies the displacement superposition value to the reconstructed personalized voice avatar geometry grid.

9. The voice image face expression collection method based on a visual large model according to claim 1, characterized in that, Before presenting the voice avatar synchronized with the target interactive object on the display unit in the S4 step, the S4 step also includes: The vehicle terminal extracts the face average skin color value of the target interactive object from the video stream data and analyzes the red, green, and blue color components; The vehicle terminal assigns the red, green, and blue color components to the diffuse reflection color attribute of the personalized voice avatar geometry grid; The vehicle terminal rasterizes the personalized voice image geometric grid by using a programmable shading pipeline, and calculates pixel color values in combination with the illumination influence of a virtual light source.

10. The voice image face expression collection method based on a visual large model according to claim 3, characterized in that, The S1 step further includes the following before the vehicle terminal sends the compressed video stream data to the cloud server: The vehicle terminal divides the compressed video stream data into multiple data payload units, and adds packet header information to each data payload unit. The packet header information includes a vehicle unique identification code, timestamp synchronization information, and a data packet sequence number.