AI generation method and system of vehicle-mounted digital human, electronic equipment and storage medium

By generating digital human images in the cloud and rendering them on the vehicle, the problem of fixed digital human images in vehicles failing to meet diverse user needs has been solved, thus achieving personalized digital human generation and efficient interaction.

CN121600153APending Publication Date: 2026-03-03CHINA FAW CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511782825.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-30
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing in-vehicle digital humans use a fixed design and cannot flexibly generate different forms, making it difficult to meet the diverse needs of users.

Method used

By acquiring user interaction input information from the vehicle, a personalized 3D digital human model is generated using a cloud-based AI image generation model and 3D reconstruction network. This model is then stored and rendered on the vehicle, supporting incremental updates.

Benefits of technology

It enables personalized in-vehicle digital human generation, reduces the computing load on the in-vehicle end, improves the interaction response speed and stability, and simplifies the update operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600153A_ABST
    Figure CN121600153A_ABST
Patent Text Reader

Abstract

The invention discloses an AI generation method and system of a vehicle-mounted digital person, electronic equipment and a storage medium, and relates to the technical field of vehicle-mounted digital persons, and the method comprises the steps: obtaining user interaction input information containing a voice instruction, image input and a digital person in a vehicle-mounted system, and generating interaction auxiliary information; sending the image to a cloud end, and generating digital human multi-angle images, textures and feature pattern data through an AI image generation model; a grid model is generated through a three-dimensional reconstruction network, a standard skeleton system and expression driving parameters are bound, driving association data is generated, and a drivable digital human model is obtained; the model and matched data are sent to a vehicle-mounted terminal for storage, and cloud incremental updating is supported; the vehicle-mounted end generates an action and expression control instruction driving model based on user interaction input and real-time vehicle control information; and finally rendering in real time and displaying on a vehicle-mounted screen. According to the method, the personalized vehicle-mounted digital human is generated according to user input, and the problems that a traditional digital human image is fixed and diverse requirements of users are difficult to meet are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle-mounted digital human technology, and in particular to AI generation methods, systems, electronic devices and storage media for vehicle-mounted digital humans. Background Technology

[0002] With the continuous evolution of smart cockpit technology, the intelligence level of in-vehicle human-machine interaction is constantly improving. Digital humans, possessing human-like interactive capabilities, are gradually becoming the core carrier connecting users and in-vehicle systems. Currently, in-vehicle digital humans are widely used in scenarios such as navigation guidance, voice dialogue, and vehicle status prompts. The industry generally uses AI technology, 3D modeling, and real-time rendering technology to build the basic capabilities of digital humans. Some automakers have already deployed digital human functions in multi-screen interconnected models, enabling digital humans to respond to user voice commands and visually display basic information, such as providing navigation and turning prompts and broadcasting vehicle energy consumption data through preset digital human images.

[0003] Existing digital humans use fixed image designs and cannot flexibly generate different forms, which makes it difficult to meet the diverse needs of users.

[0004] Therefore, there is a need for AI generation methods, systems, electronic devices, and storage media for in-vehicle digital humans. This solution uses user interaction input in the vehicle to drive the generation of a drivable 3D digital human model in the cloud. The model is then stored, driven, and rendered on the vehicle side, enabling personalized in-vehicle digital human generation. This improves upon the problem that traditional in-vehicle digital humans have fixed appearances and cannot meet the diverse needs of users. Summary of the Invention

[0005] The purpose of this invention is to provide an AI generation method, system, electronic device, and storage medium for in-vehicle digital humans, at least to solve the problem that existing digital humans adopt a fixed image design and cannot flexibly generate different forms, which makes it difficult to meet the diverse needs of users.

[0006] This invention provides the following solution:

[0007] According to one aspect of the present invention, an AI generation method for in-vehicle digital humans is provided, comprising the following steps:

[0008] Acquire user interaction input information from the vehicle system, the user interaction input information including voice command information, image input information, and auxiliary information related to digital human generation or interaction;

[0009] The user interaction input information is sent to the cloud, where an AI image generation model is used to generate digital human multi-angle image data, corresponding texture data, and feature pattern data based on the user interaction input information.

[0010] In the cloud, the multi-angle image data, texture data and feature pattern data of the digital human are input into the 3D reconstruction network to generate a 3D mesh model of the digital human. The standard skeleton system and expression driving parameters are bound to the 3D mesh model, and the corresponding driving association data is generated to obtain a drivable 3D digital human model.

[0011] The driveable 3D digital human model and its associated data are sent to the vehicle terminal, where the 3D digital human model and its associated data are stored locally and incremental updates based on cloud update commands are supported.

[0012] On the vehicle-mounted terminal, based on user interaction input information and real-time vehicle control status information, a pre-set action library and driving model are called to generate corresponding action control instructions and expression control parameters, and action-driven processing and expression-driven processing are performed on the three-dimensional digital human model respectively.

[0013] The 3D digital human model, after being processed by the drive, is rendered in real time and displayed on the vehicle's display screen.

[0014] Furthermore, the acquisition of user interaction input information from the in-vehicle system includes:

[0015] Receive input information provided by the user via voice or image;

[0016] Receive input information provided by the user through touch operation or operation command;

[0017] Acquire vehicle control status information related to digital human generation or interaction;

[0018] The received voice, image, touch, operation commands and vehicle control status information are uniformly encapsulated to form user interactive input information.

[0019] Furthermore, the generation of multi-angle image data, corresponding texture data, and feature pattern data of the digital human includes:

[0020] The generation style or template of the digital human is determined based on user interaction input information;

[0021] Multi-angle image data of the digital human is generated according to the style or template. The size of the digital human is fixed according to the preset vehicle adaptation standard, and the style is selected from cartoon, anthropomorphic animal or real-person simulation type.

[0022] Texture information is extracted from the multi-angle image data to form corresponding texture data;

[0023] Extract digital human appearance features from the multi-angle image data or texture information to generate feature pattern data;

[0024] The multi-angle image data, texture data, and feature pattern data are uniformly formatted.

[0025] Furthermore, the obtained drivable 3D digital human model includes:

[0026] Multi-angle image data of digital humans are input into a 3D reconstruction network to generate a 3D mesh model containing facial and body geometry;

[0027] Texture data is mapped onto the surface of the 3D mesh model, and feature pattern data is applied to the corresponding areas to form a complete appearance;

[0028] A standardized skeletal system is bound to the 3D mesh model, including body skeleton and joint information, and weights are assigned to each part of the mesh;

[0029] Generate facial expression control parameters for the facial mesh model;

[0030] The skeletal binding information, facial expression parameters, and texture data are integrated to form a related data set;

[0031] Output the final, drivable 3D digital human model.

[0032] Furthermore, the step of performing local storage on the 3D digital human model and its associated data, and supporting incremental updates based on cloud-based update commands, includes:

[0033] Receives a driveable 3D digital human model and its associated data downloaded from the cloud;

[0034] The received data is parsed and classified according to model files, texture files, bone weight data, facial expression parameter dictionary, and motion library.

[0035] The parsed data is saved to the vehicle's local storage space, and corresponding data indexes or mapping relationships are established.

[0036] The received data is labeled with version information or a timestamp;

[0037] When an update command is issued from the cloud, the incremental data corresponding to the update command is extracted;

[0038] Apply incremental data to local storage to update or replace existing models or supporting data;

[0039] After the update is completed, perform a data integrity check.

[0040] Furthermore, the generation of corresponding motion control commands and facial expression control parameters includes:

[0041] Analyze user interaction input information collected from the vehicle terminal to extract control intentions related to the digital human's actions and expressions;

[0042] Obtain real-time vehicle control status information to determine the execution conditions for actions and expressions;

[0043] Based on the parsed control intent, the corresponding action template is matched in the preset action library to determine the action type and motion parameters;

[0044] The motion template is mapped to the standard skeletal system to generate the rotation angle, position change and motion transition parameters of the skeletal joints;

[0045] Facial expression driving parameters are calculated based on user interaction input information, and the driving parameters include control values ​​for eyebrow, eye and mouth expression units;

[0046] The motion control commands and facial expression control parameters are organized into a data format that can be directly used to drive 3D digital human models.

[0047] Furthermore, the real-time rendering of the driven 3D digital human model includes:

[0048] S1. Load the 3D digital human model and its texture, material, facial parameters and skeletal weight data after motion-driven and expression-driven processing;

[0049] S2. Based on the motion control instructions and bone weight mapping, perform vertex transformation calculations on the model's bone nodes;

[0050] S3. Adjust facial expression units according to expression control parameters and update the 3D mesh shape;

[0051] S4. Perform lighting, material and texture mapping on the model surface, and calculate pixel color and reflection effect;

[0052] S5. Convert the 3D model coordinates into 2D screen projection coordinates and output pixel data to the display screen;

[0053] S6. Repeat steps S1-S5 to perform continuous rendering.

[0054] According to a second aspect of the present invention, an AI generation system for in-vehicle digital humans is provided, comprising:

[0055] The information acquisition module is used to acquire user interaction input information in the vehicle system. The user interaction input information includes voice command information, image input information, and auxiliary information related to digital human generation or interaction.

[0056] The cloud-based data generation module is used to send the user interaction input information to the cloud, and generate digital human multi-angle image data, corresponding texture data and feature pattern data based on the user interaction input information through an AI image generation model in the cloud.

[0057] The cloud-based modeling and binding module is used to input the multi-angle image data, texture data, and feature pattern data of the digital human into the 3D reconstruction network in the cloud to generate a 3D mesh model of the digital human, bind a standard skeleton system and expression driving parameters to the 3D mesh model, and generate corresponding driving association data to obtain a drivable 3D digital human model.

[0058] The storage update module is used to send the driveable 3D digital human model and its supporting data to the vehicle terminal, perform local storage on the 3D digital human model and its supporting data on the vehicle terminal, and support incremental updates based on cloud update commands.

[0059] The drive control module is used to call the preset action library and drive model on the vehicle terminal based on user interaction input information and real-time vehicle control status information, generate corresponding action control instructions and expression control parameters, and perform action drive processing and expression drive processing on the three-dimensional digital human model respectively.

[0060] The rendering and display module is used to render the 3D digital human model after driving processing in real time on the vehicle terminal and display it on the vehicle terminal's display screen.

[0061] According to three aspects of the present invention, an electronic device is provided, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0062] The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the AI ​​generation method for the in-vehicle digital human.

[0063] According to four aspects of the present invention, a computer-readable storage medium is provided, comprising: storing a computer program executable by an electronic device, wherein when the computer program is run on the electronic device, the electronic device causes the electronic device to perform the steps of an AI generation method for an in-vehicle digital human.

[0064] The above solution achieves the following beneficial technical effects:

[0065] This application obtains user interaction input information from the vehicle and sends it to the cloud. Through AI image generation, 3D reconstruction and parameter binding, a driveable 3D digital human model is obtained. Then, through storage, driving and rendering on the vehicle, a personalized in-vehicle digital human can be generated according to user input. This improves the problem that traditional in-vehicle digital humans adopt a fixed image design, cannot flexibly generate different forms, and cannot meet the diverse needs of users.

[0066] This application completes the heavy computation of digital human generation, modeling and parameter binding in the cloud, while the vehicle terminal only performs lightweight driving and rendering, reducing the computing load on the vehicle terminal, ensuring the speed of interactive response, and improving the problem of high interactive latency and insufficient stability caused by the local execution of heavy computation of traditional vehicle digital human.

[0067] This application enables the local storage of digital human models and related data on the vehicle-mounted terminal, supports incremental updates in the cloud, simplifies the update operation and reduces the amount of data transmitted, and improves the problem that the traditional full update of vehicle-mounted digital human models involves a large amount of data, takes a long time, and affects normal use by users. Attached Figure Description

[0068] Figure 1 This is a flowchart of an AI generation method for in-vehicle digital humans provided in one or more embodiments of the present invention.

[0069] Figure 2 This is an architecture diagram of an AI generation system for in-vehicle digital humans provided in one or more embodiments of the present invention.

[0070] Figure 3 This is an electronic device structural block diagram of an AI generation method for in-vehicle digital humans provided in one or more embodiments of the present invention. Detailed Implementation

[0071] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0072] Figure 1 This is a flowchart of an AI generation method for in-vehicle digital humans provided in one or more embodiments of the present invention.

[0073] like Figure 1 The AI ​​generation method for in-vehicle digital humans shown includes the following steps:

[0074] Acquire user interaction input information from the in-vehicle system. This user interaction input information includes voice command information, image input information, and auxiliary information related to digital human generation or interaction.

[0075] Obtaining user interaction input information from the in-vehicle system includes:

[0076] Receive input information provided by the user via voice or image;

[0077] Receive input information provided by the user through touch operation or operation command;

[0078] Acquire vehicle control status information related to digital human generation or interaction;

[0079] The received voice, image, touch, operation commands and vehicle control status information are uniformly encapsulated to form user interactive input information.

[0080] Specifically, the system receives voice commands recorded by the user through the in-vehicle voice acquisition device. This information originates from the user's description or selection of the in-vehicle digital human's form and style. It also receives image input information uploaded by the user through the in-vehicle touchscreen or facial images captured by the in-vehicle camera; this image input information provides a reference for the digital human's appearance. Furthermore, it receives touch operation information and operation commands triggered by the user through physical buttons on the steering wheel or virtual buttons on the in-vehicle central control screen. This information is used to help confirm the digital human's generation preferences or interaction triggering needs. Finally, it acquires vehicle control status information related to the digital human's generation or interaction through in-vehicle sensors and the control system. This includes vehicle speed, battery status, and door open / close status. Vehicle speed is collected by the onboard speed sensor, battery status is obtained by the onboard battery management system, and door open / close status is fed back by the door status sensor. The above-mentioned voice command information, image input information, touch operation information, operation command information, and vehicle control status information are uniformly packaged and processed, and integrated according to the preset data format specifications to form a unified user interaction input information. This user interaction input information is subsequently sent to the cloud, providing input basis for the cloud to generate multi-angle image data, corresponding texture data, and feature pattern data of digital human through AI image generation model.

[0081] In this embodiment, user interaction input information is sent to the cloud, where an AI image generation model is used to generate digital human multi-angle image data, corresponding texture data, and feature pattern data based on the user interaction input information.

[0082] The generation of digital human multi-angle image data, corresponding texture data, and feature pattern data includes:

[0083] The generation style or template of the digital human is determined based on user interaction input information;

[0084] Multi-angle image data of the digital human is generated according to the style or template. The size of the digital human is fixed according to the preset vehicle adaptation standard. The style is selected from cartoon, anthropomorphic animal or real-life simulation type.

[0085] Texture information is extracted from multi-angle image data to form corresponding texture data;

[0086] Extract digital human appearance features from multi-angle image data or texture information to generate feature pattern data;

[0087] The data, texture data, and feature pattern data from multiple angles are formatted in a unified manner.

[0088] Specifically, the previously generated user interaction input information is sent to the cloud. Using this user interaction input information as input, the cloud uses an AI image generation model to perform data generation operations related to the digital human. First, based on the voice command information and image input information in the user interaction input information, the generation style or template of the digital human is determined. The generation style T is selected from cartoon T1, anthropomorphic animal T2, or live-action simulation T3. The size of the digital human is fixed according to a preset vehicle adaptation standard S0, which is set according to the resolution and visual presentation requirements of the vehicle display screen. Then, according to the determined style T or template and the fixed size S0, the AI ​​image generation model generates multi-angle image data I of the digital human. The multi-angle image data I includes images of the digital human from at least three perspectives, such as the front, side, and 45° angle. Subsequently, surface texture information is extracted from the multi-angle image data I to form texture data T corresponding to the multi-angle image data I. ex Then, from multi-angle image data I or texture data T ex Extracting appearance features such as hairstyle and clothing from the digital human to generate feature pattern data P; finally, according to the data format requirements that the 3D reconstruction network can recognize, processing multi-angle image data I and texture data T... ex The feature pattern data P is uniformly formatted to form a structured data set D={I,T}. ex The structured dataset D will then be input into a cloud-based 3D reconstruction network to generate a 3D mesh model of the digital human.

[0089] In this embodiment, multi-angle image data, texture data and feature pattern data of the digital human are input into the 3D reconstruction network in the cloud to generate a 3D mesh model of the digital human. The standard skeleton system and expression driving parameters are bound to the 3D mesh model, and the corresponding driving association data is generated to obtain a drivable 3D digital human model.

[0090] Obtaining a driveable 3D digital human model includes:

[0091] Multi-angle image data of digital humans are input into a 3D reconstruction network to generate a 3D mesh model containing facial and body geometry;

[0092] Texture data is mapped onto the surface of a 3D mesh model, and feature pattern data is applied to the corresponding areas to form a complete appearance;

[0093] A standardized skeletal system is bound to the 3D mesh model, including body skeleton and joint information, and weights are assigned to each part of the mesh;

[0094] Generate facial expression control parameters for the facial mesh model;

[0095] Integrate skeletal binding information, facial expression parameters, and texture data to form a related data set;

[0096] Output the final, drivable 3D digital human model.

[0097] Specifically, a 3D reconstruction network is invoked in the cloud, using multi-angle image data, texture data, and feature pattern data of the digital human previously processed by an image generation model as input. First, the multi-angle image data of the digital human is input into the 3D reconstruction network, which reconstructs the geometric information of the multi-view images to generate a 3D mesh model containing the geometry of the digital human's face and body. Next, the texture data is mapped onto the surface of the 3D mesh model according to the correspondence between image coordinates and mesh vertices. Simultaneously, the feature pattern data is applied to the corresponding appearance areas of the mesh model, forming a complete 3D mesh model. Then, a standardized skeletal system is bound to the 3D mesh model. This standardized skeletal system includes information on the body's skeleton and joints, and the skeletal connections are set according to the laws of human kinematics. A weight allocation algorithm is used to assign weights to the vertices of each part of the mesh, obtaining the weight value W of the i-th bone to the j-th mesh vertex. ij Next, for the facial mesh model, facial expression control parameters are generated based on the basic expression requirements in the in-vehicle voice interaction scenario. These parameters cover the control values ​​of expression units such as eyebrows, eyes, and mouth. Then, the skeletal binding information, expression control parameters, and texture data are integrated to form a driving-related data set, where the skeletal binding information includes a standardized skeletal system and weight values. Finally, based on the complete 3D mesh model and the driving-related data set, the final drivable 3D digital human model is output. This drivable 3D digital human model is then sent to the in-vehicle terminal for local storage and action-driven processing and expression-driven processing based on user interaction input information and real-time vehicle control status information.

[0098] In this embodiment, the drivable 3D digital human model and its supporting data are sent to the vehicle terminal, where the 3D digital human model and its supporting data are stored locally and incremental updates based on cloud update commands are supported.

[0099] Local storage is performed on the 3D digital human model and its associated data, and incremental updates based on cloud-based update commands are supported, including:

[0100] Receives a driveable 3D digital human model and its associated data downloaded from the cloud;

[0101] The received data is parsed and classified according to model files, texture files, bone weight data, facial expression parameter dictionary, and motion library.

[0102] The parsed data is saved to the vehicle's local storage space, and corresponding data indexes or mapping relationships are established.

[0103] The received data is labeled with version information or a timestamp;

[0104] When an update command is issued from the cloud, the incremental data corresponding to the update command is extracted;

[0105] Apply incremental data to local storage to update or replace existing models or supporting data;

[0106] After the update is completed, perform a data integrity check.

[0107] Specifically, it receives a driveable 3D digital human model M downloaded from the cloud. drv and related data D drv The driveable three-dimensional digital human model M drv and related data D drv This refers to structured data generated in the cloud after 3D reconstruction network modeling, skeleton rigging, and parameter integration; the received data is processed according to the model file F. model Texture file F tex Skeletal weight data D w Expression parameter dictionary D exp and motion library L act The system parses and classifies data based on its functional attributes and file format characteristics; it saves the parsed data to the vehicle's local storage space and creates corresponding data indexes for quick data location and retrieval; and it marks the received data with version information V or a timestamp T. stamp timestamp T stamp Records the time information of the data storage moment; when the cloud issues an update command C... update At that time, the corresponding incremental data D is extracted based on the difference identifier carried in the update instruction. inc Incremental data D inc This refers to the difference between locally stored data and the latest data in the cloud; incremental data D inc Applied to local storage, for the original model file F model Texture file F tex Skeletal weight data D w Expression parameter dictionary D exp or action library L act Update or replace any parts that differ from the original data; after the update is complete, perform data integrity verification and calculate the updated local dataset D. local The hash value Hcalc = Hash(D) local And compare it with the verification hash value H sent from the cloud. recv A comparison is performed to confirm data integrity; the complete data set D stored locally on the vehicle terminal. local ={F model,F tex D w D exp ,L act}, D local For M drv With D drv The structured results after parsing and classification are then used by the vehicle-mounted terminal to generate corresponding motion control commands and facial expression control parameters based on user interaction input information and real-time vehicle control status information, and then to perform motion-driven processing and facial expression-driven processing on the 3D digital human model.

[0108] In this embodiment, based on user interaction input information and real-time vehicle control status information, the vehicle terminal calls the preset action library and driving model to generate corresponding action control instructions and expression control parameters, and performs action-driven processing and expression-driven processing on the three-dimensional digital human model respectively.

[0109] The corresponding motion control instructions and facial expression control parameters include:

[0110] Analyze user interaction input information collected from the vehicle terminal to extract control intentions related to the digital human's actions and expressions;

[0111] Obtain real-time vehicle control status information to determine the execution conditions for actions and expressions;

[0112] Based on the parsed control intent, the corresponding action template is matched in the preset action library to determine the action type and motion parameters;

[0113] The motion template is mapped to the standard skeletal system to generate the rotation angle, position change and motion transition parameters of the skeletal joints;

[0114] The facial expression driving parameters are calculated based on user interaction input information. The driving parameters include control values ​​for the expression units of eyebrows, eyes, and mouth.

[0115] The motion control commands and facial expression control parameters are organized into a data format that can be directly used to drive 3D digital human models.

[0116] Specifically, the in-vehicle system uses previously generated user interaction input information and real-time vehicle control status information as input. User interaction input information includes voice commands, image input, and auxiliary interaction information. Real-time vehicle control status information includes vehicle speed, battery status, and other vehicle operation data collected by in-vehicle sensors and the control system. First, the user interaction input information is parsed, and through semantic recognition and feature extraction, the control intent related to the digital human's actions and expressions is extracted. Next, the real-time vehicle control status information is acquired, and the execution conditions for actions and expressions are determined based on data such as vehicle speed and battery status. Then, according to the control intent, a corresponding action template is matched from a pre-set action library, and the action type and motion parameters are determined based on the action template. Finally, the action template is mapped to a standard skeletal system, and the rotation angle θ of the skeletal joints is calculated and generated. ij Position change and motion transition parameters, where θ ij Let be the rotation angle of the i-th joint around the j-th axis; simultaneously, based on the emotional inclination and command requirements in the user's interactive input information, calculate facial expression driving parameters, which include eyebrow control values, eye control values, and mouth control values; finally, organize the rotation angle, position change, and motion transition parameters of the skeletal joints into motion control commands, and organize the facial expression driving parameters into expression control parameters, both of which adopt a data format that can be directly used for driving the 3D digital human model; the generated motion control commands and expression control parameters are subsequently used to perform motion-driven processing and expression-driven processing on the 3D digital human model stored on the vehicle, respectively, to realize the interactive response of the digital human in the vehicle scene.

[0117] In this embodiment, the three-dimensional digital human model after driving processing is rendered in real time on the vehicle terminal and displayed on the vehicle terminal's display screen.

[0118] Real-time rendering of the processed 3D digital human model includes:

[0119] S1. Load the 3D digital human model and its texture, material, facial parameters and skeletal weight data after motion-driven and expression-driven processing;

[0120] S2. Based on the motion control instructions and bone weight mapping, perform vertex transformation calculations on the model's bone nodes;

[0121] S3. Adjust facial expression units according to expression control parameters and update the 3D mesh shape;

[0122] S4. Perform lighting, material and texture mapping on the model surface, and calculate pixel color and reflection effect;

[0123] S5. Convert the 3D model coordinates into 2D screen projection coordinates and output pixel data to the display screen;

[0124] S6. Repeat steps S1-S5 to perform continuous rendering.

[0125] Specifically, real-time rendering is performed on the 3D digital human model processed by motion and expression control on the in-vehicle device. First, the motion- and expression-driven 3D digital human model, along with its associated textures, materials, expression parameters, and skeletal weights, are loaded. This data is a collection of digital human-related data previously received and stored on the in-vehicle device. Next, based on the previously generated motion control commands and skeletal weight mapping, vertex transformation calculations are performed on the skeletal nodes of the 3D digital human model, causing the model's skeletal movement to synchronously update the mesh vertex positions. Then, based on the previously generated expression control parameters, the model's facial expression units are adjusted, updating the 3D mesh shape to present the corresponding expression. Finally, the adjusted 3D digital human model is rendered... The model undergoes lighting processing, material property configuration, and texture mapping. The color values ​​and reflection effects of each pixel on the model's surface are calculated to enhance visual realism. Then, the processed 3D model coordinates are converted into 2D projection coordinates adapted to the in-vehicle display screen, and the calculated pixel data is output to the in-vehicle display screen. Finally, the above steps of loading, vertex transformation, expression adjustment, surface processing, coordinate transformation, and pixel output are executed in a loop to achieve continuous rendering of the 3D digital human model. The rendered digital human dynamic image presented on the in-vehicle display screen is used to show users the digital human's actions and facial expressions in the in-vehicle scene, meeting the visual presentation requirements of in-vehicle human-computer interaction and fitting the scenario setting for subsequent use in multi-screen vehicles.

[0126] Figure 2 This is an architecture diagram of an AI generation system for in-vehicle digital humans provided in one or more embodiments of the present invention.

[0127] like Figure 2 The AI-generated in-vehicle digital human system shown includes:

[0128] The information acquisition module is used to acquire user interaction input information from the in-vehicle system. The user interaction input information includes voice command information, image input information, and auxiliary information related to digital human generation or interaction.

[0129] The cloud-based data generation module is used to send user interaction input information to the cloud. In the cloud, an AI image generation model generates multi-angle image data, corresponding texture data, and feature pattern data of the digital human based on the user interaction input information.

[0130] The cloud-based modeling and binding module is used to input multi-angle image data, texture data, and feature pattern data of the digital human into the 3D reconstruction network in the cloud to generate a 3D mesh model of the digital human. It binds the standard skeleton system and expression driving parameters to the 3D mesh model and generates the corresponding driving association data to obtain a drivable 3D digital human model.

[0131] The storage update module is used to send the drivable 3D digital human model and its supporting data to the vehicle terminal, perform local storage of the 3D digital human model and its supporting data on the vehicle terminal, and support incremental updates based on cloud update commands.

[0132] The drive control module is used to call the preset motion library and drive model on the vehicle end based on user interaction input information and real-time vehicle control status information, generate corresponding motion control instructions and expression control parameters, and perform motion drive processing and expression drive processing on the 3D digital human model respectively.

[0133] The rendering and display module is used to render the 3D digital human model after driving processing in real time on the vehicle terminal and display it on the vehicle terminal's display screen.

[0134] In the context of intelligent vehicle cockpits, in-vehicle interaction systems often only provide virtual avatars with fixed expressions or limited movements. They cannot generate digital humans with personalized appearances and natural interaction capabilities based on real-time user voice, images, or vehicle control status. This results in a limited in-vehicle interaction experience, failing to meet the emotional interaction needs across various scenarios. To address these issues, the AI ​​generation system for in-vehicle digital humans provided by this invention is employed, with the structure as follows: Figure 2 As shown. The specific implementation process of this system is as follows:

[0135] The information acquisition module receives input information from users through voice, images, touch, and operation commands, and combines it with vehicle control status data to encapsulate it into user interactive input information, thereby achieving a unified expression of the user's intentions within the vehicle.

[0136] The cloud-based data generation module sends the encapsulated user interaction input information to the cloud and generates multi-angle image data, texture data, and feature pattern data of the digital human based on the AI ​​image generation model, thereby generating data for constructing the digital human appearance according to user preferences.

[0137] The cloud-based modeling and binding module generates a 3D mesh model of a digital human based on a 3D reconstruction network and binds it with a standard skeletal system and facial expression driving parameters. At the same time, it generates driving association data to obtain a 3D digital human model with executable actions and facial expression driving.

[0138] The storage update module sends the driveable 3D digital human model and its supporting data to the vehicle terminal for local storage. When it receives an update command from the cloud, it performs incremental updates to maintain the dynamic availability and real-time update capability of the digital human model on the vehicle terminal.

[0139] Based on user interaction input information and real-time vehicle control status, the drive control module calls the preset action library and drive model to generate action control commands and facial expression control parameters, enabling the digital human to generate corresponding actions and expressions according to the in-vehicle scene status and user intentions, thus achieving natural interaction.

[0140] The rendering and display module performs real-time rendering of the 3D digital human model after driver processing and displays it on the in-vehicle screen, realizing the real-time presentation effect of the digital human in the in-vehicle environment, so that users can obtain a continuous and natural interactive experience.

[0141] Through the above steps, the system can dynamically generate a digital human image with a personalized appearance, natural movements, and linkage with the vehicle control status in various interactive scenarios during vehicle operation, thereby achieving a more immersive and responsive interactive experience in the intelligent cockpit environment.

[0142] Figure 3 This is an electronic device structural block diagram of an AI generation method for in-vehicle digital humans provided in one or more embodiments of the present invention.

[0143] like Figure 3 As shown, this application provides an electronic device, including: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0144] The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the AI ​​generation method for the in-vehicle digital human.

[0145] This application also provides a computer-readable storage medium storing a computer program executable by an electronic device, which, when run on the electronic device, causes the electronic device to perform the steps of an AI generation method for an in-vehicle digital human.

[0146] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0147] The electronic device comprises a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on the operating system. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory. The operating system can be any one or more computer operating systems that control the electronic device through processes, such as Linux, Unix, Android, iOS, or Windows. Furthermore, in this embodiment of the invention, the electronic device can be a smartphone, tablet computer, or other handheld device, or a desktop computer, portable computer, or other electronic device; there is no particular limitation in this embodiment.

[0148] In this embodiment of the invention, the executing entity for electronic device control can be an electronic device itself, or a functional module within an electronic device capable of calling and executing a program. The electronic device can obtain the firmware corresponding to the storage medium. This firmware is provided by the supplier, and different storage media may have the same or different firmware; no limitation is made here. After obtaining the firmware corresponding to the storage medium, the electronic device can write this firmware into the storage medium; specifically, it burns the firmware corresponding to the storage medium into the storage medium. The process of burning the firmware into the storage medium can be implemented using existing technology, and will not be elaborated upon in this embodiment of the invention.

[0149] Electronic devices can also obtain reset commands corresponding to the storage media. The reset commands corresponding to the storage media are provided by the supplier. The reset commands corresponding to different storage media can be the same or different, and no restrictions are imposed here.

[0150] At this time, the storage medium of the electronic device is a storage medium on which the corresponding firmware has been written. The electronic device can respond to the reset command corresponding to the storage medium on which the corresponding firmware has been written, thereby resetting the storage medium on which the corresponding firmware has been written according to the reset command. The process of resetting the storage medium according to the reset command can be implemented by existing technology and will not be described in detail in this embodiment of the invention.

[0151] For ease of description, the above devices are described separately by function as various units and modules. Of course, in implementing this application, the functions of each unit and module can be implemented in one or more software and / or hardware.

[0152] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the meaning consistent with their meaning in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined.

[0153] For the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0154] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An AI generation method for in-vehicle digital humans, characterized in that, Includes the following steps: Acquire user interaction input information from the vehicle system, the user interaction input information including voice command information, image input information, and auxiliary information related to digital human generation or interaction; The user interaction input information is sent to the cloud, where an AI image generation model is used to generate digital human multi-angle image data, corresponding texture data, and feature pattern data based on the user interaction input information. In the cloud, the multi-angle image data, texture data and feature pattern data of the digital human are input into the 3D reconstruction network to generate a 3D mesh model of the digital human. The standard skeleton system and expression driving parameters are bound to the 3D mesh model, and the corresponding driving association data is generated to obtain a drivable 3D digital human model. The driveable 3D digital human model and its associated data are sent to the vehicle terminal, where the 3D digital human model and its associated data are stored locally and incremental updates based on cloud update commands are supported. On the vehicle-mounted terminal, based on user interaction input information and real-time vehicle control status information, a pre-set action library and driving model are called to generate corresponding action control instructions and expression control parameters, and action-driven processing and expression-driven processing are performed on the three-dimensional digital human model respectively. The 3D digital human model, after being processed by the drive, is rendered in real time and displayed on the vehicle's display screen.

2. The AI ​​generation method for in-vehicle digital humans according to claim 1, characterized in that, The acquisition of user interaction input information from the vehicle system includes: Receive input information provided by the user via voice or image; Receive input information provided by the user through touch operation or operation command; Acquire vehicle control status information related to digital human generation or interaction; The received voice, image, touch, operation commands and vehicle control status information are uniformly encapsulated to form user interactive input information.

3. The AI ​​generation method for in-vehicle digital humans according to claim 1, characterized in that, The generated digital human multi-angle image data, corresponding texture data, and feature pattern data include: The generation style or template of the digital human is determined based on user interaction input information; Multi-angle image data of the digital human is generated according to the style or template. The size of the digital human is fixed according to the preset vehicle adaptation standard, and the style is selected from cartoon, anthropomorphic animal or real-person simulation type. Texture information is extracted from the multi-angle image data to form corresponding texture data; Extract digital human appearance features from the multi-angle image data or texture information to generate feature pattern data; The multi-angle image data, texture data, and feature pattern data are uniformly formatted.

4. The AI ​​generation method for in-vehicle digital humans according to claim 1, characterized in that, The obtained driveable 3D digital human model includes: Multi-angle image data of digital humans are input into a 3D reconstruction network to generate a 3D mesh model containing facial and body geometry; Texture data is mapped onto the surface of the 3D mesh model, and feature pattern data is applied to the corresponding areas to form a complete appearance; A standardized skeletal system is bound to the 3D mesh model, including body skeleton and joint information, and weights are assigned to each part of the mesh; Generate facial expression control parameters for the facial mesh model; The skeletal binding information, facial expression parameters, and texture data are integrated to form a related data set; Output the final, drivable 3D digital human model.

5. The AI ​​generation method for in-vehicle digital humans according to claim 1, characterized in that, The step of performing local storage on the 3D digital human model and its associated data, and supporting incremental updates based on cloud update commands, includes: Receives a driveable 3D digital human model and its associated data from the cloud; The received data is parsed and classified according to model files, texture files, bone weight data, facial expression parameter dictionary, and motion library. The parsed data is saved to the vehicle's local storage space, and corresponding data indexes or mapping relationships are established. The received data is labeled with version information or a timestamp; When an update command is issued from the cloud, the incremental data corresponding to the update command is extracted; Apply incremental data to local storage to update or replace existing models or supporting data; After the update is completed, perform a data integrity check.

6. The AI ​​generation method for in-vehicle digital humans according to claim 1, characterized in that, The generation of corresponding motion control commands and facial expression control parameters includes: Analyze user interaction input information collected from the vehicle terminal to extract control intentions related to the digital human's actions and expressions; Obtain real-time vehicle control status information to determine the execution conditions for actions and expressions; Based on the parsed control intent, the corresponding action template is matched in the preset action library to determine the action type and motion parameters; The motion template is mapped to the standard skeletal system to generate the rotation angle, position change and motion transition parameters of the skeletal joints; Facial expression driving parameters are calculated based on user interaction input information, and the driving parameters include control values ​​for eyebrow, eye and mouth expression units; The motion control commands and facial expression control parameters are organized into a data format that can be directly used to drive 3D digital human models.

7. The AI ​​generation method for in-vehicle digital humans according to claim 1, characterized in that, The real-time rendering of the processed 3D digital human model includes: S1. Load the 3D digital human model and its texture, material, facial parameters and skeletal weight data after motion-driven and expression-driven processing; S2. Based on the motion control instructions and bone weight mapping, perform vertex transformation calculations on the model's bone nodes; S3. Adjust facial expression units according to expression control parameters and update the 3D mesh shape; S4. Perform lighting, material and texture mapping on the model surface, and calculate pixel color and reflection effect; S5. Convert the 3D model coordinates into 2D screen projection coordinates and output pixel data to the display screen; S6. Repeat steps S1-S5 to perform continuous rendering.

8. An AI-generated in-vehicle digital human system, characterized in that, The AI ​​generation method for in-vehicle digital humans according to any one of claims 1-7 includes: The information acquisition module is used to acquire user interaction input information in the vehicle system. The user interaction input information includes voice command information, image input information, and auxiliary information related to digital human generation or interaction. The cloud-based data generation module is used to send the user interaction input information to the cloud, and generate digital human multi-angle image data, corresponding texture data and feature pattern data based on the user interaction input information through an AI image generation model in the cloud. The cloud-based modeling and binding module is used to input the multi-angle image data, texture data, and feature pattern data of the digital human into the 3D reconstruction network in the cloud to generate a 3D mesh model of the digital human, bind a standard skeleton system and expression driving parameters to the 3D mesh model, and generate corresponding driving association data to obtain a drivable 3D digital human model. The storage update module is used to send the driveable 3D digital human model and its supporting data to the vehicle terminal, perform local storage on the 3D digital human model and its supporting data on the vehicle terminal, and support incremental updates based on cloud update commands. The drive control module is used to call the preset action library and drive model on the vehicle terminal based on user interaction input information and real-time vehicle control status information, generate corresponding action control instructions and expression control parameters, and perform action drive processing and expression drive processing on the three-dimensional digital human model respectively. The rendering and display module is used to render the 3D digital human model after driving processing in real time on the vehicle terminal and display it on the vehicle terminal's display screen.

9. An electronic device, characterized in that, include: The processor, communication interface, memory, and communication bus are connected, with the processor, communication interface, and memory communicating with each other via the communication bus. The memory stores a computer program that, when executed by a processor, causes the processor to perform the steps of the AI ​​generation method for an in-vehicle digital human as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The device stores a computer program executable by an electronic device, which, when run on the electronic device, causes the electronic device to perform the steps of the AI ​​generation method for an in-vehicle digital human as described in any one of claims 1-7.