Vehicle-mounted virtual human generation method and system, vehicle, vehicle machine and aerial imaging device
Patent Information
- Application Number
- CN202411634683.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2044-11-15
AI Technical Summary
[0004]有鉴于此,本发明提供了一种车载虚拟人生成方法、系统、车辆、车机及空中成像设备,以解决传统的车载虚拟人严重依赖于车机屏幕,导致渲染效果差、车机耗能高的问题
[0075] By acquiring communication information from at least one business module within the vehicle's infotainment system, virtual human imaging parameters reflecting the driver's intentions and vehicle movement trends are determined. Communication packets are then generated based on these parameters and sent to an aerial imaging device. The aerial imaging device decodes the communication packets and renders them according to the obtained virtual human imaging parameters, generating an in-vehicle virtual human with excellent rendering quality. This approach eliminates the need to consume the vehicle's rendering performance, reducing energy consumption, and the virtual human imaging is independent of the vehicle's screen, resulting in a more immersive user experience. Furthermore, this invention can utilize communication information from different business modules within the vehicle for virtual human imaging, leading to richer imaging effects.
Smart Images

Figure CN119484733B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent driving technology, specifically to a method, system, vehicle, vehicle-mounted infotainment system, and aerial imaging equipment for generating in-vehicle virtual humans. Background Technology
[0002] In-cabin features such as intelligent driving, voice assistants, premium audio systems, and cinematic movie theaters have become crucial areas for enhancing the driving experience. Global automakers are committed to creating the ultimate intelligent cockpit experience, leveraging new technologies to develop in-vehicle infotainment systems that significantly improve driving safety.
[0003] Currently, in-vehicle virtual humans, as a smart cockpit feature, can provide users with an ultimate driving experience. However, traditional in-vehicle virtual humans can only rely on the vehicle's infotainment screen for display and rendering, and the rendering process is limited by the vehicle's performance, which can easily lead to poor rendering effects and high energy consumption. Summary of the Invention
[0004] In view of this, the present invention provides a method, system, vehicle, vehicle-mounted system and aerial imaging device for generating in-vehicle virtual humans, in order to solve the problem that traditional in-vehicle virtual humans rely heavily on the vehicle-mounted screen, resulting in poor rendering effects and high energy consumption of the vehicle-mounted system.
[0005] In a first aspect, the present invention provides a method for generating a virtual human in a vehicle, applied to an in-vehicle infotainment system, the method comprising:
[0006] Acquire communication information generated by at least one service module in the vehicle's infotainment system;
[0007] Based on communication information, determine the imaging parameters of the virtual human;
[0008] Communication packets are generated based on the virtual human imaging parameters and sent to the aerial imaging equipment. The aerial imaging equipment then parses the communication packets to obtain the virtual human imaging parameters and renders them to generate an in-vehicle virtual human.
[0009] Beneficial effects: By acquiring communication information from at least one business module through the vehicle's infotainment system, virtual human imaging parameters reflecting the driver's intentions and vehicle driving trends are determined. Communication packets are then generated based on these parameters and sent to an aerial imaging device. After decoding the communication packets, the aerial imaging device renders the virtual human according to the obtained imaging parameters, generating an in-vehicle virtual human with excellent rendering quality. This eliminates the need to consume the vehicle's rendering performance, reducing energy consumption. Furthermore, the virtual human imaging is independent of the vehicle's screen, enhancing user immersion. In addition, this invention can utilize communication information from different vehicle business modules for virtual human imaging, resulting in richer imaging effects.
[0010] In one alternative implementation, determining virtual human imaging parameters based on communication information includes:
[0011] The communication information is filtered to obtain interactive data sent by at least one service module; the service module includes an aerial imaging equipment setting module, a voice module, and a navigation module.
[0012] Based on the business instructions corresponding to the interactive data, the virtual human imaging parameters are obtained; among them, the virtual human imaging parameters include imaging trigger type, imaging area, imaging volume, virtual human model parameters, virtual human action parameters and / or virtual human voice parameters.
[0013] Beneficial effects: By extracting and identifying the communication information of various business modules of the vehicle, business instructions that can reflect the driver's intention and the vehicle's driving trend can be generated. Based on the business instructions, the imaging trigger type, imaging area, imaging volume, virtual human model, and the actions and voices that the virtual human needs to perform can be determined, thereby realizing the interaction between the virtual human and the user and enhancing the driving experience.
[0014] In one optional implementation, the service instruction includes an imaging parameter setting instruction, the virtual human action parameters include the user-set initial virtual human action, and the virtual human voice parameters include the user-set initial virtual human voice; based on the service instruction corresponding to the interaction data, the virtual human imaging parameters are obtained, including:
[0015] Based on the interactive data sent by the aerial imaging equipment setting module, the imaging parameter setting instructions issued by the user for the aerial imaging equipment are obtained.
[0016] Based on the imaging parameter setting instructions, the imaging trigger type, imaging area, imaging volume, virtual human model parameters, initial virtual human actions, and / or initial virtual human speech set by the user for the virtual human are obtained.
[0017] Beneficial effects: In this embodiment of the invention, users can issue imaging parameter setting commands through the aerial imaging device setting module to set the imaging trigger type, imaging area, imaging volume, virtual human model parameters, virtual human action parameters and / or virtual human voice parameters of the virtual human, thereby improving the richness of the imaging and the user experience.
[0018] In one optional implementation, the business instruction further includes a voice instruction, the virtual human action parameters further include a target action corresponding to the voice instruction, and the virtual human voice parameters further include a target voice corresponding to the voice instruction; based on the business instruction corresponding to the interaction data, the virtual human imaging parameters are obtained, further including:
[0019] The voice commands issued by the user are obtained by extracting and recognizing the interactive data sent by the voice module.
[0020] Based on the preset mapping relationship between voice commands and virtual human imaging parameters, the target action and / or target speech corresponding to the voice command are identified.
[0021] Beneficial effects: In this embodiment of the invention, the vehicle system analyzes the user's intent through the interactive data sent by the voice module, obtains the user's voice commands, and queries the pre-stored mapping relationship between voice commands and virtual human imaging parameters to determine the target action that the virtual human needs to perform and / or the target voice that needs to be broadcast, so as to facilitate the rendering and projection of the aerial imaging equipment, thereby realizing the interaction between the vehicle system's language module and the virtual human, and improving the user's immersive experience during the driving process.
[0022] In one optional implementation, the business instruction further includes a navigation instruction, the virtual human action parameters further include a target action corresponding to the navigation instruction, and the virtual human voice parameters further include a target voice corresponding to the navigation instruction; based on the business instruction corresponding to the interaction data, the virtual human imaging parameters are obtained, further including:
[0023] Based on the interactive data sent by the navigation module, navigation instructions are obtained;
[0024] Based on the preset mapping relationship between navigation instructions and virtual human imaging parameters, the target action and / or target voice corresponding to the navigation instructions are identified.
[0025] Beneficial effects: In this embodiment of the invention, the vehicle system analyzes the driving trend of the vehicle through the interactive data sent by the navigation module to obtain navigation instructions. Then, based on the pre-stored mapping relationship between the navigation instructions and the virtual human imaging parameters, it queries the target action and / or target voice corresponding to the navigation instructions, which facilitates the rendering and projection by the aerial imaging equipment. This enables the interaction between the vehicle system navigation module and the virtual human, which helps to improve the immersive experience of the user during the driving process.
[0026] In one optional implementation, generating a communication packet based on virtual human imaging parameters and sending the communication packet to an aerial imaging device includes:
[0027] The virtual human imaging parameters are byte-encoded to generate a communication body, and a communication header is generated based on the byte length of the communication body and the version information of the preset communication protocol.
[0028] The communication packet is obtained based on the communication header and communication body. The communication packet is sent to the airborne imaging device according to the preset communication protocol so that the airborne imaging device can parse the communication packet to obtain the virtual human imaging parameters. Based on the virtual human imaging parameters, the initial state machine is updated. The updated state machine is used to render the model to obtain the target virtual human model. The target virtual human model is then projected to generate the vehicle-mounted virtual human.
[0029] Beneficial effects: The virtual human imaging parameters are byte-encoded according to the preset communication protocol to generate a communication body, and a communication header is generated according to the byte length of the communication body and the protocol version information. The communication packet containing the communication header and the communication body is sent to the aerial imaging equipment, which facilitates the aerial imaging equipment to unpack the packet, avoids reading erroneous data, and improves the accuracy of information transmission.
[0030] In one optional implementation, sending a communication packet to an aerial imaging device according to a preset communication protocol includes:
[0031] Send a synchronization message to the airborne imaging equipment to request the establishment of a connection;
[0032] After receiving the synchronization confirmation message returned by the airborne imaging device based on the synchronization message, the system sends a first confirmation message to the airborne imaging device to confirm the establishment of the connection and establishes a connection with the airborne imaging device.
[0033] Send communication packets to aerial imaging equipment;
[0034] After detecting that the communication packet has been sent, a termination message is sent to the air imaging device to request the closure of the connection;
[0035] After receiving a second acknowledgment message from the airborne imaging device to indicate that a termination message has been received, and a termination acknowledgment message returned based on the termination message, the system sends a third acknowledgment message to the airborne imaging device to confirm the closure of the connection, and disconnects from the airborne imaging device.
[0036] Beneficial effects: In this embodiment of the application, by sending communication packets to the aerial imaging equipment according to a preset communication protocol, the reliability, integrity and accuracy of data transmission are improved, and there is no need to develop different communication standards for different aerial imaging equipment manufacturers, which reduces project maintenance costs and development difficulty.
[0037] Secondly, this invention provides a method for generating a vehicle-mounted virtual human, applied to aerial imaging equipment, the method comprising:
[0038] Receive communication packets sent by the vehicle-mounted system; wherein, the communication packets are generated by the vehicle-mounted system based on the virtual human imaging parameters corresponding to the communication information obtained by at least one service module;
[0039] The communication packets are parsed to obtain the virtual human imaging parameters;
[0040] The virtual human is generated by rendering based on the virtual human imaging parameters.
[0041] Beneficial effects: By acquiring communication information from at least one business module through the vehicle's infotainment system, virtual human imaging parameters reflecting the driver's intentions and vehicle driving trends are determined. Communication packets are then generated based on these parameters and sent to an aerial imaging device. After decoding the communication packets, the aerial imaging device renders the virtual human according to the obtained imaging parameters, generating an in-vehicle virtual human with excellent rendering quality. This eliminates the need to consume the vehicle's rendering performance, reducing energy consumption. Furthermore, the virtual human imaging is independent of the vehicle's screen, enhancing user immersion. In addition, this invention can utilize communication information from different vehicle business modules for virtual human imaging, resulting in richer imaging effects.
[0042] In one alternative implementation, rendering is performed based on virtual human imaging parameters to generate an in-vehicle virtual human, including:
[0043] The initial state machine is updated based on the virtual human imaging parameters; the virtual human imaging parameters include imaging trigger type, imaging area, imaging volume, virtual human model parameters, virtual human action parameters and / or virtual human voice parameters.
[0044] The updated state machine is used for model rendering to obtain the target virtual human model;
[0045] The target virtual human model is projected to generate an in-vehicle virtual human.
[0046] Beneficial effects: By using the virtual human imaging parameters sent by the vehicle-mounted unit, the initial state machine of the aerial imaging device is updated, enabling the updated state machine to perform graphic drawing and projection according to the virtual human imaging parameters, resulting in good rendering effects. This eliminates the need for the vehicle-mounted unit to perform imaging rendering, thus reducing the energy consumption of the vehicle-mounted unit.
[0047] In one alternative implementation, the updated state machine is used for model rendering to obtain the target virtual human model, including:
[0048] The updated state machine is used for geometry drawing, texture mapping, lighting and shadow processing, color processing, depth testing, and template testing to obtain the target virtual human model.
[0049] Beneficial effects: The updated state machine is used for geometric rendering to obtain the virtual human's outline. Texture mapping is then applied to enrich the virtual human's color, bumpiness, and roughness, improving realism. Next, lighting and color processing are performed to determine the illumination color of each voxel in the imaging area. Finally, the rendered results can be stored in a frame buffer before subsequent projection, reducing image flicker and improving rendering stability. Furthermore, depth testing and stencil testing are used to filter noise and interference, further enhancing the realism of the rendering.
[0050] In one alternative implementation, after projecting the target virtual human model to generate the in-vehicle virtual human, the method further includes:
[0051] Based on the virtual human motion parameters and / or virtual human speech parameters, the target motion and / or target speech are obtained;
[0052] Control the in-vehicle virtual human to perform target actions and / or broadcast target speech.
[0053] Beneficial effects: Based on the virtual human's motion parameters and / or voice parameters, the system identifies the target actions that the virtual human needs to perform and / or the target voices that need to be broadcast, and controls the in-vehicle virtual human to perform the target actions and / or broadcast the target voices, thereby realizing the interaction process between the virtual human and the driver and passengers, and improving the driver and passengers' driving experience and immersion.
[0054] In one alternative implementation, after projecting the target virtual human model to generate the in-vehicle virtual human, the method further includes:
[0055] If no communication packet is received from the vehicle's infotainment system within a preset time, the virtual human in the vehicle will be controlled to perform a preset standby action and / or broadcast a preset standby voice message.
[0056] Beneficial effects: In this embodiment, if no new communication packets are acquired within a preset time, that is, if the virtual human does not interact with the user, the virtual human is controlled to enter a leisure state, and the in-vehicle virtual human is controlled to perform some leisure standby actions. Standby voice can also be played to further enhance the realism of the in-vehicle virtual human.
[0057] In one alternative implementation, after projecting the target virtual human model to generate the in-vehicle virtual human, the method further includes:
[0058] New virtual human imaging parameters are obtained based on the new communication packets sent by the vehicle's infotainment system.
[0059] The in-vehicle virtual human was adjusted based on the new virtual human imaging parameters.
[0060] Beneficial effects: In this embodiment, the vehicle-to-vehicle (V2V) machine requests the aerial imaging device to adjust parameters such as the imaging area, imaging size, and displayed virtual human model. It will also issue new business instructions to adjust the virtual human's motion and voice parameters. After receiving the new communication packet sent by the V2V, the aerial imaging device will adjust according to the corresponding parameters to achieve real-time adjustment of the virtual human and enhance the user experience.
[0061] In one alternative implementation, the vehicle-mounted virtual human is adjusted according to the new virtual human imaging parameters, including:
[0062] Adjust the projection duration of the in-vehicle virtual human based on the new imaging trigger type;
[0063] And / or, adjust the projection position of the in-vehicle virtual human based on the new imaging area.
[0064] Beneficial effects: In this embodiment of the invention, users are allowed to adjust the projection duration and projection position of the in-vehicle virtual human, thereby improving the user experience.
[0065] Thirdly, the present invention provides an in-vehicle virtual human generation system, including: an in-vehicle system and an aerial imaging device;
[0066] The vehicle-mounted system acquires communication information generated by at least one business module within the system; based on the communication information, it determines the virtual human imaging parameters; based on the virtual human imaging parameters, it generates a communication packet and sends the communication packet to the aerial imaging device.
[0067] The aerial imaging equipment receives communication packets sent by the vehicle-mounted unit; it parses the communication packets to obtain virtual human imaging parameters; and it renders the virtual human based on these parameters to generate an in-vehicle virtual human.
[0068] Beneficial effects: By acquiring communication information generated by at least one business module through the vehicle's infotainment system, virtual human imaging parameters that can reflect the driver's intention and vehicle driving trend can be determined. Then, communication packets are generated based on the virtual human imaging parameters and sent to the airborne imaging equipment. After decoding the communication packets, the airborne imaging equipment renders them according to the obtained virtual human imaging parameters to generate an in-vehicle virtual human with good rendering effect. This does not consume the rendering performance of the vehicle's infotainment system, reducing the vehicle's energy consumption. Moreover, the virtual human imaging does not depend on the vehicle's infotainment system screen, resulting in a stronger sense of user immersion.
[0069] In one alternative implementation, the aerial imaging equipment is mounted in the passenger seat, rear seat, and / or driver's cab of the vehicle.
[0070] Beneficial effects: By installing aerial imaging equipment in multiple areas inside the vehicle, the imaging range is expanded, allowing virtual human figures to be displayed in multiple areas inside the vehicle, which helps to improve the driving experience for multiple passengers.
[0071] Fourthly, the present invention provides a vehicle including the in-vehicle virtual human generation system of the third aspect above or any corresponding embodiment thereof.
[0072] Fifthly, the present invention provides a vehicle infotainment system, which includes a first memory and a first processor, the first memory and the first processor being communicatively connected to each other, the first memory storing first computer instructions, and the first processor executing the first computer instructions to perform the vehicle virtual human generation method of the first aspect or any corresponding embodiment described above.
[0073] In a sixth aspect, the present invention provides an aerial imaging device, which includes a second memory and a second processor, the second memory and the second processor being communicatively connected to each other, the second memory storing second computer instructions, and the second processor executing the vehicle-mounted virtual human generation method of the second aspect or any corresponding embodiment thereof by executing the second computer instructions.
[0074] The beneficial effects of this invention are as follows:
[0075] By acquiring communication information from at least one business module within the vehicle's infotainment system, virtual human imaging parameters reflecting the driver's intentions and vehicle movement trends are determined. Communication packets are then generated based on these parameters and sent to an aerial imaging device. The aerial imaging device decodes the communication packets and renders them according to the obtained virtual human imaging parameters, generating an in-vehicle virtual human with excellent rendering quality. This approach eliminates the need to consume the vehicle's rendering performance, reducing energy consumption, and the virtual human imaging is independent of the vehicle's screen, resulting in a more immersive user experience. Furthermore, this invention can utilize communication information from different business modules within the vehicle for virtual human imaging, leading to richer imaging effects. Attached Figure Description
[0076] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0077] Figure 1 This is a structural block diagram of an in-vehicle virtual human generation system according to an embodiment of the present invention;
[0078] Figure 2 This is a schematic diagram of the installation position of an aerial imaging device according to an embodiment of the present invention;
[0079] Figure 3 This is a schematic diagram of the interaction process of an in-vehicle virtual human generation system according to an embodiment of the present invention;
[0080] Figure 4 This is a schematic diagram of another interaction process of the in-vehicle virtual human generation system according to an embodiment of the present invention;
[0081] Figure 5 This is a schematic diagram of an aerial imaging process according to an embodiment of the present invention;
[0082] Figure 6 This is a schematic diagram of a TCP communication process according to an embodiment of the present invention;
[0083] Figure 7 This is a communication flowchart of a vehicle-mounted system and an aerial imaging device according to an embodiment of the present invention;
[0084] Figure 8 This is a schematic diagram of a virtual human rendering image according to an embodiment of the present invention;
[0085] Figure 9 This is a schematic diagram of another virtual human rendering image according to an embodiment of the present invention;
[0086] Figure 10 This is a schematic diagram of the architecture of an in-vehicle virtual human generation system according to an embodiment of the present invention;
[0087] Figure 11 This is a structural block diagram of a vehicle according to an embodiment of the present invention;
[0088] Figure 12 This is a schematic diagram of the structure of the vehicle system according to an embodiment of the present invention;
[0089] Figure 13 This is a schematic diagram of the structure of an aerial imaging device according to an embodiment of the present invention. Detailed Implementation
[0090] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0091] Currently, in a period without significant advancements in automotive engine technology, enriching cabin features such as intelligent driving, voice assistants, premium audio systems, and cinematic movie theaters have become crucial areas for enhancing the driving experience. Global automakers are investing heavily in creating the ultimate intelligent cockpit experience, redeveloping traditional in-vehicle system modules using new technologies. For example, they are using 3D rendering technology to transform 2D car desktops into 3D desktops, allowing drivers to use vehicle functions in a more intuitive and WYSIWYG manner. Another example is Advanced Driver Assistance Systems (ADAS), which utilize 3D world reconstruction technology to render the surrounding environment on the in-vehicle screen, significantly improving driving safety.
[0092] In-vehicle virtual humans, as a smart cockpit feature, can provide users with an ultimate driving experience. However, traditional in-vehicle virtual humans have some problems: 1) They can only rely on the vehicle's infotainment screen for display and cannot be displayed off-screen; 2) The rendering effect is related to the screen size. The effect is poor on small screens, making it difficult to see clearly and lacking in detail; large screens have better effects, but consume more resources; 3) Due to the limitations of the vehicle's infotainment system, running multiple high-resource modules (such as voice and navigation) at the same time can cause the image rendering program to lag; 4) They are prone to rendering area conflicts with other apps, such as the virtual human image obscuring map navigation content; 5) Because the image is only displayed on a 2D screen, it cannot interact with specific objects in the cabin (such as air conditioning vents).
[0093] The implementation of a new scientific technology often brings about significant changes to related industries. Air Imaging Technology (or Free-space Imaging Technology) is an advanced display technology that can display images in the air without a physical medium (such as a screen or projection screen). This technology utilizes the principles of light reflection, refraction, or scattering, using specific optical devices to suspend images in the air, creating a visual effect of suspension. Air Imaging Technology directly brings many scenes previously only seen in science fiction movies into people's real lives, with applications covering almost every aspect of life. If it could be used in a car's cockpit, one or more realistic 3D virtual human figures could be projected onto multiple areas of the cabin as needed. If the virtual human figures could be customized by the car owner, given specific voices and movements, and provide comprehensive driving assistance throughout the journey, it would undoubtedly greatly enhance the driving experience, making users feel as if they were in a futuristic scenario.
[0094] Therefore, this invention provides a vehicle-mounted virtual human generation scheme, which transmits communication data collected by the vehicle to an aerial imaging device. The aerial imaging device generates a vehicle-mounted virtual human based on medium-free aerial imaging technology, and projects one or more realistic 3D virtual human images into multiple areas within the cabin as needed, thereby greatly enhancing the experience of drivers and passengers.
[0095] According to an embodiment of the present invention, an in-vehicle virtual human generation system 100 is provided, including: a vehicle-mounted infotainment system 101 and an aerial imaging device 102. The vehicle-mounted infotainment system 101 can act as a client, and the aerial imaging device 102 can act as a server. The client and server communicate according to a preset communication protocol. It should be noted that the vehicle-mounted infotainment system 101, i.e., the client, can include various modules of the vehicle-mounted infotainment system (provided that the vehicle-mounted infotainment system has access to an upper-layer application development SDK), such as an aerial imaging device setting module, a voice module, a navigation module, and a sensor module, etc. The specific configuration can be determined according to actual application requirements, and the present invention is not limited thereto.
[0096] The vehicle-mounted unit 101 acquires communication information generated by at least one business module in the vehicle-mounted unit; determines virtual human imaging parameters based on the communication information; generates a communication packet based on the virtual human imaging parameters, and sends the communication packet to the aerial imaging device 102.
[0097] The aerial imaging device 102 receives communication packets sent by the vehicle-mounted unit 101; it parses the communication packets to obtain virtual human imaging parameters; and it renders the virtual human based on the virtual human imaging parameters to generate an in-vehicle virtual human.
[0098] The specific working principles and processes of the vehicle-mounted unit 101 and the aerial imaging device 102 are described in the relevant descriptions of the method embodiments below, and will not be repeated here.
[0099] In some alternative implementations, the aerial imaging device 102 can be mounted in the passenger seat, rear seat, and / or dashboard of the vehicle. The installation of the aerial imaging device 102 is typically performed by the vehicle manufacturer as an original equipment manufacturer (OEM). Figure 2 As shown, the possible installation areas are: the front passenger seat, the rear seats, and the dashboard. In this embodiment, the images projected onto the front passenger seat and the rear seats will be approximately the size of a real person, while the image projected onto the dashboard will be approximately 30cm*30cm*30cm.
[0100] In this embodiment, the aerial imaging device 102 employs two imaging strategies: selective imaging and on-demand imaging. The driver can switch between these two strategies at any time during the journey. Selective imaging means the vehicle's central control system 101 automatically traverses multiple imaging areas and selects the first empty area without passengers for projection, resulting in potentially different imaging areas each time. On-demand imaging, on the other hand, assigns a unique identifier ID to each imaging area, allowing the driver to select a specific area for projection via the vehicle's central control system, regardless of whether passengers are present in that area.
[0101] This embodiment expands the imaging range by installing aerial imaging equipment in multiple areas inside the vehicle, enabling virtual human figures to be displayed in multiple areas inside the vehicle, which helps to improve the driving experience for multiple drivers and passengers.
[0102] In some optional implementations, the imaging triggering mechanism of the aerial imaging device 102 can be divided into two types: constant projection and on-demand projection. The driver can switch between these two mechanisms at any time while driving. Constant projection means that as soon as the driver starts the vehicle, the vehicle's central control system will project an image according to the imaging area strategy described above, and the projection will remain until the vehicle is turned off or the driver actively changes the imaging area strategy through the central control system. At this time, if the virtual human does not interact with humans (such as responding to voice commands, navigation commands, or chatting with passengers or the driver), the virtual human will enter a leisure state, performing some leisurely standby actions and playing standby voice prompts.
[0103] Specifically, on-demand projection means that the image will only be projected when there is a need to interact with the virtual human (such as responding to voice commands, navigation commands, or chatting and greeting passengers or drivers). The image will be projected according to the imaging area strategy described above. Once the interaction is completed, the virtual human will disappear and wait for the next opportunity to interact.
[0104] The in-vehicle virtual human generation system provided in this embodiment of the invention utilizes the vehicle's infotainment system 101 to obtain communication information generated by at least one business module, thereby determining virtual human imaging parameters that can reflect the driver's intention and the vehicle's driving trend. Then, based on the virtual human imaging parameters, a communication packet is generated and sent to the aerial imaging device 102. After decoding the communication packet, the aerial imaging device 102 renders it according to the obtained virtual human imaging parameters to generate an in-vehicle virtual human. The rendering effect is good, which does not consume the rendering performance of the vehicle's infotainment system, reducing the vehicle's energy consumption. Moreover, the virtual human imaging does not depend on the vehicle's infotainment system screen, resulting in a stronger sense of user immersion.
[0105] According to an embodiment of the present invention, an embodiment of a method for generating a vehicle-mounted virtual human is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0106] This embodiment provides a method for generating in-vehicle virtual humans, which can be used for, for example... Figure 1 The vehicle-mounted unit 101 and the aerial imaging device 102 shown are shown. Figure 3 This is a schematic diagram of the interaction process of an in-vehicle virtual human generation system according to an embodiment of the present invention. The vehicle-mounted system 101 is used to execute steps S101 to S103, and the aerial imaging device 102 is used to execute steps S201 to S203. The specific interaction process between the vehicle-mounted system 101 and the aerial imaging device 102 is as follows:
[0107] Step S101: Obtain communication information generated by at least one service module in the vehicle infotainment system.
[0108] Specifically, the vehicle infotainment system can include multiple service modules, as long as it can obtain the communication information of each service module. Service modules can be voice modules, navigation modules, sensor modules, parking modules, multimedia modules, ADAS modules, radar sensing modules, braking modules, etc. All of these service modules can communicate with the vehicle infotainment system, and the vehicle infotainment system will uniformly arbitrate and process the communication information of these service modules.
[0109] Step S102: Determine the virtual human imaging parameters based on the communication information.
[0110] Specifically, the vehicle machine will identify, extract, and process the communication information sent by various business modules, analyze the user's intentions and the vehicle's driving trends, and obtain virtual human imaging parameters that can reflect the user's intentions or the vehicle's driving trends. In this way, imaging can be performed based on the virtual human imaging parameters.
[0111] For example, the voice module of the vehicle system will collect the driver's speech through the microphone and recognize the driver's intention through the pre-installed AI voice recognition module. The navigation module of the vehicle system will obtain the vehicle's driving trend based on the car's location and the map navigation route. The vehicle system will further analyze the virtual human's imaging parameters, such as the actions the virtual human needs to perform and the voice it needs to emit, based on the driver's intention and the vehicle's driving trend.
[0112] Step S103: Generate a communication packet based on the virtual human imaging parameters and send the communication packet to the aerial imaging device.
[0113] Specifically, since there may be multiple manufacturers of aerial imaging equipment, each with different hardware and development SDKs, without a unified communication standard, each project would need to develop different communication protocols for different aerial imaging equipment manufacturers. This would undoubtedly increase project maintenance costs and development complexity. Therefore, this invention provides a pre-defined communication protocol for a unified business layer independent of the aerial imaging equipment. Communication between the vehicle-mounted system and the aerial imaging equipment is achieved according to this pre-defined protocol, enabling the vehicle-mounted system to send virtual human imaging parameters to the aerial imaging equipment. Furthermore, this invention can utilize communication information from different business modules of the vehicle for virtual human imaging, resulting in richer imaging effects.
[0114] Step S201: Receive the communication packet sent by the vehicle's infotainment system.
[0115] Since there may be multiple manufacturers of aerial imaging equipment, each with different hardware and SDK development methods, without a unified communication standard, each project would need to develop different communication protocols for different aerial imaging equipment manufacturers. This would undoubtedly increase project maintenance costs and development difficulty. Therefore, this embodiment of the invention implements communication between the vehicle-mounted system and the aerial imaging equipment according to a preset communication protocol, enabling the aerial imaging equipment to receive communication packets sent by the vehicle-mounted system.
[0116] Step S202: Parse the communication packet to obtain the virtual human imaging parameters.
[0117] Specifically, after receiving a communication packet, the aerial imaging equipment will unpack the packet according to the communication header and communication body, decode it, and obtain detailed virtual human imaging parameters.
[0118] Step S203: Render the virtual human based on the virtual human imaging parameters to generate an in-vehicle virtual human.
[0119] Specifically, a 3D graphics rendering library can be pre-installed in the aerial imaging device to draw complex graphics in the air. The aerial imaging device performs rendering according to the virtual human imaging parameters and projects to generate a virtual human.
[0120] The in-vehicle virtual human generation method provided in this embodiment obtains communication information generated by at least one business module through the vehicle's infotainment system, thereby determining virtual human imaging parameters that reflect the driver's intention and vehicle driving trend. Then, communication packets are generated based on the virtual human imaging parameters and sent to an aerial imaging device. After decoding the communication packets using the aerial imaging device, the virtual human is rendered according to the obtained virtual human imaging parameters to generate an in-vehicle virtual human. The rendering effect is good, thus reducing the vehicle's infotainment system's energy consumption without consuming its rendering performance. Furthermore, the virtual human imaging does not rely on the vehicle's infotainment system screen, resulting in a stronger user immersion experience.
[0121] This embodiment provides a method for generating in-vehicle virtual humans, which can be used for, for example... Figure 1 The vehicle-mounted unit 101 and the aerial imaging device 102 shown are shown. Figure 4 This is a schematic diagram of the interaction process of an in-vehicle virtual human generation system according to an embodiment of the present invention. The vehicle-mounted system 101 is used to execute steps S301 to S305, and the aerial imaging device 102 is used to execute steps S401 to S203. The specific interaction process between the vehicle-mounted system 101 and the aerial imaging device 102 is as follows:
[0122] Step S301: Obtain communication information generated by at least one service module in the vehicle's infotainment system. For details, please refer to [link / reference]. Figure 3 Step S101 of the illustrated embodiment will not be described again here.
[0123] Step S302: Filter the communication information to obtain interactive data sent by at least one service module; wherein the service module includes an aerial imaging equipment setting module, a voice module, and a navigation module.
[0124] Specifically, such as Figure 5 As shown, vehicles can be equipped with hardware devices such as 360-degree cameras, infrared devices, and microphones to collect communication information such as external environmental information, so that various business modules of the vehicle system (such as voice modules, navigation modules, etc.) can send the corresponding communication information to the vehicle system.
[0125] Specifically, since some communication information does not contain data suitable for imaging, it is first necessary to filter the communication information to identify the interactive data from each business module that can be used to interact with the virtual human. This filtering can be based on keyword type, the type of recipient of the communication information, etc. If the communication information contains preset keywords and a recipient, the corresponding information is treated as imageable interactive data, facilitating the analysis of the virtual human's imaging parameters and improving data processing efficiency.
[0126] Step S303: Obtain virtual human imaging parameters based on the business instructions corresponding to the interactive data.
[0127] The virtual human imaging parameters may include imaging trigger type, imaging area, imaging volume, virtual human model parameters, virtual human action parameters and / or virtual human voice parameters, which can be set according to the actual scenario.
[0128] In this embodiment of the invention, the types and descriptions of virtual human imaging parameters are shown in Table 1 below:
[0129] Table 1
[0130]
[0131] It should be noted that the virtual human's voice parameters can include the virtual human's timbre, volume, and the content of the voice broadcast, which can be set according to the actual application scenario.
[0132] In one optional implementation, the service instruction may include a voice instruction, the virtual human action parameters may include a target action corresponding to the voice instruction, and the virtual human voice parameters may include a target voice corresponding to the voice instruction. First, the user's voice instruction can be obtained by extracting and recognizing the interaction data sent by the voice module. Then, based on a preset mapping relationship between the voice instruction and the virtual human imaging parameters, the target action and / or target voice corresponding to the voice instruction is confirmed.
[0133] Specifically, since the interactive data sent by the voice module may contain some background noise and other interference data, key information in the interactive data can be extracted and identified through keyword extraction algorithms and filtering algorithms to filter out background noise, thereby accurately identifying the target user's intent and obtaining language commands.
[0134] For example, the vehicle's voice module collects the driver's speech through the microphone and identifies the driver's intent through the AI voice recognition module. For instance, if the driver says, "Set the air conditioning temperature to 25 degrees," the business layer of the vehicle's client voice module will generate the voice command "Set the air conditioning temperature to 25 degrees." The vehicle's system pre-stores a mapping relationship between voice commands and virtual human imaging parameters. The system then queries this pre-stored mapping relationship based on the voice command to obtain the corresponding target action, such as "pointing the left hand towards the air conditioning vent."
[0135] In this embodiment of the invention, the vehicle system analyzes the user's intent through the interactive data sent by the voice module, obtains the user's voice commands, and queries the pre-stored mapping relationship between voice commands and virtual human imaging parameters to determine the target action that the virtual human needs to perform and / or the target voice that needs to be broadcast, so as to facilitate the rendering and projection of the aerial imaging equipment, thereby realizing the interaction between the vehicle system's language module and the virtual human and improving the user's immersive experience during driving.
[0136] In some optional implementations, the service instructions may include navigation instructions, the virtual human motion parameters may include the target action corresponding to the navigation instructions, and the virtual human voice parameters may include the target voice corresponding to the navigation instructions. The vehicle system first obtains the navigation instructions based on the interaction data sent by the navigation module. Then, based on the preset mapping relationship between navigation instructions and virtual human imaging parameters, it confirms the target action and / or target voice corresponding to the navigation instructions.
[0137] For example, the vehicle navigation module can identify in advance that a right turn is required at the next intersection based on the car's location and the map navigation route. The corresponding navigation instruction could be "turn right at the next intersection". The vehicle system can then retrieve the target action "pointing right to the right" and the target language "turn right at the next intersection" based on the pre-stored mapping relationship between navigation instructions and virtual human imaging parameters.
[0138] In this embodiment of the invention, the vehicle-mounted system analyzes the vehicle's driving trend through interactive data sent by the navigation module to obtain navigation instructions. Then, based on the pre-stored mapping relationship between navigation instructions and virtual human imaging parameters, it queries the target action and / or target voice corresponding to the navigation instructions, which facilitates rendering and projection by the aerial imaging equipment. This enables interaction between the vehicle-mounted system navigation module and the virtual human, thereby improving the user's immersive experience during driving.
[0139] In some optional implementations, the service instructions may include imaging parameter setting instructions, virtual human action parameters including user-set initial virtual human actions, and virtual human voice parameters including user-set initial virtual human voice. The vehicle-mounted system can obtain the imaging parameter setting instructions issued by the user for the aerial imaging equipment based on the interactive data sent by the aerial imaging equipment setting module. Based on the imaging parameter setting instructions, the system obtains the imaging trigger type, imaging area, imaging volume, virtual human model parameters, initial virtual human actions, and / or initial virtual human voice set by the user for the virtual human.
[0140] For example, an aerial imaging device setting module, such as an operating app, can be provided at the vehicle infotainment application layer. Users can use this app to select the imaging area, imaging volume, and specific virtual human model, and preset the initial actions and / or voice of the virtual human. Simultaneously, the app generates corresponding imaging parameter setting instructions, allowing the vehicle infotainment system to determine the virtual human imaging parameters based on these instructions. Alternatively, the vehicle infotainment system can also determine the imaging area, imaging volume, and specific virtual human model parameters independently based on the vehicle's interior space occupancy, improving imaging flexibility and interior space utilization.
[0141] In some optional implementations, users can first set the virtual human imaging parameters through the aerial imaging device setting module, such as setting the imaging area, imaging trigger type, and virtual human model used, so that the aerial imaging device can initially generate a virtual human. During the subsequent driving process, the vehicle system can further combine the voice commands issued by the user through the voice module and the corresponding navigation commands from the navigation module to redetermine the actions and / or voices that the virtual human needs to perform. This enables the virtual human to provide action and / or voice feedback to the user's intentions or the vehicle's driving intentions during the driving process, thereby improving the user experience.
[0142] In this embodiment of the invention, users are allowed to issue imaging parameter setting commands through the aerial imaging device setting module to set the imaging trigger type, imaging area, imaging volume, virtual human model parameters, virtual human action parameters and / or virtual human voice parameters of the virtual human, thereby improving the richness of imaging and the user experience.
[0143] In some optional implementations, the mapping relationship between voice commands, navigation commands, and virtual human imaging parameters can be stored not only in the vehicle's infotainment system but also in the aerial imaging device. Accordingly, the vehicle's infotainment system can directly use voice commands, navigation commands, and other business commands as specific virtual human action parameters and virtual human voice parameters. The aerial imaging device will then query according to the preset mapping relationship to determine the action or voice message the virtual human needs to perform. For example, for the business command "set the air conditioning temperature to 25 degrees Celsius," the vehicle's infotainment system can package and send it to the aerial imaging device according to a preset communication protocol. After receiving the communication packet, the aerial imaging device will decode it and, based on the preset mapping relationship between business commands and virtual human actions, determine that the virtual human's action is "pointing its left hand towards the air conditioning vent." As another example, for the business command "play news," the vehicle's infotainment system can directly send this business command as a virtual human voice parameter to the aerial imaging device. The aerial imaging device will then access the corresponding news app interface according to the preset mapping relationship between business commands and virtual human voice to obtain the voice content the virtual human needs to broadcast.
[0144] In this embodiment of the invention, by extracting and identifying the communication information of various business modules of the vehicle, business instructions that can reflect the driver's intention and the vehicle's driving trend are generated. In order to determine the imaging trigger type, imaging area, imaging volume, virtual human model, and the actions and voice to be performed by the virtual human according to the business instructions, the interaction between the virtual human and the user is realized, and the driving experience is enhanced.
[0145] Step S304: Byte-encode the virtual human imaging parameters to generate a communication body, and generate a communication header based on the byte length of the communication body and the version information of the preset communication protocol.
[0146] Specifically, the basic format of the preset communication protocol is a communication header + communication body. The communication body mainly consists of specific business data, such as virtual human imaging parameters. During transmission, this data is actually encoded into a byte stream. The specific business format can be common data formats such as JSON and XML, or it can be a custom data format. This embodiment of the invention is not limited to this.
[0147] Furthermore, the communication header will at least contain the version information of the communication protocol and the byte length of the communication body. The version information is mainly used for protocol upgrades and compatibility strategies between different protocol versions. The byte length of the communication body refers to the byte length of the specific business data in the communication stream. This data is mainly used for packet splicing and reassembly in the communication stream to avoid reading data that does not belong to the current message.
[0148] Step S305: Obtain the communication packet based on the communication header and communication body, and send the communication packet to the air imaging device according to the preset communication protocol.
[0149] Specifically, a communication packet is synthesized according to the basic format of communication header + communication body. In subsequent communication, the vehicle or aerial imaging equipment reads a specific length of byte data from the communication stream according to the byte length of the communication header. Then, the business layer decodes the communication body to obtain common data formats (json, xml), thus completing a communication.
[0150] In some optional implementations, the vehicle-mounted unit sends a synchronization message to the airborne imaging device to request the establishment of a connection. After receiving a synchronization confirmation message returned by the airborne imaging device based on the synchronization message, the vehicle-mounted unit sends a first confirmation message to the airborne imaging device to confirm the establishment of the connection, thus establishing a connection with the airborne imaging device. Then, the vehicle-mounted unit sends communication packets to the airborne imaging device. When it detects that the communication packet transmission is complete, it sends a termination message to the airborne imaging device to request the closure of the connection. Finally, after receiving a second confirmation message from the airborne imaging device indicating receipt of the termination message and a termination confirmation message returned based on the termination message, the vehicle-mounted unit sends a third confirmation message to the airborne imaging device to confirm the closure of the connection, thus disconnecting from the airborne imaging device.
[0151] In this embodiment, the communication standard can be built based on the Transmission Control Protocol (TCP). Figure 6 This is a basic flowchart of TCP communication. TCP is a connection-oriented, reliable, stream-based communication protocol widely used for data transmission in the Internet and local area networks. The principle of using TCP to implement communication can be divided into the following steps (steps A to C):
[0152] Step A: TCP establishes a connection through a three-way handshake. The process of establishing a TCP connection is called a three-way handshake, and its process is shown in steps a1 to a3:
[0153] Step a1, First handshake (client sends SYN).
[0154] Specifically, the client (i.e., the vehicle-mounted system) sends a SYN (Synchronize) message to the server (i.e., the aerial imaging device) to indicate that it wants to establish a connection and sets an initial sequence number, where the message segment is: SYN = 1, sequence number = x.
[0155] Step a2, the second handshake (the server responds with SYN-ACK).
[0156] Specifically, after receiving a SYN packet from the client, the server replies with a SYN-ACK packet (synchronization acknowledgment packet) to indicate its agreement to establish a connection and also sets an initial sequence number. The packet contains the following parameters: SYN = 1, ACK = 1, sequence number = y, and acknowledgment number = x + 1.
[0157] Step a3, the third handshake (client responds with ACK).
[0158] Specifically, after receiving the SYN-ACK packet from the server, the client replies with an ACK packet (the first acknowledgment packet) to confirm the connection establishment. In this packet, ACK = 1, sequence number = x + 1, and acknowledgment number = y + 1. At this point, the three-way handshake is complete, and the TCP connection is established.
[0159] Step B, Data Transmission. After establishing a TCP connection, the client and server can transmit data. TCP provides full-duplex communication, allowing both parties to send and receive data simultaneously. The data transmission process includes:
[0160] Step b1, data fragmentation and reassembly: Data is divided into multiple segments or communication packets for transmission, and the receiver reassembles the segments in sequence according to the sequence number.
[0161] Step b2, Flow Control: TCP uses a sliding window mechanism to control the data transmission rate of the sender in order to prevent network congestion.
[0162] Step b3, Reliable Transmission: Reliable data transmission is ensured through an acknowledgment (ACK) mechanism and a retransmission mechanism. After sending data, the sender waits for acknowledgment from the receiver. If no acknowledgment is received within a certain time, the sender will retransmit the data.
[0163] Step C: TCP four-way handshake to close the connection. The process of closing a TCP connection is called a four-way handshake, and its steps are as follows:
[0164] Step c1, first handshake (client sends FIN).
[0165] Specifically, the client sends a FIN (Finish) message (termination message) to indicate that it wishes to close the connection and stop sending data. The message segment contains: FIN = 1, sequence number = u.
[0166] Step c2, the second handshake (server responds with ACK).
[0167] Specifically, after receiving the FIN packet, the server replies with an ACK packet (second acknowledgment packet) to acknowledge receipt of the close request. The packet contains: ACK = 1, sequence number = v, and acknowledgment number = u+1.
[0168] Step c3, the third handshake (the server sends FIN).
[0169] Specifically, after processing all the data, the server sends a FIN message (Termination Acknowledgment message) to indicate its agreement to close the connection and to stop sending data. The message segment contains: FIN = 1, sequence number = w.
[0170] Step c4, the fourth handshake (client responds with ACK). After receiving the server's FIN packet, the client replies with an ACK packet (third acknowledgment packet), indicating confirmation of the close request. In this packet, ACK = 1, sequence number = u+1, and acknowledgment number = w+1. At this point, the four-way handshake is complete, and the TCP connection is closed.
[0171] In this embodiment of the application, by sending communication packets to the aerial imaging equipment according to a preset communication protocol, the reliability, integrity and accuracy of data transmission are improved, and there is no need to develop different communication standards for different aerial imaging equipment manufacturers, which reduces project maintenance costs and development difficulty.
[0172] In this embodiment of the invention, the virtual human imaging parameters are byte-encoded according to a preset communication protocol to generate a communication body, and a communication header is generated according to the byte length of the communication body and the protocol version information. The communication packet containing the communication header and the communication body is sent to the aerial imaging device, which facilitates the aerial imaging device to unpack the packet, avoids reading erroneous data, and improves the accuracy of information transmission.
[0173] Step S401: Receive the communication packet sent by the vehicle's infotainment system. For details, please refer to [link / reference]. Figure 3 Step S201 of the illustrated embodiment will not be described again here.
[0174] Step S402: Parse the communication packet to obtain the virtual human imaging parameters. For details, please refer to [link to relevant documentation]. Figure 3 Step S202 of the illustrated embodiment will not be described again here.
[0175] Step S403: Update the initial state machine based on the virtual human imaging parameters.
[0176] like Figure 7 As shown, this embodiment of the invention provides a rendering imaging interface driver layer in an aerial imaging device. The rendering imaging interface driver layer is a three-dimensional graphics rendering library that runs on a mediumless aerial imaging device. It provides a set of functions for drawing complex graphics in the air.
[0177] Specifically, the rendering and imaging interface driver layer is a state machine whose behavior depends on its current state. The current state includes the current color, texture, matrix, etc., and these states can be set and modified through relevant functions. All driver layer states are stored in a context environment, which can be created and managed by the upper-layer application, and the rendering area for imaging can be defined within the context environment.
[0178] Furthermore, the rendering and imaging interface driver layer uses a graphics pipeline to process graphics data. This pipeline includes multiple stages, from vertex processing to fragment processing (also known as voxel processing). Each stage can be controlled by programming (shaders). Its main stages include: defining model vertex data, compiling vertex shaders, coordinate transformation (from model space to projection space), rasterization (which generates voxels in projection space, similar to 2D pixels), and compiling fragment shaders (which mainly calculate lighting and determine the color of each voxel). For details, please refer to the description of the relevant technology, which will not be repeated here.
[0179] In this embodiment, the initial state machine can be updated using virtual human imaging parameters and functions in the 3D graphics rendering library, thereby using the updated state machine for rendering and imaging.
[0180] Step S404: Render the model using the updated state machine to obtain the target virtual human model.
[0181] Specifically, when rendering using the rendering and imaging interface driver layer, the driver layer context environment is first created and initialized. Then, using the virtual human imaging parameters and functions in the 3D graphics rendering library, the initial state machine is updated, the projection area is set (defining the projection position and size of the solid), the projection matrix and view matrix are set, and vertex data is defined (vertex data defines the shape of the geometry and can include information such as position, color, normal, and texture coordinates). Next, the shader is written and compiled (the shader is a small program that runs on the aerial imaging device and is used to control the processing of each voxel). Finally, the rendering loop is performed (usually, rendering commands are executed in a loop, including clearing the screen, drawing geometry, swapping buffers, etc.), to obtain the target virtual human model.
[0182] It's important to note that the projection matrix is primarily a transformation matrix used to project a 3D model onto a specific area of space. The view matrix can also be categorized into translation, rotation, and scaling. The view matrix is the inverse of the model matrix representing the observer's transformation within the world, obtained by treating the observer as a model. If the observer's coordinates are translated by (tx, ty, tz), the view matrix is as follows. It can be seen that if the view matrix is considered as the model matrix of the entire world, it's equivalent to the entire world being translated by (-tx, -ty, -tz):
[0183]
[0184] If the observer rotates by an angle θ around the z-axis, the view matrix is as follows, which is equivalent to the entire world rotating by -θ degrees around the z-axis:
[0185]
[0186] If the observer is scaled down by a factor of s in three directions, the resulting view matrix is as follows, which is equivalent to the entire world being magnified by a factor of s:
[0187]
[0188] This invention utilizes the virtual human imaging parameters sent by the vehicle-mounted unit to update the initial state machine of the aerial imaging device. The updated state machine can then perform graphic drawing and projection according to the virtual human imaging parameters, resulting in good rendering effects. This eliminates the need for the vehicle-mounted unit to perform imaging rendering, thus reducing its energy consumption.
[0189] In some alternative implementations, the updated state machine can be used for geometry rendering, texture mapping, lighting and shading, color processing, depth testing, and template testing to obtain the target virtual human model.
[0190] Specifically, geometry drawing is the basic drawing command provided by the rendering and imaging interface driver layer, such as drawArrays and drawElements, used to draw geometric objects such as points, lines, and triangles; texture mapping is used to apply image textures to the surface of geometric objects, and it supports a variety of texture operations, including texture loading, texture binding, setting texture parameters, and texture filtering; lighting and shadow processing achieves realistic lighting and shadow effects through shaders and lighting models (such as the Phong model); color processing is used to control semi-transparency effects by mixing different color values; frame buffer (FBO) is used for off-screen rendering, allowing the rendering results to be stored in textures or render buffers instead of being directly displayed on the imaging area; depth testing is used to hide occluded objects; and stencil testing is used for complex clipping and blending effects.
[0191] For example, to render a character image, you can follow these steps:
[0192] Step d1 first provides a 3D model file such as .fbx, .obj, .gltf, etc. These files are composed of many three-dimensional points (Vector3), which form the outline of the character and are a necessary input for the geometry drawing process.
[0193] Step d2 provides several image textures (base color map, normal map, roughness map, etc.). The base color map determines the color of the character model, such as black hair and yellow skin. The normal map represents the surface roughness, lighting, etc., of the model. The roughness map indicates where the model's surface is smooth and where it is rough. Texture mapping can perfectly overlay these texture maps onto the model's surface. These maps, as optional inputs to the rendering pipeline, improve the realism of the rendered results.
[0194] In step d3, after receiving the above input, the aerial imaging device performs lighting calculations on the virtual human model based on algorithms for generating lighting and shadows (such as PBR algorithm, Phone algorithm), obtaining the lighting color of each voxel in the imaging area (which can be understood as the color we ultimately see), producing the same lighting effect as reality. After the lighting calculation, each voxel in the imaging area has a color. Some culling may be required through depth testing and stencil testing. For example, when viewed from the front of the character, voxels that are obscured, such as internal organs and the back, are not visible. These colors need to be culled to prevent interference with the rendering results, and also to save performance.
[0195] Step d4: Through the above steps, a model with a stereoscopic rendering effect will be obtained, or multiple 3D voxel color arrays. These color arrays cannot be projected directly, but are first stored in a frame buffer (FBO). When the imaging device refreshes the imaging area next time, the contents of this frame buffer will be projected. This reduces image flicker and achieves a stable rendering effect. In this way, the aerial imaging device completes the rendering and imaging process.
[0196] In this embodiment, an updated state machine is used to draw the geometry, obtaining the outline of the virtual human. Then, texture mapping is performed to enrich the virtual human's color, bumpiness, and roughness, improving its realism. Next, lighting and color processing are performed to determine the illumination color of each voxel in the imaging area. Finally, the rendered effect is stored in a frame buffer before subsequent projection, reducing image flicker and improving rendering stability. Furthermore, depth testing and stencil testing are used to filter noise and interference, further enhancing the realism of the rendering.
[0197] Step S405: Project the target virtual human model to generate a vehicle-mounted virtual human.
[0198] Specifically, the aerial imaging equipment projects the target virtual human model onto the target area according to parameters such as imaging area and imaging volume, generating an in-vehicle virtual human. The aerial imaging equipment can be mounted in the passenger seat, rear seats, and / or driver's seat of the vehicle, etc. (See again...) Figure 7The vehicle-mounted system can generate communication packets containing virtual human imaging parameters based on the communication information of different business modules, and send them to the rendering interfaces of multiple aerial imaging devices according to the standard TCP communication protocol. Each aerial imaging device uses its own rendering interface driver layer to perform imaging rendering, thereby projecting the vehicle-mounted virtual human in different areas of the vehicle, such as the passenger seat and the rear seats.
[0199] Step S406: Based on the virtual human's action parameters and / or virtual human's voice parameters, obtain the target action and / or target voice, and control the vehicle-mounted virtual human to perform the target action and / or broadcast the target voice.
[0200] For example, the vehicle's voice module collects the driver's speech through the microphone and identifies the driver's intent through the AI voice recognition module. For instance, if the driver says, "Set the air conditioning temperature to 25 degrees," the vehicle's voice module's business layer will generate a communication header plus a byte array containing "Set the air conditioning temperature to 25 degrees." Then, it will call the communication interface to send this message to the aerial imaging device. The aerial imaging device will receive and unpack the message, decoding it to obtain the specific executable message content. Depending on the product requirements, this message will cause the virtual human to play a gesture of pointing to the air conditioning vent with its left hand and showing "OK" with its right hand, indicating that the operation has been completed, giving the driver a concrete interactive feedback.
[0201] For example, during driving, cars rely on many sensors and cameras to perceive their surroundings. A simple scenario is turning right at an upcoming intersection. The navigation module, based on the car's position and the map navigation route, anticipates the need to turn right at the next intersection. A virtual human needs to play specific actions and sounds to prompt the driver to move to the right lane in advance. The vehicle navigation module then generates a communication header plus a communication body containing "turn right at the next intersection." The vehicle navigation service layer calls the communication interface to send this message to the aerial imaging device. The server-side aerial imaging device receives the message, decodes it to obtain the intent to play the specific actions and sounds, performs rendering, and completes the interaction.
[0202] In this embodiment, by controlling the in-vehicle virtual human to perform actions and / or broadcast voice according to the virtual human's action parameters and / or voice parameters, the interaction process between the virtual human and the driver / passengers is realized, thereby improving the driver / passengers' driving experience and immersion.
[0203] In some optional implementations, if no communication packet is received from the vehicle-mounted system within a preset time, the in-vehicle virtual human is controlled to perform preset standby actions and / or play preset standby voice messages. For example, if the aerial imaging device's imaging trigger mechanism is always-on display, meaning the aerial system in the vehicle-mounted system continuously projects when the vehicle is started, and if the vehicle-mounted system does not send a new communication packet within a preset time, i.e., the virtual human does not interact with the user, then the virtual human is controlled to enter a leisure state, performing some leisurely standby actions and playing standby voice messages to further enhance the realism of the in-vehicle virtual human.
[0204] In some alternative implementations, the aerial imaging equipment may also obtain new virtual human imaging parameters based on new communication packets sent by the vehicle-mounted system, and adjust the vehicle-mounted virtual human according to the new virtual human imaging parameters.
[0205] Specifically, the projection duration of the in-vehicle virtual human can be adjusted based on the new imaging trigger types. These types mainly include constant projection and on-demand projection. Constant projection corresponds to a projection duration from vehicle system startup to shutdown; that is, the aerial imaging device will continuously run and image until the vehicle system is shut down. On-demand projection allows users to set the projection duration and trigger timing themselves. Furthermore, the projection position of the in-vehicle virtual human can be adjusted based on the new imaging area.
[0206] In this embodiment of the invention, users are allowed to adjust the projection duration and projection position of the in-vehicle virtual human, thereby improving the user experience.
[0207] For example, the vehicle can request the aerial imaging device to adjust parameters such as the image size and the displayed virtual human model. It can also issue new business instructions to adjust the virtual human's motion and voice parameters. After receiving the new communication packet sent by the vehicle, the aerial imaging device will adjust according to the corresponding parameters to achieve real-time adjustment of the virtual human and enhance the user experience.
[0208] The in-vehicle virtual human generation method provided in this embodiment addresses the pain points of traditional virtual humans by utilizing medium-free aerial imaging to realize a virtual human driving assistant. This invention customizes a standard communication protocol between the aerial imaging device and the vehicle's infotainment system using TCP, and establishes rendering and imaging interfaces running on both the vehicle and aerial imaging ends. The vehicle's infotainment system is responsible for collecting data, status, and other communication information from various vehicle business modules, and transmitting the collected data and status to the aerial imaging device via the standard communication protocol. The aerial imaging device then displays images in specific areas through the rendering and imaging interface, such as... Figure 8 and Figure 9 As shown. Furthermore, the vehicle's infotainment system can directly access the rendering and imaging interface to directly control the aerial imaging equipment to display images or change image states.
[0209] This invention also provides an SDK for upper-layer application development, such as... Figure 10 As shown, this part is located in the JNI interface layer and is entirely developed in C++. This satisfies the high-speed interface call requirements of upper-layer applications, while also making it easier for the JNI interface layer to call the hardware capabilities of the aerial imaging device (because the underlying capabilities provided by imaging device manufacturers are usually also provided to the outside world using C++). The JNI interface layer contains the implementation of the communication standard and most of the interface calls of the rendering layer, and encapsulates the interfaces needed for the development of upper-layer applications.
[0210] This invention provides an SDK for upper-layer application development, primarily integrating communication standards and rendering interface capabilities. This allows upper-layer developers to focus solely on their business requirements without needing to understand the principles of aerial imaging equipment, significantly improving development efficiency.
[0211] This invention provides a standard rendering and imaging interface, primarily used to bridge the upper-layer application's call requirements and the driver layer of the aerial imaging equipment's rendering capabilities. This allows devices from different aerial imaging manufacturers to use this standard for aerial image rendering. In this invention, the aerial imaging equipment is mounted on the area of the cockpit that needs to be projected. Automobile manufacturers can rationally place the aerial imaging equipment in the correct position according to the cockpit layout of their vehicle models, and it also provides a fine-tuning function for the projection position.
[0212] Compared to traditional technologies, this invention does not consume the rendering performance of the vehicle's infotainment system. The system only needs to handle communication with the aerial imaging equipment and the interactive development of rendering effects. In contrast, traditional virtual human rendering requires rendering on the vehicle's infotainment system, which is very resource-intensive. Furthermore, the virtual human of this invention can interact with any object in the real world, while traditional virtual humans only exist on the vehicle's infotainment screen and cannot leave the screen, thus preventing them from interacting with real-world objects.
[0213] This invention also provides a vehicle, such as... Figure 11 As shown, the vehicle has the following features: Figure 1 The vehicle-mounted virtual human generation system 100 shown.
[0214] Please see Figure 12 , Figure 12 This is a schematic diagram of the structure of a vehicle infotainment system provided in an optional embodiment of the present invention, such as... Figure 12As shown, the vehicle infotainment system includes one or more first processors 10, a first memory 20, and interfaces for connecting the various components, including high-speed interfaces and low-speed interfaces. The various components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the vehicle infotainment system, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple devices can be connected, each providing some of the necessary operations (e.g., as a server array, a set of blade servers, or a multiprocessor system). Figure 12 Take a first processor 10 as an example.
[0215] The first processor 10 may be a central processing unit, a network processor, or a combination thereof. The first processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GPRS), or any combination thereof.
[0216] The first memory 20 stores instructions executable by at least one first processor 10 to cause the at least one first processor 10 to perform the method shown in the above embodiments.
[0217] The first memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the vehicle system. Furthermore, the first memory 20 may include high-speed random access memory and may also include non-transient memory, such as at least one disk storage device, flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the first memory 20 may optionally include memory remotely located relative to the first processor 10, and these remote memories can be connected to the vehicle system via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0218] The first memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the first memory 20 may also include a combination of the above types of memory.
[0219] The vehicle infotainment system also includes a first communication interface 30 for communicating with other devices or communication networks.
[0220] Please see Figure 13 , Figure 13 This is a schematic diagram of the structure of an aerial imaging device provided in an optional embodiment of the present invention, such as... Figure 13 As shown, the vehicle infotainment system includes one or more second processors 40, a second memory 50, and a second communication interface 60. For details on the specific working principles of the second processor 40, the second memory 50, and the second communication interface 60, please refer to the description of the above vehicle infotainment system embodiment, which will not be repeated here.
[0221] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0222] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0223] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for generating a virtual human in a vehicle, characterized in that, Applied to in-vehicle infotainment systems, the method includes: Acquire communication information generated by at least one service module in the vehicle's infotainment system; Based on the communication information, the virtual human imaging parameters are determined; A communication packet is generated based on the virtual human imaging parameters, and the communication packet is sent to the aerial imaging device so that the aerial imaging device can parse the communication packet to obtain the virtual human imaging parameters, and render the virtual human based on the virtual human imaging parameters to generate a vehicle-mounted virtual human; Determining the virtual human imaging parameters based on the communication information includes: The communication information is filtered to obtain interactive data sent by at least one service module; wherein, the service module includes an aerial imaging equipment setting module, a voice module, and a navigation module; Based on the business instructions corresponding to the interactive data, virtual human imaging parameters are obtained; wherein, the virtual human imaging parameters include imaging trigger type, imaging area, imaging volume, virtual human model parameters, virtual human action parameters and / or virtual human voice parameters; The step of generating a communication packet based on the virtual human imaging parameters and sending the communication packet to the aerial imaging device includes: The virtual human imaging parameters are byte-encoded to generate a communication body, and a communication header is generated according to the byte length of the communication body and the version information of the preset communication protocol. The communication packet is obtained based on the communication header and communication body. The communication packet is sent to the aerial imaging device according to the preset communication protocol so that the aerial imaging device can parse the communication packet to obtain the virtual human imaging parameters. First, the rendering imaging interface driver layer context environment is created and initialized. The rendering imaging interface driver layer is a three-dimensional graphics rendering library that runs on the mediumless aerial imaging device. Based on the virtual human imaging parameters and the functions in the three-dimensional graphics rendering library, the initial state machine is updated, and the projection area, projection matrix and view matrix are set, and vertex data is defined. Then, the shader is written and compiled, and the updated state machine is used to render the model to obtain the target virtual human model. The target virtual human model is then projected to generate the vehicle-mounted virtual human. The step of rendering the model using the updated state machine to obtain the target virtual human model includes: The updated state machine is used for geometry drawing, texture mapping, lighting and shadow processing, color processing, depth testing, and template testing to obtain the target virtual human model.
2. The method according to claim 1, characterized in that, The business instruction includes an imaging parameter setting instruction; the virtual human action parameters include the user-set initial virtual human action; and the virtual human voice parameters include the user-set initial virtual human voice. Obtaining the virtual human imaging parameters based on the business instruction corresponding to the interaction data includes: Based on the interactive data sent by the aerial imaging equipment setting module, the imaging parameter setting command issued by the user for the aerial imaging equipment is obtained; Based on the imaging parameter setting instructions, the imaging trigger type, imaging area, imaging volume, virtual human model parameters, initial virtual human actions, and / or initial virtual human speech set by the user for the virtual human are obtained.
3. The method according to claim 2, characterized in that, The business instruction further includes a voice instruction, the virtual human action parameters further include a target action corresponding to the voice instruction, and the virtual human voice parameters further include a target voice corresponding to the voice instruction; obtaining the virtual human imaging parameters based on the business instruction corresponding to the interaction data further includes: The voice commands issued by the user are obtained by extracting and recognizing the interactive data sent by the voice module. Based on the preset mapping relationship between voice commands and virtual human imaging parameters, the target action and / or target voice corresponding to the voice command is confirmed.
4. The method according to claim 2, characterized in that, The business instructions also include navigation instructions, the virtual human action parameters also include target actions corresponding to the navigation instructions, and the virtual human voice parameters also include target voice corresponding to the navigation instructions; obtaining virtual human imaging parameters based on the business instructions corresponding to the interaction data further includes: Based on the interactive data sent by the navigation module, navigation instructions are obtained; Based on the preset mapping relationship between navigation instructions and virtual human imaging parameters, the target action and / or target voice corresponding to the navigation instructions are identified.
5. The method according to claim 1, characterized in that, Sending the communication packet to the aerial imaging device according to a preset communication protocol includes: Send a synchronization message to the aerial imaging device to request the establishment of a connection; After receiving the synchronization confirmation message returned by the airborne imaging device based on the synchronization message, the system sends a first confirmation message to the airborne imaging device to confirm the establishment of the connection, thereby establishing a connection with the airborne imaging device. The communication packet is sent to the aerial imaging device; After the communication packet is detected to have been sent, a termination message is sent to the air imaging device to request the closure of the connection. After receiving a second confirmation message from the airborne imaging device indicating receipt of the termination message and a termination confirmation message returned based on the termination message, a third confirmation message is sent to the airborne imaging device to confirm the closure of the connection, thereby disconnecting the connection with the airborne imaging device.
6. A method for generating a virtual human in a vehicle, characterized in that, The method, applied to aerial imaging equipment, includes: The vehicle receives communication packets sent by the vehicle-mounted system; wherein the communication packets are generated by the vehicle-mounted system based on the virtual human imaging parameters corresponding to the communication information obtained by at least one service module. The communication packets are parsed to obtain the virtual human imaging parameters; Rendering is performed based on the virtual human imaging parameters to generate an in-vehicle virtual human; The step of rendering based on the virtual human imaging parameters to generate an in-vehicle virtual human includes: First, create and initialize the rendering and imaging interface driver layer context environment; the rendering and imaging interface driver layer is a 3D graphics rendering library that runs on a mediumless aerial imaging device; Based on the virtual human imaging parameters and functions in the 3D graphics rendering library, the initial state machine is updated, and the projection area, projection matrix, and view matrix are set, vertex data is defined. Then, shaders are written and compiled. The virtual human imaging parameters include imaging trigger type, imaging area, imaging volume, virtual human model parameters, virtual human action parameters, and / or virtual human voice parameters. These parameters are obtained by filtering the communication information to obtain interactive data sent by at least one service module, and based on the corresponding service instructions. The service modules include an aerial imaging equipment setting module, a voice module, and a navigation module. The updated state machine is used for model rendering to obtain the target virtual human model; The target virtual human model is projected to generate an in-vehicle virtual human; The step of rendering the model using the updated state machine to obtain the target virtual human model includes: The updated state machine is used for geometry drawing, texture mapping, lighting and shadow processing, color processing, depth testing, and template testing to obtain the target virtual human model.
7. The method according to claim 6, characterized in that, After projecting the target virtual human model to generate an in-vehicle virtual human, the method further includes: Based on the virtual human action parameters and / or virtual human voice parameters, the target action and / or target voice are obtained; Control the in-vehicle virtual human to perform the target action and / or broadcast the target voice.
8. The method according to claim 6, characterized in that, After projecting the target virtual human model to generate an in-vehicle virtual human, the method further includes: If no communication packet is received from the vehicle's infotainment system within a preset time, the virtual human in the vehicle will be controlled to perform a preset standby action and / or broadcast a preset standby voice message.
9. The method according to claim 6, characterized in that, After projecting the target virtual human model to generate an in-vehicle virtual human, the method further includes: New virtual human imaging parameters are obtained based on the new communication packets sent by the vehicle's infotainment system. The vehicle-mounted virtual human is adjusted based on the new virtual human imaging parameters.
10. The method according to claim 9, characterized in that, The step of adjusting the vehicle-mounted virtual human based on the new virtual human imaging parameters includes: Adjust the projection duration of the in-vehicle virtual human according to the new imaging trigger type; And / or, adjust the projection position of the in-vehicle virtual human according to the new imaging area.
11. A vehicle-mounted virtual human generation system, characterized in that, include: Vehicle-mounted and aerial imaging equipment; The vehicle-mounted system acquires communication information generated by at least one service module within the vehicle-mounted system. Based on the communication information, the virtual human imaging parameters are determined; A communication packet is generated based on the virtual human imaging parameters, and the communication packet is sent to the aerial imaging device; The vehicle-mounted system is also used to filter the communication information to obtain interactive data sent by at least one service module; wherein, the service module includes an aerial imaging equipment setting module, a voice module, and a navigation module; and obtains virtual human imaging parameters according to the service instructions corresponding to the interactive data; wherein, the virtual human imaging parameters include imaging trigger type, imaging area, imaging volume, virtual human model parameters, virtual human action parameters, and / or virtual human voice parameters; The aerial imaging device receives communication packets sent by the vehicle-mounted unit; parses the communication packets to obtain virtual human imaging parameters; and renders the virtual human based on the virtual human imaging parameters to generate an in-vehicle virtual human. The aerial imaging device is further used to first create and initialize the rendering imaging interface driver layer context environment; the rendering imaging interface driver layer is a 3D graphics rendering library running on the media-free aerial imaging device; based on the virtual human imaging parameters and functions in the 3D graphics rendering library, the initial state machine is updated, and the projection area, projection matrix and view matrix are set, vertex data is defined, and then the shader is written and compiled; wherein, the virtual human imaging parameters include imaging trigger type, imaging area, imaging volume, virtual human model parameters, virtual human action parameters and / or virtual human voice parameters; the updated state machine is used to render the model to obtain the target virtual human model; the target virtual human model is projected to generate a vehicle-mounted virtual human; The aerial imaging device is also used to perform geometry drawing, texture mapping, lighting and shadow processing, color processing, depth testing, and template testing using the updated state machine to obtain a target virtual human model.
12. The system according to claim 11, characterized in that, The aerial imaging equipment is mounted in the passenger seat, rear seat, and / or driver's seat of the vehicle.
13. A vehicle, characterized in that, The in-vehicle virtual human generation system includes any one of claims 11 to 12.
14. A vehicle infotainment system, characterized in that, The vehicle infotainment system includes a first memory and a first processor; The first memory and the first processor are interconnected. The first memory stores first computer instructions. The first processor executes the vehicle-mounted virtual human generation method according to any one of claims 1 to 5 by executing the first computer instructions.
15. An aerial imaging device, characterized in that, The aerial imaging device includes a second memory and a second processor; The second memory and the second processor are interconnected. The second memory stores second computer instructions. The second processor executes the vehicle-mounted virtual human generation method according to any one of claims 6 to 10 by executing the second computer instructions.
Citation Information
Patent Citations
In-car 3D holographic projection device
CN105679209A
Cabin interaction method and system, vehicle and computer storage medium
CN115681178A