Interaction methods, devices, electronic equipment and storage media for in-vehicle digital humans

By recognizing the emotions and identities of passengers in the in-vehicle system, playing matching audio and video files and synchronizing with rhythm, the problem of lack of emotional interaction in in-vehicle music playback is solved, achieving an immersive user experience and emotional companionship.

CN122493880APending Publication Date: 2026-07-31BEIJING YINGZHI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING YINGZHI TECH CO LTD
Filing Date
2026-04-21
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing in-car music playback functions lack emotional interaction, failing to perceive and respond to user emotions, resulting in a dull music experience.

Method used

By identifying the emotional information and identity tags of occupants in the smart cockpit, audio and video files matching the emotional information are played, and the in-vehicle digital human responds synchronously with the rhythm of the music, creating an immersive interactive feedback.

Benefits of technology

It enhances the immersive experience of in-vehicle human-machine interaction, alleviates driving loneliness, and strengthens the emotional companionship experience in the cabin.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493880A_ABST
    Figure CN122493880A_ABST
Patent Text Reader

Abstract

This application proposes an interaction method, device, electronic device, and storage medium for an in-vehicle digital human, relating to the field of in-vehicle human-machine interaction technology. The interaction method for the in-vehicle digital human includes: determining the emotional information and identity tags of occupants in a smart cockpit; determining the interaction information of the in-vehicle digital human based on the occupants' emotional information and identity tags; playing an audio-visual file matching the emotional information based on a playback request; acquiring the occupants' body movements and the audio rhythm of the audio-visual file; and responding to a successful rhythmic match between the body movements and the audio rhythm by providing a rhythmic response to the in-vehicle digital human's body according to the audio rhythm, forming an immersive interactive feedback that alleviates the loneliness of driving and enhances the emotional companionship experience in the cockpit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle-mounted human-machine interaction technology, and in particular to an interaction method, device, electronic device and storage medium for a vehicle-mounted digital human. Background Technology

[0002] Existing in-car music playback functions simply play music without any emotional interaction between the car system and the user; when playing music, the digital assistant usually just sits quietly in the corner of the screen without any response.

[0003] When listening to music, users experience emotional resonance, move to the rhythm, or feel joy or sadness. However, existing in-car infotainment systems cannot sense user feelings or provide any emotional response, resulting in a mediocre in-car music experience. Summary of the Invention

[0004] This application aims to at least partially address one of the technical problems in the related art.

[0005] Therefore, the first objective of this application is to propose an interaction method for in-vehicle digital humans to enhance the immersive experience of in-vehicle human-computer interaction.

[0006] The second objective of this application is to propose an interactive device for an in-vehicle digital human.

[0007] The third objective of this application is to propose an electronic device.

[0008] The fourth objective of this application is to provide a computer-readable storage medium.

[0009] The fifth objective of this application is to provide a computer program product.

[0010] To achieve the above objectives, a first aspect of this application proposes an interaction method for an in-vehicle digital human, comprising:

[0011] Determine the emotional information and identity tags of occupants in the smart cockpit; Based on the occupant's emotional information and identity tags, the interaction information of the in-vehicle digital human is determined, wherein the interaction information includes at least a playback request for playing an audio or video file that matches the emotional information. Based on the playback request, play the audio and video file that matches the emotional information; Acquire the occupant's body movements and the audio beats of the audio and video files; In response to a successful rhythmic match between the body movements and the audio beat, the in-vehicle digital human's body responds rhythmically to the audio beat.

[0012] To achieve the above objectives, a second aspect of this application provides an interactive device for an in-vehicle digital human, comprising: The identification module is used to determine the emotional information and identity tags of the occupants in the smart cockpit; The determination module is used to determine the interaction information of the in-vehicle digital human based on the occupant's emotional information and identity tag, wherein the interaction information includes at least a playback request for playing an audio or video file that matches the emotional information; A playback module is used to play an audio or video file that matches the emotional information based on the playback request; The acquisition module is used to acquire the occupant's body movements and the audio beats of the audio and video files; The response module is used to respond to a successful rhythmic match between the body movement and the audio beat by providing a rhythmic response to the body of the in-vehicle digital human according to the audio beat.

[0013] To achieve the above objectives, a third aspect of this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method described in the first aspect embodiment.

[0014] To achieve the above objectives, a fourth aspect of this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the method described in the first aspect embodiment.

[0015] To achieve the above objectives, a fifth aspect of this application provides a computer program product including a computer program that, when executed by a processor, implements the method described in the first aspect.

[0016] The in-vehicle digital human interaction method, device, electronic device, and storage medium provided in this application obtain more accurate in-vehicle digital human interaction information by determining the emotional information and identity tags of the occupants in the smart cockpit. Based on the playback requests in the interaction information, audio and video files are played to achieve personalized audio and video recommendations. After the audio and video files are played, the occupants' body movements and the audio rhythm of the audio and video files are obtained. The body movements and audio rhythm are matched. If the rhythm matching is successful, it means that the occupants are moving in rhythm with the audio and video files. At this time, the in-vehicle digital human is controlled to respond rhythmically according to the audio rhythm, forming an immersive interactive feedback, alleviating the loneliness of driving, and enhancing the emotional companionship experience in the cockpit.

[0017] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0018] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A flowchart illustrating an in-vehicle digital human interaction method provided in an embodiment of this application; Figure 2 A flowchart illustrating an interaction method for an in-vehicle digital human provided in an embodiment of this application; Figure 3 A flowchart illustrating an interaction method for an in-vehicle digital human provided in an embodiment of this application; Figure 4 A logical schematic diagram of an interaction method for an in-vehicle digital human provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an interactive device for an in-vehicle digital human provided in an embodiment of this application. Detailed Implementation

[0019] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0020] The following description, with reference to the accompanying drawings, outlines an interaction method, apparatus, electronic device, and storage medium for an in-vehicle digital human based on embodiments of this application.

[0021] Figure 1 This is a flowchart illustrating an in-vehicle digital human interaction method provided in an embodiment of this application. Figure 1 As shown, the method includes the following steps: S101 determines the emotional information and identity tags of the occupants in the smart cockpit.

[0022] The intelligent cockpit is a driving space system that integrates various information technologies (IT) and artificial intelligence technologies to create a brand-new integrated digital platform in the vehicle, providing drivers with an intelligent experience and promoting driving safety.

[0023] In some embodiments, facial images of occupants can be collected and analyzed using cameras inside the smart cockpit to determine the current emotional information of the occupants. This emotional information is used to characterize the user's current facial emotions, such as happiness, anger, sadness, calmness, and fatigue.

[0024] In some embodiments, the facial images of the occupants can be analyzed to determine the facial features of the occupants in the facial images. Based on the facial features, the facial features are compared with the facial features of users in the stored identity database to determine the identity label corresponding to the facial features when the similarity is greater than the similarity threshold. The identity label may include the occupant's nickname, name, and occupant's preference information, etc.

[0025] It is understandable that if no matching identity tag is found in the stored identity database, the current occupant may be a user with a new identity, and the identity tag may be an unknown identity.

[0026] S102 determines the interaction information of the in-vehicle digital human based on the occupant's emotional information and identity tags.

[0027] The interactive information includes at least a playback request for an audio or video file that matches the emotional information.

[0028] For example, assuming that the occupant's emotional information indicates that the user is sad, the playback request could be to push healing audio and video files to the occupant; when pushing healing audio and video files, preference information in the identity tag can be referenced to prioritize recommending audio and video files of the user's favorite singers or preferred types, and the generated playback request is more in line with the user's emotional information and identity tag, thereby improving the user experience as interactive information of the in-vehicle digital human.

[0029] S103, based on the playback request, plays the audio and video file that matches the emotional information.

[0030] Understandably, the playback request includes an audio / video file that matches the emotional information, and that audio / video file is then played.

[0031] S104, acquire the occupant's body movements and the audio beats of the audio and video files.

[0032] In some embodiments, user motion videos can be collected using cameras inside the smart cockpit, and user body movements can be determined based on these videos.

[0033] Optionally, the rhythm of the current audio and video file can be analyzed in real time to obtain the audio beat of the video clip file, such as the number of beats per minute (BPM) of the audio and video file.

[0034] S105 responds to the successful rhythm matching of body movements and audio beats, and provides rhythmic responses to the body of the in-vehicle digital human according to the audio beats.

[0035] Optionally, the rhythm of the body movement and the rhythm of the audio beat can be matched. For example, the first rhythm feature of the body movement and the second rhythm feature of the audio beat can be obtained, and the similarity between the first rhythm feature and the second rhythm feature can be calculated. If the similarity is greater than a preset similarity threshold, the rhythm matching is determined to be successful. If the similarity is less than or equal to the similarity threshold, the rhythm matching is determined to be unsuccessful.

[0036] In some embodiments, the rhythm of bodily movements may be based on rhythmic behaviors such as nodding, swaying, and tapping to the music.

[0037] When the body movements and the rhythm of the audio beat are successfully matched, it means that the current occupant is moving to the rhythm of the music. At this time, the body of the in-vehicle digital human is responded to with rhythm according to the audio beat, that is, the in-vehicle digital human is driven to move synchronously and sway with the rhythm of the audio beat, presenting a sense of companionship as listening to audio and video with the occupant, thus improving the user experience.

[0038] In this embodiment, by determining the emotional information and identity tags of the occupants in the smart cockpit, more accurate and suitable interactive information from the in-vehicle digital human is provided to the occupants. Based on the playback requests in the interactive information, audio and video files that match the current occupants' emotional information and identity tags are played, realizing personalized and contextualized audio and video recommendations. After playing the audio and video files, the occupants' body movements and the audio rhythm of the audio and video files are obtained. The body movements and audio rhythm are matched. If the rhythm matching is successful, it means that the occupants are moving in rhythm with the audio and video files. At this time, the in-vehicle digital human is controlled to respond rhythmically according to the audio rhythm, forming an immersive interactive feedback, alleviating the loneliness of driving, and enhancing the emotional companionship experience in the cockpit.

[0039] Based on the above embodiments, Figure 2 This is a flowchart illustrating an interaction method for an in-vehicle digital human provided in an embodiment of this application. Figure 2 As shown, the method includes the following steps: S201 determines the emotional information and identity tags of the occupants in the smart cockpit.

[0040] In some embodiments, at least one of the occupant's facial information and audio information may be acquired; the facial information may be acquired based on a camera inside the smart cockpit, and the audio information may be acquired through a microphone inside the smart cockpit; it is understood that if the occupant does not speak, there is no audio information.

[0041] Optionally, facial expression recognition is performed on the facial information to determine the first emotional information; tone recognition is performed on the audio information to determine the second emotional information. In this embodiment, facial expression recognition can be performed on the facial information based on a pre-trained expression recognition model, and the facial image features can be extracted to identify emotional states such as happiness, sadness, anger, and fatigue; tone recognition can be performed on the audio information based on a pre-trained tone recognition model, and the user's tone and emotion can be judged by analyzing the spectrum, prosody, and sound intensity of the speech signal; the model structure and training process of the expression recognition model and tone recognition model are common existing technologies, and will not be described in detail here.

[0042] In some embodiments, the first emotion information and the second emotion information may include the probability of the occupant's corresponding emotion. For example, the first emotion information may include "70% calm and 30% happy" for the current occupant, and the second emotion information may include "60% calm and 40% happy" for the occupant. The emotion with a higher probability in the emotion information is the occupant's facial emotion.

[0043] Optionally, the occupant's emotional information can be determined based on at least one of the first emotional information and the second emotional information. It is understood that if there is no audio information, the first emotional information is determined to be the occupant's emotional information. If there is audio information, the first emotional information and the second emotional information can be fused. In this embodiment, the probability of the same emotion in the first emotional information and the second emotional information is averaged. For example, the first emotional information "calm 70%, happy 30%" and the second emotional information "calm 60%, happy 40%" are fused to obtain the occupant's emotional information as "calm 65%, happy 35%".

[0044] Understandably, when judging a passenger's final facial emotion based on their emotional information of "65% calm and 35% happy", the current passenger's facial emotion is determined to be calm.

[0045] In some embodiments, audio information can be matched with historical user audio in an audio library to obtain a first matching probability; facial information can be matched with historical user faces in a face library to obtain a second matching probability; historical user audio in the audio library refers to the audio of the user currently recorded by the smart cockpit, and historical user faces in the face library refers to the faces of the user currently recorded by the smart cockpit. The first matching probability can be the matching probability of the current audio information with one or more historical user audios, and the second matching probability can be the matching probability of the current facial information with one or more historical user faces.

[0046] Furthermore, the first matching probability and the second matching probability are fused. In this embodiment, the fusion of the first matching probability and the second matching probability can be calculated by averaging the first matching probability and the second matching probability. For example, for the same user A, the corresponding first matching probability is 0.8 and the corresponding second matching probability is 0.6. Then, the fusion result of the first matching probability and the second matching probability is 0.7.

[0047] Furthermore, the target matching user is determined based on the fusion result, and the preference information of the target matching user is obtained as the identity label of the passenger; in this embodiment, the user with the highest probability among all probabilities greater than a preset probability threshold in the fusion result is the target matching user, and the preference information of the target matching user is used as the identity label of the passenger.

[0048] S202, Obtain audio and video tags that match the emotional information.

[0049] Audio and video tags are labels used to identify the content and type of audio and video files. For example, music file A is tagged as healing and soothing, while music file B is tagged as cheerful and fast-paced.

[0050] For example, when the emotional information indicates happiness, the corresponding audio and video tags can be cheerful, relaxed, and positive; when the emotional information indicates sadness, the corresponding audio and video tags can be soothing, gentle, and soft music; when the emotional information indicates anger, the corresponding audio and video tags can be relaxing, stress-relieving, and gentle; when the emotional information indicates calmness, the corresponding audio and video tags can be soft music and gentle; and when the emotional information indicates fatigue, the corresponding audio and video tags can be brisk, rhythmic, and invigorating.

[0051] S203, determine the occupant's music preference information based on the identity tag.

[0052] Optionally, the identity tag may include the passenger's music preference information, which may include information such as the user's favorite singers, favorite audio and video types, and favorite music genres.

[0053] S204: Based on music preference information and audio / video tags, determine the matching audio / video files.

[0054] For example, assuming the audio / video tags are soothing, gentle, and light music, then based on the passenger's favorite singers in the music preference information, the audio / video files of soothing, gentle, and light music sung by that singer are selected as the matching audio / video files.

[0055] In some embodiments, the interactive information may also include a greeting; the corresponding greeting content can be determined based on the passenger's emotional information. For example, when a passenger is feeling sad, the greeting content could be, "You seem a little unhappy today. Would you like to listen to some songs that will cheer you up?"

[0056] In some embodiments, a specific title for an occupant can also be determined based on the occupant's identity tag, such as the occupant's name or a nickname set by the occupant.

[0057] Furthermore, a greeting can be generated based on the content of the greeting and the specific title. For example, the greeting could be, "Little Y, you seem a little unhappy today. Would you like to listen to some songs that will cheer you up?"

[0058] In some embodiments, the corresponding playback tone can be determined based on the passenger's emotional information; a greeting can be generated and played based on the playback tone; for example, when the passenger is sad, the playback tone can be gentle and calm; when the passenger is happy, the playback tone can be cheerful and pleasant.

[0059] In some embodiments, expression-driven information matching emotional information can also be obtained, and the expression of the in-vehicle digital human can be switched based on the expression-driven information; for example, when a user is feeling sad, the matched expression-driven information can be information corresponding to concern and comfort, and the expression of the digital human can be switched to a concern or comforting expression based on the expression-driven information to enhance the emotional comfort effect on the occupants.

[0060] S205, based on the playback request, play the audio and video file that matches the emotional information.

[0061] In this application embodiment, the implementation method of step S205 can be implemented in any of the various embodiments of this disclosure, and no limitation is made here, nor will it be described in detail.

[0062] S206, acquire the occupant's body movements and the audio beats of audio and video files.

[0063] In this application embodiment, the implementation method of step S206 can be implemented in any of the various embodiments of this disclosure, and no limitation is made here, nor will it be described in detail.

[0064] S207 responds to the successful rhythm matching of body movements and audio beats, and provides rhythmic responses to the body of the in-vehicle digital human according to the audio beats.

[0065] In some embodiments, the first temporal rhythm features of the occupant can be extracted based on bodily movements using a temporal characteristic extraction algorithm. The first temporal rhythm features represent the rhythmic features of the occupant's movements in a continuous time dimension.

[0066] Obtain the second temporal rhythm feature corresponding to the audio beat. The second temporal rhythm feature is consistent with the first temporal rhythm feature in the temporal dimension. Calculate the similarity between the first and second temporal rhythm features to obtain the similarity between them.

[0067] If the similarity is greater than or equal to a preset similarity threshold, the rhythm matching is considered successful.

[0068] If the similarity is less than a preset similarity threshold, the rhythm matching is determined to have failed.

[0069] When the rhythm is successfully matched, the in-vehicle digital human body responds rhythmically to the music beat; optionally, the beat point and beat frequency of the audio beat can be extracted; based on the beat point and beat frequency, the in-vehicle digital human is controlled to respond rhythmically, so that the in-vehicle digital human sways or nods in accordance with the beat point and beat frequency, showing a sense of companionship while listening to the same music.

[0070] In this embodiment, by using at least one of the facial and audio information of the occupants in the smart cockpit, the occupants' emotional information and identity tags are obtained, improving the accuracy of occupant emotion and identity judgment in the smart cockpit. Audio and video tags are determined based on the emotional information, and music preference information is determined based on the identity tags. Audio and video files are determined based on the audio and video tags and music preference information, improving the adaptability of recommended content. After playing the audio and video files, the rhythm is matched according to the occupants' body movements and the audio beat of the audio and video files. When the occupants follow the rhythm of the music, the in-vehicle digital human is controlled to make a synchronous rhythmic response, making the smart cockpit more warm and human-like.

[0071] Based on the above embodiments, Figure 3 This is a flowchart illustrating an interaction method for an in-vehicle digital human provided in an embodiment of this application. Figure 3 As shown, the method includes the following steps: S301 determines the emotional information and identity tags of the occupants in the smart cockpit.

[0072] In this application embodiment, the implementation method of step S301 can be implemented in any of the various embodiments of this disclosure, and no limitation is made here, nor will it be described in detail.

[0073] S302, Obtain audio and video tags that match the emotional information.

[0074] In this application embodiment, the implementation method of step S302 can be implemented in any of the various embodiments of this disclosure, and no limitation is made here, nor will it be described in detail.

[0075] S303 determines the occupant's music preference information based on their identity tag.

[0076] In this application embodiment, the implementation method of step S303 can be implemented in any of the various embodiments of this disclosure, and no limitation is made here, nor will it be described in detail.

[0077] S304, an audio / video file matched based on music preference information and audio / video tags.

[0078] In this application embodiment, the implementation method of step S304 can be implemented in any of the various embodiments of this disclosure, and no limitation is made here, nor will it be described in detail.

[0079] S305, based on the playback request, plays the audio and video file that matches the emotional information.

[0080] In this application embodiment, the implementation method of step S305 can be implemented in any of the various embodiments of this disclosure, and no limitation is made here, nor will it be described in detail.

[0081] S306, acquire the occupant's body movements and the audio beats of audio and video files.

[0082] In this application embodiment, the implementation method of step S306 can be implemented in any of the various embodiments of this disclosure, and no limitation is made here, nor will it be described in detail.

[0083] S307 responds to the successful rhythm matching of body movements and audio beats, and provides rhythmic responses to the body of the in-vehicle digital human according to the audio beats.

[0084] In this application embodiment, the implementation method of step S307 can be implemented in any of the various embodiments of this disclosure, and no limitation is made here, nor will it be described in detail.

[0085] S308 controls the ambient lighting to change according to the audio beat.

[0086] In some embodiments, the beat points of the audio beats can be extracted; the ambient lights can be controlled to blink based on the beat points; for example, the lights blink rapidly during fast tempos and breathe slowly during slow tempos.

[0087] Optionally, the audio / video type of the current audio / video file can also be obtained; the corresponding ambient light color can be determined based on the audio / video type; for example, the ambient light color corresponding to a cheerful audio / video type is a warm tone, and the ambient light color corresponding to a melancholic audio / video type is a cool tone.

[0088] Alternatively, the ambient light color can be determined based on emotional information. For example, the ambient light color is warm when you are happy and cool when you are sad, which can create an immersive space with multi-dimensional audiovisual interaction.

[0089] In this embodiment, by using at least one of the occupants' facial and audio information within the smart cockpit, the occupants' emotional information and identity tags are obtained, improving the accuracy of occupant emotion and identity judgment within the smart cockpit. Audio and video tags are determined based on emotional information, and music preference information is determined based on identity tags. Audio and video files are jointly determined based on audio and video tags and music preference information, improving the adaptability of recommended content. After playing the audio and video files, rhythm matching is performed based on the occupants' body movements and the audio beat of the audio and video files. When the occupants move to the rhythm of the music, the in-vehicle digital human is controlled to synchronously respond rhythmically, making the smart cockpit more warm and human-like. Furthermore, the ambient light color and flashing rhythm can be controlled according to the audio and video type and emotional information, enhancing the immersiveness and atmosphere of the music and environment linkage within the smart cockpit, further improving the emotional experience and entertainment interaction effect of the smart cockpit.

[0090] Figure 4 This is a logical schematic diagram of an interaction method for an in-vehicle digital human provided in an embodiment of this application. Facial information is acquired through a camera, and audio information is acquired through a microphone. The facial and audio information are analyzed, and the music beat (BPM) is extracted in real time. The facial, audio, and music beat information are then fused and analyzed to obtain a unified driving signal. This signal controls the in-vehicle digital human's facial expressions and body movements, controls the ambient light color and beat synchronization, and recommends music based on emotion-matched playlists, enabling users to experience a synchronized empathetic experience across video, environment, and auditory aspects.

[0091] To achieve the above embodiments, this application also proposes an interactive device for an in-vehicle digital human.

[0092] Figure 5 This is a schematic diagram of the structure of an interactive device for an in-vehicle digital human, provided as an embodiment of this application. Figure 5 As shown, the in-vehicle digital human interaction device 500 includes: The identification module 501 is used to determine the emotional information and identity tags of the occupants in the smart cockpit; The determination module 502 is used to determine the interaction information of the in-vehicle digital human based on the occupant's emotional information and identity tags, wherein the interaction information includes at least a playback request for playing an audio or video file that matches the emotional information; The playback module 503 is used to play audio and video files that match the emotional information based on the playback request; The acquisition module 504 is used to acquire the occupant's body movements and the audio beats of audio and video files; The response module 505 is used to respond to the successful rhythm matching of body movements and audio beats, and to provide rhythmic responses to the body of the in-vehicle digital human according to the audio beats.

[0093] Furthermore, in one possible implementation of this application embodiment, the interactive information further includes a greeting, and the determining module 502 is further used for: Determine the appropriate greeting content based on the passenger's emotional state; The specific title of the passenger is determined based on the passenger's identification tag; Generate a greeting based on the content of the greeting and the specific title.

[0094] Furthermore, in one possible implementation of this application embodiment, the playback module 503 is also used for: Determine the appropriate tone of voice based on the passengers' emotional information; Generate a greeting based on the tone of the speech and play it.

[0095] Furthermore, in one possible implementation of this application embodiment, the device 500 is also used for: Control the ambient lighting to change according to the audio beat.

[0096] Furthermore, in one possible implementation of this application embodiment, the device 500 is also used for: Extract the beat points of the audio beat; The ambient lights blink based on the beat point.

[0097] Furthermore, in one possible implementation of this application embodiment, the device 500 is also used for: Get the audio / video type of the current audio / video file; The ambient light color is determined based on the audio / video type. Alternatively, the ambient light color can be determined based on emotional information.

[0098] Furthermore, in one possible implementation of this application embodiment, the identification module 501 is used for: Obtain at least one of the occupant's facial information and audio information; Facial information is used for expression recognition to determine the primary emotional information; Tone recognition is performed on audio information to determine secondary emotional information; The occupant's emotional information is determined based on at least one of the first emotional information and the second emotional information.

[0099] Furthermore, in one possible implementation of this application embodiment, the identification module 501 is also used for: The audio information is matched with historical user audio in the audio library to obtain the first matching probability; The facial information is matched with the faces of historical users in the facial database to obtain the second matching probability; The first matching probability and the second matching probability are fused together, and the target matching user is determined based on the fusion result; Obtain the preference information of the target matching user as the identity label of the passenger.

[0100] Furthermore, in one possible implementation of this application embodiment, the determining module 502 is used for: Obtain audio and video tags that match emotional information; Determine passengers' music preferences based on their identity tags; Based on music preference information and audio / video tags, determine the matching audio / video files.

[0101] Furthermore, in one possible implementation of this application embodiment, the playback module 503 is also used for: Acquire expression-driven information that matches emotional information, and switch the expressions of the in-vehicle digital human based on the expression-driven information.

[0102] Furthermore, in one possible implementation of this application embodiment, the response module 505 is also used for: Based on body movements, extract the first temporal rhythm characteristics of the occupants; Obtain the second temporal rhythm features corresponding to the audio beat; Similarity calculation is performed between the first temporal rhythm features and the second temporal rhythm features; If the similarity is greater than or equal to a preset similarity threshold, the rhythm matching is determined to be successful; If the similarity is less than a preset similarity threshold, the rhythm matching is determined to have failed.

[0103] Furthermore, in one possible implementation of this application embodiment, the response module 505 is also used for: Extract the beat point and beat frequency of the audio beat; Based on the beat point and beat frequency, the vehicle-mounted digital human is controlled to respond rhythmically.

[0104] It should be noted that the foregoing explanation of the interaction method embodiment for the in-vehicle digital human also applies to the interaction device of the in-vehicle digital human in this embodiment, and will not be repeated here.

[0105] In this embodiment, by using at least one of the facial and audio information of the occupants in the smart cockpit, the occupants' emotional information and identity tags are obtained, improving the accuracy of occupant emotion and identity judgment in the smart cockpit. Audio and video tags are determined based on emotional information, and music preference information is determined based on identity tags. Audio and video files are jointly determined based on audio and video tags and music preference information, improving the adaptability of recommended content. After playing the audio and video files, the rhythm is matched according to the occupants' body movements and the audio beat of the audio and video files. When the occupants follow the rhythm of the music, the in-vehicle digital human is controlled to make a synchronous rhythmic response, making the smart cockpit more warm and human-like. Furthermore, the ambient light color and flashing rhythm can be controlled according to the audio and video type and emotional information, enhancing the immersiveness and atmosphere of the music and environment linkage in the smart cockpit, further improving the emotional experience and entertainment interaction effect of the smart cockpit.

[0106] To implement the above embodiments, this application also proposes an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments.

[0107] To implement the above embodiments, this application also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the foregoing embodiments.

[0108] To implement the above embodiments, this application also proposes a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the foregoing embodiments.

[0109] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0110] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.

[0111] This application is intended to provide an implementation scheme for users to selectively prevent the use or access to their personal information data. Specifically, this disclosure is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information is de-identified to protect user privacy.

[0112] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0113] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0114] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0115] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0116] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0117] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0118] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0119] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. An in-vehicle digital human interaction method, characterized by, The method includes: Determine the emotional information and identity tags of occupants in the smart cockpit; Based on the occupant's emotional information and identity tags, the interaction information of the in-vehicle digital human is determined, wherein the interaction information includes at least a playback request for an audio or video file that matches the emotional information. Based on the playback request, play the audio and video file that matches the emotional information; Acquire the occupant's body movements and the audio beats of the audio and video files; In response to a successful rhythmic match between the body movements and the audio beat, the in-vehicle digital human's body responds rhythmically to the audio beat.

2. The method of claim 1, wherein, The interactive information also includes a greeting, and the method further includes: Based on the emotional information of the passengers, determine the corresponding greeting content; The specific title of the passenger is determined based on the passenger's identification tag; Generate a greeting based on the content of the greeting and the specific title.

3. The method of claim 2, wherein, Before playing the audio / video file matching the emotion information based on the playback request, the method further includes: Based on the emotional information of the passengers, determine the corresponding playback tone; The greeting is generated and played according to the tone of the playback.

4. The method according to claim 1, characterized in that, The method further includes: The ambient lighting is controlled to change according to the audio beat.

5. The method according to claim 4, characterized in that, The control of the ambient lighting to change according to the audio beat includes: Extract the beat points of the audio beats; The ambient light is controlled to blink based on the beat point.

6. The method according to claim 5, characterized in that, The method further includes: Get the audio / video type of the current audio / video file; The corresponding ambient light color is determined based on the audio / video type; Alternatively, the corresponding ambient light color can be determined based on the emotional information.

7. The method according to claim 1, characterized in that, Determine the emotional information of occupants within the smart cockpit, including: Obtain at least one of the occupant's facial information and audio information; Facial information is used for expression recognition to determine the first emotion information; The audio information is subjected to tone recognition to determine the second emotion information; The occupant's emotional information is determined based on at least one of the first emotional information and the second emotional information.

8. The method according to claim 7, characterized in that, Identify the occupants' identities within the smart cockpit, including: The audio information is matched with historical user audio in the audio library to obtain a first matching probability; The facial information is matched with the faces of historical users in the facial database to obtain a second matching probability; The first matching probability and the second matching probability are fused together, and the target matching user is determined based on the fusion result; The preference information of the target matching user is obtained as the identity label of the passenger.

9. The method according to claim 1, characterized in that, The process of determining the interaction information of the in-vehicle digital human based on the occupant's emotional information and identity tags includes: Obtain audio and video tags that match the emotional information; The occupant's music preference information is determined based on the identity tag; Based on the music preference information and the audio / video tags, the matching audio / video files are determined.

10. The method according to claim 9, characterized in that, The step of playing an audio / video file matching the emotion information based on the playback request further includes: Obtain expression-driven information that matches the emotional information, and switch the expressions of the in-vehicle digital human based on the expression-driven information.

11. The method according to claim 1, characterized in that, Rhythmic matching of the body movements and the audio beat includes: Based on the described body movements, the first temporal rhythm features of the occupants are extracted; Obtain the second temporal rhythm features corresponding to the audio beat; The similarity between the first temporal rhythm feature and the second temporal rhythm feature is calculated. If the similarity is greater than or equal to a preset similarity threshold, the rhythm matching is determined to be successful; If the similarity is less than a preset similarity threshold, rhythm matching is determined to have failed.

12. The method according to claim 1, characterized in that, The step of responding to the rhythm of the in-vehicle digital human's body according to the audio beat includes: Extract the beat point and beat frequency of the audio beat; Based on the beat point and the beat frequency, the vehicle-mounted digital human is controlled to perform rhythmic responses.

13. An interactive device for a vehicle-mounted digital human, characterized in that, The device includes: The identification module is used to determine the emotional information and identity tags of the occupants in the smart cockpit; The determination module is used to determine the interaction information of the in-vehicle digital human based on the occupant's emotional information and identity tag, wherein the interaction information includes at least a playback request for playing an audio or video file that matches the emotional information; A playback module is used to play an audio or video file that matches the emotional information based on the playback request; The acquisition module is used to acquire the occupant's body movements and the audio beats of the audio and video files; The response module is used to respond to a successful rhythmic match between the body movement and the audio beat by providing a rhythmic response to the body of the in-vehicle digital human according to the audio beat.

14. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-12.

16. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1-12.