Human-computer interaction method and device and intelligent cockpit

By adjusting the virtual character's posture based on the music's progress and user behavior data, the problem of poor interactivity of virtual characters was solved, resulting in more vivid interaction and a better user experience.

CN116788182BActive Publication Date: 2026-01-16PATEO CONNECT (NANJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210271619.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-18
Publication Date
2026-01-16
Estimated Expiration
2042-03-18

AI Technical Summary

Technical Problem

Existing commercial virtual characters have poor interactivity, resulting in a poor user experience, difficulty in tracking user progress in real time, or high costs.

Method used

By adjusting the virtual avatar's posture data based on the music's progress and user behavior data, including semantic recognition and image recognition technologies, the virtual avatar's posture is adjusted in real time to match the user's singing progress and body movements.

Benefits of technology

It enhances the interactivity and immersive experience of virtual avatars, and strengthens their realism and interactivity with users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116788182B_ABST
    Figure CN116788182B_ABST
Patent Text Reader

Abstract

The application provides a human-computer interaction method and device and an intelligent cockpit, relates to virtual technology, and particularly relates to a human-computer interaction method on an intelligent cockpit of a vehicle, comprising the following steps: determining a virtual image corresponding to a currently played music based on the currently played music; determining first posture data corresponding to a current time from posture data of the virtual image based on a progress time of the music; adjusting a display state of the first posture data of the virtual image in response to behavior data of a user at the current time; and / or controlling the virtual image to change from the first posture data to second posture data in response to the behavior data of the user at the current time.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to virtual technology, in particular to a human-computer interaction method and device and intelligent cockpit. BACKGROUND

[0002] A virtual image is a display mode of a virtual character, commonly seen in comics, TV series, movies, etc., and is an image that does not exist in reality. In recent years, virtual characters have frequently appeared in life, such as the application of virtual images such as "Hatsune Miku" and "Luo Tianyi" in the music field, or the application of virtual anchors in live broadcast, etc. However, the display of such commercial virtual characters is difficult, and the interactivity of these virtual images is poor, resulting in poor user experience. SUMMARY

[0003] The present application provides a human-computer interaction method and device to solve the problem of poor human-computer interaction.

[0004] According to the first aspect of the present application, a human-computer interaction method of the present application capable of achieving the aforementioned objects and other objects and advantages comprises the following steps: determining a virtual image corresponding to a currently played music based on the music; determining first attitude data corresponding to a current time from attitude data of the virtual image based on a progress time of the music; adjusting a display state of the first attitude data of the virtual image in response to behavior data of the user at the current time; and / or controlling the virtual image to change from the first attitude data to second attitude data in response to the behavior data of the user at the current time.

[0005] In an embodiment of the present application, the human-computer interaction method further comprises: determining a performance stage of the music based on the progress time; and controlling the virtual image to display in a third attitude data in response to the performance stage being a prelude stage and / or a coda stage.

[0006] In an embodiment of the present application, the attitude data is collected from at least one of a singer, a performer, a composer, and a lyricist of the music; and / or is pre-set based on a preference of the user. For example, the projection data of the virtual image is pre-set from another real image, including face pinching of the virtual image and / or system default; wherein the face pinching includes pre-setting in a real picture generation mode.

[0007] In an embodiment of the present application, determining the first attitude data of the current time from the attitude data of the virtual image comprises: determining the first attitude data based on the singing mode.

[0008] In the embodiments of the present application, the determining the first posture data of the current moment from the posture data of the virtual image comprises: taking the current moment as a starting point, obtaining a melody of the music within a preset time, and determining the first posture data based on the melody.

[0009] In the embodiments of the present application, the determining the first posture data of the current moment from the posture data of the virtual image comprises: taking the current moment as a starting point, obtaining a melody of the music within a preset time, and determining the first posture data based on the melody.

[0010] In the embodiments of the present application, the adjusting the display state of the first posture data of the virtual image in response to the behavior data of the user of the current moment comprises: collecting voice data of the user for semantic recognition, determining a singing progress of the user of the current moment based on the result of semantic recognition, and delaying or reducing the time length of displaying the first posture data when the singing progress and the current progress of the music are inconsistent.

[0011] In the embodiments of the present application, the controlling the virtual image to change from the first posture data to the second posture data in response to the behavior data of the user of the current moment comprises: collecting real-time body data of the user, and updating the first posture data to the second posture data to cater to the user when the real-time body data matches the preset body data.

[0012] In the embodiments of the present application, the preset body data comprises at least one of face rotation, hand waving, and body shaking.

[0013] In the embodiments of the present application, the first posture data is customized based on the preset body data.

[0014] According to a second aspect of the present application, a human-computer interaction device is disclosed, comprising: a virtual image determination unit configured to determine a virtual image corresponding to a currently played music based on the music; a first posture data determination unit configured to determine first posture data corresponding to a current moment from a series of posture data of the virtual image based on a progress moment of the music; a posture adjustment unit configured to adjust a display state of the first posture data of the virtual image in response to behavior data of the user of the current moment, and / or control the virtual image to change from the first posture data to second posture data in response to the behavior data of the user of the current moment.

[0015] According to a third aspect of the present application, an intelligent cockpit is disclosed, comprising: a human-computer interaction device, a music playing device, and a vehicle-mounted sound device.

[0016] In the embodiment of the present application, the intelligent cockpit further comprises an audio acquisition device, which is in communication connection with the human-computer interaction device, and is configured to provide voice data of the user to the human-computer interaction device in real time.

[0017] In the embodiment of the present application, the audio acquisition device comprises a microphone.

[0018] In the embodiment of the present application, the intelligent cockpit further comprises an image acquisition device, which is in communication connection with the human-computer interaction device, and is configured to provide limb data of the user to the human-computer interaction device in real time.

[0019] In the embodiment of the present application, the image acquisition device comprises a vehicle-mounted camera or a vehicle-mounted radar.

[0020] In the embodiment of the present application, the intelligent cockpit further comprises a singing evaluation device, which is in communication connection with the human-computer interaction device, and is configured to quantitatively evaluate a singing level of the user based on audio information provided by the audio acquisition device.

[0021] In the embodiment of the present application, the intelligent cockpit quantitatively evaluates the singing level of the user based on the audio information provided by the audio acquisition device, which comprises at least one of sound intensity, tone color, tonality, and frequency.

[0022] In the embodiment of the present application, the intelligent cockpit further comprises a singing evaluation device, which quantitatively evaluates the singing level of the user based on the limb data of the user provided by the image acquisition device.

[0023] In the embodiment of the present application, the singing evaluation device further comprises at least one of a system score, a like score, and a reward score.

[0024] In the embodiment of the present application, the score form after the song is played is adjusted based on the playing mode.

[0025] According to a fourth aspect of the present disclosure, a vehicle is disclosed, comprising an intelligent cockpit.

[0026] According to a fifth aspect of the present disclosure, a computer readable storage medium is disclosed, which stores a computer program, and the computer program is executed by a processor to implement the steps of the human-computer interaction method.

[0027] The present disclosure determines a virtual image corresponding to the music based on the currently played music, determines first posture data corresponding to the current time from the posture data of the virtual image based on the progress time of the music, adjusts the display state of the first posture data of the virtual image in response to the behavior data of the user at the current time, and / or controls the virtual image to change from the first posture data to second posture data in response to the behavior data of the user at the current time. Thus, the posture of the virtual image is added in the song and can be projected in the real world, enhancing the realism of the virtual image in the singing process. At the same time, the posture of the virtual image can be adjusted according to the real-time input behavior data of the user, making the virtual image more lively, increasing the interactive interest of singing, and improving the immersive experience of the user. BRIEF DESCRIPTION OF DRAWINGS

[0028] The above and other objects, features and advantages of the present disclosure will become more apparent from the following description of embodiments of the present disclosure, taken in conjunction with the accompanying drawings, in which:

[0029] Figure 1 An application scenario diagram of the human-computer interaction method, device, medium and program product according to embodiments of the present disclosure is schematically shown.

[0030] Figure 2 A flowchart of the human-computer interaction method according to embodiments of the present disclosure is schematically shown.

[0031] Figure 3a A simple diagram of the music progress timeline according to embodiments of the present disclosure is schematically shown.

[0032] Figure 3b A flowchart of adjusting the first posture data of the virtual image according to embodiments of the present disclosure is schematically shown.

[0033] Figure 3c A flowchart of updating the first posture data of the virtual image to the second posture data according to embodiments of the present disclosure is schematically shown.

[0034] Figure 4 A structural block diagram of the human-computer interaction device according to embodiments of the present disclosure is schematically shown.

[0035] Figure 5 A structural block diagram of the intelligent cockpit according to embodiments of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0036] The following description is presented to enable any person skilled in the art to practice the application as claimed. Preferred embodiments are presented in the following description only as examples and modifications can be made by persons skilled in the art having the benefit of this disclosure. The present application is directed to overcoming the foregoing problems and others.

[0037] It should be understood that various steps in the method implementations of the disclosure can be performed in different order and / or in parallel. Additionally, the method implementations can include additional steps and / or omit the steps shown. The scope of the disclosure is not limited in this regard.

[0038] It is to be noted that like-numbered terms and / or components throughout the figures denote like elements, and thus, upon initial introduction of a term in a figure, subsequent description of a similar item can not be repeated in a subsequent figure. In the description of the present application, the terms "first", "second", "third", "fourth" and the like, designate only to distinguish between similar objects, and can not imply a relative importance.

[0039] The names of the messages or information exchanged between the devices in the embodiments of the disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0040] In the prior art, it is difficult to display commercial virtual characters, either real-time following the progress of the user or at a high cost, which causes a variety of technical problems such as poor interaction of virtual images and poor user experience.

[0041] The steps of the human-computer interaction method provided by the present application include: determining a virtual image corresponding to a currently played music based on the music; determining first posture data corresponding to a current time from posture data of the virtual image based on a progress time of the music; adjusting a display state of the first posture data of the virtual image in response to behavior data of the user at the current time; and / or controlling the virtual image to change from the first posture data to second posture data in response to the behavior data of the user at the current time. The posture of the virtual image can be adjusted according to the behavior data input by the user in real time, making the virtual image more lively, increasing the interactive interest of singing, and improving the immersive experience of the user.

[0042] It should be noted that the human-computer interaction method provided by the embodiments of the present application relates to virtual technology, and detailed technical solutions can be used in indoor singing devices, outdoor singing devices, for example, including live concerts and related aspects, etc., and is especially suitable for use in vehicle intelligent cockpit.

[0043] The projection data of the other singer interacting with the user in the music playing, wherein the other singer can be but is not limited to a virtual character user, that is, the object matched by the real user can be a real user or a virtual character.

[0044] Figure 1 An application scenario of a human-computer interaction method, apparatus, medium and program product according to an embodiment of the present disclosure is schematically shown.

[0045] As Figure 1 shown, the application scenario according to this embodiment can include a car terminal 101, a user 102 and a server 103. A network 104 is used to provide a communication link medium between the car terminal 101, the user 102 and the server 103. The network 104 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0046] The user can log in an account applied for on the car terminal 101 to interact with the server 103 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the car terminal 101, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0047] The user 102 operates the car terminal 101, including logging in an account, selecting a song to sing, uploading a song to sing, etc.

[0048] The server 103 can be a server providing various services, such as a background management server providing support for the user 102 to browse applications using the car terminal 101 (only as an example). The background management server can analyze and process received requests or other instructions, and feed back the processing results (such as obtaining the lyrics of the selected song and the corresponding virtual character posture data according to the user request) to the car terminal 101.

[0049] The human-computer interaction method according to an embodiment of the present disclosure will be described in detail below based on the scenario described. Figure 1 Figure 2 , Figure 3a- Figure 3c

[0050] Figure 2 ​​The flowchart of the human-computer interaction method according to an embodiment of the present disclosure is shown schematically, including: in step S210, determining a virtual image corresponding to a currently played music based on the music; in step S220, determining first posture data corresponding to a current time from posture data of the virtual image based on a progress time of the music; in step S230, adjusting a display state of the first posture data of the virtual image in response to behavior data of a user at the current time; and / or controlling the virtual image to change from the first posture data to second posture data in response to the behavior data of the user at the current time.

[0051] Before step S210, a singing mode is further selected, and the first posture data is determined based on the singing mode; the singing mode includes a PK mode and a chorus mode, and the first posture data displayed by the virtual image is different in different modes. For example, when the lyrics of the music “Youth Training Manual” are “follow me, left hand and right hand, a slow motion” are played, if the user selects the PK mode, the first posture data controls the virtual image to display left and right nods, and if the user selects the chorus mode, the first posture data controls the virtual image to display left and right hands drawing a circle in turn.

[0052] In step S210, for a played music selected by the user, a virtual image corresponding to the music is determined. Specifically, the virtual image corresponding to the music can be determined in the following steps: Figure 3a In step S210, for a played music selected by the user, a virtual image corresponding to the music is determined. Specifically, the virtual image corresponding to the music can be determined in the following steps: Figure 3a A simple diagram of a music progress timeline according to an embodiment of the present disclosure is shown schematically. The music progress includes a prelude stage 301, a music stage 302, and an ending stage 303. The virtual image displayed in the music stage 302 is controlled by the first posture data or the second posture data, and the virtual image displayed in the prelude stage 301 and the ending stage 303 is controlled by the third posture data.

[0053] In step S220, the first posture data corresponding to the current time is determined from the posture data of the virtual image based on the progress time of the music. Determining the first posture data corresponding to the current time from the posture data of the virtual image includes: taking the current time as a starting point, obtaining a tune of the music within a preset time, and determining the first posture data based on the tune. For example, taking 0:00 of the music “Youth Training Manual” as a starting point, the tune of the music within 5 seconds belongs to the youth series, and the first posture data belongs to the youth series.

[0054] In the embodiment of the present application, the first posture data of the virtual image at the current time is determined from the posture data of the virtual image, which includes: taking the current time as the starting point, obtaining the lyrics of the music within a preset time, and determining the first posture data based on the semantic analysis result of the lyrics. For example, taking 0:00 of the music "Youth Manual" as the starting point, the lyrics of the music within 10 seconds are "follow me left hand right hand a slow motion, right hand left hand slow motion replay", and the semantic analysis result of the lyrics is "left hand and right hand respectively circle", so the first posture data of the virtual image is that the left hand and the right hand respectively circle.

[0055] In actual life, when a user sings a selected music, the user cannot always keep pace with the progress of the music playing. Compared with the currently played music, the user may sing faster or slower. According to the actual singing progress of the user, the duration of the first posture data of the virtual image is adjusted in time, so that the user can obtain a better experience in the singing process.

[0056] In the embodiment of the present application, the virtual image corresponding to the music is determined based on the currently played music, the first posture data of the virtual image at the current time is determined from the posture data of the virtual image based on the progress time of the music, the semantic recognition of the voice data of the user is performed, the singing progress of the user at the current time is determined based on the result of the semantic recognition, and when the singing progress and the current progress of the music are inconsistent, the duration of the display of the first posture data is delayed or reduced.

[0057] Figure 3b A flowchart for adjusting the first posture data of the virtual image is provided. In step S310, real-time voice data of a user is collected. In step S311, semantic recognition of the collected real-time voice data of the user is performed. Then the result of the semantic recognition is compared with the lyrics of the selected music to determine the singing progress of the user, that is, step S312. In step S313, it is judged whether the current singing progress of the user is consistent with the progress of the actually played music. If the singing progress of the user is inconsistent with the progress of the actually played music, step S314 is entered to delay / reduce the duration of the display of the first posture data. If the singing progress of the user is consistent with the progress of the actually played music, the duration of the display of the first posture data is not changed. Next, step S316 is entered to continue playing the first posture data of the virtual image.

[0058] In step S311, the step of semantic recognition includes but is not limited to extracting text feature information from the collected real-time voice data of the user, and identifying the predicted semantic information of the collected real-time voice data of the user based on the text feature information of the collected real-time voice data of the user. In the process of semantic recognition, a semantic recognition model is used, including semantic understanding technology and semantic analysis technology therein.

[0059] The semantic recognition model of the present application can be constructed based on any artificial neural network structure available for semantic recognition, for example, the semantic recognition model can be implemented based on BERT (Bidirectional Encoder Representation from Transformers), RNN (Recurrent Neural Network), CNN (Convolutional Neural Networks), RCNN (Rich feature hierarchies Convolutional Neural Network), etc., and there is no limitation thereto.

[0060] In Figure 3a For the embodiment of the present application, step S314 is further described. In one possible embodiment, t represents the current progress of the music playing, and t1 represents the current singing progress of the user. At this time, the current progress of the music playing is inconsistent with the current singing progress of the user, and the current progress of the music playing is faster than the current singing progress of the user. Then, the time length of the first pose data of the virtual image displayed at the t progress is delayed, and the delayed time is t-t1, so as to enable the virtual image displayed at the current music playing progress to be consistent with the singing progress of the user, thereby improving the experience of the user.

[0061] In another possible embodiment, t represents the current progress of the music playing, and t2 represents the current singing progress of the user. At this time, the current progress of the music playing is inconsistent with the current singing progress of the user, and the current progress of the music playing is slower than the current singing progress of the user. Then, the time length of the first pose data of the virtual image displayed at the t progress is shortened, and the shortened time is t2-t, so as to enable the virtual image displayed at the current music playing progress to be consistent with the singing progress of the user, thereby improving the experience of the user.

[0062] In step S230, the virtual image is controlled to be changed from the first pose data to the second pose data in response to the behavior data of the user at the current time.

[0063] Figure 3cA flowchart for updating first pose data of a virtual image to second pose data is provided. In step S320, real-time body data of a user is collected. In step S321, image recognition is performed on the collected real-time body data of the user, and then the result of the image recognition is compared with preset body data to determine whether the real-time body data is consistent with the preset body data. If the collected real-time body data is consistent with the preset body data, then in step S322, the first pose data is updated to the second pose data; if the collected real-time body data is not consistent with the preset body data, then in step S323, the first pose data of the virtual image is continued to be played.

[0064] In step S321, the image recognition can be, but is not limited to, real-time recognition of human body actions based on infrared images. After the collected real-time body data is preprocessed, a dynamic skeleton feature of the preprocessed image is extracted by using an infrared image human body pose extraction network to obtain a human body dynamic skeleton feature map to be recognized. Next, a region of interest of the human body dynamic skeleton feature map is intercepted as an action recognition grid input sequence, the intercepted human body dynamic skeleton feature map is adjusted in pixels, an action recognition network is used to classify and predict the action of the human body dynamic skeleton feature map after the adjustment, and then the human body dynamic skeleton feature map is compared with preset body data. The image recognition can be, but is not limited to, action recognition based on a depth image. Existing tools, such as Microsoft Kinect SDK, are used to directly obtain human body joint or skeleton information, and then a traditional pattern recognition algorithm is used for recognition; or image features are extracted from the collected real-time body data.

[0065] In an embodiment of the present application, a virtual image corresponding to a music piece is determined based on a currently played music piece; first pose data corresponding to a current time is determined from pose data of the virtual image based on a progress time of the music piece; real-time body data of a user is collected; and when the real-time body data matches preset body data, the first pose data is updated to second pose data to cater to the user.

[0066] In an embodiment of the present application, a virtual image corresponding to a music piece is determined based on a currently played music piece; first pose data corresponding to a current time is determined from pose data of the virtual image based on a progress time of the music piece; real-time body data of a user is collected; and when the real-time body data matches preset body data, the first pose data is updated to second pose data to cater to the user. Figure 3a In an embodiment of the present application, a virtual image corresponding to a music piece is determined based on a currently played music piece; first pose data corresponding to a current time is determined from pose data of the virtual image based on a progress time of the music piece; real-time body data of a user is collected; and when the real-time body data matches preset body data, the first pose data is updated to second pose data to cater to the user. Figure 3cThe process is further described. In one possible embodiment, t3 represents the current user's singing progress, at which time the virtual image is displayed with the first pose data at t3. In step S320, real-time body data of the user at t3 is collected, and image recognition is performed on the collected real-time body data of the user, the image recognition being to extract human dynamic skeleton features of the real-time body data. Then, the extracted human dynamic skeleton features of the real-time body data are compared with human dynamic skeleton features of preset body data, if the extracted human dynamic skeleton features of the real-time body data are inconsistent with the human dynamic skeleton features of the preset body data, the first pose data is continuously played; if the extracted human dynamic skeleton features of the real-time body data are consistent with the human dynamic skeleton features of the preset body data, the original first pose data is updated to second pose data.

[0067] In an embodiment of the present application, the preset body data includes at least one of face rotation, hand waving, and body shaking.

[0068] In an embodiment of the present application, after starting to play the selected music, a progress time of the music determines a stage of playing the music. After determining that the playing stage of the music is the coda stage 303 based on the progress time, the coda stage 303 has no lyric singing, and the virtual image is controlled to be displayed to the user in third pose data. For example, when the playing progress of the song is the progress time of the coda, the virtual image is in a bowing pose for a period of time (for example, 8 seconds). It should be noted that the third pose data of the virtual image used for display in the coda stage 303 can be preset according to the song, including the display pose and display duration of the virtual character, can also be preset according to the user's preference, or collected from one of the singer, performer, composer, and lyricist of the music.

[0069] In an embodiment of the present application, after starting to play the selected music, a progress time of the music determines a stage of playing the music. After determining that the playing stage of the music is the coda stage 303 based on the progress time, the coda stage 303 has no lyric singing, and the virtual image is controlled to be displayed to the user in third pose data. For example, when the playing progress of the song is the progress time of the coda, the virtual image is in a bowing pose for a period of time (for example, 8 seconds). It should be noted that the third pose data of the virtual image used for display in the coda stage 303 can be preset according to the song, including the display pose and display duration of the virtual character, can also be preset according to the user's preference, or collected from one of the singer, performer, composer, and lyricist of the music.

[0070] Figure 4 The structure block diagram of the human-computer interaction device according to an embodiment of the present application is schematically shown.

[0071] As Figure 4As shown, the human-computer interaction device 400 of the embodiment includes a virtual image determination unit 401, a first posture data determination unit 402, and a posture adjustment unit 403.

[0072] The virtual image determination unit 401 is configured to determine a virtual image corresponding to a currently played music based on the music.

[0073] The first posture data determination unit 402 is configured to determine first posture data corresponding to a current time from a series of posture data of the virtual image based on a progress time of the music.

[0074] The posture adjustment unit 403 is configured to adjust a display state of the first posture data of the virtual image in response to behavior data of a user at the current time, and / or control the virtual image to transition from the first posture data to second posture data in response to the behavior data of the user at the current time.

[0075] In the embodiment of the present application, in the virtual image determination unit 401, virtual images corresponding to the first posture data, the second posture data, and the third posture data are determined based on a currently played music. In the first posture data determination unit 402, relevant information about the first posture data is retrieved, the first posture data of the virtual image corresponding to the music is determined based on the currently played music, and then the retrieved first posture data is determined. In the posture adjustment unit 403, real-time body data of a user is collected, image recognition is performed on the collected real-time body data of the user, and then the result of the image recognition is compared with preset body data. According to the comparison result, it is determined whether to update the first posture data to the second posture data or continue to play the first posture data of the virtual image. It should be noted that after it is determined to continue to play the first posture data of the virtual image, the first posture data determination unit 402 should be returned to retrieve relevant information, and then determine the first posture data of the virtual image based on the currently played music.

[0076] Figure 5 An illustrative structural block diagram of an intelligent cockpit according to an embodiment of the present application is shown.

[0077] As shown in FIG. 4, the intelligent cockpit 500 includes a human-computer interaction device 501, a vehicle-mounted audio device 502, a music playing device 503, and other devices 504. Figure 5 As shown, the intelligent cockpit 500 includes a human-computer interaction device 501, a vehicle-mounted audio device 502, a music playing device 503, and other devices 504.

[0078] The vehicle-mounted audio device 502 is in communication connection with the music playing device 503, and is configured to play the music out loud.

[0079] The music playing device 503 is in communication connection with the human-computer interaction device 501, and is configured to play music according to a request of a user.

[0080] In the embodiment of the present application, the other device 504 can include but is not limited to an audio acquisition device, which is in communication connection with the human-computer interaction device 501, and is used to provide the voice data of the user to the human-computer interaction device 501 in real time. For example, when the user sings, a microphone is used to record and save the song sung by the user, and then the song is sent to the human-computer interaction device 501, so as to determine or adjust the first posture data of the virtual image corresponding to the song. The quantified evaluation of the singing level of the user based on the audio information provided by the audio acquisition device includes that the audio information includes at least one of sound intensity, tone, tonality, and frequency.

[0081] In another embodiment of the present application, the audio acquisition device is used to provide the voice data of the user to the virtual image determination unit 401, the first posture data determination unit 402, and the posture adjustment unit 403 in real time.

[0082] In the embodiment of the present application, the other device 504 can include but is not limited to an image acquisition device, which is in communication connection with the human-computer interaction device 501, and is used to provide the body data of the user to the human-computer interaction device 501 in real time. Optionally, the image acquisition device 501 includes a vehicle-mounted camera or a vehicle-mounted radar. Preferably, the image acquisition device is a vehicle-mounted camera.

[0083] In the embodiment of the present application, the other device 504 can include but is not limited to a singing evaluation device, which is in communication connection with the human-computer interaction device, and is used to quantitatively evaluate the singing level of the user based on the audio information provided by the audio acquisition device. The singing level of the user is quantitatively evaluated based on the body data of the user provided by the image acquisition device. The quantified evaluation includes matching the body data of the user with the first posture data in the song, and determining the singing level of the user based on the coincidence degree of the matching.

[0084] The singing evaluation device further includes at least one of a system score, a like score, and a reward score. For example, after the user finishes singing, the system performs a weighted score on the song sung by the user, and according to the proportion of the main song and the refrain, a final score is obtained by weighted calculation of the score of each sentence.

[0085] The embodiment of the present application provides a computer readable storage medium, which stores computer instructions, and when the computer instructions run on a computer, the computer executes the method of human-computer interaction.

[0086] The computer readable storage medium described above can be implemented by any type of volatile or nonvolatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special purpose computer.

[0087] Corresponding structures, acts, deeds, and all elements of all devices or steps of the method described in the disclosure are intended to be encompassed by the term "means" or "step" plus function without regard to structure. The description of the various preferred embodiments of the disclosure has been presented for purposes of illustration and description, and is not intended to be exhaustive or limited to the disclosure in the form disclosed. The disclosure is susceptible to various modifications and alternative forms, and should be construed and interpreted in the context of the disclosure. The terminology used by the inventor or publications identified herein, and by others, is intended to be interpreted in an inclusive, and not an exclusive, sense. The use of "including", "comprising", or "having" and variations thereof herein is meant to encompass the items listed thereafter and equivalents thereof as well as additional items.

Claims

1. A human-computer interaction method, comprising the steps of: determining a virtual image corresponding to a currently played music based on the music; determining first posture data corresponding to a current time from posture data of the virtual image based on a progress time of the music; adjusting a display state of the first posture data of the virtual image in response to behavior data of a user at the current time; and / or controlling the virtual image to change from the first posture data to second posture data in response to the behavior data of the user at the current time; wherein the adjusting the display state of the first posture data of the virtual image in response to the behavior data of the user at the current time comprises: collecting voice data of the user for semantic recognition, and determining a singing progress of the user at the current time based on a result of the semantic recognition; delaying or reducing a time length of displaying the first posture data when the singing progress and a current progress of the music are inconsistent. 2.The method of claim 1, further comprising: determining a performance stage of the music based on the progress time; controlling the virtual image to display in third posture data in response to the performance stage being a prelude stage and / or a coda stage. 3.The method of claim 2, the posture data is collected from at least one of a singer, a performer, a composer and a lyricist of the music; and / or is pre-set based on a preference of the user. 4.The method of claim 1, wherein the determining the first posture data corresponding to the current time from the posture data of the virtual image comprises: determining the first posture data based on a singing mode. 5.The method of claim 1, wherein the determining the first posture data corresponding to the current time from the posture data of the virtual image comprises: acquiring a tune of the music within a preset time from the current time, and determining the first posture data based on the tune; and / or acquiring lyrics of the music within the preset time from the current time, and determining the first posture data based on a semantic analysis result of the lyrics. 6.The method of claim 1, wherein the controlling the virtual image to change from the first posture data to the second posture data in response to the behavior data of the user at the current time comprises: collecting real-time body data of the user; determining that the real-time body data matches preset body data, and updating the first posture data to the second posture data to cater to the user. 7.The method of claim 6, the preset body data comprises at least one of face turning, hand waving and body shaking. 8.The method of claim 7, the first posture data is customized based on the preset body data. 9.A human-computer interaction device, comprising: a virtual image determination unit configured to determine a virtual image corresponding to a currently played music based on the music; a first posture data determination unit configured to determine first posture data corresponding to a current time from a series of posture data of the virtual image based on a progress time of the music. ​ The posture adjusting unit is configured to adjust a display state of the first posture data of the virtual image in response to behavior data of the user at the current time, and / or control the virtual image to transition from the first posture data to the second posture data in response to the behavior data of the user at the current time. The adjusting of the display state of the first posture data of the virtual image in response to the behavior data of the user at the current time includes: collecting voice data of the user for semantic recognition, determining a singing progress of the user at the current time based on a result of the semantic recognition, and delaying or reducing a display duration of the first posture data when the singing progress and a current progress of the music are inconsistent.

10. An intelligent cabin, comprising: The human-computer interaction device in claim 9; A music playing device in communication connection with the human-computer interaction device, configured to play music according to a request of a user; A vehicle-mounted audio device in communication connection with the music playing device, configured to play out the music.

11. The intelligent cabin in claim 10, further comprising: An audio collecting device in communication connection with the human-computer interaction device, configured to provide voice data of the user to the human-computer interaction device in real time.

12. The intelligent cabin in claim 11, The audio collecting device is a microphone.

13. The intelligent cabin in claim 11, further comprising: An image collecting device in communication connection with the human-computer interaction device, configured to provide limb data of the user to the human-computer interaction device in real time.

14. The intelligent cabin in claim 13, The image collecting device includes a vehicle-mounted camera or a vehicle-mounted radar.

15. The intelligent cabin in claim 13, further comprising: A singing evaluation device in communication connection with the human-computer interaction device, and configured to quantitatively evaluate a singing level of the user based on audio information provided by the audio collecting device.

16. The intelligent cabin in claim 15, the quantitatively evaluating the singing level of the user based on the audio information provided by the audio collecting device includes: The audio information includes at least one of sound intensity, tone color, tonality, and frequency.

17. The intelligent cabin in claim 15, further comprising: The singing evaluation device quantitatively evaluates the singing level of the user based on the limb data of the user provided by the image collecting device.

18. The intelligent cabin in claim 17, the singing evaluation device quantitatively evaluates the singing level of the user based on the limb data of the user provided by the image collecting device includes: Matching the limb data of the user with first posture data in the music; Determining the singing level of the user based on a coincidence degree of the matching.

19. The intelligent cabin in claim 15, the singing evaluation device further includes at least one of a system score, a like score, and a reward score.

20. A human-computer interaction device, comprising: A virtual image determining unit configured to determine a virtual image corresponding to a currently played music based on the currently played music. The first attitude data determination unit is configured to determine first attitude data corresponding to the current time from a series of attitude data of the virtual image based on the progress time of the music; The attitude adjustment unit is configured to adjust the display state of the first attitude data of the virtual image in response to the behavior data of the user at the current time; and / or control the virtual image to change from the first attitude data to second attitude data in response to the behavior data of the user at the current time; wherein the adjusting the display state of the first attitude data of the virtual image in response to the behavior data of the user at the current time comprises: collecting voice data of the user for semantic recognition, determining the singing progress of the user at the current time based on the result of the semantic recognition; and when the singing progress and the current progress of the music are inconsistent, delaying or reducing the time length of displaying the first attitude data.

21. The human-computer interaction device according to claim 20, further comprising: An audio acquisition device configured to provide the virtual image determination unit, the first attitude data determination unit and the attitude adjustment unit with voice data of the user in real time.

22. The human-computer interaction device according to claim 21, The audio acquisition device is a microphone.

23. The human-computer interaction device according to claim 20 or 21, further comprising: An image acquisition device configured to provide the attitude adjustment unit with limb data of the user in real time.

24. The human-computer interaction device according to claim 23, The image acquisition device is a vehicle-mounted camera or a vehicle-mounted sensor.

25. An intelligent cabin, comprising: The human-computer interaction device according to any one of claims 20-24.

26. A vehicle, comprising: The intelligent cabin according to any one of claims 10-19; or the intelligent cabin according to claim 25.

27. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by a processor to implement the steps of the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Virtual image interaction method in virtual-reality scene

    CN107422862A

  • Limb interaction method and system based on virtual person

    CN108416420A