Cross-box interaction method based on virtual digital human, medium and equipment

By generating virtual digital people based on the boxes and user portraits, the problem of difficulty in cross-box interaction is solved, real-time chorus and personalized entertainment experience are achieved, and the quality of user interaction in places such as KTV is improved.

CN120653103APending Publication Date: 2025-09-16FUJIAN KAIMI NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510504189.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In multi-private room environments such as KTVs and bars, it is difficult for users to achieve convenient and real-time cross-private room interaction and chorus, resulting in a very limited entertainment experience and poor chorus quality.

Method used

By generating virtual digital people based on the boxes and user portraits, audio data mixing and display synchronization across boxes are achieved, combined with lighting effect adjustment to provide a personalized interactive experience.

Benefits of technology

It enables real-time chorus interaction across boxes, improves communication efficiency and entertainment experience, and enhances user immersion and interactive fun.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653103A_ABST
    Figure CN120653103A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-box interaction method based on virtual digital humans, a medium and equipment, and the method comprises the steps: obtaining box portraits of a first box and a second box and user group portraits in the box portraits, and generating the corresponding virtual digital humans; then the first virtual digital person initiates a chorus request containing a box ID and track information to a second virtual digital person, adds a chorus track after receiving response information of the second virtual digital person, and enters a chorus mode after receiving a chorus trigger instruction; and in the mode, the audio acquisition devices respectively acquire sound data of the boxes where the audio acquisition devices are located and mutually receive the sound data, and the corresponding virtual digital humans perform sound mixing on the sound data to obtain sound mixing data and play the sound mixing data through the audio playing devices in the respective boxes. According to the scheme, cross-box real-time chorus interaction can be realized, the communication efficiency of cross-box interaction is improved, and the user experience of a multi-box entertainment scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of box interaction, and in particular to a cross-box interaction method based on virtual digital humans, a computer-readable storage medium, and an electronic device. Background Art

[0002] In traditional entertainment venues with multiple private rooms, such as KTVs and bars, each room is typically relatively independent and lacks effective interaction. Patrons in different rooms can only engage in entertainment activities within their own private rooms, making it difficult to communicate and collaborate with patrons in other rooms, resulting in a significantly limited entertainment experience.

[0003] For example, in a chorus scenario, if a customer wishes to sing with others in other rooms, existing technologies often fail to provide a convenient, real-time interactive channel. Traditional methods may require manual communication and coordination with other rooms, which is not only inefficient but also subject to time and space constraints, making it impossible to fulfill the chorus request in a timely manner. Moreover, even if coordination is successful, it is difficult to ensure real-time sound transmission, which significantly reduces the quality of the chorus. Summary of the Invention

[0004] To this end, it is necessary to provide a cross-box interaction method based on virtual digital humans to solve the problems of communication difficulties, low efficiency, poor chorus quality, etc. that users have when performing cross-box interactive chorus in the existing technology.

[0005] To achieve the above objectives, in a first aspect, the present application provides a cross-box interaction method based on a virtual digital human, which is applicable to an interactive scenario of multiple boxes, wherein the multiple boxes include a first box and a second box, the first box is provided with a first audio acquisition device and a first audio playback device, and the second box is provided with a second audio acquisition device and a second audio playback device;

[0006] The method comprises the following steps:

[0007] Obtain a first private room portrait corresponding to the first private room and a first user group portrait corresponding to the first user group in the first private room, and generate a first virtual digital human based on the first private room portrait and the first user group portrait; and obtain a second private room portrait corresponding to the second private room and a second user group portrait corresponding to the second user group in the second private room, and generate a second virtual digital human based on the second private room portrait and the second user group portrait;

[0008] Initiate a chorus request to the second virtual digital human through the first virtual digital human, wherein the chorus request includes the ID information of the first box and the chorus repertoire information;

[0009] receiving the response information of the second virtual digital human to the chorus request, adding the chorus repertoire information to the song sequence, and entering the chorus mode after receiving the chorus trigger instruction;

[0010] In chorus mode, the first audio acquisition device acquires the first sound data in the first box and receives the second sound data in the second box acquired by the second audio acquisition device. The first virtual digital human mixes the first sound data and the second sound data to obtain first mixed data and plays it through the first audio playback device; or, the second audio acquisition device acquires the second sound data in the second box and receives the first sound data in the first box acquired by the first audio acquisition device. The second virtual digital human mixes the second sound data and the first sound data to obtain second mixed data and plays it through the second audio playback device.

[0011] Furthermore, a first display device is provided in the first box, and a second display device is provided in the second box. The method includes:

[0012] In chorus mode, the first virtual digital human image and the second virtual digital human image are synchronously displayed on the first display device and / or the second display device;

[0013] Collect the first singing data of the singer in the first box, and adjust the first virtual digital human image displayed on the first display device and / or the second display device according to the first singing data, or collect the second singing data of the singer in the second box, and adjust the second virtual digital human image displayed on the first display device and / or the second display device according to the second singing data, the first singing data and / or the second singing data including any one or more of the singer's corresponding body movements, expressions, timbre, pitch, and rhythm.

[0014] Furthermore, the first box is provided with a first lighting device, and the second box is provided with a second lighting device. The method includes:

[0015] In the chorus mode, the first virtual digital human and / or the second virtual digital human adaptively generates first adjustment information according to the song parameters corresponding to the chorus track information;

[0016] The first virtual digital human determines a first adjustment factor based on the first box portrait and the first user group portrait, generates second adjustment information based on the first adjustment factor and the first adjustment information, and adjusts the operating parameters of the first lighting device based on the second adjustment information; or, the second virtual digital human determines a second adjustment factor based on the second box portrait and the second user group portrait, generates third adjustment information based on the second adjustment factor and the first adjustment information, and adjusts the operating parameters of the second lighting device based on the third adjustment information.

[0017] Furthermore, a first lighting device and a first image acquisition device are further provided in the first box, and the method includes:

[0018] The first image acquisition device acquires first behavior data of a user in the first box, and the first audio acquisition device acquires third sound data in the first box;

[0019] The first virtual digital human determines the scene mode of the current first box based on the first behavior data and / or the third sound data. When it is determined that the learned scene mode belongs to the preset scene mode, the first virtual digital human adjusts the operating parameters of the first lighting device and / or the first audio playback device according to the lighting adjustment strategy corresponding to the preset scene mode, and determines the dissemination level according to the first user group portrait;

[0020] The first virtual digital person determines the range of boxes that need to be disseminated based on the dissemination level and the location of the first box, and transmits the first message to the virtual digital people in all boxes within the box range, so that after receiving the first message, the virtual digital people in these boxes control the song in their own box to stop playing and adjust the operating parameters of the lighting equipment in their own box according to the preset interaction strategy.

[0021] Furthermore, a first display device is provided in the first box, and the method includes:

[0022] The first virtual digital person receives the blessing message sent by the virtual digital person in the box range, and presents the virtual digital person who initiates the blessing message, the blessing action and the blessing message through the first display device.

[0023] Furthermore, the method includes:

[0024] When it is detected that the playback resumption condition is met, the virtual digital person corresponding to the box in the song playback pause state controls the audio playback device of its own box to resume song playback.

[0025] Furthermore, the method includes:

[0026] Initiate a game co-play request to the second virtual digital human through the first virtual digital human, wherein the game co-play request includes the ID information of the first box and co-play game information;

[0027] receiving a response message from the second virtual digital human to the game co-play request, wherein the response message includes the ID of the user of the second box who agrees to participate in the co-play;

[0028] Enter the game co-play mode, and the IDs of users participating in the co-play game will be displayed synchronously in the first and second boxes. The game will be hosted by the first virtual digital human in the first box or by the second virtual digital human in the second box. When the game end conditions are met, game settlement information will be generated and displayed.

[0029] Furthermore, the method includes:

[0030] After the chorus is finished, the first virtual digital human stores the audio data of the chorus;

[0031] The first virtual digital person receives a friend-adding request from a user in the first box, sends the friend-adding request to the second virtual digital person in the second box, and upon receiving feedback information from the second virtual digital person regarding the friend-adding request, establishes a cross-box social group and sends the audio data of the chorus track to the cross-box social group, wherein the social group includes the ID information of all users participating in the chorus.

[0032] In a second aspect, the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the cross-box interaction method based on a virtual digital human as described in the first aspect of the present application is implemented.

[0033] In a third aspect, the present application also provides an electronic device comprising a processor and a storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by the processor, the cross-box interaction method based on a virtual digital human as described in the first aspect of the present application is implemented.

[0034] Different from the existing technology, the cross-box interaction method, medium, and device based on virtual digital humans described in the above technical solution include: first obtaining the box portraits of the first and second boxes and the portraits of the user groups within them, and generating corresponding virtual digital humans; then the first virtual digital human initiates a chorus request containing the box ID and song information to the second virtual digital human, and adds the chorus song after receiving the response information from the second virtual digital human, and enters the chorus mode after receiving the chorus trigger instruction; in this mode, the audio acquisition devices respectively collect the sound data of the boxes in which they are located and receive it from each other, and the corresponding virtual digital humans mix the sound data, obtain the mixed data, and play it through the audio playback devices in their respective boxes. The above solution can realize real-time cross-box chorus interaction, improve the communication efficiency of cross-box interaction, and enhance the user experience of multi-box entertainment scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 Flowchart of a cross-box interaction method based on virtual digital humans according to a first exemplary embodiment of the present invention;

[0036] Figure 2 Flowchart of a cross-box interaction method based on virtual digital humans according to a second exemplary embodiment of the present invention;

[0037] Figure 3 Flowchart of a cross-box interaction method based on virtual digital humans according to a third exemplary embodiment of the present invention;

[0038] Figure 4 Flowchart of a cross-box interaction method based on virtual digital humans according to a fourth exemplary embodiment of the present invention;

[0039] Figure 5 Flowchart of a cross-box interaction method based on virtual digital humans according to a fifth exemplary embodiment of the present invention;

[0040] Figure 6 Flowchart of a cross-box interaction method based on virtual digital humans according to a sixth exemplary embodiment of the present invention;

[0041] Figure 7 A schematic diagram of a module of an electronic device according to the present invention;

[0042] Reference numerals:

[0043] 10. Electronic equipment;

[0044] 101. Processor;

[0045] 102. Storage medium. DETAILED DESCRIPTION

[0046] In order to explain in detail the possible application scenarios, technical principles, specific solutions that can be implemented, and the purpose and effects of this application, the following is a detailed description of the specific embodiments listed in conjunction with the accompanying drawings. The embodiments described herein are only used to more clearly illustrate the technical solutions of this application and are therefore only examples and are not intended to limit the scope of protection of this application.

[0047] References to "embodiments" herein mean that the specific features, structures, or characteristics described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the word "embodiment" in various places in the specification does not necessarily refer to the same embodiment, nor does it particularly limit its independence or relevance to other embodiments. In principle, in this application, as long as there are no technical contradictions or conflicts, the various technical features mentioned in the embodiments can be combined in any manner to form a corresponding implementable technical solution.

[0048] Unless otherwise defined, the technical terms used herein have the same meanings as those generally understood by those skilled in the art to which this application belongs; the use of relevant terms herein is only for describing specific embodiments and is not intended to limit this application.

[0049] In the description of this application, the term "and / or" is used to describe a logical relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A exists, B exists, and both A and B exist. In addition, the character " / " in this document generally indicates that the objects before and after are in a logical "or" relationship.

[0050] In this application, terms such as "first" and "second" are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual quantity, priority or sequence relationship between these entities or operations.

[0051] Without further limitations, in this application, the words "include", "comprise", "have" or other similar expressions used in sentences are intended to cover non-exclusive inclusion. These expressions do not exclude the presence of additional elements in the process, method or product that includes the elements, so that the process, method or product that includes a series of elements may include not only those limited elements, but also other elements that are not explicitly listed, or also include elements that are inherent to such process, method or product.

[0052] Consistent with the understanding in the Examination Guidelines, in this application, expressions such as "greater than," "less than," and "exceed" are understood to exclude the number itself; expressions such as "above," "below," and "within" are understood to include the number itself. Furthermore, in the description of the embodiments of this application, "multiple" means more than two (including two), and similar expressions related to "multiple" are also understood in this manner, such as "multiple groups," "multiple times," etc., unless otherwise specifically defined.

[0053] To facilitate further understanding of the solution of this application, the following definitions are given for the terms involved in the solution of this application:

[0054] In this application, a private room can be a KTV room, a cinema, a game room, an immersive experience center, a multi-functional conference room, etc. The private room portrait includes the room's space and layout information (such as the room area, seat distribution and arrangement, whether there is a separate bathroom, etc.), facility and equipment information (such as audio equipment, karaoke system, display device configuration, etc.), decoration style information (such as color matching), usage record information (such as booking frequency, peak usage time), and cost-related information (such as charging standards, additional charges).

[0055] In this application, the user group profile includes at least one user profile, which includes basic user information (such as age, gender, region, occupation, membership status, etc.), music preferences, and behavioral characteristics (such as the frequency of requesting various types of songs). The reserved box themes include birthday themes, Valentine's Day themes, anime themes, rock themes, etc.

[0056] In this application, lighting equipment, audio playback equipment, display equipment, image acquisition equipment and audio acquisition equipment can be set up in the box.

[0057] The lighting equipment in the box includes basic lighting equipment, atmosphere-creating lighting equipment, effect-enhancing lighting equipment, etc. Audio playback equipment includes audio sound-generating equipment (such as speakers, microphones), audio signal processing equipment (such as amplifiers, mixers, sound effect processors), audio playback source equipment (such as karaoke machines), and audio transmission equipment (such as audio cables or wireless transmission equipment). Display equipment includes video display equipment, LED light curtains, projectors, etc. Image acquisition equipment can be cameras distributed at different angles in the box, such as panoramic cameras, depth cameras, thermal imaging cameras, etc. Audio acquisition equipment can be microphones distributed at different locations in the box, such as dynamic microphones, condenser microphones, array microphones, and wireless microphones.

[0058] In the first aspect, the present application provides a cross-box interaction method based on virtual digital people, which is suitable for interactive scenarios of multiple boxes. The multiple boxes include a first box and a second box. The first box is provided with a first audio acquisition device and a first audio playback device, and the second box is provided with a second audio acquisition device and a second audio playback device.

[0059] like Figure 1 As shown, the method includes the following steps:

[0060] Step S101: Obtain a first private room portrait corresponding to the first private room and a first user group portrait corresponding to the first user group in the first private room, and generate a first virtual digital human based on the first private room portrait and the first user group portrait; and obtain a second private room portrait corresponding to the second private room and a second user group portrait corresponding to the second user group in the second private room, and generate a second virtual digital human based on the second private room portrait and the second user group portrait;

[0061] Step S102: Initiating a chorus request to the second virtual digital human through the first virtual digital human, wherein the chorus request includes the ID information of the first box and the chorus repertoire information;

[0062] Step S103: receiving the response information of the second virtual digital human to the chorus request, adding the chorus repertoire information to the song sequence, and entering the chorus mode after receiving the chorus trigger instruction;

[0063] Step S104: In the chorus mode, the first audio acquisition device acquires the first sound data in the first box, and receives the second sound data in the second box acquired by the second audio acquisition device. The first virtual digital human mixes the first sound data and the second sound data to obtain first mixed data and plays it through the first audio playback device; or, the second audio acquisition device acquires the second sound data in the second box, and receives the first sound data in the first box acquired by the first audio acquisition device. The second virtual digital human mixes the second sound data and the first sound data to obtain second mixed data and plays it through the second audio playback device.

[0064] In step S101, the image and style of the generated first virtual digital persona can be determined based on a comprehensive consideration of the user's preferences, the private room's decorative style, and the theme information of the current private room reservation. For example, if the user profiles of the first user group are mostly male, the image of the first virtual digital persona can be set to female. Conversely, if the user profiles of the first user group are mostly female, the image of the first virtual digital persona can be set to male, and the clothing designed for the first virtual digital persona matches the decorative style of the merchant or private room. The image and style of the second virtual digital persona are generated in a similar manner to the first virtual digital persona and will not be further described here.

[0065] In step S102, the ID information of the first box serves as the unique identifier of the first box, which can ensure that when the chorus request is transmitted to the second box, the user of the second box can learn the initiating box of the chorus request and the chorus repertoire information through the second virtual digital person in the second box.

[0066] In step S103, when the second virtual digital human in the second private room receives the chorus request, it can ask the user in the second private room to confirm whether to agree to the chorus request through voice or pop-up window. If a confirmation reply is received, a response message to the chorus request will be generated and sent to the first virtual digital human in the first private room. The chorus trigger instruction can be triggered by a conversation between the user in the first private room and the first virtual digital human. For example, when the first virtual digital human receives the user in the first private room saying "enter chorus mode" or "start chorus", it is considered to have triggered the chorus trigger instruction. It can also be triggered by the user manually clicking the start chorus button on the screen.

[0067] In step S104, in chorus mode, the audio capture devices in both rooms operate simultaneously. The first audio capture device collects first sound data from the first room, while simultaneously receiving second sound data from the second room, collected by the second audio capture device. Conversely, the second audio capture device performs a similar operation. The first and second virtual digital humans respectively mix the collected sound data to generate first and second mixed data, which are then played back through the audio playback devices in their respective rooms, achieving real-time chorus.

[0068] Furthermore, the sound data of the first box collected by the first audio collection device includes voice data of conversations between users, singing data of users singing songs, and sound data expressing users' emotions (such as laughter, crying, screaming, sighing, etc.). Preferably, in chorus mode, the system will first denoise and clean the sound data of the first box collected by the first audio collection device, and only retain the singing data of the users singing songs (i.e., the first sound data) and mix it with the singing data of the users singing songs from the second box that has also undergone denoising (i.e., the second sound data), so as to present a better singing effect in chorus mode.

[0069] This application generates virtual digital people based on the box portrait and user group portrait, which can provide a personalized interactive experience for each box and user. For example, chorus tracks can be recommended based on the user's music preferences, and the mixing effect can be adjusted according to the sound characteristics of the box. Through the above solution, the isolation between traditional boxes can be broken. Users in different boxes can easily initiate and participate in chorus activities. In addition, by using audio collection, transmission and mixing technology, real-time interaction of sound between the two boxes is realized, making users feel as if they are singing in the same space, improving the quality and immersion of the chorus.

[0070] In some embodiments, as Figure 2 As shown, a first display device is further provided in the first box, and a second display device is provided in the second box, and the method includes:

[0071] Step S201: In chorus mode, the first virtual digital human image and the second virtual digital human image are synchronously displayed on the first display device and / or the second display device;

[0072] Step S202: Collect the first singing data of the singer in the first box, and adjust the first virtual digital human image displayed on the first display device and / or the second display device according to the first singing data, or collect the second singing data of the singer in the second box, and adjust the second virtual digital human image displayed on the first display device and / or the second display device according to the second singing data. The first singing data and / or the second singing data include any one or more of the singer's corresponding body movements, expressions, timbre, pitch, and rhythm.

[0073] In this embodiment, when displaying the virtual digital human image, the display device will present it on the screen based on the received virtual digital human image data through graphics rendering technology, including details such as the virtual digital human's appearance, clothing, posture, etc., so that users in both boxes can see the virtual digital human image corresponding to the other box in real time.

[0074] In this embodiment, the first audio acquisition device and the first audio playback device can be used to respectively collect the behavior data and sound data of the singer in the first box to obtain the first singing data, and the second audio acquisition device and the second audio playback device can be used to respectively collect the behavior data and sound data of the singer in the second box to obtain the second singing data.

[0075] Taking the collection of the first singing data as an example, the body movements and expressions of the singer can be captured through the camera in the first box. After the camera obtains the video stream, it uses image recognition technology to detect and track the human body joints, facial features, etc. in the video, thereby parsing the body movement and expression data. For the collection of timbre, pitch and rhythm, the first audio acquisition device is used to convert the current sound signal of the singer into an electrical signal, and after analog-to-digital conversion, the sound data is obtained, and the sound data is transmitted to the background analysis model. The analysis model uses an audio analysis algorithm to extract timbre features (such as harmonic structure, etc.), pitch (through fundamental frequency detection algorithm) and rhythm (based on beat detection algorithm) from the audio data.

[0076] After collecting the first or second performance data, the system adjusts the virtual human's image according to preset mapping rules. If the singer's body movements are detected to be large, the mapping rules may be set to increase the amplitude of the virtual human's arm swings; if the voice is bright, the virtual human's facial gloss may be adjusted accordingly. This adjusted data is then transmitted again via the network to the first and second display devices, which then re-render the virtual human's image, presenting real-time changes that match the singer's performance.

[0077] This solution adjusts the virtual human's image in real time based on the singer's body movements, facial expressions, timbre, and other performance data. This allows users to intuitively experience how their own singing impacts the virtual human's performance, adding a touch of fun to the interaction. This approach allows for more lively interactions between users in different rooms, for example, by using exaggerated body movements to elicit amusing reactions from the virtual human, enhancing the interactive experience. Furthermore, since each singer's performance data is unique, adjustments to the virtual human's image based on this data can achieve personalized presentation. Different timbres, pitches, and rhythms give the virtual human a distinct "singing style," satisfying users' demand for a personalized entertainment experience.

[0078] In some embodiments, as Figure 3 As shown, a first lighting device is further provided in the first box, and a second lighting device is provided in the second box. The method includes:

[0079] Step S301: In the chorus mode, the first virtual digital human and / or the second virtual digital human adaptively generates first adjustment information according to song parameters corresponding to the chorus track information;

[0080] Step S302: The first virtual digital human determines a first adjustment factor based on the first box portrait and the first user group portrait, generates second adjustment information based on the first adjustment factor and the first adjustment information, and adjusts the operating parameters of the first lighting device based on the second adjustment information; or, the second virtual digital human determines a second adjustment factor based on the second box portrait and the second user group portrait, generates third adjustment information based on the second adjustment factor and the first adjustment information, and adjusts the operating parameters of the second lighting device based on the third adjustment information.

[0081] In this embodiment, the system pre-stores relevant song parameters for various choral tracks, such as tempo, emotional tone (cheerful, lyrical, passionate, etc.), climax position, etc. After the chorus mode is activated, the first virtual digital human and / or the second virtual digital human obtain information about the current choral track, conduct an in-depth analysis of its song parameters, and then generate the first adjustment information.

[0082] For example, if the song has a fast tempo and a cheerful emotional tone, the first virtual digital human and / or the second virtual digital human may determine that a brighter and more frequently flashing lighting effect is needed to match it, and accordingly generate first adjustment information, which includes a range of light brightness values, color tendencies (such as warm tones), a flashing frequency range, etc.

[0083] After generating the first adjustment information, the first virtual digital person will refer to the first box portrait and the first user group portrait. The first box portrait covers the size of the box space, decoration style (such as simple style, retro style, etc.), existing lighting layout, etc.; the first user group portrait includes the user age distribution, music preference style, etc. If the first box space is large, the decoration is modern and simple, and most of the users are young people who prefer pop music, the first virtual digital person will combine these factors to determine a suitable first adjustment factor. The first adjustment factor can be used as a correction coefficient for the first adjustment information. For example, considering the large space of the box, the light brightness adjustment range is appropriately increased; because the user prefers pop music, the color selection highlights more vibrant colors, etc.

[0084] Similarly, the second virtual digital person determines the second adjustment factor based on the second box portrait (including box size, decoration features, basic lighting settings) and the second user group portrait (age level, music preferences) to adapt to the specific situation of the second box.

[0085] The first virtual digital human applies the first adjustment factor to the first adjustment information, generating the second adjustment information using specific calculation rules. For example, if the initial brightness adjustment value in the first adjustment information is 50%, and the first adjustment factor is calculated to be 1.2 based on the room and user conditions, the brightness adjustment value in the second adjustment information can be set to 60%. The first virtual digital human then transmits the second adjustment information to the controller of the first lighting device via the network. The controller of the first lighting device adjusts the operating parameters of the first lighting device, such as brightness, color, and flashing frequency, according to the second adjustment information, achieving real-time changes in the lighting effects.

[0086] The second virtual digital human performs a similar operation, combining the second adjustment factor with the first adjustment information to generate third adjustment information, and sends it to the controller of the second lighting device, thereby adjusting the operating parameters of the second box lighting device to match the lighting effect with the chorus repertoire, box environment and user group.

[0087] This solution allows for real-time lighting adjustments based on the rhythm and emotional characteristics of the chorus track. For example, enhancing lighting brightness and flickering effects during climaxes creates an immersive atmosphere that aligns closely with the music, allowing users to immerse themselves more deeply in the chorus and enhance their entertainment experience. Adjustment factors are generated based on the profiles of both the private rooms and user groups, ensuring that lighting adjustments take into account the unique environment of each private room and the preferences of each user group. Different private rooms and users with different musical preferences can each receive lighting effects tailored to their needs, enhancing the personalization of entertainment services.

[0088] In some embodiments, as Figure 4 As shown, the first box is further provided with a first lighting device and a first image acquisition device, and the method includes:

[0089] Step S401: a first image acquisition device acquires first behavior data of a user in a first box, and a first audio acquisition device acquires third sound data in the first box;

[0090] Step S402: The first virtual digital human determines the current scene mode of the first private room based on the first behavior data and / or the third sound data. If the determined scene mode belongs to a preset scene mode, the first virtual digital human adjusts the operating parameters of the first lighting device and / or the first audio playback device according to the lighting adjustment strategy corresponding to the preset scene mode, and determines the dissemination level based on the first user group portrait.

[0091] Step S403: The first virtual human determines the range of rooms to be disseminated based on the dissemination level and the location of the first room. It then transmits the first message to the virtual human beings in all rooms within the range. Upon receiving the first message, the virtual human beings in these rooms stop playing the song in their own rooms and adjust the operating parameters of the lighting equipment in their own rooms according to the preset interaction strategy. The first message includes the current scene mode of the first room.

[0092] Furthermore, a first display device is provided in the first box, and the method includes: a first virtual digital person receives a blessing message sent by a virtual digital person in the box range, and presents the virtual digital person who initiates the blessing message, the blessing action, and the blessing message through the first display device.

[0093] In this embodiment, there are multiple first image capture devices, which can be set up in different locations within the first private room to ensure full coverage of the private room space and capture various user behaviors. The video stream images captured by the first image capture device can be analyzed in real time using an image recognition algorithm. For example, the image recognition algorithm can identify the user's body movements, such as whether they are holding a birthday cake, throwing streamers, kneeling on one knee with flowers, and other special movements, as well as the user's facial expressions, such as whether they are filled with celebratory joy, and convert this information into first behavior data, providing a key basis for subsequent scene judgment.

[0094] The first audio capture device uses a highly sensitive microphone, clearly capturing all sounds within the first private room. Using sound signal processing technology, the sound signals are converted into digital third-party sound data. This data includes users' cheers, singing, and conversations. Using an audio feature extraction algorithm, it identifies characteristic sound elements of the scene, such as the melody of a birthday song or wedding greetings.

[0095] When determining the scene mode, the first virtual digital human uses a data analysis model to comprehensively analyze the first behavioral data and the third sound data. This data analysis model pre-stores behavioral and sound feature templates for various preset scene modes. For example, the birthday scene template includes the melody of the birthday song and the action of holding a cake; the wedding scene template includes the melody of the wedding march, the sound of the newlyweds taking vows, and related celebratory actions. When the input data matches a template to a certain degree, the scene mode currently in the first private room is determined.

[0096] Once the scene mode is determined, the first virtual digital human immediately retrieves the corresponding lighting adjustment strategy and audio playback device adjustment strategy from the preset rule library (which contains the mapping relationship between scene modes and interaction strategies). For a birthday scene, the lighting adjustment strategy may be to switch the lights to a warm yellow, dim the brightness and set it to a flashing effect to create a romantic atmosphere; the audio playback device in the first box may pause the current song and play birthday greeting songs. At the same time, the first virtual digital human can also determine the box communication level based on the portrait of the first user group, taking into account factors such as the user's consumption habits and membership level. The corresponding box communication level may be higher for user groups with strong consumption power and high membership level. The higher the box communication level, the wider the spread of the information that the first box is currently in the current preset scene mode.

[0097] The first virtual digital human defines the communication range based on the communication level and the location of the first private room. For birthday scenes, since the communication level is relatively low, location information, such as the entertainment venue's network topology or positioning system, can be combined to determine the nearby private rooms as the communication range. Wedding scenes, on the other hand, have a higher communication level. If the entertainment venue is a chain KTV, the network architecture and merchant information management system can be used to determine all private rooms under the same merchant as the communication range.

[0098] The first virtual digital person then transmits a first message, containing the current scene mode of the first private room (such as a birthday or wedding) and data related to the first private room (such as user characteristics within the first private room, the first private room's ID information, etc.), via the network to the virtual digital persons in all private rooms within the transmission range. After receiving the first message, the virtual digital persons in these rooms first control the pause of the song in their own private room to avoid sound interference. Then, they adjust the operating parameters of the lighting equipment according to the preset interaction strategy. For example, the lights in the adjacent private room switch to warm tones similar to the first room to create an overall celebratory atmosphere. In the case of a wedding, users in other private rooms can choose to send gifts, record blessing videos, etc. through the operation interface of their own private room virtual person or the display device. These blessings or gifts recorded by other private room users are transmitted to the first virtual digital person in the first private room via the network, and the first virtual digital person presents these gifts and blessing videos on the display device in the first private room.

[0099] Through precise data collection and intelligent analysis, this solution can quickly and accurately determine the scene mode within the private room and automatically adjust the equipment within the room to create an entertainment environment that suits the scene, greatly enhancing the user experience and immersion in the room. For the first private room in different scene modes, the communication level is determined based on the user group profile, and differentiated communication ranges are set to achieve personalized service. A high communication level provides high-value users with a wider interactive experience, meeting the needs of different users, improving user satisfaction and loyalty to the entertainment venue, and thus creating more business opportunities for merchants, such as increased user consumption for a better interactive experience. After triggering a specific scene mode, users can receive real-time blessings and attention from other private rooms. This large-scale cross-private room linkage creates a warm celebratory atmosphere, enhances the social nature of the entertainment venue, and effectively improves the user's interactive experience.

[0100] In some embodiments, the method includes: when it is detected that the playback resumption condition is met, the virtual digital person corresponding to the box in the song playback pause state controls the audio playback device of its own box to resume song playback.

[0101] In this embodiment, playback resumption conditions can be pre-set, and these conditions may be triggered by time, events, or user actions. For example, a fixed time threshold, such as 5 minutes, can be set. When the duration after the scene mode is triggered reaches this threshold, the playback resumption condition is met. Alternatively, the playback resumption condition can be based on a specific event, such as when a user in the first private room completes a birthday wish or a key step in a wedding ceremony. The playback resumption condition can also be set based on user actions, such as when a user in the first private room or a private room within the transmission range manually issues a resume command through the private room control panel.

[0102] The virtual human in the box can interact with relevant data interfaces to obtain data such as time information, event status feedback, and user operation instructions. For example, the virtual human communicates with the time management module to obtain the duration since the current scene mode was triggered; receives signals from the event monitoring system indicating the completion of key events; and monitors instructions from the user operation interface through the box control network. Once a set condition is detected, the virtual human immediately captures this information and prepares to perform subsequent recovery operations.

[0103] When the virtual digital human detects that the playback resumption conditions are met, it quickly generates an instruction to resume song playback. The instruction contains clear control information, such as the audio device identifier that specifies the audio device to resume playback (to ensure accurate control in a multi-audio device environment), the song track information to resume playback (in case the current playlist has been updated), and the playback position information (if the song has been played to a specific progress, it needs to continue playing from that position). The virtual digital human sends this instruction to the controller of the audio playback device in its own box through the network communication protocol. After receiving the instruction, the audio playback device controller in the box immediately parses the instruction. According to the analysis results, it locates the position where the song playback is interrupted, reads the song data, and resumes playback of the song. During the entire process, the audio playback device will feedback the playback status information to the controller, and the controller will then transmit this information back to the virtual digital human so that the virtual digital human can confirm whether the recovery operation has been successfully completed.

[0104] Through the above method, it is possible to avoid the situation where the song playing in other boxes is interrupted for a long time due to the triggering of the scene mode of the first box, thereby ensuring the continuity of the entertainment activities.

[0105] In some embodiments, as Figure 5 As shown, the method includes:

[0106] Step S501: Initiate a game co-play request to the second virtual digital human through the first virtual digital human, wherein the game co-play request includes the ID information of the first box and the co-play game information;

[0107] Step S502: receiving the response information of the second virtual digital human to the game play request, wherein the response information includes the user ID of the second box who agrees to participate in the game play;

[0108] Step S503: Entering the game co-play mode, the IDs of the users who are co-playing the game are displayed synchronously in the first and second boxes, and the game is hosted by the first virtual digital human in the first box or by the second virtual digital human in the second box. When the game end conditions are met, the game settlement information is generated and displayed.

[0109] In this embodiment, when the second virtual human receives a request from the first virtual human, it parses the request, extracts the first private room ID and game information, and displays a play invitation to the user in the second private room via a display device or voice prompt. The user then chooses whether to participate in the second private room. If the user in the second private room agrees to participate, the second virtual human obtains the user ID information of the user who agrees to participate, packages this user ID information into a response message, and returns the response message to the first virtual human. After receiving the response message, the first virtual human extracts the user ID information of the second private room user who agrees to participate. Then, the first and second virtual human beings simultaneously display the ID information of all users participating in the play on the display devices in their respective private rooms.

[0110] Based on preset rules (such as the default hosting by the inviting party's virtual human or alternating hosting by the virtual human beings in the two rooms), the first virtual human in the first room or the second virtual human in the second room is determined to host the game. The hosting virtual human obtains detailed information such as the game flow and rules from the game server or local game resource library and guides the user through the game through voice and text prompts. During the game, user operation data (such as user answers in quiz games and user action commands in action games) is collected in real time and synchronized between the two rooms via the network to ensure consistent game status. When the game meets the preset end conditions (such as reaching the specified time, completing a specific task, or answering all questions), the first or second virtual human generates game settlement information based on the game rules and user operation data, including each user's score, ranking, and completion time. After the settlement information is generated, it is displayed simultaneously in the first and second rooms through interaction with the display device, allowing users to intuitively understand the game results.

[0111] This solution breaks down the boundaries between private rooms, providing users with a brand-new cross-room gaming experience. This enriches the gameplay of entertainment venues and meets users' diverse entertainment needs. For example, while traditional private room entertainment is limited to singing, cross-room gaming adds more interactive game items, enhancing the attraction of entertainment venues.

[0112] In some embodiments, as Figure 6 As shown, the method includes:

[0113] Step S601: After the chorus song is finished, the first virtual digital human stores the audio data of the chorus song;

[0114] Step S602: The first virtual digital human receives a friend-adding request from a user in the first box, sends the friend-adding request to the second virtual digital human in the second box, and upon receiving feedback from the second virtual digital human regarding the friend-adding request, establishes a cross-box social group, and sends the audio data of the chorus track to the cross-box social group, where the social group includes the ID information of all users participating in the chorus.

[0115] In this embodiment, when a user in the first private room sends a friend request through voice interaction with the first virtual human, the first virtual human captures this operation signal and first obtains the ID of the user initiating the request and relevant information about the target second private room. It then packages the friend request information, which includes key information such as the initiator's ID and the request content. Then, via a network communication protocol, it sends the request to the second virtual human in the second private room.

[0116] After receiving the friend request, the second virtual digital person displays the request to the target user in the second box through a display device or voice prompt. The target user gives feedback of agreement or rejection. If agreed, the second virtual digital person packages the feedback information, including the target user ID, and sends it back to the first virtual digital person. After receiving the feedback information, the first virtual digital person confirms that both parties agree to add friends, and then begins to create a social group across boxes. The ID information of all users participating in the chorus is obtained, and these user ID information is integrated into the newly created social group. At the same time, the first virtual digital person retrieves the corresponding chorus track audio data from the previously stored audio data and uploads it to the file storage area of ​​the social group so that group members can access and download it at any time.

[0117] This solution provides users in different rooms with a convenient way to add friends, breaking down the social barriers between rooms. The interactive foundation established through choral activities further fosters social connections between users, meeting their need to expand their social circles during entertainment and enhancing their social experience in entertainment venues.

[0118] In some embodiments, users can also set permissions for their private rooms by interacting with the virtual digital people in the room or by operating on the operating interface of the display device. For example, interactive elements such as switch buttons or drop-down menus containing various permission options can be set on the operating interface of the display device. For example, for "Do you agree to open the portrait of this private room?", a switch button is set, and turning it on means agreeing, and turning it off means rejecting; for options such as "Do you agree to communicate, sing, or play games in other private rooms?", a drop-down menu can be used to provide options such as "Agree", "Reject", and "Only specific rooms can request" for users to choose.

[0119] When a virtual digital person in a private room (the initiator) initiates a request to another private room (such as a chorus request), the virtual digital person will first obtain the permission setting data of the target private room from the central server. After obtaining the permission data, the virtual digital person will determine the permission of the target private room for the requested type. If the target private room is set to reject this type of request, the virtual digital person of the initiator will exclude the target private room from the list of requested private rooms and will not send a request to it. If the target private room is set to agree or has other specific conditions, the virtual digital person of the initiator will package and send the request according to the corresponding situation.

[0120] When the virtual human in the target room receives a request from another room, it will first obtain the permission settings for that room from the central server. It will then determine whether to respond to the request based on the permission settings. If the permission is set to "yes," the virtual human will display the request information to the user in the room, who will then decide whether to accept the request. If the permission is set to "reject," the virtual human will simply discard the request without displaying it to the user in the room or sending any response to the initiator.

[0121] Through this solution, users can flexibly set the permissions for their private rooms to respond to different types of requests based on their needs and preferences, fully protecting their privacy and autonomy. For example, when engaging in private activities, users can deny all external requests to ensure that their private room activities are not disturbed; however, when they want to interact, they can grant corresponding permissions and enjoy the fun of cross-private room entertainment.

[0122] In a second aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the cross-box interaction method based on virtual digital humans as described in the first aspect of the present invention.

[0123] The computer-readable storage medium may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories.

[0124] The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface storage, an optical disc, or a compact disc read-only memory (CD ROM); the magnetic surface storage may be a magnetic disk storage or a magnetic tape storage.

[0125] The volatile memory may be a random access memory (RAM) that is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronized dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM). The computer-readable storage medium described in the embodiments of the present invention is intended to include these and any other suitable types of memory.

[0126] like Figure 7 As shown, in the third aspect, the present invention provides an electronic device 10, including a processor 101 and a storage medium 102, on which a computer program is stored. When the computer program is executed by the processor, the cross-box interaction method based on virtual digital people as described in the first aspect of the present invention is implemented.

[0127] In some embodiments, the processor can be implemented through software, hardware, firmware or a combination thereof, and can use at least one of a circuit, a single or multiple application-specific integrated circuits (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor, so that the processor can execute some or all of the steps or any combination of the steps in the cross-box interaction method based on virtual digital humans described in the various embodiments of the present application.

[0128] It should be noted that although the above embodiments have been described herein, this does not limit the scope of patent protection of the present invention. Therefore, based on the innovative concept of the present invention, changes and modifications to the embodiments described herein, or equivalent structural or equivalent process transformations made using the contents of the present invention's specification and drawings, and direct or indirect application of the above technical solutions to other related technical fields, are all included in the scope of patent protection of the present invention.

Claims

1. A cross-box interaction method based on virtual digital humans, characterized in that: Applicable to an interactive scenario of multiple boxes, the multiple boxes include a first box and a second box, the first box is provided with a first audio acquisition device and a first audio playback device, and the second box is provided with a second audio acquisition device and a second audio playback device; The method comprises the following steps: Obtain a first private room portrait corresponding to the first private room and a first user group portrait corresponding to the first user group in the first private room, and generate a first virtual digital human based on the first private room portrait and the first user group portrait; and obtain a second private room portrait corresponding to the second private room and a second user group portrait corresponding to the second user group in the second private room, and generate a second virtual digital human based on the second private room portrait and the second user group portrait; Initiate a chorus request to the second virtual digital human through the first virtual digital human, wherein the chorus request includes the ID information of the first box and the chorus repertoire information; receiving the response information of the second virtual digital human to the chorus request, adding the chorus repertoire information to the song sequence, and entering the chorus mode after receiving the chorus trigger instruction; In chorus mode, the first audio acquisition device acquires the first sound data in the first box and receives the second sound data in the second box acquired by the second audio acquisition device. The first virtual digital human mixes the first sound data and the second sound data to obtain first mixed data and plays it through the first audio playback device; or, the second audio acquisition device acquires the second sound data in the second box and receives the first sound data in the first box acquired by the first audio acquisition device. The second virtual digital human mixes the second sound data and the first sound data to obtain second mixed data and plays it through the second audio playback device.

2. The cross-box interaction method based on virtual digital humans according to claim 1, characterized in that: A first display device is further provided in the first box, and a second display device is provided in the second box. The method includes: In chorus mode, the first virtual digital human image and the second virtual digital human image are synchronously displayed on the first display device and / or the second display device; Collect the first singing data of the singer in the first box, and adjust the first virtual digital human image displayed on the first display device and / or the second display device according to the first singing data, or collect the second singing data of the singer in the second box, and adjust the second virtual digital human image displayed on the first display device and / or the second display device according to the second singing data, the first singing data and / or the second singing data including any one or more of the singer's corresponding body movements, expressions, timbre, pitch, and rhythm.

3. The cross-box interaction method based on virtual digital humans according to claim 1, characterized in that: The first compartment is further provided with a first lighting device, and the second compartment is provided with a second lighting device. The method includes: In the chorus mode, the first virtual digital human and / or the second virtual digital human adaptively generates first adjustment information according to the song parameters corresponding to the chorus track information; The first virtual digital human determines a first adjustment factor based on the first box portrait and the first user group portrait, generates second adjustment information based on the first adjustment factor and the first adjustment information, and adjusts the operating parameters of the first lighting device based on the second adjustment information; or, the second virtual digital human determines a second adjustment factor based on the second box portrait and the second user group portrait, generates third adjustment information based on the second adjustment factor and the first adjustment information, and adjusts the operating parameters of the second lighting device based on the third adjustment information.

4. The cross-box interaction method based on virtual digital humans according to claim 1, characterized in that: The first box is also provided with a first lighting device and a first image acquisition device, and the method includes: The first image acquisition device acquires first behavior data of a user in the first box, and the first audio acquisition device acquires third sound data in the first box; The first virtual digital human determines the scene mode of the current first box based on the first behavior data and / or the third sound data. When it is determined that the learned scene mode belongs to the preset scene mode, the first virtual digital human adjusts the operating parameters of the first lighting device and / or the first audio playback device according to the lighting adjustment strategy corresponding to the preset scene mode, and determines the dissemination level according to the first user group portrait; The first virtual digital human determines the range of boxes that need to be disseminated based on the dissemination level and the location of the first box, and transmits the first message to the virtual digital humans in all boxes within the box range, so that after receiving the first message, the virtual digital humans in these boxes control the songs in their own boxes to stop playing, and adjust the operating parameters of the lighting equipment in their own boxes according to the preset interaction strategy. The first message includes the scene mode currently in which the first box is located.

5. The cross-box interaction method based on virtual digital humans according to claim 4 is characterized in that: A first display device is also provided in the first box, and the method includes: The first virtual digital person receives the blessing message sent by the virtual digital person in the box range, and presents the virtual digital person who initiates the blessing message, the blessing action and the blessing message through the first display device.

6. The cross-box interaction method based on virtual digital humans according to claim 4 is characterized in that: The method comprises: When it is detected that the playback resumption condition is met, the virtual digital person corresponding to the box in the song playback pause state controls the audio playback device of its own box to resume song playback.

7. The cross-box interaction method based on virtual digital humans according to claim 1, characterized in that: The method comprises: Initiate a game co-play request to the second virtual digital human through the first virtual digital human, wherein the game co-play request includes the ID information of the first box and co-play game information; receiving a response message from the second virtual digital human to the game co-play request, wherein the response message includes the ID of the user of the second box who agrees to participate in the co-play; Enter the game co-play mode, and the IDs of users participating in the co-play game will be displayed synchronously in the first and second boxes. The game will be hosted by the first virtual digital human in the first box or by the second virtual digital human in the second box. When the game end conditions are met, game settlement information will be generated and displayed.

8. The cross-box interaction method based on virtual digital humans according to claim 1, characterized in that: The method comprises: After the chorus is finished, the first virtual digital human stores the audio data of the chorus; The first virtual digital person receives a friend-adding request from a user in the first box, sends the friend-adding request to the second virtual digital person in the second box, and upon receiving feedback information from the second virtual digital person regarding the friend-adding request, establishes a cross-box social group and sends the audio data of the chorus track to the cross-box social group, wherein the social group includes the ID information of all users participating in the chorus.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the cross-box interaction method based on virtual digital humans as described in any one of claims 1 to 8 is implemented.

10. An electronic device, characterized in that: It includes a processor and a storage medium, wherein a computer program is stored on the storage medium, and when the computer program is executed by the processor, the cross-box interaction method based on virtual digital humans as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Multimedia interaction method, related device, equipment and storage medium

    CN111741370A

  • Networking chorus method based on virtual image and terminal

    CN114494537A

  • Intelligent debugging method and system for adaptive scene of combined atmosphere lamp

    CN117241445A

  • Loudspeaker box light control method and control system

    CN119584388A

  • Interaction method, apparatus, device, and storage medium

    US20240428777A1