Home music system, device and apparatus supporting multi-person interaction

CN122531397APending Publication Date: 2026-08-07SOYO TECH DEV CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOYO TECH DEV CO LTD
Filing Date
2026-07-06
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]目前,传统家庭音乐系统在交互功能上存在一些局限,仅能支持单用户语音交互,无法适配家庭多成员同时使用的场景

Benefits of technology

可以看出,本申请中所描述的支持多人互动的家庭音乐系统、装置及设备,通过为每个房间配置独立音箱与控制模块,各控制模块互联互通且区分主控与子控,可避免多用户同时操作造成的指令冲突,适配多房间多用户使用场景;通过主控模块定位互动用户所处房间,结合用户身份与语音确定互动意图,实现多用户针对性语音互动,打破传统系统的单向交互局限,从而有效提高了家庭音乐系统对多人互动场景的适配性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531397A_ABST
    Figure CN122531397A_ABST
Patent Text Reader

Abstract

The application discloses a family music system, device and equipment supporting multi-person interaction, which comprises a sound box and a control module. The sound box is used for playing audio in a corresponding room in a room. The first control module is used for receiving a first voice of a first user. The master control module is used for determining first user identity information corresponding to the first voice based on a preset voiceprint recognition technology. The first voice interaction intention and second user identity information are determined according to the first user identity information and the first voice. Voice information of users in the a rooms is acquired, the voice information is recognized based on the preset voiceprint recognition technology, and room information is obtained. The first voice interaction intention is realized according to the room information, the first control module and the a sound boxes. The application improves the adaptability of the family music system to the multi-person interaction scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio control technology, and in particular to a home music system, device, and equipment that supports multi-user interaction. Background Technology

[0002] With the development of smart home technology, home music systems have become an important home device, effectively enhancing the auditory experience of home users.

[0003] Currently, traditional home music systems have limitations in interactive functions, supporting only single-user voice interaction and failing to adapt to scenarios where multiple family members use the system simultaneously. When multiple people issue voice commands at the same time, the system is prone to response confusion and command conflicts, and it cannot achieve voice interaction and collaborative operation between multiple users based on the music system. Its interactive adaptability is poor, making it difficult to meet the personalized and diverse needs of multiple family members using the system together.

[0004] Therefore, how to improve the adaptability of home music systems to multi-person interactive scenarios has become an urgent problem to be solved. Summary of the Invention

[0005] This application provides a home music system, device, and equipment that supports multi-user interaction, improving the adaptability of the home music system to multi-user interaction scenarios.

[0006] In a first aspect, embodiments of this application provide a home music system supporting multi-user interaction. The system is installed in a target home, which includes *a* rooms. The system includes *a* speakers and *a* control modules, with each control module corresponding to one speaker. Each room contains one speaker and one control module. The *a* control modules are interconnected. Each *a* control module includes one main control module and *a-1* sub-control modules; *a* is an integer greater than 1. The a speakers are used to play audio in the corresponding rooms of the a rooms; A first control module is used to receive a first voice message from a first user; the first user is in a first room; the first room is any one of the a rooms; the control module is the control module in the first room; The main control module is used to determine the first user identity information corresponding to the first voice based on a preset voiceprint recognition technology; determine the first voice interaction intention and the second user identity information based on the first user identity information and the first voice; acquire the voice information of users in the a rooms, recognize the voice information based on the preset voiceprint recognition technology, and obtain the room information of the user corresponding to the second user identity information; and realize the first voice interaction intention based on the room information, the first control module, and the a speakers.

[0007] Secondly, embodiments of this application provide a home music device that supports multi-user interaction, including the system described in the first aspect.

[0008] Thirdly, embodiments of this application provide a home music device that supports multi-user interaction, including the apparatus as described in the second aspect, or the system as described in the first aspect.

[0009] Implementing this application will have the following beneficial effects: As can be seen, the home music system, device, and equipment described in this application that supports multi-user interaction, by configuring independent speakers and control modules for each room, with each control module interconnected and distinguishing between master and sub-controllers, can avoid command conflicts caused by simultaneous operation by multiple users and adapt to multi-room, multi-user usage scenarios; by locating the room where the interactive user is located through the master control module, and determining the interactive intent by combining the user's identity and voice, targeted voice interaction for multiple users can be realized, breaking the one-way interaction limitations of traditional systems, thereby effectively improving the adaptability of the home music system to multi-user interactive scenarios. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the background art, the accompanying drawings used in the embodiments of this application or the background art will be described below.

[0011] Figure 1 This is an application scenario diagram of a home music system supporting multi-user interaction provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a home music system supporting multi-user interaction provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a control module provided in an embodiment of this application; Figure 4 This is a flowchart of a method for determining first user identity information provided in an embodiment of this application; Figure 5 This is a flowchart of a method for determining room information provided in an embodiment of this application; Figure 6 This is a schematic diagram of a voice channel provided in an embodiment of this application; Figure 7 This is a schematic diagram of a voice group provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a home music device that supports multi-user interaction, provided in an embodiment of this application. Detailed Implementation

[0012] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0013] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0014] It should be understood that the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this document indicates that the preceding and following related objects are in an "or" relationship. In the embodiments of this application, "multiple" refers to two or more.

[0015] In the embodiments of this application, "at least one item" or its similar expression refers to any combination of these items, including any combination of a single item or a plurality of items. "One or more" means one or more, while "multiple" means two or more. For example, "at least one item" of a, b, or c can represent the following seven cases: a, b, c; a and b; a and c; b and c; a, b, and c. Each of a, b, and c can be an element or a set containing one or more elements.

[0016] In this application, the term "connection" refers to various connection methods, such as direct connection or indirect connection, to achieve communication between devices. This application does not impose any limitations on this.

[0017] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0018] The following describes the relevant content, concepts, meanings, technical issues, technical solutions, and beneficial effects involved in the embodiments of this application.

[0019] First, let me explain some of the technical terms or phrases used in this application: Home music system: A system installed in the target home to provide audio playback, voice interaction, and multi-user interactive services to each room of the home.

[0020] Preset voiceprint recognition technology: a technology that identifies and distinguishes users by extracting voiceprint features (e.g., Mel-spectral coefficients, fundamental frequencies, etc.) from speech signals.

[0021] Voice channel: A communication link used to transmit user voice signals and audio data.

[0022] Voice group: A virtual communication set consisting of multiple user terminals connected to a network, supporting real-time voice communication and message exchange among members within the group.

[0023] With the rapid iteration and popularization of smart home technology, home music systems have gradually become an indispensable smart device in modern homes. Breaking the limitations of traditional single-speaker use, they provide multi-scenario, personalized audio playback services, effectively enriching family life and enhancing the auditory experience and convenience of family members. Currently, most home music systems on the market have basic functions such as audio playback, volume adjustment, and track switching. Some mid-to-high-end systems also support multi-room music customization settings, allowing users to set different playback content and volume levels for different rooms in the home according to their own needs, meeting the auditory needs of family members in different spaces.

[0024] However, in real-world home applications, home music systems are typically used by multiple family members. The need for simultaneous use and interaction among multiple users is increasingly frequent, but traditional home music systems currently suffer from significant limitations in their interactive design, making them ill-suited for such multi-user scenarios. Specifically, traditional home music systems often employ a single interaction mode, supporting only one-way voice interaction between a single user and the system. They cannot accommodate the simultaneous actions initiated by multiple family members. For example, if a family member wants to share the currently playing music with someone in another room or engage in voice communication with other members, traditional home music systems cannot provide the corresponding interactive channels and functional support, requiring the use of additional communication devices. This results in a limited range of interactive scenarios and poor adaptability for home music systems.

[0025] Therefore, how to improve the adaptability of home music systems to multi-person interactive scenarios has become an urgent problem to be solved.

[0026] Please see Figure 1 , Figure 1 This is an application scenario diagram of a home music system supporting multi-user interaction provided in an embodiment of this application. It can be seen that the target family includes four rooms: Room A, Room B, Room C, and Room D. Each room is equipped with a control module (S1, S2, S3, S4) and a speaker (L1, L2, L3, L4). Each control module corresponds to a speaker and is used to provide audio playback and voice interaction services in its respective room. The control modules are interconnected, with one control module serving as the master control module and the others as sub-control modules, collectively forming a distributed control network. Users can initiate voice commands through the corresponding control module in any room to achieve voice interaction between multiple users and collaborative audio control across multiple rooms.

[0027] Please see Figure 2 , Figure 2 This is a schematic diagram of a home music system supporting multi-user interaction provided in an embodiment of this application; the system is installed in a target home, and the target home includes a rooms; as shown Figure 2 As shown, the system includes: *a* speakers and *a* control modules, with each control module corresponding to one speaker; each room has one speaker and one control module; the *a* control modules are interconnected; please refer to [link / reference]. Figure 3 , Figure 3 This is a schematic diagram of the structure of *a* control modules provided in an embodiment of this application. It can be seen that the *a* control modules include one main control module and *a-1* sub-control modules; *a* is an integer greater than 1; where: The a speakers are used to play audio in the corresponding rooms of the a rooms; In this embodiment of the application, the speaker can be any of the following: portable speaker, wall-mounted speaker, in-wall speaker, etc., and is not limited thereto.

[0028] Each of the a control modules can integrate a voice unit to collect user voice and audio data of the surrounding environment of the room. The voice unit can include any of the following: microphone, microphone array, pickup module, audio acquisition chip, etc., without limitation.

[0029] In a specific embodiment, each speaker is deployed in a corresponding room to play audio content in that room, including but not limited to music, voice intercom data, system prompts, etc.

[0030] In some embodiments, 'a' can be equal to 4, and the target household contains 4 rooms (rooms A, B, C, and D), corresponding to the deployment of speakers L1, L2, L3, and L4. Speaker L1 in room A plays popular music for users to use while relaxing in the living room; Speaker L2 in room B plays children's stories for the children in the bedroom to listen to; Speaker L3 in room C plays white noise to help users sleep; The L4 speaker in room D plays news and information to meet the information needs of the study user.

[0031] Each speaker plays different content independently without interfering with each other, enabling personalized audio services in multiple rooms.

[0032] In some embodiments, users may want the same song to play synchronously throughout the entire house: The main control module sends synchronous playback commands to speakers L1, L2, L3, and L4. Speakers L1, L2, L3, and L4 play the same music synchronously in their respective rooms, keeping the audio track timing consistent. Users can adjust the volume or switch tracks in any room using the corresponding control module, and all room speakers will respond synchronously to achieve a whole-room background music effect.

[0033] In some embodiments, a user in room B can share the currently playing song with a user in room C: The main control module obtains the song information currently being played by speaker L2 in room B and sends a playback command to speaker L3 in room C; Speaker L3 started playing the song in room C, allowing the user in room C to listen to the shared content; The speakers in rooms A and B continue playing the original content, while only the speaker in room C changes the track, achieving precise point-to-point music sharing.

[0034] In this way, by using a one-to-one correspondence between speakers and rooms, and allowing audio to be played independently in each room, it is possible to achieve independent control and personalized playback of audio in multiple rooms, meet the different listening needs of different family members, and avoid mutual interference between audio in different rooms.

[0035] A first control module is used to receive a first voice message from a first user; the first user is in a first room; the first room is any one of the a rooms; the control module is the control module in the first room; In a specific embodiment, the audio signal in the first room can be collected in real time by the voice unit in the first control module, and the voice signal issued by the first user can be extracted and identified to form the corresponding first voice.

[0036] It should be explained that the first control module can be a main control module or a sub-control module; if the first control module is a sub-control module, then the first control module can send the first voice to the main control module.

[0037] The main control module is used to determine the first user identity information corresponding to the first voice based on a preset voiceprint recognition technology; determine the first voice interaction intention and the second user identity information based on the first user identity information and the first voice; acquire the voice information of users in the a rooms, recognize the voice information based on the preset voiceprint recognition technology, and obtain the room information of the user corresponding to the second user identity information; and realize the first voice interaction intention based on the room information, the first control module, and the a speakers.

[0038] In this embodiment of the application, the preset voiceprint recognition technology can be preset in advance or defaulted. Specifically, the preset voiceprint recognition technology can be any of the following: a recognition method based on speech feature parameters, a voiceprint recognition method based on deep learning, a recognition method based on Gaussian mixture model, etc., without limitation.

[0039] In a specific embodiment, the main control module can analyze the first voice using preset voiceprint recognition technology to obtain the first user identity information; then, it can analyze the first voice based on the first user identity information to obtain the first voice interaction intent and the second user identity information; then, it can control the aforementioned a control modules to collect the voices of users in the corresponding rooms of a rooms to obtain voice information, and the main control module receives all the voice information collected by these a control modules; next, it can identify the voice information based on preset voiceprint recognition technology to obtain the room information of the user corresponding to the second user identity information; finally, it can realize the first voice interaction intent based on the room information, the first control module, and a speakers.

[0040] In this way, the main control module identifies the first user's identity through voiceprint recognition, combines the interactive intent with the second user's identity through voice parsing, and then identifies the voice in each room to accurately locate the room where the second user is located. This enables accurate differentiation and automatic matching of multiple users, avoiding misidentification and accidental triggering.

[0041] Furthermore, based on room information, control modules, and speakers, interactive intentions can be executed to transmit voice and audio to the target room, enhancing the relevance, privacy, and real-time nature of the interaction. This method eliminates the need for manual selection of objects or rooms, automatically completing the entire process, improving system intelligence and response efficiency, and enhancing the user experience, adaptability, and stability in multi-user, multi-room scenarios.

[0042] Optional, please refer to Figure 4 , Figure 4 This is a flowchart of a method for determining first user identity information provided in an embodiment of this application. It can be seen that, in determining the first user identity information corresponding to the first voice based on preset voiceprint recognition technology, the main control module is specifically used to execute... Figure 4The steps shown are as follows: S11. Perform speech preprocessing on the first speech to obtain the first speech data; S12. Extract the voiceprint from the first speech data based on the preset voiceprint recognition technology to obtain the first voiceprint; S13. Match the first voiceprint with the voiceprints in the preset voiceprint library to obtain the first matching result; S14. When the first matching result includes a successful match, determine the first user identity information based on the first matching result and the preset voiceprint database.

[0043] In this embodiment, the preset voiceprint library can be preset in advance or defaulted; speech preprocessing can include at least one of the following: noise reduction, filtering, echo cancellation, amplitude normalization, speech endpoint detection, etc., which are not limited here; the first matching result includes any one of the following: matching successful, matching failed; each user identity information can include at least one of the following: user name, user identifier, user permission level, user preference information, family role, etc., which are not limited here.

[0044] In a specific embodiment, the first speech is preprocessed to obtain the first speech data. Specifically, the first speech can first be denoised and filtered to suppress environmental noise, circuit interference and echo interference; then speech endpoint detection is performed to remove silent segments and invalid audio segments, and retain the valid speech part; then the valid speech part is adjusted for gain and normalized for amplitude to make the speech signal strength within a preset reasonable range; then the valid speech part can be converted into a digital signal to obtain the first speech data.

[0045] Then, the voiceprint in the first speech data can be extracted based on the preset voiceprint recognition technology to obtain the first voiceprint. For example, the preset voiceprint recognition technology can be a recognition method based on speech feature parameters. The preset voiceprint recognition technology is used to perform frame segmentation and windowing processing on the first speech data, extract the voiceprint feature parameters in the speech data, and normalize and vector encode the extracted voiceprint feature parameters to obtain the first voiceprint for identity matching. The voiceprint feature parameters can include any of the following: Mel frequency cepstral coefficients, linear prediction cepstral coefficients, perceptual linear prediction coefficients, fundamental frequency, formant features, spectral envelope features, etc., which are not limited here.

[0046] Furthermore, the first voiceprint is matched with voiceprints in a preset voiceprint library to obtain a first matching result. Specifically, the first voiceprint can be compared with the voiceprints of each user in the preset voiceprint library to obtain multiple similarity values. The maximum similarity value among these multiple similarity values ​​is determined. If the maximum similarity value is greater than or equal to a preset similarity threshold, the matching is considered successful; otherwise, the matching is considered unsuccessful, thus obtaining the first matching result. The preset similarity threshold can be preset in advance or defaulted.

[0047] When the first matching result includes a successful match, the first user identity information is determined based on the first matching result and the preset voiceprint library. Specifically, the first matching result may also include the user voiceprint that was successfully matched in the preset voiceprint library. The first user identity information is determined based on the user voiceprint. For example, a preset mapping relationship between voiceprints and user identity information may be stored in advance, and the first user identity information corresponding to the successfully matched user voiceprint is determined based on the mapping relationship.

[0048] When the first matching result includes a matching failure, it is determined that the first voice did not match the corresponding user in the preset voiceprint library, the first voice is determined to be the voice of an unknown user, and preset abnormal handling operations are performed, such as: outputting an abnormal prompt message to indicate that the user has not registered or the identity cannot be recognized; refusing to perform subsequent voice interaction operations; or, initiating the user voiceprint registration process to add the voiceprint corresponding to the first voice to the preset voiceprint library.

[0049] Thus, by preprocessing the first speech to obtain the first speech data, noise, interference signals, and invalid audio segments can be effectively filtered out, improving speech quality. Based on the preset voiceprint recognition technology, the voiceprint is extracted and matched with the preset voiceprint database, which can accurately identify the user's identity and improve the accuracy and reliability of identity confirmation. When the match is successful, the user's identity information is determined according to the voiceprint database, which can realize fast and automatic user identity verification, avoid misidentification and misoperation, improve system security and interaction efficiency, and provide a reliable identity foundation for subsequent voice interaction.

[0050] Optionally, in determining the first voice interaction intent and the second user identity information based on the first user identity information and the first voice, the main control module is specifically used for: S21. Determine the first text data corresponding to the first voice; S22. Perform semantic recognition on the first text data to obtain the first voice interaction intent and the first interaction object information; S23. Determine the first user's first permission information based on the first user's identity information; S24. Based on the first permission information and the first voice interaction intent, perform permission verification on the first user to obtain a first verification result; the first verification result includes any one of the following: verification successful, verification failed; S25. When the first verification result includes the verification success, determine the second user identity information based on the first interactive object information; S26. When the first verification result includes the verification failure, generate a first prompt voice; control the speaker in the first room to play the first prompt voice.

[0051] In the embodiments of this application, each voice interaction intent may include any of the following: music request, calling a specific user, multi-person intercom, system settings, etc., without limitation.

[0052] In a specific embodiment, the first text data corresponding to the first speech can be determined first. Specifically, the first speech is processed by Automatic Speech Recognition (ASR) to convert the speech into text, resulting in a text sequence corresponding to the first speech, i.e., the first text data. Then, semantic recognition is performed on the first text data to obtain the first speech interaction intent and the first interaction object information. Specifically, the first text data can be segmented, part-of-speech tagging, syntactic analysis, and semantic understanding to extract intent keywords and object keywords. The intent keywords are matched according to a preset intent library to determine the first speech interaction intent. The object keywords such as user identifier, user name, and user role are matched according to a preset object library to obtain the first interaction object information.

[0053] For example, suppose the first text data is "I want to talk to my dad". We perform word segmentation, part-of-speech tagging, syntactic analysis, and semantic understanding on this first text data, and extract the intent keywords as "find" and "talk", and the object keyword as "dad". Based on the preset intent library, the intent keywords "find" and "speak" are matched to determine the first voice interaction intent as calling the specified user; Based on the keyword "father" in the preset object library, it is matched with the user identity information in the preset object library to determine the first interactive object information as the information corresponding to the identity of "father", such as the user identifier of "father".

[0054] Furthermore, based on the first user's identity information, the first user's first permission information can be determined. Specifically, a pre-stored mapping relationship between user identity information and permission information can be used to determine the first permission information corresponding to the first user's identity information. Then, based on the first permission information and the first voice interaction intent, the first user's permissions can be verified to obtain a first verification result. Specifically, the intent operation permission corresponding to the first voice interaction intent can be determined first, and the intent operation permission can be matched and compared with the first permission information. If the intended operation is within the permitted scope of the first permission information, the verification is successful. If the intended operation permission is not within the allowed range of the first permission information, the verification will fail.

[0055] For example, if the first user is a child and the intended operation permission is "music on demand," which is within the allowed range, the verification will succeed; if the intended operation permission is "system settings," which is outside the allowed range, the verification will fail.

[0056] If the first verification result is successful, the information of the first interactive object can be directly identified as the identity information of the second user.

[0057] If the first verification result includes verification failure, a first prompt voice message is generated to indicate insufficient permissions. For example, the first prompt voice message could be "Sorry, you do not have permission to perform this operation." Then, the first prompt voice message can be sent to the speaker in the first room, and the speaker can be controlled to play the first prompt voice message to remind the first user that the operation is invalid.

[0058] Thus, by converting the first speech into text and performing semantic recognition, the interactive intent and object can be accurately analyzed, improving the intelligence of the interaction and the accuracy of recognition. In addition, by determining and verifying the corresponding permission information based on the user's identity, fine-grained control of operation permissions can be achieved, improving system security.

[0059] Optionally, the main control module is located in the main room among the a rooms; the voice information includes a first voice and a-1 voices; please refer to [link / reference]. Figure 5 , Figure 5 This is a flowchart of a method for determining room information provided in an embodiment of this application; it can be seen that, in the step of obtaining the voice information of users in the a rooms, recognizing the voice information based on the preset voiceprint recognition technology, and obtaining the room information of the user corresponding to the second user identity information, the main control module is specifically used to execute Figure 5 The steps shown are as follows: S31. Collect the first voice in the main room, and control the a-1 sub-control modules to collect the a-1 voices in the a rooms other than the main room; S32. Based on the preset voiceprint recognition technology, the first speech and the a-1 speech are recognized to obtain b voiceprints; there are no duplicate voiceprints among the b voiceprints; b is a positive integer. S33. Determine the b user identity information corresponding to the b voiceprints; each voiceprint corresponds to one user identity information. S34. Based on the b user identity information, the a rooms, and the a-1 voice messages, determine the room information of the user corresponding to the second user identity information.

[0060] In this embodiment of the application, the voice unit integrated inside the main control module can collect the first voice in the main room. Furthermore, the main control module can also send collection instructions to a-1 sub-control modules to control the a-1 sub-control modules to collect a-1 voices from a-1 rooms other than the main room. Then, the a-1 sub-control modules can send the collected a-1 voices to the main control module.

[0061] The main control module uses a preset voiceprint recognition technology to recognize the first speech and a-1 speech samples, resulting in b voiceprints. Specifically, voiceprints can be extracted from the first speech and a-1 speech samples to obtain a voiceprints. Then, each of the a voiceprints is paired and matched for similarity to determine if there are duplicate voiceprints. If the similarity between any two voiceprints is greater than or equal to a preset similarity threshold, they are identified as duplicate voiceprints and are merged and removed. If the similarity is less than the preset similarity threshold, they are identified as independent voiceprints from different users. Finally, b voiceprints without duplicate voiceprints are obtained.

[0062] Furthermore, we can determine the b user identity information corresponding to the b voiceprints. Specifically, we can match each of the b voiceprints with voiceprints in a preset voiceprint database to obtain b matching results. Based on these b matching results, we can determine the b user identity information. For example, for each matching result: If the matching result is successful, the corresponding user identity information is determined based on the successfully matched voiceprint and the above-preset mapping relationship between voiceprint and user identity information. If the matching result is a failure, the user's identity information is marked as an unregistered stranger.

[0063] In this way, we can obtain b user identity information; Finally, based on b user identity information, a rooms, and a-1 voice messages, the room information of the user corresponding to the second user identity information can be determined.

[0064] In this way, by the master control module and multiple slave control modules cooperating to collect voices in multiple rooms, it is possible to comprehensively and synchronously obtain voice data in all areas of the whole house, avoiding omission of user information; by performing voiceprint recognition on multiple channels of voices and removing duplicates, valid voiceprints without duplicates can be obtained, excluding interference caused by multiple collections of the same user, and improving the recognition accuracy; by matching the user identity according to the voiceprint and establishing the corresponding relationship among voice, room and user, it is possible to quickly and accurately locate the room where the target user is located, realizing automatic perception of the user's position, and effectively improving the intelligence level and use convenience of voice interaction.

[0065] Optionally, in terms of determining the room information where the user corresponding to the second user identity information is located according to the b user identity information, the a rooms and the a - 1 voices, the master control module is specifically configured to: S41. Match the second user identity information with each user identity information in the b user identity information to obtain a second matching result; the second matching result includes any one of the following: matching successfully, matching failed; S42. When the second matching result includes matching successfully, determine the room information where the user corresponding to the second user identity information is located according to the second matching result, the a rooms and the a - 1 voices; S43. When the second matching result includes matching failed, determine the room information where the user corresponding to the second user identity information is located as empty.

[0066] In the embodiment of the present application, the second user identity information can be matched with each user identity information in the b user identity information based on a preset matching algorithm to obtain a second matching result, where the preset matching algorithm can include at least one of the following: character matching method, content comparison method, identification consistency judgment method, etc., which are not limited herein.

[0067] In some embodiments, the preset matching algorithm can be the character matching method. Assume that the second user identity information is "dad", and the b user identity information are respectively "dad", "mom", "grandpa"; using the character matching method, the string of the second user identity information is completely and identically matched with the strings of the b user identity information one by one: Perform character matching on "dad" and "dad", and the two strings are exactly the same, so the matching is successful; Perform character matching on "dad" and "mom", and the two strings are different, so the matching fails; Perform character matching on "dad" and "grandpa", and the two strings are different, so the matching fails.

[0068] In this way, the second matching result can be obtained.

[0069] When the second matching result includes a successful match, the room information of the user corresponding to the second user identity information is determined based on the second matching result, a rooms, and a-1 voices. Specifically, the second matching result may also include the target user identity information that was successfully matched among b user identity information. Based on the target voiceprint corresponding to the target user identity information, the target voice corresponding to the target voiceprint in a-1 voices is determined, the target room corresponding to the target voice in a rooms is determined, and the room information of the target room is obtained, that is, the room information of the user corresponding to the second user identity information.

[0070] The room information may include at least one of the following: room identifier, room name, room type (e.g., living room, master bedroom, secondary bedroom), room location, room area, etc., without limitation.

[0071] It should be explained that the target user identity information may include one or more identity information items, and correspondingly, the target voice may include one or more voice items.

[0072] In some embodiments, assuming the second user identity information is "father", if only one target user identity information that matches "father" is found among b user identity information, then the target user identity information is one identity information; based on the target voiceprint corresponding to the target user identity information, a voice corresponding to the target voiceprint is determined among a-1 voices, and the room to which the voice belongs is determined as the room where "father" is located, so as to obtain the corresponding room information.

[0073] In some embodiments, assuming the second user identity information includes "father" and "mother", the target user identity information that matches "father" and "mother" is searched among b user identity information. The target user identity information contains two identity information (one is the identity information of "father" and the other is the identity information of "mother"). Based on the two voiceprints corresponding to each of the two identity information, two voices that correspond one-to-one with the two voiceprints are determined from a-1 voices. Based on the collection room corresponding to each voice, the target room where "father" and "mother" are located are determined respectively to obtain the corresponding room information.

[0074] If the second matching result includes a failed match, the room information of the user corresponding to the second user identity information will be set to empty.

[0075] In this way, by matching the second user's identity information with the identity information of b users one by one, it is possible to accurately determine whether the user exists within the current collection range, avoiding invalid location operations; when the match is successful, the room where the user is located is quickly determined by combining room and voice information, ensuring accurate execution of interactive commands; when the match fails, the room information is cleared, realizing standardized handling of abnormal scenarios and avoiding system errors or misjudgments.

[0076] Optionally, the second user identity information includes: the identity information of the second user; the room information includes information about the second room; the second room is the room where the second user is located; in terms of realizing the first voice interaction intent based on the room information, the first control module, and the a speakers, the main control module is specifically used for: S51. Determine the first speaker corresponding to the first room and the second speaker corresponding to the second room among the a speakers; the first speaker corresponds to the first control module and the second speaker corresponds to the second control module; S52. Based on the first speaker and the second speaker, establish a first voice channel between the first control module and the second control module; the first voice channel is used to realize the first voice interaction intent.

[0077] In this embodiment of the application, the first speaker corresponding to the first room and the second speaker corresponding to the second room are determined among a speakers. Specifically, a preset mapping relationship between rooms and speakers can be stored in advance, and the first speaker corresponding to the first room and the second speaker corresponding to the second room are determined based on the mapping relationship. Then, based on the first and second speakers, a first voice channel is established between the first control module and the second control module. Specifically, the main control module can send a channel establishment command to the first and second control modules. Then, the first control module initiates a first voice channel establishment request to the second control module according to the device identifiers and communication addresses of the first and second speakers. After receiving the first voice channel establishment request, the second control module completes handshake verification and link negotiation with the first control module. After successful verification, the first voice channel between the first and second control modules is established. The first voice channel is a bidirectional voice data transmission link used to transmit voice data and voice control commands to realize the first voice interaction intent.

[0078] In this way, by establishing a one-to-one correspondence between rooms, speakers, and control modules, the devices that need to be used for voice interaction can be accurately located, ensuring directional transmission of voice data. By establishing a dedicated primary voice channel between corresponding control modules, point-to-point dedicated voice interaction between rooms can be achieved. The voice transmission is stable and has strong anti-interference capabilities, effectively improving the privacy and reliability of voice interaction.

[0079] Optionally, the second user identity information includes: c identity information corresponding to c users, each identity information corresponding to one user, where c is an integer greater than 1; the room information includes information about d rooms where the c users are located; d is a positive integer less than or equal to a; in terms of realizing the first voice interaction intent based on the room information, the first control module, and the a speakers, the main control module is specifically used for: S61. Determine the first speaker among the a speakers that corresponds to the first room, and the d speakers that correspond to the d rooms; the first speaker corresponds to the first control module; the d speakers correspond to the d control modules; S62. Based on the first speaker and the d speakers, establish voice channels between the first control module and each of the d control modules to obtain d voice channels; the d voice channels are used to realize the first voice interaction intent; or, based on the first speaker and the d speakers, establish voice groups between the first control module and the d control modules; the voice groups are used to realize the first voice interaction intent.

[0080] In this embodiment of the application, the first speaker corresponding to the first room among the a speakers and the d speakers corresponding to the d rooms can be determined firstly. Specifically, the first speaker corresponding to the first room and the d speakers corresponding to the d rooms can be determined according to the above-mentioned preset mapping relationship between rooms and speakers.

[0081] Next, based on the first speaker and d speakers, voice channels can be established between the first control module and each of the d control modules, resulting in d voice channels. Specifically, the main control module can send channel establishment commands to the first control module and the d control modules respectively. Then, the first control module initiates channel establishment requests to the d control modules respectively based on the device information of the first speaker and the d speakers. After receiving the request, the d control modules complete handshake verification and link negotiation with the first control module. After successful verification, a dedicated voice channel is established between the first control module and each of the d control modules, resulting in d voice channels. Based on these d voice channels, real-time communication between the first room and the d rooms can be achieved. The system enables real-time voice communication, voice forwarding, or collaborative playback to fulfill the first voice interaction intent. Alternatively, based on the first speaker and d speakers, a voice group is established between the first control module and d control modules. Specifically, the main control module can create a voice group and assign a group identifier based on the device information of the first speaker and d speakers. Then, the first control module and d control modules can be added to the voice group to complete the group configuration. A group communication link is established between the control modules within the group, supporting group broadcasting, intra-group forwarding, and synchronous playback of voice data. Based on this voice group, group intercom, whole-room voice broadcasting, and synchronous voice interaction between the first room and d rooms can be realized to fulfill the first voice interaction intent.

[0082] In some embodiments, after establishing a voice channel (or voice group), audio chat content can be played through the speakers in the room corresponding to the voice channel (or voice group); at the same time, the volume of the original audio currently being played by the speakers in the corresponding room is reduced (or muted) to ensure that the audio chat content can be played clearly.

[0083] Thus, by establishing point-to-point voice channels, stable one-to-one voice interaction can be achieved, avoiding interference and improving the privacy and reliability of the interaction; or by establishing voice groups, broadcast-style voice interaction can be achieved, enabling one-to-many and many-to-many communication, meeting user needs such as whole-house intercom and simultaneous broadcasting. These two methods can be flexibly switched, compatible with diverse voice interaction scenarios, effectively improving the flexibility and applicability of voice interaction.

[0084] Please see Figure 6 , Figure 6 This is a schematic diagram of a voice channel provided in an embodiment of this application; it can be seen that, Figure 6 It includes a first control module, a second control module, and a third control module, wherein: The first control module, corresponding to the speaker in the first room, is used to receive voice commands from the first user; The second control module corresponds to the speaker in the second room, which is the room where the second user is located; The third control module corresponds to the speaker in the third room, which is the room where the third user is located.

[0085] Based on the voice interaction needs of the first and second users, a first voice channel is established between the first and second control modules. This channel is a two-way voice data transmission link used to realize real-time voice intercom, audio forwarding, or collaborative playback between the first and second rooms. Similarly, based on the voice interaction needs of the first and third users, a second voice channel is established between the first and third control modules. This channel is a two-way voice data transmission link independent of the first voice channel used to realize real-time voice interaction between the first and third rooms.

[0086] Please see Figure 7 , Figure 7 This is a schematic diagram of a voice group provided in an embodiment of this application; it can be seen that, based on the interactive needs of multi-room intercom, the main control module can add the first control module, the second control module, and the third control module to the same voice group to complete the group configuration; The voice group is an intra-group broadcast communication unit. The first control module, the second control module, and the third control module are all connected to the voice group, forming a one-to-many and many-to-many voice interaction link. Voice data collected by any control module within the group can be synchronously broadcast to other control modules within the group via the voice group, and played by the speakers in the corresponding rooms, enabling synchronous voice interaction between multiple rooms.

[0087] Optionally, after determining that the room information of the user corresponding to the second user identity information is empty, the main control module is further specifically used for: S71. Generate a second prompt voice; S72. Control the speaker in the first room to play the second prompt voice; S73. Obtain the first user's first feedback instruction in response to the second prompt voice; the first feedback instruction includes any one of the following: stop interaction, continue interaction; S74. When the first feedback instruction includes stopping interaction, stop implementing the first voice interaction intent; S75. When the first feedback instruction includes continuing interaction, determine the user terminal of the user corresponding to the second user identity information; call the user terminal using a preset calling method to realize the first voice interaction intention.

[0088] In this embodiment of the application, the preset calling method can be preset in advance or defaulted. The preset calling method can include any of the following: telephone calling method, network calling method, application calling method, etc., which are not limited here.

[0089] In a specific embodiment, after determining that the room information of the user corresponding to the second user identity information is empty, it is determined that no voice data of the user corresponding to the second user identity information has been collected. Based on this, the main control module generates a second prompt voice. The second prompt voice is used to indicate that no interactive user has been found, so as to provide feedback to the user who initiated the voice interaction request. For example, the second prompt voice can be "Dear xx user, no xx interactive user has been found".

[0090] Then, the main control module can send the second prompt voice to the speaker in the first room and control the speaker in the first room to play the second prompt voice; next, it can obtain the first user's first feedback command in response to the second prompt voice. Specifically, the first control module can collect the first user's voice information in real time, perform voice recognition and semantic analysis on the collected voice information, and obtain the first feedback command.

[0091] When the first feedback instruction includes stopping the interaction, the main control module generates an interaction termination instruction and sends the interaction termination instruction to the first control module so that the first control module stops implementing the first voice interaction intent.

[0092] When the first feedback instruction includes continuing interaction, the user terminal corresponding to the second user identity information is determined. Specifically, a preset mapping relationship between user identity information and terminals can be stored in advance, and the user terminal corresponding to the second user identity information is determined based on the mapping relationship. The user terminal can include any of the following: mobile phone, computer, tablet, smartwatch, etc., without limitation. Finally, the user terminal can be called using a preset calling method to realize the first voice interaction intention.

[0093] In some embodiments, the preset calling method can be a telephone calling method. When the first feedback instruction includes continuing interaction, the main control module can determine the user's telephone number corresponding to the second user's identity information based on the pre-stored mapping relationship between the user's identity information and the telephone number. The first user accesses the mobile communication network through the communication device authorized by the first user and initiates a call to the user's telephone number. After the call is connected, the first user can talk to the user corresponding to the second user's identity information, thereby realizing the first voice interaction intention.

[0094] In some embodiments, the preset calling method can be a network calling method. The main control module can determine the user IP address corresponding to the second user identity information based on the pre-stored mapping relationship between the user identity information and the IP address. Based on the VoIP communication protocol, an IP voice call link between the first control module and the user terminal is established through a cloud server relay (or a point-to-point direct connection). Based on the IP voice call link, real-time transmission of voice data is realized to complete the first voice interaction intent.

[0095] VoIP, or Voice over IP (VoIP) communication protocol, is a general term for a class of communication protocols that convert voice signals into IP data packets and transmit them in real time through IP networks such as the Internet and local area networks.

[0096] In some embodiments, the preset calling method can be an application calling method. The main control module can determine the user application account corresponding to the second user identity information based on the pre-stored mapping relationship between the user identity information and the application account. The main control module sends a call request to the user application account through the application calling SDK on the device side (main control module). After the user terminal answers the call through the application client, an in-application voice call channel is established between the device side and the user terminal to realize the first voice interaction intention.

[0097] Among them, the application call SDK is an in-application communication function module pre-integrated into the device, used to send call requests and establish voice links based on application accounts.

[0098] In some embodiments, if a first user in a first room inputs a first voice command to request a voice conversation with a second user, and a third user in a second room inputs a second voice command to request a voice conversation with a fourth user; the main control module can parse the first and second voice commands respectively to determine that the voice interaction object corresponding to the first user is the second user, and the voice interaction object corresponding to the third user is the fourth user; detect whether the second user and the fourth user are in the target household to obtain the target detection result; If the target detection result is that both the second user and the fourth user are in the target household, the main control module can establish a first voice channel between the first control module and the control module of the speaker corresponding to the second user, and at the same time, establish a second voice channel between the second control module and the control module of the speaker corresponding to the fourth user. The first voice channel enables voice dialogue between the first user and the second user, and the second voice channel enables voice dialogue between the third user and the fourth user. The two sets of voice interactions are executed in parallel without interfering with each other.

[0099] If the target detection result is: the second user is in the target home and the fourth user is not in the target home; then the main control module can establish a first voice channel between the first control module and the control module of the speaker corresponding to the second user to realize voice dialogue between the first user and the second user; at the same time, the user terminal of the fourth user is called using a preset calling method to realize voice dialogue between the third user and the fourth user. If the target detection result is: the second user is not in the target home, and the fourth user is in the target home; then the main control module can establish a first voice channel between the third control module and the control module of the speaker corresponding to the fourth user, so as to realize the voice dialogue between the third user and the fourth user; at the same time, the user terminal of the second user is called using a preset calling method to realize the voice dialogue between the first user and the second user. If the target detection result is that neither the second user nor the fourth user is in the target household, the main control module can use a preset calling method to call the user terminal of the second user and the user terminal of the fourth user respectively, so as to realize the voice dialogue between the first user and the second user, and the voice dialogue between the third user and the fourth user.

[0100] In some embodiments, the main control module is further specifically used for: S81. Obtain the second voice of the third user and the third voice of the fourth user at the current moment; S82. Based on the preset voiceprint recognition technology, determine the third user identity information corresponding to the second voice and the fourth user identity information corresponding to the third voice; S83. Determine the second voice interaction intent based on the second voice and the third user identity information; S84. Determine the third voice interaction intent based on the third voice and the fourth user identity information; S85. When there is no conflict between the second voice interaction intention and the third voice interaction intention, determine the current interaction period corresponding to the current moment; perform permission verification on the second voice interaction intention based on the current interaction period and the third user identity information to obtain a second verification result; perform permission verification on the third voice interaction intention based on the current interaction period and the fourth user identity information to obtain a third verification result; based on the second verification result and the third verification result, determine whether to implement the second voice interaction intention and the third voice interaction intention.

[0101] In this embodiment, when a third user is in the third room, the third control module in the third room can collect the third user's second voice at the current moment and send it to the main control module; similarly, when a fourth user is in the third room, the fourth control module in the fourth room can collect the third voice and send it to the main control module; wherein, the third room and the fourth room are both rooms in a group of rooms.

[0102] Then, the second and third voices can be processed using preset voiceprint recognition technology to obtain the third user identity information and the fourth user identity information. Next, the interaction intent of the second voice is determined based on the second voice and the third user identity information. Specifically, the method for obtaining the interaction intent of the second voice can be the same as the method for obtaining the interaction intent of the first voice, which will not be repeated here. Similarly, the interaction intent of the third voice can also be determined based on the third voice and the fourth user identity information.

[0103] Then, the second and third voice interaction intentions can be analyzed to determine if there is a conflict between them. Specifically, the main control module can be configured with a preset intention conflict judgment rule base. Different types of voice interaction intentions are respectively bound to the audio playback resources, voice channel resources, and system permission resources they occupy. The main control module identifies the intention type of the second and third voice interaction intentions and matches them with the resource types they occupy. If the two types of voice interaction intentions have resource preemption or mutual permission exclusion, it is determined that there is a conflict between them; if the two types of voice interaction intentions occupy resources independently without preemption, it is determined that there is no conflict between them. For example: Suppose the second voice interaction intent is to call a specific user A (requiring the establishment of a point-to-point dedicated voice channel), and the third voice interaction intent is a multi-person intercom (requiring the formation of a multi-room voice group, and user A is included among these multiple people); since user A is included in the scope of both the point-to-point dedicated voice channel and the multi-room voice group, the two types of intents will simultaneously compete for the limited voice channel scheduling resources within the home, making it impossible to establish two independent and non-interfering voice interaction links for user A in parallel. Therefore, the main control module determines that the second voice interaction intent and the third voice interaction intent conflict.

[0104] For example, suppose a third user initiates a second voice interaction in the second room with the intent to request music, and a fourth user initiates a third voice interaction in the third room with the intent to call a specific user B; the two belong to different rooms, occupy independent audio resources and independent voice channels, do not preempt each other, are determined to have no conflict, and can be executed in parallel.

[0105] When there is no conflict between the second and third voice interaction intentions, the current interaction time period corresponding to the current moment is determined. Specifically, the main control module has multiple preset interaction time periods. The current moment is compared with each interaction time period one by one to match the time period to which the current moment belongs, thereby determining the current interaction time period. For example, if the multiple interaction time periods are: daytime period: 07:00~19:00, nighttime period: 19:00~07:00 the next day, and assuming the current moment is 9:30, then it falls within the daytime period, that is, the current interaction time period is the daytime period. Then, the second voice interaction intention can be verified for permissions based on the current interaction time period and the third user identity information to obtain the second verification. Specifically, based on the mapping relationship between the user identity information and permission information mentioned above, the second permission information corresponding to the third user's identity information can be determined. Combining the second permission information with the time period constraint rules corresponding to the current interaction time period, it can be verified whether the third user has the permission to execute the second voice interaction intent within the current interaction time period, and a corresponding first verification result can be generated. If both permission and time period constraint rules are satisfied, the first verification result is a successful verification; if either condition is not satisfied, the first verification result is a failed verification. Similarly, permission verification of the third voice interaction intent can also be performed based on the current interaction time period and the fourth user's identity information to obtain a third verification result.

[0106] Finally, based on the second and third verification results, it can be determined whether the second and third voice interaction intentions have been implemented, as follows: If both the second and third verification results are successful, then the second and third voice interaction intentions can be realized respectively. If the second verification result is successful and the third verification result is unsuccessful, then only the second voice interaction intent is executed and implemented, while the prompt voice indicating insufficient permissions for the fourth user is played. If the second verification result is verification failure and the third verification result is verification success, then only the third voice interaction intent is executed and implemented, and the prompt voice indicating insufficient permissions for the third user is played. If both the second and third verification results fail, the second and third voice interaction intentions will not be executed, and the corresponding permission-insufficient prompt voices will be played to the third and fourth users respectively.

[0107] In this way, by simultaneously collecting voice data from multiple users and accurately identifying users through voiceprint recognition, the system analyzes each user's voice interaction intent. It first predicts whether there are any conflicts in the interaction intents; if no conflicts exist, it performs dual permission verification based on the current interaction time and user identity, and then makes a unified decision on whether to execute the corresponding interaction intent based on the verification results. This approach not only avoids resource contention issues caused by concurrent multi-user interactions but also enables refined permission control based on different family member identities and time periods, preventing unauthorized operations. Simultaneously, it ensures compliant and orderly responses to multiple voice interaction intents in conflict-free scenarios, improving the standardization, security, and intelligence of family voice interaction.

[0108] In some embodiments, when there is a conflict between the second voice interaction intent and the third voice interaction intent, the main control module is further configured to: S91. Determine the first family role corresponding to the third user identity information and the second family role corresponding to the fourth user identity information; S92. Determine the first priority corresponding to the first family role and the second priority corresponding to the second family role; S93. When the first priority and the second priority are the same, determine the first initiation time corresponding to the second voice interaction intention, and determine the second initiation time corresponding to the third voice interaction intention; determine a first implementation order according to the first initiation time and the second initiation time; implement the second voice interaction intention and the third voice interaction intention according to the first implementation order. S94. When the first priority and the second priority are not the same, determine the second implementation order according to the first priority and the second priority; implement the second voice interaction intention and the third voice interaction intention according to the second implementation order.

[0109] In this embodiment, each user identity information may include a family role. The first family role can be extracted from the third user identity information, and similarly, the second family role can be extracted from the fourth user identity information. Then, the first priority corresponding to the first family role and the second priority corresponding to the second family role can be determined. Specifically, a preset mapping relationship between family roles and priorities can be stored in advance, and the first priority corresponding to the first family role and the second priority corresponding to the second family role can be determined based on the mapping relationship. For example, the priorities can be set from high to low according to the identity level of the family roles: the priority of the elderly role (grandfather, grandmother) is the highest > the priority of the parent adult role (father, mother) is the second highest > the priority of the child role (son, daughter) is the lowest.

[0110] When the first and second priorities are the same, the first initiation time corresponding to the second voice interaction intent is determined. The earliest acquisition time of the second voice can be obtained and used as the first initiation time. Then, the second initiation time corresponding to the third voice interaction intent is determined. Similarly, the earliest acquisition time of the third voice can be obtained and used as the second initiation time. Further, the first implementation order can be determined based on the first and second initiation times. Specifically, the time sequence of the first and second initiation times is compared, and the execution order of the two voice interaction intents, i.e., the first implementation order, is determined according to the rule of first-come-first-served. Then, the second and third voice interaction intents can be implemented sequentially according to the first implementation order.

[0111] When the first priority and the second priority are different, the second implementation order is determined according to the first priority and the second priority. Specifically, they can be sorted by priority, with the voice interaction intents corresponding to the family roles with higher priority executed first and the ones with lower priority executed later, thus determining the second implementation order. Finally, the second voice interaction intent and the third voice interaction intent are implemented in sequence according to the second implementation order.

[0112] In this way, by first matching family roles based on user identity and setting corresponding priorities, a two-layer arbitration strategy is adopted, which prioritizes high-priority users and queues users with the same priority according to the time of initiation. When priorities are different, the execution order is determined according to the role level, and when priorities are the same, the requests are responded to in order of initiation time. This can efficiently resolve resource conflicts of concurrent voice intentions from multiple users, which not only conforms to the usage habits of family role hierarchy, but also ensures fair and orderly scheduling of requests from users of the same level, thereby improving the regularity and intelligent scheduling effect of multi-room voice interaction.

[0113] In summary, the home music system for multi-user interaction described in this application avoids command conflicts caused by simultaneous operation by multiple users by configuring independent speakers and control modules for each room, with interconnected control modules and a distinction between master and sub-controllers, thus adapting to multi-room, multi-user scenarios. Furthermore, by locating the room where the interactive user is located through the master control module and determining the interaction intent based on the user's identity and voice, targeted voice interaction for multiple users is achieved, breaking the one-way interaction limitations of traditional systems and effectively improving the adaptability of the home music system to multi-user interactive scenarios.

[0114] This application embodiment also provides a control method for a home music system, which is applied to the home music system supporting multi-user interaction described in the above embodiment. The method includes the following steps: Audio is played in the corresponding rooms of the a rooms through the a speakers; The first user's voice is received through the first control module; the first user is in the first room; the first room is any one of the a rooms; the control module is the control module in the first room; The main control module determines the first user identity information corresponding to the first voice based on preset voiceprint recognition technology; determines the first voice interaction intention and the second user identity information based on the first user identity information and the first voice; obtains the voice information of users in the a rooms, identifies the voice information based on the preset voiceprint recognition technology, and obtains the room information of the user corresponding to the second user identity information; and realizes the first voice interaction intention based on the room information, the first control module, and the a speakers.

[0115] In specific implementations, the control method of the home music system described in the embodiments of this application can also execute other implementation methods described in the home music system supporting multi-person interaction provided in the embodiments of the present invention, which will not be repeated here.

[0116] This application also provides a computer-readable storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of the control method for a home music system as described in the above method embodiments, wherein the computer includes a home music system that supports multi-user interaction.

[0117] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of the control method for a home music system as described in the above method embodiments. The computer program product can be a software installation package, and the computer includes a home music system supporting multi-user interaction.

[0118] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a home music device supporting multi-user interaction provided in an embodiment of this application. As can be seen, the device includes a home music system supporting multi-user interaction as described in the above embodiment.

[0119] This application also provides a home music device that supports multi-user interaction, wherein the device includes a home music system that supports multi-user interaction as described in the above embodiments, or includes a home music device that supports multi-user interaction as described in the above embodiments.

[0120] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0121] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0122] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.

[0123] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

[0124] The steps of the methods or algorithms described in the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in RAM, flash memory, ROM, EPROM, electrically erasable programmable read-only memory (EEPROM), registers, hard disk, portable hard disk, read-only optical disk (CD-ROM), or any other form of storage medium well known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Furthermore, the ASIC can reside in a terminal device or management device. Alternatively, the processor and storage medium can exist as discrete components in the terminal device or management device.

[0125] Those skilled in the art will recognize that, in one or more of the examples above, the functions described in the embodiments of this application can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated.

[0126] The aforementioned computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media.

[0127] The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).

[0128] The modules / units included in the various devices and products described in the above embodiments can be software modules / units, hardware modules / units, or a combination of both. For example, for devices and products applied to or integrated into a chip, all modules / units can be implemented using hardware methods such as circuits, or at least some modules / units can be implemented using software programs that run on a processor integrated within the chip, while the remaining (if any) modules / units can be implemented using hardware methods such as circuits. For devices and products applied to or integrated into a chip module, all modules / units can be implemented using hardware methods such as circuits. Different modules / units can be located in the same component (e.g., chip, circuit module, etc.) or different components of the chip module, or at least some modules / units can be implemented using hardware methods such as circuits. The implementation is achieved through a software program that runs on the processor integrated within the chip module. The remaining modules / units (if any) can be implemented using hardware methods such as circuits. For various devices and products applied to or integrated into terminal equipment, each of their modules / units can be implemented using hardware methods such as circuits. Different modules / units can be located in the same component (e.g., chip, circuit module, etc.) or different components within the terminal equipment. Alternatively, at least some modules / units can be implemented through a software program that runs on the processor integrated within the terminal equipment, while the remaining modules / units (if any) can be implemented using hardware methods such as circuits.

[0129] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the embodiments of this application. It should be understood that the above descriptions are merely specific embodiments of the embodiments of this application and are not intended to limit the protection scope of the embodiments of this application. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solutions of the embodiments of this application should be included within the protection scope of the embodiments of this application.

Claims

1. A home music system supporting multi-user interaction, characterized in that, The system is installed in a target home, which includes *a* rooms. The system includes *a* speakers and *a* control modules, with each control module corresponding to one speaker. Each room contains one speaker and one control module. The *a* control modules are interconnected. Each *a* control module includes one main control module and *a-1* sub-control modules; *a* is an integer greater than 1. The a speakers are used to play audio in the corresponding rooms of the a rooms; A first control module is used to receive a first voice message from a first user; the first user is in a first room; the first room is any one of the a rooms; the control module is the control module in the first room; The main control module is used to determine the first user identity information corresponding to the first voice based on preset voiceprint recognition technology; determine the first voice interaction intention and the second user identity information based on the first user identity information and the first voice; acquire the voice information of users in the a rooms, recognize the voice information based on the preset voiceprint recognition technology, and obtain the room information of the user corresponding to the second user identity information; and realize the first voice interaction intention based on the room information, the first control module, and the a speakers. Specifically, in determining the first voice interaction intent and the second user identity information based on the first user identity information and the first voice, the main control module is used for: Determine the first text data corresponding to the first speech; Semantic recognition is performed on the first text data to obtain the first voice interaction intent and the first interaction object information; Based on the first user's identity information, determine the first user's first permission information; Based on the first permission information and the first voice interaction intent, the first user's permission is verified to obtain a first verification result; the first verification result includes any of the following: verification successful, verification failed; When the first verification result includes the verification success, the second user identity information is determined based on the first interaction object information; When the first verification result includes the verification failure, a first prompt voice is generated; the speaker in the first room is controlled to play the first prompt voice.

2. The system as described in claim 1, characterized in that, In determining the first user identity information corresponding to the first voice based on preset voiceprint recognition technology, the main control module is specifically used for: The first speech is preprocessed to obtain the first speech data; The first voiceprint is obtained by extracting the voiceprint from the first speech data based on the preset voiceprint recognition technology. The first voiceprint is matched with the voiceprints in the preset voiceprint library to obtain the first matching result; When the first matching result includes a successful match, the first user's identity information is determined based on the first matching result and the preset voiceprint database.

3. The system as described in claim 1 or 2, characterized in that, The main control module is located in the main room of the a rooms; the voice information includes a first voice and a-1 voices; In acquiring the voice information of users in the a rooms, recognizing the voice information based on the preset voiceprint recognition technology, and obtaining the room information of the user corresponding to the second user identity information, the main control module is specifically used for: Collect the first voice in the main room, and control the a-1 sub-control modules to collect the a-1 voices in the a rooms other than the main room; Based on the preset voiceprint recognition technology, the first speech and the a-1 speech are recognized to obtain b voiceprints; there are no duplicate voiceprints among the b voiceprints; b is a positive integer; Determine the b user identity information corresponding to the b voiceprints; each voiceprint corresponds to one user identity information. Based on the b user identity information, the a rooms, and the a-1 voice messages, determine the room information of the user corresponding to the second user identity information.

4. The system as described in claim 3, characterized in that, In determining the room information of the user corresponding to the second user identity information based on the b user identity information, the a rooms, and the a-1 voice messages, the main control module is specifically used for: The second user identity information is matched with each of the b user identity information to obtain a second matching result; the second matching result includes any of the following: matching successful, matching failed; When the second matching result includes a successful match, the room information of the user corresponding to the second user identity information is determined based on the second matching result, the a rooms and the a-1 voices; If the second matching result includes a matching failure, the room information of the user corresponding to the second user identity information will be set to empty.

5. The system as described in claim 3, characterized in that, The second user identity information includes: the identity information of the second user; the room information includes information about the second room; the second room is the room where the second user is located; In terms of realizing the first voice interaction intent based on the room information, the first control module, and the a speakers, the main control module is specifically used for: Identify the first speaker corresponding to the first room and the second speaker corresponding to the second room from among the a speakers; the first speaker corresponds to the first control module, and the second speaker corresponds to the second control module; Based on the first speaker and the second speaker, a first voice channel is established between the first control module and the second control module; the first voice channel is used to realize the first voice interaction intent.

6. The system as described in claim 3, characterized in that, The second user identity information includes: c identity information corresponding to c users, each identity information corresponding to one user, where c is an integer greater than 1; the room information includes information about d rooms where the c users are located; where d is a positive integer less than or equal to a. In terms of realizing the first voice interaction intent based on the room information, the first control module, and the a speakers, the main control module is specifically used for: Identify the first speaker among the a speakers that corresponds to the first room, and the d speakers that correspond to the d rooms; the first speaker corresponds to the first control module; the d speakers correspond to the d control modules; Based on the first speaker and the d speakers, voice channels are established between the first control module and each of the d control modules, resulting in d voice channels; these d voice channels are used to realize the first voice interaction intent; or... Based on the first speaker and the d speakers, a voice group is established between the first control module and the d control modules; the voice group is used to realize the first voice interaction intent.

7. The system as described in claim 1 or 2, characterized in that, The main control module is also specifically used for: Get the second voice of the third user and the third voice of the fourth user at the current moment; Based on the preset voiceprint recognition technology, the third user identity information corresponding to the second voice and the fourth user identity information corresponding to the third voice are determined. The second voice interaction intent is determined based on the second voice and the third user identity information; The third voice interaction intent is determined based on the third voice and the fourth user identity information; When there is no conflict between the second voice interaction intention and the third voice interaction intention, the current interaction period corresponding to the current moment is determined; Based on the current interaction time period and the third user's identity information, the second voice interaction intent is validated for permissions to obtain a second validation result. Based on the current interaction time period and the fourth user identity information, the third voice interaction intent is validated for permission, and a third validation result is obtained. Based on the second verification result and the third verification result, determine whether the second voice interaction intent and the third voice interaction intent are realized.

8. A home music device supporting multi-user interaction, characterized in that, Includes the system as described in any one of claims 1-7.

9. A home music device supporting multi-user interaction, characterized in that, Includes the apparatus as described in claim 8, or includes the system as described in any one of claims 1-7.