Broadcasting control method and system of intelligent loudspeaker box

Through multi-microphone analysis, voice recognition and broadcast content binding technology, the delay and confusion problems of traditional speakers when handling multiple user commands in complex environments are solved, achieving more efficient and accurate response.

CN120108389APending Publication Date: 2025-06-06SHENZHEN MTN ELECTRONICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411436012.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Traditional speakers are difficult to deal with the situation where multiple users issue commands at the same time in complex environments, resulting in delays and confusion in responses, and cannot dynamically adapt to changes in the user environment.

Method used

Audio signals are collected through multiple microphones, sound intensity and time difference are analyzed, sound source direction is judged; the characteristics of sound signals are extracted, and sound signals from different sources are separated; each separate audio signal is voiced, broadcast content is generated, and it is bound to the response area, and the report is broadcasted through the smart speaker in the bound area.

Benefits of technology

It realizes accurate identification of the direction and location of multiple sound sources in a complex environment, dynamically divides spatial areas, optimizes response priority, and improves the operation efficiency and response correlation of speakers when multiple people are instructed simultaneously.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108389A_ABST
    Figure CN120108389A_ABST
Patent Text Reader

Abstract

The invention is suitable for the field of intelligent sound box control, and provides a broadcast control method and system for an intelligent sound box, and the system comprises a sound collection and direction judgment module, a sound feature extraction and separation module, a voice recognition and content analysis module, a region division and priority judgment module, and a broadcast content generation and region binding module. According to the scheme, voice input in multiple directions can be received at the same time, and the direction of each sound source can be accurately recognized. The capability remarkably improves the operation efficiency of the sound box in a complex environment, and can meet the requirement that multiple persons send out instructions at the same time. Through analysis of the sound intensity and the time difference, the sound box can accurately identify the position of each sound source, and the space is divided into different areas. According to the dynamic space management mode, the sound box can be intelligently adapted according to the actual environment, and the accuracy of sound source identification and the correlation of response are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of smart speaker control, and in particular relates to a smart speaker broadcast control method and system. Background Art

[0002] As a device that integrates functions such as voice recognition, audio playback, and smart home control, smart speakers have received widespread attention and application in recent years. In the in-car environment, the background of smart speaker control technology mainly stems from the development trend of automobile intelligence and interconnection. With the continuous advancement of automotive electronic technology, more and more vehicles are equipped with advanced driver assistance systems (ADAS) and in-vehicle entertainment information systems. These systems require efficient and intuitive user interaction methods, while traditional physical button controls often cannot meet safety and convenience requirements in complex driving environments. Therefore, voice control technology came into being, allowing drivers to achieve functions such as navigation, music playback, and phone calls through simple voice commands. The application of this smart speaker not only improves driving safety, but also facilitates users' multi-tasking capabilities during driving, and promotes the progress of smart travel.

[0003] Traditional speakers can usually only process single sound source input and cannot effectively handle scenarios where multiple users issue commands at the same time. It is also difficult to perform segmented command recognition in complex environments. When faced with multiple commands, it is easy to cause response delays and confusion. It also relies on static spatial division and cannot dynamically adapt to changes in the user environment. Summary of the invention

[0004] The purpose of the present invention is to provide a broadcast control method for an intelligent speaker, aiming to solve the technical problems existing in the prior art identified in the background technology.

[0005] The present invention is implemented as follows: a broadcast control method for an intelligent speaker, the method comprising:

[0006] Collect audio signals through each microphone, analyze the sound intensity received by each microphone, calculate the time difference between different microphones receiving the sound, and determine the direction of the sound source;

[0007] Extract the characteristics of each sound signal and separate the sound signals from different sources;

[0008] Performing speech recognition on each separated audio signal, extracting and analyzing the content of the recognized speech input;

[0009] Based on the source location of the sound, the space is divided into different areas, and the corresponding sound source of each area is recorded. Based on the identified content and sound source location, the response priority and corresponding area of ​​each request are determined;

[0010] Corresponding broadcast content is generated according to the requests of different identified areas, and the broadcast content is bound to the response area, and the broadcast content is broadcast through the smart speaker in the bound area.

[0011] As a further solution of the present invention, the analysis of the sound intensity received by each microphone, the calculation of the time difference between different microphones receiving the sound, and the determination of the source direction of the sound specifically include:

[0012] Collect surrounding audio signals through all microphones at the same time, and analyze the intensity of the audio signal received by each microphone;

[0013] Measure the time difference between the sound reaching different microphones, and combine the sound intensity and the results of time difference analysis to determine the specific source direction of the sound.

[0014] As a further solution of the present invention, the extracting the features of each sound signal and separating the sound signals from different sources specifically includes:

[0015] Extract features from the collected audio signals, separate the audio signals from different sound sources, and identify all the sound source signals;

[0016] Perform a quality assessment on the separated sound source signal to determine whether the serial number clarity meets the recognition requirements;

[0017] For the sound source signals that meet the recognition requirements, they are stored in the recognized storage memory and the signals are directly transmitted. For the sound source signals that do not meet the recognition requirements, they are stored in the unrecognized storage memory and the unrecognizable signals are transmitted.

[0018] As a further solution of the present invention, performing speech recognition on each separated audio signal, extracting and analyzing the content of the speech input, specifically includes:

[0019] Identify each sound source signal transmitted in the identified storage, convert the audio signal into text, and identify the speech content therein;

[0020] Natural language processing is used to parse the recognized text results, including understanding the intent and extracting information.

[0021] As a further solution of the present invention, the space is divided into different areas, and the corresponding sound source of each area is recorded, and the response priority and the corresponding area of ​​each request are determined according to the identified content and the location of the sound source, which specifically includes:

[0022] By analyzing the sound intensity and time difference, the relative position of each sound source signal in space is identified;

[0023] According to the identified sound source location and the actual environment, the entire environment space is divided into several areas;

[0024] Record the sound source information corresponding to each divided area, associate the sound source position with the area, and establish a mapping relationship;

[0025] The priority of each request is determined based on the identified content and the location of the sound source, and the distance to the sound source.

[0026] Based on the priority evaluation, the response area corresponding to each sound source request is determined.

[0027] As a further solution of the present invention, the generating of corresponding broadcast content, binding the broadcast content with the response area, and broadcasting the broadcast content through the smart speaker in the bound area specifically includes:

[0028] Generate corresponding broadcast information and control instructions according to the recognized voice input content;

[0029] Binding the generated announcement content and control instructions to the area mapped by the sound source position;

[0030] The control instructions are passed to the smart speakers in the corresponding bound area, and the generated broadcast control instructions are executed by the smart speakers to output the broadcast content to the bound area.

[0031] Another object of the present invention is to provide a broadcast control system for an intelligent speaker, the system comprising:

[0032] The sound collection and direction determination module is used to collect audio signals through various microphones, analyze the sound intensity received by each microphone, calculate the time difference between different microphones receiving the sound, and determine the direction of the sound source;

[0033] Sound feature extraction and separation module, used to extract the features of each sound signal and separate sound signals from different sources;

[0034] A speech recognition and content analysis module, used to perform speech recognition on each separated audio signal, extract and analyze the content of the recognized speech input;

[0035] The area division and priority judgment module is used to divide the space into different areas based on the source location of the sound, and record the corresponding sound source of each area. According to the identified content and sound source location, the response priority of each request and the corresponding area are determined;

[0036] The broadcast content generation and area binding module is used to generate corresponding broadcast content according to the requests of different identified areas, bind the broadcast content to the response area, and broadcast the broadcast content through the smart speaker in the bound area.

[0037] As a further solution of the present invention, the area division and priority determination module includes:

[0038] A sound source localization analysis unit, used to identify the relative position of each sound source signal in space by analyzing the sound intensity and time difference;

[0039] The environment area division unit is used to divide the entire environment space into several areas according to the identified sound source position and the actual environment;

[0040] A sound source mapping recording unit is used to record the sound source information corresponding to each divided area, by associating the sound source position with the area and establishing a mapping relationship;

[0041] A request priority evaluation unit, used to determine the response priority of each request based on the identified content and the sound source location, taking the distance from the sound source location as a judgment rule;

[0042] The response area decision unit is used to determine the response area corresponding to each sound source request based on the priority evaluation.

[0043] As a further solution of the present invention, the broadcast content generation and area binding module includes:

[0044] A broadcast information generating unit, used to generate corresponding broadcast information and control instructions according to the recognized voice input content;

[0045] A content and area binding unit, used to bind the generated broadcast content and control instructions to the area mapped by the sound source position;

[0046] The instruction transmission and execution unit is used to transmit the control instruction to the smart speaker in the corresponding bound area, execute the generated broadcast control instruction through the smart speaker, and output the broadcast content to the bound area.

[0047] The beneficial effects of the present invention are:

[0048] The solution can simultaneously receive voice input from multiple directions and accurately identify the direction of each sound source. This capability significantly improves the operating efficiency of the speaker in complex environments and can meet the needs of multiple people issuing commands at the same time. Through the analysis of sound intensity and time difference, the speaker can accurately identify the location of each sound source and divide the space into different areas. This dynamic space management method enables the speaker to intelligently adapt according to the actual environment, improving the accuracy of sound source identification and the relevance of the response. The system sets the response priority of the request based on the distance between the sound source location and the user. This design ensures that the speaker can give priority to requests that are closer to the user, making the user experience smoother and more natural and reducing waiting time. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 A flowchart of a method for controlling the broadcast of an intelligent speaker provided by an embodiment of the present invention;

[0050] Figure 2 A flowchart for analyzing the sound intensity received by each microphone, calculating the time difference between different microphones receiving the sound, and determining the source direction of the sound provided by an embodiment of the present invention;

[0051] Figure 3 A flowchart for extracting features of each sound signal and separating sound signals from different sources provided by an embodiment of the present invention;

[0052] Figure 4 A flowchart of performing speech recognition on each separated audio signal, extracting and analyzing the content of the recognized speech input provided by an embodiment of the present invention;

[0053] Figure 5 A flowchart for dividing a space into different areas, recording the corresponding sound source of each area, and determining the response priority of each request and the corresponding area according to the identified content and the location of the sound source provided in an embodiment of the present invention;

[0054] Figure 6 A flowchart of generating corresponding broadcast content provided by an embodiment of the present invention, binding the broadcast content with a response area, and broadcasting the broadcast content through a smart speaker in the bound area;

[0055] Figure 7 A structural block diagram of a broadcast control system of an intelligent speaker provided in an embodiment of the present invention;

[0056] Figure 8 A structural block diagram of a region division and priority determination module provided in an embodiment of the present invention;

[0057] Fig. 9 This is a structural block diagram of the broadcast content generation and area binding module provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0059] It is understood that the terms "first", "second", etc. used in this application may be used herein to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, without departing from the scope of this application, a first xx script may be referred to as a second xx script, and similarly, a second xx script may be referred to as a first xx script.

[0060] Figure 1 A flowchart of a method for controlling the broadcast of an intelligent speaker provided by an embodiment of the present invention, such as Figure 1 As shown, the method includes:

[0061] S100, collecting audio signals through various microphones, analyzing the sound intensity received by each microphone, calculating the time difference between different microphones receiving the sound, and determining the source direction of the sound;

[0062] like Figure 2 As shown, the analysis of the sound intensity received by each microphone, calculating the time difference between different microphones receiving the sound, and determining the source direction of the sound specifically includes:

[0063] S110, collecting surrounding audio signals through all microphones at the same time, and performing strength analysis on the audio signal received by each microphone;

[0064] S120, measuring the time difference of the sound reaching different microphones, and combining the sound intensity and the time difference analysis result to determine the specific source direction of the sound.

[0065] This step will determine the direction of the sound source. First, use the multiple microphones configured in the smart speaker to collect audio signals from the surrounding environment at the same time. Each microphone receives sound independently and converts it into an electrical signal. Next, the audio signal received by each microphone is analyzed for intensity. Specifically, the amplitude information of the audio signal is extracted to determine the intensity of the sound received by each microphone. Then, by recording the time it takes for the sound to reach each microphone, the time difference between different microphones is calculated. This can be calculated by understanding the speed of sound wave propagation and combining the order in which the sound waves arrive at the microphones. Combine the results of sound intensity and time difference analysis to determine the specific source direction of the sound. It is worth mentioning that when there are multiple sound sources at the same time, the sound source separation algorithm can be used to separate the audio signals of different sound sources to ensure that the direction and intensity of each sound source can be accurately identified.

[0066] By combining the dual analysis methods of sound intensity and time difference, the S100 can provide higher sound source positioning accuracy than traditional single microphone or simple sound source localization technology, especially in complex environments. In addition, the S100 can process inputs from multiple sound sources at the same time, which enables the smart speaker to accurately judge and respond to multiple requests from different directions in a noisy environment, which is an advantage that existing technologies do not have. Furthermore, by analyzing the audio signals of the surrounding environment in real time, the smart speaker can dynamically adapt to different sound source situations and provide more flexible response capabilities in the user's actual usage scenarios. Due to the ability to accurately identify the direction of the sound source, the smart speaker can achieve more natural human-computer interaction, and the voice commands issued by users in different locations can be responded to quickly and accurately, improving the convenience and comfort of use.

[0067] S200, extracting the features of each sound signal and separating the sound signals from different sources;

[0068] like Figure 3 As shown, the extraction of the features of each sound signal and separation of sound signals from different sources specifically include:

[0069] S210, extracting features from the collected audio signals, separating the audio signals from different sound sources, and identifying all the sound source signals;

[0070] S220, performing a quality assessment on the separated sound source signal to determine whether the serial number clarity meets the recognition requirement;

[0071] S230, for the sound source signal that meets the recognition requirement, it is stored in the recognized storage memory and the signal is directly transmitted. For the sound source signal that does not meet the recognition requirement, it is stored in the unrecognized storage memory and the unrecognizable signal is transmitted.

[0072] In this step, the smart speaker first extracts features from the collected audio signals and analyzes the audio signals from different sound sources to extract key parameters that can represent the sound characteristics. Through these features, the speaker can identify all sound source signals and effectively distinguish the characteristics of different sound sources.

[0073] Next, the separated sound source signal is evaluated for quality to determine whether its clarity meets the requirements of subsequent speech recognition to ensure the validity and recognizability of the input signal. After the quality evaluation is completed, the sound source signal that meets the recognition requirements will be stored in the recognized storage repository and directly passed to the subsequent processing link to ensure the accuracy and efficiency of speech recognition. For sound source signals that do not meet the recognition requirements, they are stored in the unrecognized storage repository and feedback on unrecognizable signals is passed so that the subsequent system can be improved or the user can be reminded.

[0074] Through efficient signal feature extraction and analysis, the ability to separate speech signals in a multi-sound source environment is significantly improved, especially in a noisy environment, which can effectively reduce the interference of background noise on speech recognition. Secondly, the quality assessment mechanism ensures that only clear and quality-qualified signals enter the subsequent processing link. This process improves the accuracy of speech recognition, reduces the misrecognition rate, and enhances the user experience. In addition, the design of the unrecognized storage library can provide feedback to the system, prompting the system to continuously optimize to adapt to changing environmental conditions and improve overall performance. At the same time, when the information in the unrecognized storage library is finally broadcast through the smart speaker, voice feedback can also be provided for the information that failed to be recognized. Finally, through feature extraction and signal separation, the smart speaker can dynamically manage and respond to different sound sources, which helps to improve the user's interactive experience, allowing the system to more intelligently adapt to actual usage scenarios and provide personalized services.

[0075] S300, performing speech recognition on each separated audio signal, extracting and analyzing the content of the recognized speech input;

[0076] like Figure 4 As shown, the voice recognition is performed on each separated audio signal, and the content of the voice input is extracted and analyzed and recognized, specifically including:

[0077] S310, identifying each sound source signal transmitted in the identified storage, converting the audio signal into text, and identifying the speech content therein;

[0078] S320, using natural language processing, parsing the recognized text results, including understanding the intent and extracting information.

[0079] S400, based on the source location of the sound, the space is divided into different areas, and the corresponding sound source of each area is recorded, and the response priority and the corresponding area of ​​each request are determined according to the identified content and the location of the sound source;

[0080] like Figure 5 As shown, the space is divided into different areas, and the corresponding sound source of each area is recorded. According to the identified content and the location of the sound source, the response priority of each request and the corresponding area are determined, specifically including:

[0081] S410, identifying the relative position of each sound source signal in space by analyzing the sound intensity and time difference;

[0082] S420, dividing the entire environment space into a plurality of areas according to the identified sound source position and the actual environment;

[0083] S430, recording the sound source information corresponding to each divided area, associating the sound source position with the area, and establishing a mapping relationship;

[0084] S440, determining the response priority of each request based on the identified content and the sound source location and the distance from the sound source location as a judgment rule;

[0085] S450: Based on the priority evaluation, a response area corresponding to each sound source request is determined.

[0086] In this step, the smart speaker can first identify the relative position of each sound source signal in space by analyzing the sound intensity and time difference. Using the configuration of a multi-microphone array, the speaker can accurately measure the time difference between the sound reaching different microphones, and determine the approximate direction and distance of the sound source based on the sound intensity received by each microphone. This process lays the foundation for subsequent regional division.

[0087] Next, based on the identified sound source location, the system divides the entire ambient space into several areas. These areas are divided based on actual environmental characteristics, such as the layout of the room, furniture placement, and other factors that affect sound propagation, to ensure that the divided areas can accurately reflect the distribution of sound sources in reality. The sound source information corresponding to each divided area will be recorded and a mapping relationship will be established, allowing the speaker to quickly identify the location of the sound source in a specific area.

[0088] In addition, based on the identified content and the location of the sound source, the system will use the distance between the sound source and the user as a judgment rule to determine the response priority of each request. Specifically, sound source requests that are closer to the user will be given a higher response priority. This strategy helps improve the user experience, allowing the speaker to quickly respond to the most relevant request when the user issues a command.

[0089] Finally, based on the priority evaluation, the system determines the corresponding response area for each sound source request. In this way, the speaker can effectively manage and deploy its response strategy when multiple requests exist at the same time, ensuring that users can get timely and accurate feedback.

[0090] The use of a multi-microphone array for sound source localization and analysis enables the speaker to achieve higher-precision sound source identification in complex environments, improving performance in noisy or multi-sound source scenes. Secondly, the dynamic division strategy of the environmental space can adapt to different usage scenarios and layouts. Compared with the traditional static area division method, S400 provides a more flexible solution to ensure that the system can adapt to environmental changes in real time.

[0091] Furthermore, through the distance-based response priority judgment rules, the smart speaker can more intelligently handle requests from different directions, ensuring that the user's most recent needs are met first, thereby improving user satisfaction and interactive experience. In addition, the mapping relationship between the sound source location and the area provides solid data support for subsequent voice recognition and feedback, optimizing the overall efficiency of the system.

[0092] S500, generating corresponding broadcast content according to the requests of different identified areas, binding the broadcast content with the response area, and broadcasting the broadcast content through the smart speaker of the bound area.

[0093] like Figure 6 As shown, the generating of corresponding broadcast content, binding the broadcast content with the response area, and broadcasting the broadcast content through the smart speaker in the bound area specifically includes:

[0094] S510, generating corresponding broadcast information and control instructions according to the recognized voice input content;

[0095] S520, binding the generated announcement content and control instructions to the area mapped by the sound source position;

[0096] S530, passing the control instruction to the smart speaker in the corresponding bound area, executing the generated broadcast control instruction through the smart speaker, and outputting the broadcast content to the bound area.

[0097] In this step, the smart speaker first generates corresponding broadcast information and control instructions based on the recognized voice input content.

[0098] Next, the system binds the generated announcement content and control instructions to the area mapped to the sound source location. This binding process ensures that each piece of information and instruction can correspond to the corresponding physical space, allowing the speaker to operate accurately in the specified area. For example, if the user issues an instruction in the back seat of the car, the system will ensure that the corresponding announcement content is only played in the smart speakers on both sides of the back seat, without interfering with the speakers in other areas, and will also identify whether it is on the left or right side of the back seat, so as to issue voice announcements more accurately.

[0099] Then, the system transmits the control command to the smart speaker in the corresponding bound area. The speaker can quickly receive the command and perform the corresponding operation. At this time, the smart speaker outputs the broadcast content to the bound area according to the generated broadcast control command to achieve voice feedback.

[0100] Natural language processing technology has achieved a higher level of voice understanding and response capabilities, which can not only identify user requests, but also generate appropriate broadcast content based on the context. This is more flexible and intelligent than the traditional fixed response mechanism.

[0101] Secondly, the broadcast content and control instructions are bound to the sound source location, ensuring the accuracy and effectiveness of information transmission and reducing confusion and interference between multiple speakers. This feature enables users to have a more personalized and customized experience when using smart speakers, improving user satisfaction.

[0102] Figure 7 A structural block diagram of a broadcast control system of an intelligent speaker provided in an embodiment of the present invention, such as Figure 7 As shown, the system comprises:

[0103] The sound collection and direction determination module 100 is used to collect audio signals through various microphones, analyze the sound intensity received by each microphone, calculate the time difference between different microphones receiving the sound, and determine the source direction of the sound;

[0104] The sound feature extraction and separation module 200 is used to extract the features of each sound signal and separate the sound signals from different sources;

[0105] The speech recognition and content analysis module 300 is used to perform speech recognition on each separated audio signal, extract and analyze the content of the recognized speech input;

[0106] The area division and priority determination module 400 is used to divide the space into different areas based on the source location of the sound, and record the corresponding sound source of each area, and determine the response priority and corresponding area of ​​each request according to the identified content and sound source location;

[0107] The broadcast content generation and area binding module 500 is used to generate corresponding broadcast content according to the requests of different identified areas, bind the broadcast content to the response area, and broadcast the broadcast content through the smart speaker in the bound area.

[0108] like Figure 8 As shown, the area division and priority determination module 400 includes:

[0109] The sound source localization analysis unit 410 is used to identify the relative position of each sound source signal in space by analyzing the sound intensity and time difference;

[0110] The environment area division unit 420 is used to divide the entire environment space into several areas according to the identified sound source position and the actual environment;

[0111] The sound source mapping recording unit 430 is used to record the sound source information corresponding to each divided area, by associating the sound source position with the area and establishing a mapping relationship;

[0112] The request priority evaluation unit 440 is used to determine the response priority of each request based on the identified content and the sound source location, taking the distance from the sound source location as a judgment rule;

[0113] The response area decision unit 450 is used to determine the response area corresponding to each sound source request based on the priority evaluation.

[0114] like Fig. 9 As shown, the broadcast content generation and area binding module 500 includes:

[0115] The announcement information generating unit 510 is used to generate corresponding announcement information and control instructions according to the recognized voice input content;

[0116] The content and area binding unit 520 is used to bind the generated broadcast content and control instructions to the area mapped by the sound source position;

[0117] The instruction transmission and execution unit 530 is used to transmit the control instruction to the smart speaker in the corresponding binding area, and execute the generated broadcast control instruction through the smart speaker to output the broadcast content to the bound area.

[0118] It should be understood that, although each step in the flow chart of each embodiment of the present invention is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0119] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0120] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0121] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

[0122] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for controlling the broadcast of an intelligent speaker, characterized in that: The method comprises: Collect audio signals through each microphone, analyze the sound intensity received by each microphone, calculate the time difference between different microphones receiving the sound, and determine the direction of the sound source; Extract the characteristics of each sound signal and separate the sound signals from different sources; Performing speech recognition on each separated audio signal, extracting and analyzing the content of the recognized speech input; Based on the source location of the sound, the space is divided into different areas, and the corresponding sound source of each area is recorded. Based on the identified content and sound source location, the response priority and corresponding area of ​​each request are determined; Corresponding broadcast content is generated according to the requests of different identified areas, and the broadcast content is bound to the response area, and the broadcast content is broadcast through the smart speaker in the bound area.

2. The method according to claim 1, characterized in that The analyzing the sound intensity received by each microphone, calculating the time difference between different microphones receiving the sound, and determining the source direction of the sound specifically includes: Collect surrounding audio signals through all microphones at the same time, and analyze the intensity of the audio signal received by each microphone; Measure the time difference between the sound reaching different microphones, and combine the sound intensity and the results of time difference analysis to determine the specific source direction of the sound.

3. The method according to claim 2, characterized in that The extracting of the features of each sound signal and separating the sound signals from different sources specifically includes: Extract features from the collected audio signals, separate the audio signals from different sound sources, and identify all the sound source signals; Perform a quality assessment on the separated sound source signal to determine whether the serial number clarity meets the recognition requirements; For the sound source signals that meet the recognition requirements, they are stored in the recognized storage memory and the signals are directly transmitted. For the sound source signals that do not meet the recognition requirements, they are stored in the unrecognized storage memory and the unrecognizable signals are transmitted.

4. The method according to claim 3, characterized in that The performing speech recognition on each separated audio signal, extracting and analyzing the content of the recognized speech input, specifically includes: Identify each sound source signal transmitted in the identified storage, convert the audio signal into text, and identify the speech content therein; Natural language processing is used to parse the recognized text results, including understanding the intent and extracting information.

5. The method according to claim 4, characterized in that The space is divided into different areas, and the corresponding sound source of each area is recorded. According to the identified content and the location of the sound source, the response priority of each request and the corresponding area are determined, which specifically includes: By analyzing the sound intensity and time difference, the relative position of each sound source signal in space is identified; According to the identified sound source location and the actual environment, the entire environment space is divided into several areas; Record the sound source information corresponding to each divided area, associate the sound source position with the area, and establish a mapping relationship; The priority of each request is determined based on the identified content and the location of the sound source, and the distance to the sound source. Based on the priority evaluation, the response area corresponding to each sound source request is determined.

6. The method according to claim 4, characterized in that The generating of corresponding broadcast content, binding the broadcast content with the response area, and broadcasting the broadcast content through the smart speaker in the bound area specifically includes: Generate corresponding broadcast information and control instructions according to the recognized voice input content; Binding the generated announcement content and control instructions to the area mapped by the sound source position; The control instructions are passed to the smart speakers in the corresponding bound area, and the generated broadcast control instructions are executed by the smart speakers to output the broadcast content to the bound area.

7. A broadcast control system for an intelligent speaker, characterized in that: The system comprises: The sound collection and direction determination module is used to collect audio signals through various microphones, analyze the sound intensity received by each microphone, calculate the time difference between different microphones receiving the sound, and determine the direction of the sound source; The sound feature extraction and separation module is used to extract the features of each sound signal and separate the sound signals from different sources; A speech recognition and content analysis module, used to perform speech recognition on each separated audio signal, extract and analyze the content of the recognized speech input; The area division and priority judgment module is used to divide the space into different areas based on the source location of the sound, and record the corresponding sound source of each area. According to the identified content and sound source location, the response priority of each request and the corresponding area are determined; The broadcast content generation and area binding module is used to generate corresponding broadcast content according to the requests of different identified areas, bind the broadcast content to the response area, and broadcast the broadcast content through the smart speaker in the bound area.

8. The system according to claim 7, characterized in that The area division and priority determination module includes: A sound source localization analysis unit, used to identify the relative position of each sound source signal in space by analyzing the sound intensity and time difference; The environment area division unit is used to divide the entire environment space into several areas according to the identified sound source position and the actual environment; A sound source mapping recording unit is used to record the sound source information corresponding to each divided area, by associating the sound source position with the area and establishing a mapping relationship; A request priority evaluation unit, used to determine the response priority of each request based on the identified content and the sound source location, taking the distance from the sound source location as a judgment rule; The response area decision unit is used to determine the response area corresponding to each sound source request based on the priority evaluation.

9. The system according to claim 8, characterized in that The broadcast content generation and area binding module includes: A broadcast information generating unit, used to generate corresponding broadcast information and control instructions according to the recognized voice input content; A content and area binding unit, used to bind the generated broadcast content and control instructions to the area mapped by the sound source position; The instruction transmission and execution unit is used to transmit the control instruction to the smart speaker in the corresponding bound area, execute the generated broadcast control instruction through the smart speaker, and output the broadcast content to the bound area.