Intelligent LED screen voice wake-up method
By combining an ultrasonic transducer and a MEMS microphone array with a low-power wake-up word detection chip and a local area network arbitration protocol, the problem of false wake-up of LED screens in densely deployed multi-screen environments was solved, achieving accurate voice wake-up and low-latency interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XUZHOU KAISHIDA INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2026-05-07
- Publication Date
- 2026-07-31
AI Technical Summary
In scenarios where multiple LED screens are densely deployed, existing voice wake-up technology is prone to being falsely triggered by advertising audio playing wake-up words on adjacent screens, making it impossible to accurately wake up a specific screen and affecting the user's interactive experience. Existing improvement solutions have failed to effectively distinguish between the sound source of a living human being and the sound source of an audio playback device.
It employs an ultrasonic transducer and MEMS microphone array combined with a low-power wake-word detection chip. It generates a binary phase shift keying modulated ultrasonic spread spectrum signal for positioning and liveness authentication. Combined with a local area network arbitration protocol, it ensures unique screen wake-up and uses a dedicated pseudo-random spreading code to avoid signal interference, enabling multi-screen collaborative work.
It effectively distinguishes between the sound source of a living human and the sound source of an audio playback device, reduces the false wake-up rate, ensures unique responsiveness in multi-screen scenarios, and ensures that the end-to-end wake-up latency does not exceed 120ms, so that users perceive no delay.
Smart Images

Figure CN122493846A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice interaction technology for intelligent display devices, and more particularly to a voice wake-up method for intelligent LED screens. Background Technology
[0002] In public places such as shopping malls, airports, exhibition halls, and conference rooms, the application scenario of multiple LED screens deployed side by side or in clusters is very common. Existing LED screen voice wake-up technology generally adopts a single passive wake-up word triggering method, that is, the microphone continuously collects ambient sound, and when a preset wake-up word is recognized, it directly triggers the screen activation. This method can be used in single-screen independent deployment scenarios, but in multi-screen dense deployment scenarios, when an adjacent screen is playing an advertising TTS voice containing a wake-up word through a speaker, all surrounding screens will be triggered and woken up at the same time, forming a sound cross-triggering problem between screens. This makes it impossible for users to accurately wake up a specific screen with voice commands, which seriously affects the interactive experience.
[0003] To address the aforementioned cross-triggering problem, existing solutions mainly fall into the following categories: First, using microphone array beamforming technology to roughly estimate the direction of the sound source. However, this method can only distinguish a wide range of directions and cannot differentiate between human voices originating from the user and sounds played from adjacent screen speakers. Second, reducing false triggers by increasing the wake-up word recognition confidence threshold. However, this method sacrifices the wake-up success rate for normal users, requiring users to repeatedly utter the wake-up word. Third, introducing camera face detection to assist in determining user presence. However, this solution heavily relies on lighting conditions and involves facial data collection, posing significant privacy compliance risks and introducing additional processing delays. None of these existing solutions fundamentally solve the problem of sound source cross-triggering in multi-screen, densely deployed scenarios, nor can they effectively distinguish between live human voices and audio playback device sounds. Therefore, a novel technical solution is urgently needed. Summary of the Invention
[0004] The purpose of this invention is to provide a voice wake-up method for intelligent LED screens to solve the problems existing in the prior art.
[0005] This invention provides a method for voice wake-up of a smart LED screen, comprising the following steps:
[0006] The S100 has six ultrasonic transducers embedded horizontally at equal intervals within the lower bezel of the LED screen. The center transmission frequency of the transducers is 21kHz, and the spacing between adjacent transducers is 8cm. A MEMS microphone with a sampling rate of 48kHz and a frequency response range covering 300Hz to 22kHz is embedded in each of the four corners of the screen. The screen integrates a low-power wake-word detection chip and a main processor with an operating power consumption of less than 1mW. All LED screens in the same location are connected to the same local area network switch via wired Ethernet, and each screen is assigned a globally unique screen number.
[0007] S200: After each screen is powered on, the 127-bit maximum length sequence bound to the screen number is used as a unique pseudo-random spreading code. It is multiplied with a 21kHz sinusoidal carrier signal at a chip rate of 2kbps to generate a binary phase shift keying modulated ultrasonic spread spectrum signal, which is transmitted synchronously and continuously through 6 ultrasonic transducers with a transmission power consumption of 50mW.
[0008] S300: The microphone array receives ultrasonic echoes with a frame period of 63.5ms. It performs normalized cross-correlation operation on the echo signals of each microphone and the 127-bit pseudo-random spreading code of the screen, reads the arrival time corresponding to the position of the correlation peak, calculates the arrival time difference between adjacent microphones, and uses the least squares method to solve for the azimuth and distance of the scattering source. When the distance falls within the range of 0.5m to 3.0m and the azimuth is between -30° and +30°, feature A is determined to be passed.
[0009] After processing 16 consecutive frames of ultrasound echoes using the S400, a short-time Fourier transform is performed on the accumulated 1.0s correlation peak amplitude time-series data. The ratio of the maximum power spectral density in the 0.2Hz to 0.5Hz frequency band to the average power spectral density across the entire frequency band is calculated as the confidence level for in vivo respiration. ,when When the value is greater than 0.75, the liveness detection is successful;
[0010] The S500 low-power wake-word detection chip continuously performs keyword template matching detection on the voice signals collected by the two microphones closest to the screen. When the recognition score exceeds an internal threshold, it sends a candidate activation interrupt signal to the main processor. The main processor switches from sleep to active state within 10ms, and simultaneously records the candidate activation time. and the azimuth angle of the wake-up word sound source estimated based on the phase difference of the two microphone signals. ;
[0011] S600, main processor extracts from The wake-word speech signal was collected from four microphones within 50ms forward. The root mean square sound pressure level (RMS) of each microphone was calculated. The near-field sound pressure ratio matching score was calculated by subtracting the measured sound pressure ratio from the inverse ratio of the distance from the sound source to each microphone estimated based on the S300 positioning results. ,when When the value is greater than 0.70, the near-field verification is passed;
[0012] S700, the azimuth angle of the live scattering source calculated by S300. Recorded with S500 The directional consistency verification is passed when the absolute value of the angular difference is less than 15°. and The overall confidence score is calculated using the three scores with weights of 0.3, 0.35, and 0.35 respectively. ,Will The value, screen number, and current timestamp are encapsulated into an arbitration competition frame and broadcast to all screens within the local area network;
[0013] In S800, after each screen receives a competing frame, it enters a 50ms arbitration wait window. After the window ends... The screen with the highest value broadcasts the arbitration completion frame to all other screens and performs wake-up activation. After receiving the arbitration completion frame, the other screens suppress the wake-up response and return to S300.
[0014] The S900, after winning the arbitration, switches the screen-driven LED display module to the interactive interface, activates the full-precision voice recognition module, and the end-to-end latency from candidate activation to wake-up completion does not exceed 120ms.
[0015] Furthermore, the formulas for calculating the azimuth angle and distance of the scattering source in S300 are as follows: ,in, For the estimated azimuth angle of the scattering source, To estimate the distance from the scattering source to the center of the screen, For adjacent microphones and The time difference of arrival For the first Two-dimensional coordinates of each microphone Let be the location coordinates of the scattering source to be determined. The speed of sound is 343 m / s.
[0016] Furthermore, the confidence level of live respiration in S400 The calculation formula is: ,in, Confidence level for in vivo respiration. The power spectral density is the time-series signal of the ultrasonic echo correlation peak amplitude. For frequency variables, the unit is Hz. The total number of frequency points involved in the calculation. For the first For each frequency point, the numerator is the maximum power spectral density in the range of 0.2Hz to 0.5Hz, and the denominator is the average power spectral density across the entire frequency band.
[0017] Furthermore, the near-field sound pressure level matching score in the S600 The calculation formula is: ,in, The near-field sound pressure level ratio matching score ranges from 0 to 1. For the first The root mean square sound pressure level measured by each microphone The root mean square sound pressure level (RMS) measured by the first microphone. The sound source estimated based on the S300 localization results is located at the [missing information - likely a location or location]. The distance of one microphone, The distance from the sound source to the first microphone. This is the normalization constant, with a value of 3.
[0018] Furthermore, in S200, each screen uses a 127-bit maximum length sequence bound to its own screen number as a unique pseudo-random spreading code. The ultrasonic signals of each screen are orthogonal to each other in the spectrum. The microphone array of any screen can separate the echo of the transmitted signal from the aliased echo by performing correlation operations between the received echo signal and the pseudo-random spreading code of the screen.
[0019] Furthermore, in S800, when only the local screen sends a contention frame within the 50ms arbitration wait window and no other screen sends a contention frame, the local screen directly performs wake-up activation without waiting for the arbitration window to end.
[0020] Furthermore, the overall confidence score in S700 Based on directional consistency score, and The three terms are weighted and summed with weights of 0.3, 0.35, and 0.35 respectively. This summation is carried in the arbitration competition frame. The value serves as the sole basis for comparison in S800 multi-screen arbitration.
[0021] Furthermore, in the S400 When the value is below 0.75, the main processor determines that the sound source scattering in front is a non-living sound source, suppresses the current wake-up candidate, shuts down the main processor, and the system returns to S300 to continue listening; in S600 when When the value is below 0.70, the main processor determines that the wake-up word source is a far-field playback source, suppresses the wake-up candidate, shuts down the main processor, and the system returns to S300 to continue listening.
[0022] The beneficial effects of this invention are:
[0023] First, this invention combines active ultrasonic sound field modulation with passive voice monitoring to complete the live human presence authentication before the wake word detection is triggered. It distinguishes the voice of a live human from the voice emitted by the audio playback device at the physical signal level, fundamentally eliminating the false wake-up interference caused by the TTS voice played on adjacent screens. Compared with the existing solutions that only rely on increasing the wake word recognition threshold, the false wake-up rate is significantly reduced.
[0024] Secondly, the present invention adopts a code division multiple access ultrasonic spread spectrum system based on a 127-bit maximum length sequence dedicated to each screen, so that the ultrasonic pilot signals of different screens in the same location are orthogonal to each other in the spectrum. The microphone array of each screen can independently separate the signal of its own screen from the multi-screen superimposed ultrasonic echo by performing normalized cross-correlation operation with the pseudo-random spread spectrum code dedicated to its own screen. This achieves stable positioning and liveness detection without interference under the parallel operation of multiple screens, without the need to impose special restrictions on the screen spacing or deployment method.
[0025] Third, this invention designs a multi-screen collaborative wake-up mechanism based on a local area network confidence contention arbitration protocol. When multiple screens simultaneously meet the three conditions for joint judgment, an arbitration contention frame is broadcast and the comprehensive confidence score Q value is compared within a 50ms arbitration window. This ensures that only the screen with the highest confidence score completes wake-up activation in the same location, while the other screens automatically suppress their response and return to the listening state. This guarantees unique responsiveness in multi-screen dense scenarios at the system level, with an end-to-end wake-up latency of no more than 120ms, and no perceived delay for the user. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a schematic diagram of the hardware deployment structure of the present invention;
[0028] Figure 2 This is a schematic diagram of the ultrasonic pilot signal encoding and transmission process of the present invention;
[0029] Figure 3 This is a schematic diagram of the three-condition joint wake-up decision logic of the present invention;
[0030] Figure 4 This is a schematic diagram of the multi-screen arbitration process of the present invention. Detailed Implementation
[0031] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. It should also be noted that, to make the embodiments more comprehensive, the following embodiments are the best and preferred embodiments, and those skilled in the art can use other alternative methods to implement some well-known technologies; moreover, the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.
[0032] It should be noted that the use of terms such as "an embodiment," "an embodiment," "an exemplary embodiment," and "some embodiments" in the specification indicates that the described embodiment may include a specific feature, structure, or characteristic, but not every embodiment necessarily includes that specific feature, structure, or characteristic. Furthermore, when a specific feature, structure, or characteristic is described in connection with an embodiment, implementing such a feature, structure, or characteristic in conjunction with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the art.
[0033] Generally, terms can be understood at least partly from their use in context. For example, depending at least partly on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or a combination of features, structures, or characteristics in a plural sense. Additionally, the term "based on" can be understood not necessarily to convey an exclusive set of factors, but rather, alternatively, depending at least partly on the context, to allow for the presence of other factors that are not necessarily explicitly described.
[0034] See Figures 1 to 4 As shown
[0035] This invention provides a method for multi-screen arbitration voice wake-up of intelligent LED screens based on active ultrasonic sound field modulation and acoustic liveness pattern authentication. The invention will be further described in detail below with reference to specific embodiments.
[0036] S100 Hardware Deployment Phase
[0037] Within the bottom bezel of a single smart LED screen, six ultrasonic transducers are embedded horizontally at equal intervals. The center transmission frequency of each transducer is set to 21kHz, and the spacing between transducers is 8cm. This allows the overall ultrasonic transmission array to cover a conical effective interactive area directly in front of the screen, with a horizontal angle of ±45° and a depth of 0.5m to 3.0m. Four MEMS wideband microphones are embedded at each of the four corners of the screen. These microphones have a sampling rate of 48kHz and a frequency response range covering 300Hz to 22kHz, enabling them to simultaneously receive ultrasonic echoes and acquire audio signals. An independent low-power wake-up word detection chip is integrated inside the screen. This chip consumes less than 1mW and continuously monitors the audio signals acquired by the microphone array. When a preset wake-up word is detected, it sends a candidate activation signal to the main processor. The main processor remains in sleep mode when no candidate activation signal is received. All LED screens in the same location are connected to the same local area network switch via wired Ethernet, forming a multi-screen arbitration communication network. Each screen is assigned a globally unique screen number at the factory.
[0038] S200 Ultrasonic Pilot Signal Encoding and Transmission Stage
[0039] After the system is powered on, the ultrasonic transmitting array of each screen immediately enters continuous transmission mode. Each screen uses a 127-bit maximum-length sequence bound to its screen number as its own pseudo-random spreading code, with a chip rate set to 2kbps and a duration of 63.5ms for one frame of the pseudo-random sequence. The ultrasonic pilot signal is generated by multiplying the pseudo-random spreading code sequence with a 21kHz sinusoidal carrier signal to obtain a binary phase-shift keying modulated ultrasonic spread spectrum signal. This signal is sent to six ultrasonic transducers for synchronous transmission. The entire transmission process is continuous, with a power consumption of approximately 50mW. Because different screens use different pseudo-random spreading codes, the ultrasonic signals from each screen are orthogonal to each other in the spectrum. The microphone array of any screen can perform correlation operations with its own screen's pseudo-random code to separate the echo of its own transmitted signal from the aliased echo, without being affected by the ultrasonic signals from adjacent screens.
[0040] S300 Acoustic Live Body Shadow Feature Extraction – Ultrasonic TDOA Human Body Localization
[0041] The microphone array continuously receives ultrasonic echo signals. The main processor performs sliding window processing on the echo data from the four microphones at a frame period of 63.5ms. For each frame, the echo signal received by each microphone is subjected to normalized cross-correlation with the 127-bit pseudo-random spreading code of the local screen to obtain the time delay correlation peak of each channel. The arrival time corresponding to the position of the correlation peak is read, and the arrival time difference between microphone i and microphone j is calculated. Based on the spatial coordinates of the four microphones and the measured time difference of arrival, the azimuth and distance of the scattering source are calculated using the least squares method. The specific calculation formula is as follows:
[0042]
[0043] in, For the estimated azimuth angle of the scattering source, To estimate the distance from the scattering source to the center of the screen, For adjacent microphones and The time difference of arrival For the first Two-dimensional coordinates of each microphone Let be the location coordinates of the scattering source to be determined. The speed of sound is taken as 343 m / s. When the calculated... Falling within the range of 0.5m to 3.0m and If the temperature falls between -30° and +30°, it is determined that there is a scatterer in the effective interaction area; otherwise, feature A is determined to be unacceptable, and the system continues to monitor.
[0044] S400 Acoustic Liveness Feature Extraction – Micro-Doppler Breathing Liveness Authentication
[0045] After processing 16 frames of ultrasound echoes continuously using the S300, approximately 1.0 s of relevant peak amplitude time-series data was accumulated. A short-time Fourier transform was performed on this data to analyze the power spectral density distribution in the 0.1 Hz to 1.0 Hz frequency band. The periodic displacement amplitude of the thoracic cavity during human respiratory movement is approximately 2 mm to 10 mm, corresponding to an ultrasound Doppler frequency shift around 0.2 Hz to 0.5 Hz. The power of the maximum spectral peak within this frequency band was calculated. Average power across the entire frequency band The ratio, used as a confidence index for liveness, is calculated using the following formula:
[0046]
[0047] in, Confidence level for in vivo respiration. The power spectral density is the time-series signal of the ultrasonic echo correlation peak amplitude. For frequency variables, The total number of frequency points involved in the calculation. For the first For each frequency point, the numerator is the maximum power spectral density within the range of 0.2Hz to 0.5Hz, and the denominator is the average power spectral density across the entire frequency band. When If the value is greater than 0.75, the forward scattering source is determined to have the characteristics of a living person breathing, and the liveness authentication is passed; otherwise, it is determined to be a non-living sound source, the system suppresses the current wake-up candidate and returns to S300 to continue listening.
[0048] S500 Low-Power Wake-Up Word Detection Trigger
[0049] The low-power wake-up word detection chip continuously detects the audio signal from the two microphones closest to the screen out of four microphones based on keyword template matching. When the recognition score of the preset wake-up word exceeds an internal threshold, it immediately sends a candidate activation interrupt signal to the main processor. The main processor then switches from sleep mode to active mode with a latency of less than 10ms. The wake-up word detection chip simultaneously records the moment of this candidate activation. And the rough azimuth angle of the wake-up word sound source estimated based on the phase difference between the two microphone signals. .
[0050] S600 Acoustic Live Body Shadow Feature Extraction – Near-Field Spherical Wave Sound Pressure Ratio Verification
[0051] After the main processor is activated, extract data from the candidate activation time. The wake-up word speech signals were collected by four microphones within 50ms forward, and the root mean square sound pressure level of each microphone on the speech signal was calculated. According to the inverse square law of sound pressure ratio for near-field spherical waves, if the sound source is located in the near-field position directly in front of the screen, the ratio of sound pressure levels of each microphone should satisfy the inverse square relationship with the ratio of the distances from each microphone to the sound source. The formula for calculating the near-field sound pressure ratio matching score is as follows:
[0052]
[0053] in, The near-field sound pressure level ratio matching score ranges from 0 to 1. For the first The root mean square sound pressure level measured by each microphone The root mean square sound pressure level (RMS) measured by the first microphone. The sound source estimated based on the S300 localization results is located at the [missing information - likely a location or location]. The distance of one microphone, The distance from the sound source to the first microphone. The normalization constant is set to 3, used to limit the score to the range of 0 to 1. If the value is greater than 0.70, the wake-up word sound source is determined to be a near-field living human voice, and the near-field verification is passed; otherwise, it is determined to be a far-field sound source, the system suppresses the wake-up candidate and returns to S300.
[0054] S700 Three-Condition Joint Wake-Up Decision
[0055] Provided that the calculation results from S300 to S600 all meet the threshold, the main processor further performs directional consistency verification: the azimuth angle of the live scattering source calculated by S300 is then used for... The azimuth angle of the wake word source recorded by S500 The difference between the two angles is calculated. If the absolute value of the angle difference is less than 15°, the direction of the sound source is determined to be consistent with the direction of the living human body, and the direction consistency verification is passed. At this point, all three conditions are met: the direction consistency condition is met, and the confidence level of the living human respiration is satisfied. Over 0.75, near-field matching score If the score exceeds 0.70, the main processor will combine the three conditions to form a confidence score. The calculation is a weighted sum of the three scores, with weights set to 0.3, 0.35, and 0.35 respectively. The value, along with the screen number and the current timestamp, is encapsulated into an arbitration competition frame and broadcast to all other screens in the same location via the local area network.
[0056] S800 Multi-Screen Arbitration Phase
[0057] Upon receiving an arbitration contention frame from another screen, each screen pauses its current wake-up response and enters a 50ms arbitration wait window. After the 50ms window ends, the screen's overall confidence score is compared. Compared to the scores carried in all competing frames, if this screen If the value is the highest, this screen wins the arbitration, broadcasts the arbitration completion frame to all other screens, and immediately performs wake-up activation; if this screen... If the value is not the highest, after receiving the arbitration completion frame, this screen will suppress the wake-up response, shut down the main processor, and return to the S300 normal listening state; if only this screen sends a contention frame within the 50ms window and does not receive a contention frame from any other screen, this screen will directly perform wake-up activation without waiting.
[0058] S900 screen wake-up activation phase
[0059] The screen's main processor, which wins the arbitration, immediately drives the LED display module to switch to the interactive interface. At the same time, it activates the full-precision voice recognition module and begins to receive subsequent voice commands from the user. The screen then enters normal interactive response mode. The entire end-to-end latency from candidate activation to screen wake-up completion does not exceed 120ms.
[0060] Example 1
[0061] This embodiment uses five smart LED advertising screens installed side by side in the central hall of a shopping mall as a scenario. The five screens are numbered LED-01 to LED-05, with a spacing of 1.5m between them. They are all connected to the same switch via wired Ethernet, and the preset wake-up word for each screen is "Hello screen".
[0062] A user stands 1.2m directly in front of the LED-03, facing the LED-03, and speaks the wake-up phrase "Hello screen" at a normal speaking volume. At the same time, the LED-02 is playing an advertising TTS message containing the words "Hello screen" through its built-in speaker, at a volume similar to the user's speaking volume.
[0063] After executing S300, LED-03 was located using ultrasonic TDOA, and a continuous and stable echo scattering source was detected 1.2m directly in front of it at an azimuth angle of 3°, indicating that feature A passed. After executing S400, a significant spectral peak was detected at 0.28Hz in a 1.0s Doppler time-series analysis. The calculated value is 1.83, exceeding the 0.75 threshold, indicating successful liveness detection. During this period, the low-power wake-up word detection chip detected the "Hello Screen" wake-up word, triggering the S500 and recording the azimuth angle of the wake-up word sound source. The angle is 5°; after executing S600, the ratio of the sound pressure levels of the four microphones is in high agreement with the theoretical prediction of 1.2m near-field spherical waves. The calculated value is 0.88, exceeding the 0.70 threshold; after executing S700, the azimuth difference is 2°, less than 15°, all three conditions are met, and the overall confidence level is [not specified]. The calculated value is 0.91, and LED-03 broadcasts an arbitration contention frame to the local area network.
[0064] The LED-02 also triggered the candidate activation of the low-power chip due to the wake-up word in the advertising TTS voice. The main processor was then woken up and executed feature C calculation. Since the TTS audio came from the LED-02's own speaker, it was a screen-level near-field point source but not a human sound source. Its ultrasonic echo did not contain the periodic Doppler component of 0.2Hz to 0.5Hz. The calculated value is 0.31, which is lower than the threshold of 0.75. The liveness authentication fails, the joint decision of LED-02 fails, no contention frame is sent, the main processor shuts down and returns to the listening state.
[0065] LED-01, LED-04, and LED-05 failed due to the user's standing position being far from their effective interaction area, resulting in ultrasonic echo scattering intensity below the detection threshold. Consequently, feature A was not passed, the main processor was not activated, and they did not participate in arbitration.
[0066] Within the 50ms arbitration window, only LED-03 sent a competing frame, and the arbitration was directly won by LED-03. LED-03 then executed S900, the LED display module switched to the interactive interface, and the voice recognition module was activated. The end-to-end latency of the entire wake-up process was 87ms, with no perceived delay for the user, and the interaction proceeded normally. The other four screens were not mistakenly woken up, achieving accurate single-screen wake-up under the conditions of dense multi-screen deployment and interference from adjacent screen playback.
[0067] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.
[0068] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A smart LED screen voice wake-up method, characterized in that, Includes the following steps: The S100 has six ultrasonic transducers embedded horizontally at equal intervals within the lower bezel of the LED screen. The center transmission frequency of the transducers is 21kHz, and the spacing between adjacent transducers is 8cm. A MEMS microphone with a sampling rate of 48kHz and a frequency response range covering 300Hz to 22kHz is embedded in each of the four corners of the screen. The screen integrates a low-power wake-word detection chip and a main processor with an operating power consumption of less than 1mW. All LED screens in the same location are connected to the same local area network switch via wired Ethernet, and each screen is assigned a globally unique screen number. S200: After each screen is powered on, the 127-bit maximum length sequence bound to the screen number is used as a unique pseudo-random spreading code. It is multiplied with a 21kHz sinusoidal carrier signal at a chip rate of 2kbps to generate a binary phase shift keying modulated ultrasonic spread spectrum signal, which is transmitted synchronously and continuously through 6 ultrasonic transducers with a transmission power consumption of 50mW. S300: The microphone array receives ultrasonic echoes with a frame period of 63.5ms. It performs normalized cross-correlation operation on the echo signals of each microphone and the 127-bit pseudo-random spreading code of the screen, reads the arrival time corresponding to the position of the correlation peak, calculates the arrival time difference between adjacent microphones, and uses the least squares method to solve for the azimuth and distance of the scattering source. When the distance falls within the range of 0.5m to 3.0m and the azimuth is between -30° and +30°, feature A is determined to be passed. After processing 16 consecutive frames of ultrasound echoes using the S400, a short-time Fourier transform is performed on the accumulated 1.0s correlation peak amplitude time-series data. The ratio of the maximum power spectral density in the 0.2Hz to 0.5Hz frequency band to the average power spectral density across the entire frequency band is calculated as the confidence level for in vivo respiration. ,when When the value is greater than 0.75, the liveness detection is successful; The S500 low-power wake-word detection chip continuously performs keyword template matching detection on the voice signals collected by the two microphones closest to the screen. When the recognition score exceeds an internal threshold, it sends a candidate activation interrupt signal to the main processor. The main processor switches from sleep to active state within 10ms, and simultaneously records the candidate activation time. and the azimuth angle of the wake-up word sound source estimated based on the phase difference of the two microphone signals. ; S600, main processor extracts from The wake-word speech signal was collected from four microphones within 50ms forward. The root mean square sound pressure level (RMS) of each microphone was calculated. The near-field sound pressure ratio matching score was calculated by subtracting the measured sound pressure ratio from the inverse ratio of the distance from the sound source to each microphone estimated based on the S300 positioning results. ,when When the value is greater than 0.70, the near-field verification is passed; S700, the azimuth angle of the live scattering source calculated by S300. Recorded with S500 The directional consistency verification is passed when the absolute value of the angular difference is less than 15°. and The overall confidence score is calculated using the three scores with weights of 0.3, 0.35, and 0.35 respectively. ,Will The value, screen number, and current timestamp are encapsulated into an arbitration competition frame and broadcast to all screens within the local area network; In S800, after each screen receives a competing frame, it enters a 50ms arbitration wait window. After the window ends... The screen with the highest value broadcasts the arbitration completion frame to all other screens and performs wake-up activation. After receiving the arbitration completion frame, the other screens suppress the wake-up response and return to S300. The S900, after winning the arbitration, switches the screen-driven LED display module to the interactive interface, activates the full-precision voice recognition module, and the end-to-end latency from candidate activation to wake-up completion does not exceed 120ms. 2.The intelligent LED screen voice wake-up method of claim 1, wherein, The formulas for determining the azimuth angle and distance of the scattering source in S300 are as follows: ,in, For the estimated azimuth angle of the scattering source, To estimate the distance from the scattering source to the center of the screen, For adjacent microphones and The time difference of arrival For the first Two-dimensional coordinates of each microphone Let be the location coordinates of the scattering source to be determined. The speed of sound is 343 m / s.
3. The intelligent LED screen voice wake-up method according to claim 1, characterized in that, S400 Confidence in live respiration The calculation formula is: ,in, Confidence level for in vivo respiration. The power spectral density is the time-series signal of the ultrasonic echo correlation peak amplitude. For frequency variables, the unit is Hz. The total number of frequency points involved in the calculation. For the first For each frequency point, the numerator is the maximum power spectral density in the range of 0.2Hz to 0.5Hz, and the denominator is the average power spectral density across the entire frequency band. 4.The intelligent LED screen voice wake-up method of claim 1, wherein, S600 near-field sound pressure level matching score The calculation formula is: ,in, The near-field sound pressure level ratio matching score ranges from 0 to 1. For the first The root mean square sound pressure level measured by each microphone The root mean square sound pressure level (RMS) measured by the first microphone. The sound source estimated based on the S300 localization results is located at the [missing information - likely a location or location]. The distance of one microphone, The distance from the sound source to the first microphone. This is the normalization constant, with a value of 3. 5.The intelligent LED screen voice wake-up method of claim 1, wherein, In S200, each screen uses a 127-bit maximum length sequence bound to its own screen number as a unique pseudo-random spreading code. The ultrasonic signals of each screen are orthogonal to each other in the spectrum. The microphone array of any screen can separate the echo of the transmitted signal from the aliased echo by performing correlation operations between the received echo signal and the pseudo-random spreading code of its own screen. 6.The intelligent LED screen voice wake-up method of claim 1, wherein, In S800, if only the local screen sends a contention frame within the 50ms arbitration wait window and no other screen sends a contention frame, the local screen will directly perform wake-up activation without waiting for the arbitration window to end. 7.The intelligent LED screen voice wake-up method of claim 1, wherein, S700 Overall Confidence Score Based on directional consistency score, and The three terms are weighted and summed with weights of 0.3, 0.35, and 0.35 respectively. This summation is carried in the arbitration competition frame. The value serves as the sole basis for comparison in S800 multi-screen arbitration.
8. The intelligent LED screen voice wake-up method according to claim 1, characterized in that, S400 When the value is below 0.75, the main processor determines that the sound source scattering in front is a non-living sound source, suppresses the current wake-up candidate, shuts down the main processor, and the system returns to S300 to continue listening; in S600 when When the value is below 0.70, the main processor determines that the wake-up word source is a far-field playback source, suppresses the wake-up candidate, shuts down the main processor, and the system returns to S300 to continue listening.