Bluetooth sound box intelligent interaction method and device

By integrating the Bluetooth module, camera module, AI model module and embedded CPU module of the Bluetooth speaker, the problems of audio and video asynchrony, high computing resource consumption and low efficiency of the action template database are solved, and real-time, smooth and high-quality synchronous interaction between music and animation of the Bluetooth speaker is realized, thereby improving the user experience.

CN120672916APending Publication Date: 2025-09-19GUANGZHOU LITTLE SEA MONSTER TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510700655.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing Bluetooth speakers face problems such as audio and video asynchrony, high computing resource consumption, slow system response, and inefficient action template database when achieving real-time, smooth, and high-quality synchronous interaction between music and animation. The introduction of AI models also leads to network transmission delays and data security challenges.

Method used

By integrating Bluetooth module, camera module, AI model module and embedded CPU module, intelligent synchronous presentation of audio signals and dynamic vision is achieved. Audio delay algorithm and hardware accelerated parallel processing architecture are adopted, combined with intelligent motion recognition algorithm and dynamic 3D rendering technology to generate interaction with users. The embedded CPU module achieves precise synchronization of audio signals and animation.

Benefits of technology

It significantly improves the real-time and immersive nature of human-computer interaction, reduces the synchronization error of audio and video to the millisecond level, improves the accuracy and response speed of motion capture, and reduces system power consumption and hardware costs, achieving real-time, smooth and high-quality synchronous interaction between music and animation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672916A_ABST
    Figure CN120672916A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent sound boxes, and discloses a Bluetooth sound box intelligent interaction method and device, and the method comprises the steps: a Bluetooth module receives a Bluetooth audio signal transmitted by intelligent equipment, and transmits the Bluetooth audio signal to an embedded CPU module; the AI model module obtains an action image signal to obtain an action control signal; processing the action control signal to obtain a 3D animation action signal and sending the 3D animation action signal to the embedded CPU module; the embedded CPU module receives and decodes the Bluetooth audio signal to obtain a digital audio signal; delaying the digital audio signal to obtain a delayed digital audio signal and sending the delayed digital audio signal to the power amplifier module; obtaining a real-time video signal according to the 3D animation action signal and a preset 3D model, and sending the real-time video signal to a display module; the power amplifier module receives and processes the delayed digital audio signal and sends the delayed digital audio signal to the loudspeaker module; the loudspeaker module receives and plays the delayed digital audio signal; the display module processes the real-time video signal to obtain a real-time action picture of the 3D animation model and sends the real-time action picture to the display screen module; and the display screen module receives and displays the real-time action picture. According to the invention, low-delay fusion output of audio-visual signals can be realized, system power consumption and hardware cost are reduced, and finally real-time, smooth and high-quality synchronous interaction of music and animation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of smart speaker technology, and in particular to a method and device for smart interaction with a Bluetooth speaker. Background Art

[0002] Currently, the core technical challenge facing Bluetooth speakers with screens on the market is achieving real-time, smooth, and high-quality synchronized interaction between music and animation. Furthermore, the reception, decoding, and processing of audio signals requires extremely low latency, otherwise audio and video will become out of sync. Secondly, real-time analysis of music features and motion capture requires robust computing power to quickly extract key information. Furthermore, the generation and rendering of 3D animations is time-consuming, potentially impacting the system's real-time performance. Furthermore, minimizing computing resource consumption and improving system responsiveness while maintaining animation quality poses a daunting challenge. While the introduction of AI models can improve processing power, it also introduces new challenges with network transmission latency and data security. Finally, establishing an efficient and accurate motion template database to support rapid matching and optimization, further enhancing system performance, is a pressing issue. Resolving these technical challenges directly impacts the user experience's immersiveness and the system's practicality, significantly impacting the widespread adoption and commercialization of Bluetooth speakers. Summary of the Invention

[0003] The present application provides a method and device for intelligent interaction of Bluetooth speakers, which can realize intelligent synchronous presentation of audio signals and dynamic vision by integrating Bluetooth modules, camera modules, AI model modules and embedded CPU modules. The present invention adopts a Bluetooth module to receive audio signals from external smart devices, and optionally combines a camera module to capture user actions in real time and generates three-dimensional motion control signals through analysis by the AI ​​model module. The embedded CPU module realizes precise synchronous processing of digital audio signals through an audio delay algorithm, and at the same time drives a preset 3D model to generate an animated video that matches the user action. The present invention significantly improves the real-time and immersive nature of human-computer interaction, effectively reduces the synchronization error of audio and video to the millisecond level through a hardware-accelerated parallel processing architecture, and improves the accuracy and response speed of motion capture through an intelligent motion recognition algorithm. Finally, dynamic 3D rendering technology is used to enhance visual expression. The present invention can realize low-latency fusion output of audio and video signals, while reducing system power consumption and hardware costs, and ultimately realize real-time, smooth and high-quality synchronous interaction between music and animation.

[0004] In a first aspect, an embodiment of the present application provides a method for intelligent interaction with a Bluetooth speaker, the method comprising:

[0005] The Bluetooth module receives the Bluetooth audio signal sent by the smart device and sends it to the embedded CPU module;

[0006] The AI ​​model module acquires the motion image signal and obtains the motion control signal; processes the motion control signal, obtains the 3D animation motion signal and sends it to the embedded CPU module;

[0007] The embedded CPU module receives and decodes the Bluetooth audio signal to obtain a digital audio signal; delays the digital audio signal to obtain the delayed digital audio signal and sends it to the power amplifier module; obtains a real-time video signal based on the 3D animation action signal and the preset 3D model and sends it to the display module;

[0008] The power amplifier module receives and processes the delayed digital audio signal and sends it to the speaker module;

[0009] The speaker module receives and plays the delayed digital audio signal;

[0010] The display module processes the real-time video signal, obtains the real-time action picture of the 3D animation model and sends it to the display screen module;

[0011] The display screen module receives and displays real-time action images.

[0012] Furthermore, the method further comprises:

[0013] The embedded CPU module analyzes the digital audio signal in automatic mode, obtains the action sequence number and sends it to the AI ​​model module; the AI ​​model module obtains the 3D animation action signal based on the action sequence number and sends it to the embedded CPU module;

[0014] Alternatively, the camera module collects the user's motion image signal in camera mode and sends it to the AI ​​model module; the AI ​​model module receives and processes the motion image signal to obtain a motion control signal; processes the motion control signal to obtain a 3D animation motion signal and sends it to the embedded CPU module.

[0015] Furthermore, the method further comprises:

[0016] The AI ​​model module sends motion control signals or music rhythm and type features to the cloud server;

[0017] The cloud server inputs the motion control signal, or music rhythm and type features into the cloud AI model, obtains a 3D animation motion video stream and sends it to the smart device via a wireless network.

[0018] Furthermore, the method further comprises:

[0019] The embedded CPU module stores 3D animation motion signals and real-time video signals to obtain a motion template database.

[0020] Furthermore, the method further comprises:

[0021] After receiving the digital audio signal, the embedded CPU module samples the digital audio signal according to the resampling algorithm and performs delay processing on the sampled digital audio signal; the sampling rate of the resampling algorithm is 48kHz and the bit depth is 24bit.

[0022] Furthermore, the method further comprises:

[0023] The embedded CPU module receives and decodes the Bluetooth audio signal, obtains the digital audio signal and sends it to the AI ​​model module;

[0024] The AI ​​model module uses a convolutional neural network algorithm to extract features from digital audio signals and send them to the adaptive filter;

[0025] The adaptive filter performs audio enhancement processing on the digital audio signal after feature extraction, improves the signal-to-noise ratio of the enhanced digital audio signal to 90dB and sends it to the embedded CPU module.

[0026] Furthermore, the method further comprises:

[0027] The embedded CPU module performs weighted mixing on the enhanced digital audio signal and the original digital audio signal to obtain a weighted digital audio signal and performs delay processing to obtain a delayed digital audio signal;

[0028] The weighted mix ratio is 7:3.

[0029] Furthermore, the embedded CPU module delays the digital audio signal, including:

[0030] The embedded CPU module uses a timer with an internal clock frequency of 48kHz to perform time delay processing on the digital audio signal. The delay time is 100ms to obtain the delayed digital audio signal.

[0031] Furthermore, the power amplifier module receives and processes the delayed digital audio signal, including:

[0032] The power amplifier module uses an FIR filter algorithm to adjust the signal gain of the delayed digital audio signal to +6dB.

[0033] Furthermore, the embedded CPU module analyzes the digital audio signal in the automatic mode and obtains the action sequence number, including:

[0034] The embedded CPU module analyzes the signal data packet according to the digital audio signal to obtain the audio waveform data;

[0035] According to the audio waveform data, the Fourier transform algorithm is used to extract the frequency characteristics and determine the music rhythm characteristics;

[0036] Analyze frequency features through the pre-trained music classification model, determine the music type, and obtain the music type number;

[0037] According to the music rhythm characteristics and music type number, the preset rhythm-action mapping table is queried to obtain the corresponding action sequence number and send it to the AI ​​model module.

[0038] Furthermore, the frequency characteristics are analyzed through the pre-trained music classification model to determine the music type and obtain the music type number, including:

[0039] The pre-trained music classification model uses a convolutional neural network structure to input frequency features into the Mel-spectrogram and perform calculations to obtain a probability distribution; the music type number is determined based on the probability distribution.

[0040] Furthermore, when the music rhythm feature is BPM=128 and the music type number is EDM_003, the corresponding action sequence number in the preset rhythm-action mapping table is DANCE_SEQ_47, and the action sequence is a robotic dance action sequence.

[0041] Furthermore, the camera module collects the user's motion image signal in the camera mode, including:

[0042] The camera module collects the user's original video data at a frame rate of 30fps and compresses it using H264 encoding to obtain motion image signals.

[0043] Furthermore, the display module processes the real-time video signal to obtain the real-time action picture of the 3D animation model, including:

[0044] The display module uses bilinear interpolation to eliminate screen tearing and obtain a real-time video signal at 60fps.

[0045] Furthermore, the display screen module receives and displays real-time action images, including:

[0046] The display screen module receives real-time action images through the MIPI interface and uses PWM dimming to control the deflection of liquid crystal molecules, presenting real-time action images with a dynamic range of 1000:1 within a 5ms response time.

[0047] In a second aspect, an embodiment of the present application provides a Bluetooth speaker intelligent interactive device, the device comprising:

[0048] A Bluetooth module is used to receive Bluetooth audio signals sent by smart devices and send them to the embedded CPU module;

[0049] The AI ​​model module is used to acquire motion image signals and obtain motion control signals; process the motion control signals to obtain 3D animation motion signals and send them to the embedded CPU module;

[0050] The embedded CPU module is used to receive and decode Bluetooth audio signals to obtain digital audio signals; delay the digital audio signals to obtain delayed digital audio signals and send them to the power amplifier module; obtain real-time video signals based on 3D animation action signals and preset 3D models and send them to the display module;

[0051] A power amplifier module, used to receive and process the delayed digital audio signal and send it to the speaker module;

[0052] A speaker module, used for receiving and playing delayed digital audio signals;

[0053] The display module is used to process the real-time video signal, obtain the real-time action picture of the 3D animation model and send it to the display screen module;

[0054] The display screen module is used to receive and display real-time action images.

[0055] Furthermore, the embedded CPU module is also used to:

[0056] In automatic mode, the digital audio signal is analyzed, the action sequence number is obtained and sent to the AI ​​model module.

[0057] Furthermore, the AI ​​model module is also used to:

[0058] The 3D animation action signal is obtained according to the action sequence number and sent to the embedded CPU module.

[0059] Furthermore, the embedded CPU module is also used to:

[0060] The 3D animation action signals and real-time video signals are stored to obtain an action template database.

[0061] Furthermore, the embedded CPU module is also used to:

[0062] According to the resampling algorithm, the sampling rate of the digital audio signal is adjusted to 48kHz and the bit depth is unified to 24bit.

[0063] In summary, compared with the prior art, the technical solutions provided by the embodiments of the present application have at least the following beneficial effects:

[0064] The embodiment of the present application provides a method for intelligent interaction of Bluetooth speakers, which can realize intelligent synchronous presentation of audio signals and dynamic vision by integrating a Bluetooth module, a camera module, an AI model module and an embedded CPU module. The present invention uses a Bluetooth module to receive audio signals from external smart devices, and optionally combines a camera module to capture user actions in real time and generates three-dimensional motion control signals through analysis by the AI ​​model module. The embedded CPU module realizes accurate synchronous processing of digital audio signals through an audio delay algorithm, and at the same time drives a preset 3D model to generate an animated video that matches the user action. The present invention significantly improves the real-time and immersive nature of human-computer interaction, effectively reduces the synchronization error of audio and video to the millisecond level through a hardware-accelerated parallel processing architecture, and improves the accuracy and response speed of motion capture through an intelligent motion recognition algorithm. Finally, dynamic 3D rendering technology is used to enhance visual expression. The present invention can achieve low-latency fusion output of audio and video signals, while reducing system power consumption and hardware costs, and ultimately achieve real-time, smooth and high-quality synchronous interaction between music and animation. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 A flowchart of a Bluetooth speaker intelligent interaction method provided as an exemplary embodiment of the present application.

[0066] Figure 2 A structural diagram of a Bluetooth speaker intelligent interactive device provided as an exemplary embodiment of the present application.

[0067] Figure 3 This is a structural diagram of a Bluetooth speaker intelligent interactive device provided as another exemplary embodiment of the present application. DETAILED DESCRIPTION

[0068] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments.

[0069] Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of this application.

[0070] See Figure 1 , the embodiment of the present application provides a Bluetooth speaker intelligent interaction method, the method specifically comprising the following steps:

[0071] The Bluetooth module receives the Bluetooth audio signal sent by the smart device and sends it to the embedded CPU module.

[0072] Among them, the Bluetooth module uses the standard Bluetooth protocol to achieve seamless access to audio signals, without relying on additional hardware such as external independent decoding modules. This solves the problem of resource redundancy caused by the need for external devices in traditional Bluetooth speakers and can significantly reduce system complexity. By directly inputting the Bluetooth audio signal into the embedded CPU module, it ensures that the Bluetooth audio signal enters the subsequent processing flow in a low-latency, high-fidelity form, providing a stable Bluetooth signal source for audio and video collaborative control. At the same time, it complements the camera module, enabling the system to synchronously process audio input and user action interaction, avoiding the interaction fragmentation caused by external devices and improving the overall smoothness of interaction. In addition, the built-in Bluetooth module supports mainstream Bluetooth protocol versions and is compatible with a variety of Bluetooth speaker devices. It provides a standardized data input basis for the decoding processing of the embedded CPU module and the collaborative rendering with the display module, enhancing the compatibility and scalability of the Bluetooth speaker system.

[0073] The AI ​​model module acquires the motion image signal and obtains the motion control signal; processes the motion control signal, obtains the 3D animation motion signal and sends it to the embedded CPU module.

[0074] Among them, the AI ​​model module obtains motion image signals and obtains motion control signals; by processing the motion control signals, it can also obtain animation control signals and send them to the embedded CPU module; the embedded CPU module also includes a GPU.

[0075] Among them, the AI ​​model module uses deep learning algorithms to achieve efficient analysis of motion image signals, accurately identify body movements such as waving and jumping as standardized motion control signals, and complete the conversion from physical movements to digital signals; through secondary processing, the motion control signals are mapped into 3D animation motion signals, and the user's jumping movements can drive the virtual model to synchronously execute the corresponding animation, ensuring the accurate transmission of motion interaction; the output 3D animation motion signal can be directly used in the embedded CPU module to generate real-time video signals, realizing the integration of user movements, 3D animation and audio playback. The 3D virtual image adjusts the dance posture in real time according to the music rhythm and user movements, eliminating the delay problem caused by the separate processing of motion recognition and animation rendering in traditional solutions, and significantly improving the interactive response speed and the overall coordination of the system.

[0076] The embedded CPU module receives and decodes the Bluetooth audio signal to obtain a digital audio signal; delays the digital audio signal to obtain the delayed digital audio signal and sends it to the power amplifier module; obtains a real-time video signal based on the 3D animation action signal and the preset 3D model and sends it to the display module.

[0077] In some embodiments, the Bluetooth audio signal may also be decoded by the Bluetooth module, thereby reducing the processing complexity of the embedded CPU module and ensuring lossless decoding of the Bluetooth audio signal.

[0078] The embedded CPU module receives Bluetooth audio signals through the Bluetooth protocol stack and parses the encoded audio data stream. Based on the parsed audio data stream, the embedded CPU module uses an audio decoding algorithm to convert it into a digital audio signal. Based on the decoded digital audio signal, the embedded CPU module normalizes the sampling rate and bit depth. If the standardized digital audio signal meets the input format requirements of the AI ​​model module, a copy of the signal is stored in a buffer. The embedded CPU module also retrieves the digital audio signal from the buffer and transmits it back to the AI ​​model module via an internal bus. Based on the digital audio signal transmitted to the AI ​​model module, the AI ​​model module performs feature extraction and audio enhancement. The embedded CPU module merges the processed digital audio signal with the original digital audio signal. If the merged digital audio signal meets the input specifications of the power amplifier module, the embedded CPU module delays it. The embedded CPU module also retrieves the delayed digital audio signal and sends it to the power amplifier module via a digital audio interface. Based on the 3D animation motion signal and the preset 3D model, the embedded CPU module generates a real-time video signal and sends it to the display module, ensuring synchronization between the delayed digital audio signal and the real-time video signal.

[0079] In some embodiments, the embedded CPU module receives Bluetooth audio signals via the Bluetooth protocol stack and parses the encoded audio data stream using the A2DP protocol. The data stream format is SBC encoded with a sampling rate of 44.1kHz. Based on the parsed audio data stream, the embedded CPU module uses an SBC decoding algorithm to convert it into a digital audio signal in PCM format, maintaining a sampling rate of 44.1kHz and a bit depth of 16 bits. The embedded CPU module uses a resampling algorithm to adjust the sampling rate of the decoded digital audio signal to 48kHz and a uniform bit depth of 24 bits to meet subsequent processing requirements. If the standardized digital audio signal meets the input format requirements of the AI ​​model module (i.e., a sampling rate of 48kHz and a bit depth of 24 bits), a copy of the signal is stored in a 512KB circular buffer. The embedded CPU module obtains the digital audio signal in the buffer and transmits it to the AI ​​model module via the I2S bus at a sampling rate of 48kHz. Based on the digital audio signal transmitted to the AI ​​model module, the AI ​​model module performs feature extraction using a convolutional neural network algorithm and performs audio enhancement processing using an adaptive filter, improving the signal-to-noise ratio to 90dB. The embedded CPU module performs a weighted mix of the digital audio signal processed by the AI ​​model module with the original digital audio signal at a 7:3 ratio. If the combined digital audio signal meets the input specifications of the power amplifier module (i.e., a sampling rate of 48kHz and a bit depth of 24 bits), the embedded CPU module applies a 50ms delay using an FIR filter. The embedded CPU then transmits the delayed digital audio signal to the power amplifier module via the SPI interface at a 48kHz sampling rate, completing the final output of the audio signal.

[0080] In some embodiments, the embedded CPU module acquires a digital audio signal and time-delays the signal using an internal clock to obtain a delayed digital audio signal. The embedded CPU module then encapsulates the delayed digital audio signal and uses a standard audio transmission protocol to determine the encapsulated audio data packet. If the data packet is complete, the embedded CPU module decapsulates the data packet to obtain the delayed digital audio signal.

[0081] In some embodiments, the embedded CPU module receives a digital audio signal from an audio input interface and uses a timer with an internal clock frequency of 48kHz to perform time delay processing on the signal with a delay time of 100ms to obtain a delayed digital audio signal. The embedded CPU module then encapsulates the delayed digital audio signal and uses the I2S audio transmission protocol to segment the signal into 16-bit data frames, each containing left and right channel data. The module then determines the encapsulated audio data packets. If the data packet checksum matches, the data packet is decapsulated to obtain the delayed digital audio signal.

[0082] In some embodiments, a real-time action signal and a preset 3D model are obtained, and the correspondence between the action signal and the preset 3D model is determined. According to the correspondence between the action signal and the preset 3D model, a real-time image frame is rendered and generated. A real-time video signal is obtained by rendering the generated real-time image frame. If the real-time video signal meets the format requirements of the display module, it is transmitted to the display module to determine the signal integrity. According to the real-time video signal received by the display module, a display drive signal is processed and generated. Through the display drive signal, it is output to the display screen module to obtain the screen display content. If the user selects the automatic mode, a digital audio signal is obtained and the mapping relationship with the preset 3D model is determined. According to the mapping relationship between the digital audio signal and the preset 3D model, an action sequence number is generated and sent to the AI ​​model module. Through the action sequence number, the rendering process is updated to obtain a new real-time video signal.

[0083] In some embodiments, real-time motion signals and a preset 3D model are acquired. User motion data is collected using a motion capture device at a sampling frequency of 60Hz. Combined with the skeletal binding information of the preset 3D model, an inverse kinematics algorithm is used to calculate the correspondence between the motion signals and the model. Based on the correspondence between the motion signals and the preset 3D model, an OpenGL rendering engine is used to generate real-time image frames at a rate of 30 frames per second, with a resolution of 1920x1080 per frame. The rendered real-time image frames are compressed using the H.264 encoding algorithm to generate a real-time video signal with a bit rate of 5Mbps. If the real-time video signal meets the format requirements of the display module, it is transmitted to the display module via an HDMI interface, and a CRC checksum algorithm is used to determine signal integrity. Based on the real-time video signal received by the display module, a display drive signal is generated using VESA standards with a refresh rate of 60Hz. The display drive signal is output to the display screen, and the screen brightness is controlled using PWM dimming technology to obtain the screen display content. If the user selects automatic mode, a digital audio signal is collected using a microphone at a sampling rate of 44.1kHz. The audio spectrum is analyzed using Fourier transform to determine the mapping relationship with the preset 3D model. Based on the mapping relationship between the digital audio signal and the preset 3D model, a neural network algorithm is used to generate action sequence numbers with an output frequency of 30Hz. The action sequence numbers are used to update the rendering process and recalculate the model pose to obtain a new real-time video signal.

[0084] The power amplifier module receives and processes the delayed digital audio signal and sends it to the speaker module.

[0085] Among them, the power amplifier module performs digital signal processing on the delayed digital audio signal, and determines the optimized digital audio signal through gain adjustment; converts the optimized audio signal into an analog signal, and uses a digital-to-analog converter to obtain an analog audio signal; transmits the analog audio signal to the speaker module through the audio output interface to obtain status feedback of successful transmission.

[0086] In some embodiments, the power amplifier module performs digital signal processing on the delayed digital audio signal, employing an FIR filter algorithm to adjust the signal gain to +6dB, and thereby determines an optimized digital audio signal. The power amplifier module converts the optimized audio signal into an analog signal using a 24-bit digital-to-analog converter with a sampling rate of 96kHz to obtain the analog audio signal. The power amplifier module transmits the analog audio signal to the speaker module via an RCA audio output interface, and obtains status feedback indicating successful transmission.

[0087] The speaker module receives and plays the delayed digital audio signal.

[0088] Among them, the speaker module receives the analog audio signal, converts the signal into electroacoustic according to the characteristics of the driving unit, and obtains a playable sound wave signal; the sound wave signal is output through physical vibration, and the final audio playback effect is determined through the movement of the diaphragm.

[0089] In some embodiments, the speaker module receives an analog audio signal and, based on the characteristic of the driving unit's impedance of 8Ω, performs electroacoustic conversion on the signal to obtain a playable sound wave signal; the speaker module outputs the sound wave signal through physical vibration, and determines the final audio playback effect through the diaphragm movement frequency range of 20Hz to 20kHz.

[0090] The display module processes the real-time video signal, obtains the real-time action picture of the 3D animation model and sends it to the display screen module.

[0091] Among them, the display module quickly converts real-time video signals into real-time action pictures based on real-time rendering technology. The dance movements of the 3D virtual image can accurately restore the user's movements frame by frame, and ensure the smoothness of the movements and the continuity of the pictures; through high-precision data processing, the lossless conversion of real-time video signals to real-time action pictures is ensured, and complex limb movement details can be clearly presented, which improves the visual realism; the direct data interaction mechanism between the embedded CPU module and the display module realizes low-latency output of 3D pictures; after detecting the user's interactive action, the response picture of the 3D virtual image can be displayed immediately, avoiding the jamming problem caused by the need for multi-level signal transfer in traditional solutions; at the same time, the real-time action picture is strictly synchronized with the delayed digital audio signal, further enhancing the immersive feeling of using Bluetooth speakers.

[0092] The audio signal is received via a Bluetooth module, and the rhythm and music type in the audio data stream are analyzed to obtain a music feature signal. Based on the music feature signal, an internal algorithm is used to map the rhythm and type to a preset action sequence number to determine the 3D animation action signal. The embedded CPU module processes the correspondence between the 3D animation action signal and the preset 3D model to generate a real-time video signal. If the real-time video signal meets the preset action threshold, the locally deployed AI model module optimizes the smoothness of the action signal to obtain an optimized real-time video signal. Based on the optimized real-time video signal, the display module is used to render the real-time video signal in real time to generate a real-time action picture. The camera module collects the user's body movement video stream, analyzes the action features in the video stream, and obtains the action image signal. The display module is used to generate the real-time action picture. The display screen module receives the real-time action picture and presents the real-time action picture of the 3D animation model to obtain the final display picture.

[0093] The display screen module receives and displays real-time action images.

[0094] Among them, the display screen module presents the real-time action pictures of the dynamic 3D animation model directly on the terminal screen based on the high-response display technology, ensuring that the action details of the virtual image are completely synchronized with the action image signal. The user's waving or rotating operations can instantly trigger the virtual image's precise action feedback. The virtual image's dance posture changes or expression adjustments can seamlessly connect with the audio rhythm and maintain the consistency of audio and picture collaboration. In addition, the display screen module replaces the redundant configuration of the traditional Bluetooth speaker that requires an external screen with an integrated design. The dynamic rendering and display functions of the 3D virtual image are concentrated within the system, which not only reduces the resource consumption of multi-device collaboration, but also avoids signal attenuation or compatibility issues caused by external display devices, and ultimately achieves a stable, smooth and highly immersive interactive visual output effect, enhances the interactivity of the Bluetooth speaker, and improves the user experience.

[0095] In some embodiments, a wifi module is also included for high-speed communication with the server; the wifi module can realize wireless high-speed communication between the device and the server, getting rid of the limitations of physical cables, and can greatly improve the flexibility of deployment and the convenience of remote control.

[0096] A Bluetooth speaker intelligent interaction method provided in the above embodiment can realize the intelligent synchronous presentation of audio signals and dynamic vision through a collaborative working system integrating a Bluetooth module, a camera module, an AI model module and an embedded CPU module. The present invention uses a Bluetooth module to receive audio signals from an external smart device, optionally combines a camera to capture user actions in real time and generates three-dimensional motion control signals through AI model analysis. The embedded CPU module realizes precise synchronous processing of digital audio signals through an audio delay algorithm, and at the same time drives a preset 3D model to generate an animated video that matches the user action. The present invention significantly improves the real-time and immersive nature of human-computer interaction, effectively reduces the synchronization error of audio and video to the millisecond level through a hardware-accelerated parallel processing architecture, and improves the accuracy and response speed of motion capture through an intelligent motion recognition algorithm. Finally, dynamic 3D rendering technology is used to enhance visual expression. The present invention can achieve low-latency fusion output of audio and video signals, while reducing system power consumption and hardware costs, and ultimately achieve real-time, smooth and high-quality synchronous interaction between music and animation.

[0097] In some embodiments, the method further comprises:

[0098] The embedded CPU module analyzes the digital audio signal in automatic mode, obtains the action sequence number and sends it to the AI ​​model module; the AI ​​model module obtains the 3D animation action signal according to the action sequence number and sends it to the embedded CPU module; alternatively, the camera module collects the user's action image signal in camera mode and sends it to the AI ​​model module; the AI ​​model module receives and processes the action image signal to obtain the action control signal; processes the action control signal to obtain the 3D animation action signal and sends it to the embedded CPU module.

[0099] The camera module dynamically captures user body movements based on visual perception technology, converting interactive actions such as waving and nodding into high-precision motion image signals, providing accurate raw data input for subsequent analysis. By sending the user's motion image signals directly to the AI ​​model module, rapid parsing and processing of motion data is achieved. For example, a user's waving motion can be accurately identified and converted into a waving motion image signal, ultimately driving the 3D virtual image to respond in real time. At the same time, this module works in conjunction with other modules in the system to ensure the adaptation of user motion interaction, audio playback, and 3D animation presentation, avoiding system delays caused by external motion capture devices in traditional solutions and significantly improving the device's integration.

[0100] The camera module captures the user's motion image signals and obtains raw video data. The raw video data is transmitted to the locally configured AI model module to obtain a sequence of images to be processed. The AI ​​model module preprocesses the image sequence to obtain standardized image features. Based on the standardized image features, the AI ​​model module executes an action recognition algorithm to determine the user's motion type. The AI ​​model module maps the motion type to the motion sequence number to obtain a 3D animation motion signal. The embedded CPU module receives the 3D animation motion signal and obtains the corresponding 3D animation motion template. Using the 3D animation motion template, the embedded CPU module obtains a real-time video signal. The display module receives the real-time video signal and determines whether the image signal meets the display requirements to obtain a real-time motion image.

[0101] Specifically, in some embodiments, the camera module captures user motion image signals at a frame rate of 30 fps and compresses the raw video data using H264 encoding to reduce transmission bandwidth usage. The raw video data is transmitted to the AI ​​model module via a USB 3.0 interface and decoded by the OpenCV library into an RGB image sequence to be processed. The AI ​​model module normalizes the image sequence, adjusting the resolution to 640×480 pixels, and uses the MediaPipe pose estimation algorithm to extract the coordinates of human joints to obtain standardized image features. Based on the standardized image features, the AI ​​model module runs an LSTM-based motion classification algorithm and determines, with a 95% confidence threshold, whether the user's motion type belongs to a preset dance motion library. If the motion type matches the "waving" category in the library, the AI ​​model module queries the motion mapping table and outputs the corresponding 3D model motion number M207 as the motion sequence number. The embedded CPU module receives the M207 signal, loads a pre-set FBX format 3D animation template from the Flash memory, and parses the skeletal animation keyframe data. Using an interpolation algorithm to calculate the skeletal transformation matrix for each frame, the embedded CPU module generates a 60fps 3D animation motion signal and outputs real-time motion data including vertex coordinates. The display module invokes the Unity3D engine to render images using a cartoon-style 3D model shader, applying the Phong lighting model to generate a 1280×720 resolution real-time motion image. The real-time motion image is transmitted to the display module via an HDMI interface, where a CRC check is performed to verify data integrity, ultimately displaying the dynamic 3D character waving animation.

[0102] Among them, the embedded CPU module can also analyze digital audio signals in automatic mode, obtain the BPM of the music and send it to the AI ​​model module for processing, thereby adjusting the movement speed and rhythm in the 3D animation movement signal.

[0103] In automatic mode, the Bluetooth speaker receives digital audio signals via the Bluetooth module, parses the signal data packets, and obtains audio waveform data. Based on the audio waveform data, a Fourier transform algorithm is used to extract frequency features and determine the music rhythm characteristics. A pre-trained music classification model is used to analyze the frequency features, determine the music genre, and obtain a music genre number. Based on the music rhythm characteristics and the music genre number, a preset rhythm-action mapping table is queried to obtain the corresponding action sequence number. The action sequence number is loaded via a locally configured AI model module to generate an initial 3D animation action signal. If the timestamp of the initial 3D animation action signal does not match the rhythm characteristics, the 3D animation action signal is adjusted using a timeline synchronization algorithm to obtain a calibrated 3D animation action signal. Based on the calibrated 3D animation action signal, a 3D rendering engine is used to transform the preset 3D model into a rendered animation frame sequence. The animation frame sequence is loaded via the display module and real-time image rendering is performed to display the 3D animation on the screen. Based on user interaction settings, real-time user feedback data is obtained to update the rhythm-action mapping table and obtain an optimized action mapping relationship.

[0104] Specifically, after receiving the digital audio signal, the Bluetooth module uses an AAC decoder to parse the data packet and extract 16-bit, 44.1kHz PCM audio waveform data. A 1024-point fast Fourier transform is applied to the audio waveform data, calculating the spectral energy of each frame. The beat intervals in the 20-200Hz low-frequency range are detected, and a rhythmic feature with a BPM value of 128 is determined. The pre-trained music classification model uses a convolutional neural network structure, taking a Mel-frequency spectrogram as input and outputting a probability distribution. When the probability of the electronic dance music type reaches 85%, the type code EDM_003 is assigned. When querying the rhythm-motion mapping table, using BPM = 128 and EDM_003 as indexes, a match is found for the robotic dance motion sequence numbered DANCE_SEQ_47. After loading the motion sequence, the locally configured AI model module generates an initial 3D animation motion signal containing 23 joint rotation angles, with a timestamp interval of 468 milliseconds. When a 50-millisecond tempo deviation is detected between the 3D animation motion signal and the digital audio signal, a dynamic time warping algorithm is used to stretch the keyframes to fully align the tempo of the 3D animation motion signal with the digital audio signal. The 3D rendering engine reads the calibrated 3D animation motion signal and drives the skeletal animation system in Unity, generating an animation frame sequence with a 1024×768 resolution at 60fps. The display module receives frame data via the MIPI interface and outputs it to a 7-inch LCD screen using a hardware-accelerated renderer. If the system records that a user has remained on DANCE_SEQ_47 for more than 3 minutes, the weight coefficient for that sequence in the mapping table is increased by 0.15.

[0105] In some embodiments, the method further comprises:

[0106] The AI ​​model module sends motion control signals, or music rhythm and type features to the cloud server.

[0107] The cloud server inputs the motion control signal, or music rhythm and type features into the cloud AI model, obtains a 3D animation motion video stream and sends it to the smart device via a wireless network.

[0108] Among them, when the cloud AI model is adopted, the Bluetooth speaker receives the digital audio signal sent by the smart device through the Bluetooth WiFi module and parses the audio data in the signal; based on the parsed audio data, the rhythm characteristics and type information of the music are extracted to generate a rhythm signal and a music type number; if the rhythm signal and the music type number meet the preset conditions, the extracted rhythm signal and the music type number are transmitted to the cloud server via WiFi, and the server receives the confirmation; the cloud server calls the online deployed AI model according to the received rhythm signal and the music type number, and generates the corresponding 3D model action control signal; through the rendering module of the cloud server, the preset 3D model and the 3D model action control signal are used to render and generate a 3D animation video stream; if the video stream is successfully generated, the 3D animation video stream is transmitted to the display module of the smart device through the network to obtain a transmission completion status; the display module receives the 3D animation video stream, decodes the video stream data, and generates a playable image signal; based on the decoded image signal, the display screen is driven to play the 3D animation video stream to obtain a playback status; the user's body motion data is obtained through the camera module, converted into a motion control signal, transmitted to the cloud server, the 3D model motion is updated, and a new video stream is generated.

[0109] Specifically, the Bluetooth Wi-Fi module uses SBC encoding to parse the digital audio signal transmitted by the smart device, extracting a PCM data stream with a 16kHz sampling rate. The audio frames are analyzed using a Mel-frequency cepstral coefficient algorithm, calculating the energy spectrum and zero-crossing rate of each 100ms frame. This generates a rhythm signal with a BPM of 120, and a convolutional neural network classifier outputs a music genre number of 3 (pop music). If the BPM value is within the range of 80-160 and the music genre number is valid, a JSON packet containing the BPM and music genre number is uploaded to the cloud server via the 802.11ac protocol. The server returns an HTTP 200 status code confirming receipt. The LSTM action generation model deployed in the cloud outputs an action sequence consisting of 12 joint rotation angles, with a frame interval of 33ms, based on the parameters of BPM = 120 and music genre number 3. The rendering engine loads a preset 3D character model in FBX format and uses the Unity HDRP pipeline to render a 2560x1440 resolution video stream at 30fps. After H264 encoding, it generates a TS stream with a bitrate of 2Mbps. The transmission module uses a TCP retransmission mechanism to ensure the video stream reaches the smart device intact. The display module's hardware decoder parses the I and P frames in the TS stream and outputs RGB888-formatted images to the frame buffer. The screen driver IC synchronizes the image output at a 75Hz refresh rate. Simultaneously, the OV5640 camera captures the user's skeletal joints at 30fps. The OpenPose algorithm identifies the coordinates of 20 key points and, after Kalman filtering, generates new motion control signals for upload to the server.

[0110] In some embodiments, the method further comprises:

[0111] The embedded CPU module stores 3D animation motion signals and real-time video signals to obtain a motion template database.

[0112] Among them, the embedded CPU module obtains the real-time video signal and the corresponding 3D animation action signal, stores them in the local storage unit, and forms an initial action template data set. If the storage capacity of the initial action template data set exceeds the preset threshold, the action template data set is compressed using a data compression algorithm to obtain a compressed action template database. Based on the compressed action template database, the newly input music rhythm and type signal is obtained, and the similarity with the action template in the database is determined by a fast matching algorithm to determine the optimal matching template. The optimal matching template is fine-tuned by the embedded CPU module to generate an optimized 3D animation action signal. The optimized 3D animation action signal is re-rendered using the display module to obtain the final real-time action picture.

[0113] Specifically, the embedded CPU module encapsulates the real-time video signal and the corresponding 3D animation action signal in JSON format and writes it into the SQLite database to form an initial data set containing 500 groups of samples. When the data volume reaches the 1GB threshold, the Zstandard compression algorithm is started to compress the data to 30% of the original volume. The newly input music features are matched with the action template database through the cosine similarity algorithm to screen candidate templates with a similarity higher than 0.85. The embedded CPU calls the gradient descent algorithm to fine-tune the joint parameters of the candidate templates, and the error tolerance is controlled at ±5°. Finally, the rendering engine mixes the adjusted bone data with the texture map to output a real-time action picture with light and shadow effects.

[0114] In some embodiments, the method further comprises:

[0115] After receiving the digital audio signal, the embedded CPU module samples the digital audio signal according to the resampling algorithm and performs delay processing on the sampled digital audio signal; the sampling rate of the resampling algorithm is 48kHz and the bit depth is 24bit.

[0116] Among them, the embedded CPU module uses a resampling algorithm to uniformly adjust the sampling rate of the digital audio signal to 48kHz and fix the bit depth to 24bit, effectively improving the standardized processing capability of the digital audio signal. The 48kHz sampling rate ensures that the digital audio signal meets the general standards of mainstream audio and video equipment, enhances compatibility and avoids signal distortion caused by sampling rate differences; the 24-bit bit depth expands the audio dynamic range, significantly reduces quantization noise and improves detail restoration capabilities, making low-frequency response more accurate and high-frequency details richer. The digital audio signal with unified parameters reduces the processing complexity of the subsequent power amplifier module, reduces time domain aliasing and phase distortion, optimizes system resource usage, and provides a standardized data basis for high-fidelity audio output, further ensuring the millisecond-level synchronization accuracy of the digital audio signal and the 3D animation motion signal. In some embodiments, the sampling rate of the resampling algorithm can also be 44100Hz and the bit depth can be 16bit.

[0117] In some embodiments, the method further comprises:

[0118] The embedded CPU module receives and decodes the Bluetooth audio signal, obtains the digital audio signal and sends it to the AI ​​model module;

[0119] The AI ​​model module uses a convolutional neural network algorithm to extract features from digital audio signals and send them to the adaptive filter.

[0120] The adaptive filter performs audio enhancement processing on the digital audio signal after feature extraction, improves the signal-to-noise ratio of the enhanced digital audio signal to 90dB and sends it to the embedded CPU module.

[0121] Among them, the AI ​​model module uses a convolutional neural network algorithm to perform high-precision feature extraction on the digital audio signal, and combines it with an adaptive filter to perform audio enhancement processing on the signal, which increases the signal-to-noise ratio to more than 90dB, significantly improving the audio clarity and fidelity, thereby effectively suppressing background noise and interference signals, enhancing the spectral feature recognition capabilities of human voices and musical instruments, making the low-frequency response more stable and the high-frequency details fuller, while optimizing the accuracy of voice command recognition. The improved signal-to-noise ratio reduces signal distortion in audio transmission, ensures the synchronization and coordination of high dynamic range audio and 3D animation movements, reduces the computing load of the embedded CPU module, improves the overall energy efficiency of the system, and is suitable for immersive interactive scenarios in complex acoustic environments. In some embodiments, the signal-to-noise ratio of the enhanced digital audio signal can also be improved according to the needs of the interactive scenario, greatly improving the adaptability to the interactive scenario.

[0122] In some embodiments, the method further comprises:

[0123] The embedded CPU module performs weighted mixing on the enhanced digital audio signal and the original digital audio signal to obtain a weighted digital audio signal and performs delay processing to obtain a delayed digital audio signal.

[0124] The weighted mixing ratio is 7:3. The embedded CPU module achieves a balance between sound quality optimization and natural listening experience by mixing the enhanced digital audio signal with the original signal in a 7:3 weighted ratio. The enhanced signal accounts for 70%, which can fully preserve the high signal-to-noise ratio characteristics after noise reduction and dynamic range improvement, and enhance the clarity of human voices and key frequency bands; the original signal accounts for 30%, maintaining the original harmonic characteristics and spatial sense of the audio, avoiding sound field compression or detail loss caused by over-processing. This mixing strategy takes into account the needs of high-fidelity restoration and sound enhancement, effectively improving voice intelligibility and music layering, reducing auditory fatigue, and simplifying the complexity of the dynamic adjustment algorithm through preset ratios, reducing the real-time computing overhead of the embedded system, ensuring high-precision synchronization with 3D animation motion signals and overall system smoothness.

[0125] In some embodiments, the embedded CPU module delays the digital audio signal, including:

[0126] The embedded CPU module uses a timer with an internal clock frequency of 48kHz to perform time delay processing on the digital audio signal. The delay time is 100ms to obtain the delayed digital audio signal.

[0127] Among them, the embedded CPU module uses a timer with an internal clock frequency of 48kHz to apply a 100ms time delay to the digital audio signal, accurately matching the corresponding relationship between the audio sampling rate and the clock cycle, and effectively eliminating the accumulation of timing deviations. The 48kHz clock frequency is synchronized with the sampling rate of the original audio signal to avoid signal phase jitter or spectrum distortion caused by clock frequency mismatch, ensuring that the delay calculation error is less than 0.02ms; the 100ms delay duration adapts to the inherent processing cycle of 3D animation rendering and display in human-computer interaction scenarios, and achieves frame-level alignment of audio output and action images through a pre-compensation mechanism. While maintaining the integrity of the digital audio signal, it significantly reduces the audio and video synchronization error to a visually indistinguishable level, simplifies the multi-module collaborative scheduling logic, and improves the system's real-time response efficiency. In some embodiments, the timer can also select different internal clock frequencies according to the complexity of the digital audio signal, and perform time delay processing on the digital audio signal according to actual conditions, so that the audio and video synchronization error is eliminated, improving the Bluetooth speaker experience in various scenarios.

[0128] In some embodiments, the power amplifier module receives and processes the delayed digital audio signal, including:

[0129] The power amplifier module uses an FIR filter algorithm to adjust the signal gain of the delayed digital audio signal to +6dB.

[0130] Among them, the power amplifier module uses the FIR filter algorithm to adjust the gain of the delayed digital audio signal, accurately increasing the signal amplitude to +6dB, and optimizing the dynamic response and energy distribution of the audio output. The linear phase characteristics of the FIR filter effectively avoid phase distortion or harmonic distortion introduced by the gain adjustment, ensuring the balance between high-frequency details and low-frequency thickness; the +6dB gain amount adapts to the sensitivity threshold of the speaker module, while improving the overall loudness and avoiding signal clipping, ensuring a smooth transition of the sound pressure level. This processing strategy simultaneously enhances the clarity of speech and the layering of music, reduces transient intermodulation distortion in large dynamic audio scenes, and deeply cooperates with the preset 48kHz clock frequency of the embedded CPU module to further reduce audio and video synchronization errors, providing a low-noise, high-stability signal driving foundation for high-fidelity audio playback.

[0131] In some embodiments, the embedded CPU module analyzes the digital audio signal in automatic mode to obtain an action sequence number, including:

[0132] The embedded CPU module analyzes the signal data packet according to the digital audio signal to obtain the audio waveform data.

[0133] Based on the audio waveform data, the Fourier transform algorithm is used to extract frequency features and determine the music rhythm characteristics.

[0134] The frequency characteristics are analyzed through the pre-trained music classification model to determine the music type and obtain the music type number.

[0135] According to the music rhythm characteristics and music type number, the preset rhythm-action mapping table is queried to obtain the corresponding action sequence number and send it to the AI ​​model module.

[0136] Spectral feature analysis improves the accuracy and real-time performance of music rhythm recognition, pre-trained models enhance the generalization of music genre classification, and a rhythm-action mapping mechanism ensures the alignment of individual signals, significantly reducing manual arrangement costs. The synchronization process reduces phase deviation between action and audio, supports intelligent action adaptation in various music scenarios, enhances the dynamic expressiveness of virtual characters and 3D models, and improves user interaction immersion, while also optimizing the computing efficiency and resource usage of embedded systems.

[0137] In some embodiments, analyzing frequency features using a pre-trained music classification model to determine the music type and obtain a music type number includes:

[0138] The pre-trained music classification model uses a convolutional neural network structure to input frequency features into the Mel-spectrogram and perform calculations to obtain a probability distribution; the music type number is determined based on the probability distribution.

[0139] In some embodiments, when the music rhythm feature is BPM=128 and the music type number is EDM_003, the corresponding action sequence number in the preset rhythm-action mapping table is DANCE_SEQ_47, and the action sequence is a robotic dance action sequence.

[0140] Among them, through quantitative classification and parameterized matching, the dynamic fit between 3D animation motion signals and music content is improved, the random deviation of motion choreography is reduced, the expressiveness of virtual characters and the immersiveness of scenes are enhanced, and the real-time computing overhead of embedded systems is reduced. It is suitable for interactive entertainment or performance scenarios with high real-time requirements.

[0141] In some embodiments, the camera module collects a user's motion image signal in camera mode, including:

[0142] The camera module collects the user's original video data at a frame rate of 30fps and compresses it using H264 encoding to obtain motion image signals.

[0143] Among them, the camera module reduces data transmission bandwidth while ensuring the continuity of motion capture; the 30fps frame rate balances real-time performance and hardware resource consumption, avoiding image blur caused by high-speed motion; H.264 encoding reduces the video stream volume by more than 50% with a high compression rate, reducing the storage and transmission load of the embedded CPU module, while retaining key action details, improving the feature extraction accuracy and processing efficiency of the AI ​​model module, and ensuring low-latency action signal generation.

[0144] In some embodiments, the display module processes the real-time video signal to obtain a real-time action picture of the 3D animation model, including:

[0145] The display module uses bilinear interpolation to eliminate screen tearing and obtain a real-time video signal at 60fps.

[0146] The display module uses pixel-weighted average smoothing to enhance image edge continuity, reducing ghosting and lag in motion scenes. A high refresh rate of 60fps adapts to the persistence of human vision, enhancing the smoothness and realism of 3D animation. Bilinear interpolation achieves frame insertion with low computational complexity, avoiding GPU overload and ensuring efficient coordination between the display module and the embedded CPU, reducing audio and video synchronization deviation to less than 10ms.

[0147] In some embodiments, the display screen module receives and displays the real-time action image, including:

[0148] The display screen module receives real-time action images through the MIPI interface and uses PWM dimming to control the deflection of liquid crystal molecules, presenting real-time action images with a dynamic range of 1000:1 within a 5ms response time.

[0149] The display screen module receives real-time action images through the MIPI interface, and combines PWM dimming technology to control the liquid crystal molecules to complete the deflection response within 5ms, achieving a high dynamic range display of 1000:1. The MIPI interface ensures lossless transmission of video streams with its low power consumption and high bandwidth characteristics; the 5ms fast response eliminates the ghosting phenomenon of high-speed action images, and the 1000:1 dynamic range enhances dark field details and bright area levels, allowing the rapid switching of mechanical dance movements and the precise presentation of light and shadow effects, enhancing the user's visual immersion and the real-time nature of interactive feedback. In some embodiments, the display screen module can also receive real-time action images through other interfaces such as the LVDS interface or RGB interface. It can be compatible with signal sources of different formats, ensure the stability of high-speed transmission, and broaden the device access capabilities of the Bluetooth speaker, thereby flexibly adapting to a variety of interactive environments.

[0150] See Figure 2 Another embodiment of the present application provides a Bluetooth speaker intelligent interactive device, the device comprising:

[0151] The Bluetooth module is used to receive Bluetooth audio signals sent by the smart device and send them to the embedded CPU module.

[0152] The AI ​​model module is used to acquire motion image signals and obtain motion control signals; process the motion control signals, obtain 3D animation motion signals and send them to the embedded CPU module.

[0153] The embedded CPU module is used to receive and decode Bluetooth audio signals to obtain digital audio signals; delay the digital audio signals to obtain delayed digital audio signals and send them to the power amplifier module; obtain real-time video signals based on 3D animation action signals and preset 3D models and send them to the display module.

[0154] The power amplifier module is used to receive and process the delayed digital audio signal and send it to the speaker module.

[0155] The speaker module is used to receive and play delayed digital audio signals.

[0156] The display module is used to process the real-time video signal, obtain the real-time action picture of the 3D animation model and send it to the display screen module.

[0157] The display screen module is used to receive and display real-time action images.

[0158] See Figure 3 In some embodiments, the embedded CPU module is further configured to:

[0159] In automatic mode, the digital audio signal is analyzed, the action sequence number is obtained and sent to the AI ​​model module.

[0160] In some embodiments, the AI ​​model module is further configured to:

[0161] The 3D animation action signal is obtained according to the action sequence number and sent to the embedded CPU module.

[0162] In some embodiments, the embedded CPU module is further configured to:

[0163] The 3D animation action signals and real-time video signals are stored to obtain an action template database.

[0164] In some embodiments, the embedded CPU module is further configured to:

[0165] According to the resampling algorithm, the sampling rate of the digital audio signal is adjusted to 48kHz and the bit depth is unified to 24bit.

[0166] The specific limitations of the Bluetooth speaker intelligent interactive device provided in this embodiment can be found in the embodiment of the Bluetooth speaker intelligent interactive method described above and will not be repeated here. The various modules in the above-mentioned Bluetooth speaker intelligent interactive device can be implemented in whole or in part through software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of the above modules.

[0167] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0168] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A Bluetooth speaker intelligent interaction method, characterized in that: The method comprises: The Bluetooth module receives the Bluetooth audio signal sent by the smart device and sends it to the embedded CPU module; The AI ​​model module acquires the motion image signal and obtains the motion control signal; processes the motion control signal to obtain the 3D animation motion signal and sends it to the embedded CPU module; The embedded CPU module receives and decodes the Bluetooth audio signal to obtain a digital audio signal; delays the digital audio signal to obtain a delayed digital audio signal and sends it to the power amplifier module; obtains a real-time video signal based on the 3D animation action signal and the preset 3D model and sends it to the display module; The power amplifier module receives and processes the delayed digital audio signal and sends it to the speaker module; The speaker module receives and plays the delayed digital audio signal; The display module processes the real-time video signal to obtain a real-time action picture of the 3D animation model and sends it to the display screen module; The display screen module receives and displays the real-time action picture.

2. The Bluetooth speaker intelligent interaction method according to claim 1, characterized in that: The method further comprises: The embedded CPU module analyzes the digital audio signal in the automatic mode, obtains the action sequence number and sends it to the AI ​​model module; the AI ​​model module obtains the 3D animation action signal according to the action sequence number and sends it to the embedded CPU module; Alternatively, the camera module collects the user's motion image signal in camera mode and sends it to the AI ​​model module; the AI ​​model module receives and processes the motion image signal to obtain a motion control signal; processes the motion control signal to obtain a 3D animation motion signal and sends it to the embedded CPU module.

3. The Bluetooth speaker intelligent interaction method according to claim 2, characterized in that: The method further comprises: The AI ​​model module sends the action control signal, or the music rhythm and the type feature to a cloud server; The cloud server inputs the motion control signal, or the music rhythm and the type feature into the cloud AI model, obtains a 3D animation motion video stream and sends it to the smart device via a wireless network.

4. The Bluetooth speaker intelligent interaction method according to claim 1, characterized in that: The method further comprises: The embedded CPU module stores 3D animation motion signals and real-time video signals to obtain a motion template database.

5. The Bluetooth speaker intelligent interaction method according to claim 1, characterized in that: The method further comprises: After obtaining the digital audio signal, the embedded CPU module samples the digital audio signal according to a resampling algorithm and performs delay processing on the sampled digital audio signal; the sampling rate of the resampling algorithm is 48kHz and the bit depth is 24bit.

6. The Bluetooth speaker intelligent interaction method according to claim 1, characterized in that: The method further comprises: The embedded CPU module receives and decodes the Bluetooth audio signal, obtains a digital audio signal and sends it to the AI ​​model module; The AI ​​model module uses a convolutional neural network algorithm to extract features from the digital audio signal and sends the feature to the adaptive filter; The adaptive filter performs audio enhancement processing on the digital audio signal after feature extraction, improves the signal-to-noise ratio of the enhanced digital audio signal to 90dB, and sends the enhanced signal to the embedded CPU module.

7. The Bluetooth speaker intelligent interaction method according to claim 6, characterized in that: The method further comprises: The embedded CPU module performs weighted mixing on the enhanced digital audio signal and the original digital audio signal to obtain a weighted digital audio signal and performs delay processing to obtain a delayed digital audio signal; The weighted mixing ratio is 7:

3.

8. The Bluetooth speaker intelligent interaction method according to claim 1, characterized in that: The embedded CPU module delays the digital audio signal, comprising: The embedded CPU module uses a timer with an internal clock frequency of 48kHz to perform time delay processing on the digital audio signal, with a delay time of 100ms to obtain a delayed digital audio signal.

9. The Bluetooth speaker intelligent interaction method according to claim 1, characterized in that: The power amplifier module receives and processes the delayed digital audio signal, including: The power amplifier module uses an FIR filter algorithm to adjust the signal gain of the delayed digital audio signal to +6dB.

10. The Bluetooth speaker intelligent interaction method according to claim 2, characterized in that: The embedded CPU module analyzes the digital audio signal in the automatic mode to obtain an action sequence number, including: The embedded CPU module analyzes the signal data packet according to the digital audio signal to obtain audio waveform data; According to the audio waveform data, the Fourier transform algorithm is used to extract the frequency characteristics and determine the music rhythm characteristics; Analyze frequency features through the pre-trained music classification model, determine the music type, and obtain the music type number; According to the music rhythm characteristics and the music type number, the preset rhythm-action mapping table is queried to obtain the corresponding action sequence number and send it to the AI ​​model module.

11. The Bluetooth speaker intelligent interaction method according to claim 10, characterized in that: The method of analyzing frequency features through a pre-trained music classification model, determining the music type, and obtaining a music type number includes: The pre-trained music classification model adopts a convolutional neural network structure, inputs frequency features into a mel-spectrogram and performs calculations to obtain a probability distribution; and determines the music type number based on the probability distribution.

12. The Bluetooth speaker intelligent interaction method according to claim 11, characterized in that: When the music rhythm feature is BPM=128 and the music type number is EDM_003, the corresponding action sequence number in the preset rhythm-action mapping table is DANCE_SEQ_47, and the action sequence is a robotic dance action sequence.

13. The Bluetooth speaker intelligent interaction method according to claim 2, characterized in that: The camera module collects the user's motion image signal in the camera mode, including: The camera module collects the user's original video data at a frame rate of 30fps and compresses it using H264 encoding to obtain the action image signal.

14. The Bluetooth speaker intelligent interaction method according to claim 1, characterized in that: The display module processes the real-time video signal to obtain a real-time action picture of the 3D animation model, including: The display module uses bilinear interpolation to eliminate screen tearing and obtain a real-time video signal at 60fps.

15. The Bluetooth speaker intelligent interaction method according to claim 1, characterized in that: The display screen module receives and displays the real-time action picture, including: The display screen module receives real-time action images through the MIPI interface and uses PWM dimming to control the deflection of liquid crystal molecules, presenting real-time action images with a dynamic range of 1000:1 within a 5ms response time.

16. A Bluetooth speaker intelligent interactive device, characterized in that: The device comprises: A Bluetooth module is used to receive Bluetooth audio signals sent by smart devices and send them to the embedded CPU module; The AI ​​model module is used to acquire the motion image signal and obtain the motion control signal; process the motion control signal to obtain the 3D animation motion signal and send it to the embedded CPU module; An embedded CPU module is configured to receive and decode the Bluetooth audio signal to obtain a digital audio signal; delay the digital audio signal to obtain a delayed digital audio signal and send it to the power amplifier module; obtain a real-time video signal based on the 3D animation action signal and the preset 3D model and send it to the display module; A power amplifier module, configured to receive and process the delayed digital audio signal and send it to the speaker module; A speaker module, configured to receive and play the delayed digital audio signal; A display module is used to process the real-time video signal, obtain a real-time action picture of the 3D animation model and send it to a display screen module; The display screen module is used to receive and display the real-time action picture.

17. The Bluetooth speaker intelligent interactive device according to claim 16, characterized in that: The embedded CPU module is also used for: In automatic mode, the digital audio signal is analyzed, the action sequence number is obtained and sent to the AI ​​model module.

18. The Bluetooth speaker intelligent interactive device according to claim 16, characterized in that: The AI ​​model module is also used to: A 3D animation action signal is obtained according to the action sequence number and sent to the embedded CPU module.

19. The Bluetooth speaker intelligent interactive device according to claim 16, characterized in that: The embedded CPU module is also used for: The 3D animation action signals and real-time video signals are stored to obtain an action template database.

20. The Bluetooth speaker intelligent interactive device according to claim 16, characterized in that: The embedded CPU module is also used for: According to the resampling algorithm, the sampling rate of the digital audio signal is adjusted to 48kHz and the bit depth is unified to 24bit.