Sound system with AI (Artificial Intelligence) identification regulation and control function
By collecting data from dancers in real time through optical lenses and AI recognition modules, and automatically adjusting the volume and audio track, the system solves the problems of single volume adjustment and audio track linkage in existing smart speakers, thus improving the dance experience and auditory adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-11
- Publication Date
- 2026-04-07
AI Technical Summary
Existing smart speakers have limited volume adjustment methods, fixed sound direction, and cannot link audio tracks with dance movements. They also lack accurate recognition of dancers' movements and distances, resulting in auditory discomfort and a poor dance experience.
Using an optical lens in conjunction with an AI recognition module, the system collects real-time data on the distance, position, and movement characteristics of dancers. The main control module adjusts the overall and individual power of multiple speakers, and the audio track correction module automatically adjusts the volume and audio track to ensure that the speakers match the dance movements.
It achieves precise perception of the dancer's status, automatically adjusts the volume and audio track, improves auditory adaptability and dance coordination, reduces power consumption waste, and supports a variety of usage scenarios.
Smart Images

Figure CN121815175A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart speaker technology, and in particular to a sound system with AI recognition and control functions. Background Technology
[0002] Smart speakers, as intelligent terminal devices that integrate voice interaction, music playback, and intelligent control functions, have been widely used in various scenarios such as homes, gyms, and dance studios.
[0003] In dance practice or entertainment scenarios, users often use smart speakers to play music to accompany their dancing. However, existing smart speakers have the following shortcomings: the volume adjustment of existing smart speakers mostly relies on manual operation by the user, such as button presses, voice commands, or fixed volume modes. They cannot automatically adjust the volume according to the distance between the dancer and the speaker. When the person is close to the speaker, the volume is too loud, which can cause auditory discomfort. When the person is far away from the speaker, the volume is too low, which affects the dance experience. At the same time, most existing smart speakers are mono or fixed multi-channel sound, which cannot adjust the volume of speakers in different directions according to the real-time position of the dancer. When the person moves around the speaker while dancing, they cannot get a uniform and suitable auditory experience, and there are problems such as sound deviation and volume imbalance. Existing smart speakers play music tracks that are fixed and cannot be adjusted in real time according to the dancer's movements. When the rhythm and amplitude of the dancer's movements change, the music and dance movements are prone to become disconnected, affecting the coordination and experience of the dance. Although some existing smart devices have integrated optical recognition or AI recognition functions, they are mostly used for facial recognition and simple motion control. They cannot accurately identify the distance, position and complex movement characteristics of the dancer, and they cannot link the recognition results with the speaker's volume and track adjustment.
[0004] Therefore, there is an urgent need for a smart speaker that can use an optical lens in conjunction with AI recognition to accurately identify distance and movement, and automatically adjust the volume and adapt the audio track to solve the above-mentioned problems in existing technologies. Summary of the Invention
[0005] The technical problem to be solved by this invention is that existing smart speakers have a single volume adjustment method, fixed sound direction, inability to link audio tracks with dance movements, and lack of accurate recognition of dancers' movements and distances.
[0006] The technical solution adopted by the present invention to solve its technical problem is: an audio system with AI recognition and control function, including a main enclosure and a bottom support plate. The interior of the main enclosure is divided into a sound transmission chamber and an energy chamber by an internal partition. Multiple speaker modules are installed on the inner walls of both sides and the bottom of the sound transmission chamber. A removable lithium battery is installed on the side wall of the internal partition at the end of the energy chamber. A control component and a display panel are installed on the front of the main enclosure. An electrically controlled flip control module is set on the front of the main enclosure directly above the control component.
[0007] The upper back of the main housing is provided with a rear loading and unloading port for easy loading and unloading of detachable lithium batteries, and a flip cover plate for closing the rear loading and unloading port is hinged to the upper end of the main housing.
[0008] The top flip-up carrying frame is hinged inside the flip-up groove at the upper end of the main housing.
[0009] The main housing is mounted on the upper part of the bottom support plate via a bottom bracket. The upper surface of the bottom support plate has a downwardly recessed bottom groove. The bottom of the main housing has a bottom opening corresponding to the position of the multi-speaker module. A downwardly protruding bottom filter is installed at the lower end of the bottom opening.
[0010] The main housing has an LED light set fixedly installed on the front of the control component.
[0011] The front of the main housing is provided with a front inclined surface at the assembly end of the electronically controlled flip control module, and a flip storage slot for storing the electronically controlled flip control module is provided on the front inclined surface.
[0012] The electronically controlled flip control module includes an electronically controlled flip cover connected inside the flip storage slot, a main control module installed inside the main housing, an AI recognition module, an audio track correction module, a storage module, and an optical lens module installed inside the electronically controlled flip cover. The optical lens module is used to collect image information of the dancers, including the dancers' outlines, position coordinates, movement postures, and distance data between the dancers and the smart speaker. The AI recognition module is used to receive image information collected by the optical lens module, process the image information through a preset AI algorithm, and identify and output the real-time distance, real-time position and real-time motion feature data of the dancer; The main control module is used to receive real-time distance, real-time position, and real-time motion feature data output by the AI recognition module, adjust the overall speaker power of the multi-speaker module based on the real-time distance, and thus adjust the overall volume of the speaker; and adjust the independent power of the speakers in different positions in the multi-speaker module based on the real-time position, and thus adjust the volume of the speakers in different positions. The audio track correction module is used to receive real-time motion feature data transmitted by the main control module, and combine it with the pre-stored audio track database in the storage module to correct the currently playing music audio track in real time so that the corrected audio track matches the real-time movements of the dancers. The storage module is used to store AI recognition algorithms, audio track databases, motion feature sample libraries, and device operating parameters.
[0013] The optical lens module includes at least one depth optical lens and at least one wide-angle optical lens. The depth optical lens is used to collect distance data and three-dimensional position coordinates between the dancer and the smart speaker. The wide-angle optical lens is used to collect panoramic movement posture and contour information of the dancer. The depth optical lens and the wide-angle optical lens work together to achieve comprehensive acquisition of image information of the dancer.
[0014] The AI recognition module includes an image preprocessing unit, a distance calculation unit, a location recognition unit, and a motion feature extraction unit.
[0015] The logic of the main control module in adjusting the overall power of the multi-speaker module is as follows: a preset distance-power correspondence table is used. When the real-time distance is less than the preset near distance threshold, the overall speaker power is reduced and the overall speaker volume is decreased. When the real-time distance is greater than the preset far distance threshold, the overall speaker power is increased and the overall speaker volume is increased. When the real-time distance is between the preset near distance threshold and the preset far distance threshold, the current overall speaker power is maintained unchanged, or it is linearly adjusted according to the distance change.
[0016] The beneficial effects of this invention are: (1) The present invention uses an optical lens module in conjunction with an AI recognition module to accurately collect and identify the real-time distance, position and movement characteristics of dancers, thereby realizing intelligent perception of the user's status. (2) Based on the identification data, the main control module can automatically adjust the overall power of the multi-speaker module and the independent power of speakers in different directions, which solves the problems of traditional speaker volume adjustment relying on manual operation and uneven sound coverage, improves auditory adaptability, and also reduces power consumption waste. (3) The audio track correction module can correct the audio track in real time based on the movement characteristics of the dancers, so that the music rhythm and dance movements are highly matched, which enhances the coordination and experience of the dance. (4) The device is equipped with a detachable lithium battery, a top flip-up carrying frame and LED lights, which respectively realize the functions of flexible battery life, convenient portability and scene lighting, and can be adapted to various usage scenarios such as home and dance studio. Attached Figure Description
[0017] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0018] Figure 1 This is a schematic diagram of the structure of the present invention.
[0019] Figure 2 This is a schematic diagram of the structure of the electrically controlled flip control module in this invention.
[0020] Figure 3 This is a schematic diagram of the structure on the right side of the present invention.
[0021] Figure 4 yes Figure 3 A cross-sectional view at position AA.
[0022] Figure 5 This is a schematic diagram of the left side structure of the present invention.
[0023] Figure 6 yes Figure 5 A cross-sectional view at position BB in the middle.
[0024] Figure 7 This is the electronic control block diagram of the present invention.
[0025] 1. Main casing; 2. Bottom support plate; 3. Internal partition; 4. Sound transmission chamber; 5. Energy chamber; 6. Multi-speaker module; 7. Removable lithium battery; 8. Control components; 9. Display panel; 10. Electrically controlled flip control module; 11. Rear loading / unloading port; 12. Flip cover; 13. Flip groove; 14. Top flip carrying frame; 15. Bottom bracket; 16. Bottom recess; 17. Bottom opening; 18. Bottom filter; 19. LED light group; 20. Front inclined surface; 21. Flip storage slot; 22. Electrically controlled flip cover; 23. Main control module; 24. AI recognition module; 25. Audio track correction module; 26. Storage module; 27. Optical lens module; 28. Depth optical lens; 29. Wide-angle optical lens; 30. Image preprocessing unit; 31. Distance calculation unit; 32. Position recognition unit; 33. Motion feature extraction unit. Detailed Implementation
[0026] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0027] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "connected" and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0028] like Figures 1-7 The specific assembly process is as follows: The main housing 1 is fixed to the bottom support plate 2 by the bottom bracket 15. The bottom groove 16 of the bottom support plate 2 can enhance the stability of the equipment placement. The interior of the main housing 1 is divided into a sound transmission chamber 4 and an energy chamber 5 by an internal partition 3. Multi-speaker modules 6 are installed on the inner walls of both sides and the bottom of the sound transmission chamber 4 to achieve multi-directional sound emission. The bottom opening 17 at the bottom of the main housing 1 corresponds to the position of the multi-speaker modules 6. The bottom filter 18 installed at the bottom of the bottom opening 17 can prevent dust from entering the sound transmission chamber 4. The removable lithium battery 7 inside the energy chamber 5 is installed and removed through the rear loading and unloading port 11 on the back of the main housing 1. The flip cover 12 can close and protect the rear loading and unloading port 11. The flip groove 13 at the top of the main housing 1 contains... The top hinged top flip-up carrying frame 14 can be flipped open when the equipment needs to be carried, and stored in the flip slot 13 when not in use. The control component 8 and the display panel 9 are installed on the front of the main housing 1. The control component 8 is used for basic operations such as powering on and off and mode switching of the equipment. The display panel 9 is used to display parameters such as the volume and working mode of the equipment. The LED light group 19 around the control component 8 can be turned on for lighting as needed. A flip storage slot 21 is opened on the front inclined surface 20 of the main housing 1. The electronically controlled flip cover 22 of the electronically controlled flip control module 10 is hinged in the flip storage slot 21. When not in use, the electronically controlled flip cover 22 is closed and stored. When in use, the electronically controlled flip cover 22 is flipped open to expose the internal optical lens module 27.
[0029] The electrically controlled flip control module 10 also includes a main control module 23, an AI recognition module 24, an audio track correction module 25, a storage module 26, and a wireless communication module. The modules are electrically connected to each other via circuit boards and are integrated inside the main housing 1.
[0030] The main control module 23 uses an STM32F407 microcontroller as the core control unit of the device. It is responsible for receiving signals from each module, outputting control commands, and realizing the collaborative work of multiple modules.
[0031] In Example 1, the optical lens module 27 includes a Kinect depth lens and a wide-angle high-definition lens. Both the Kinect depth lens and the wide-angle high-definition lens are mounted on an electrically controlled flip cover 22. The Kinect depth lens is used to collect real-time distance data and three-dimensional position coordinates between the dancer and the speaker, with a collection accuracy of ±1cm. The wide-angle high-definition lens has a field of view of 120° and is used to collect panoramic movement postures and contour information of the dancer, ensuring no blind spots in the collection. The optical lens module 27 also includes a transparent acrylic lens cover and a micro stepper motor. The micro stepper motor is electrically connected to the main control module 23 and can automatically adjust the collection angle of the Kinect depth lens and the wide-angle high-definition lens according to user settings or the recognition results of the AI recognition module 24. The adjustment range is 0°-90°, ensuring full coverage of the dancing area.
[0032] The AI recognition module 24 uses an NVIDIA Jetson Nano AI development board, integrating an image preprocessing unit 30, a distance calculation unit 31, a position recognition unit 32, and a motion feature extraction unit 33. The image preprocessing unit 30 uses a Gaussian filtering algorithm to reduce noise in the original image and a perspective transformation algorithm to correct distortion, outputting a clear preprocessed image. The distance calculation unit 31 uses a binocular ranging algorithm to calculate the real-time straight-line distance between the dancer and the smart speaker based on the depth image captured by the Kinect depth camera. The position recognition unit 32... 2. The YOLOv5 target detection algorithm is adopted. Based on the human contour in the preprocessed image and combined with the three-dimensional position coordinates, the real-time position of the dancer around the smart speaker is identified, such as front, left, right and back, and the position coordinate data is output. The motion feature extraction unit 33 adopts the MediaPipe skeleton key point recognition algorithm to extract 18 limb joint key points of the dancer, such as head, shoulder, elbow, wrist, waist, knee and ankle. Based on the motion trajectory of the key points, the motion amplitude, motion rhythm and motion type, such as stretching, rotating and jumping, are calculated, and real-time motion feature data is output.
[0033] The multi-speaker module 6 includes three independently controlled full-range speakers, respectively installed on the left, right, and bottom surfaces of the inner wall of the main housing 1, forming a multi-channel sound structure. Each speaker is equipped with a TPA3116D2 power amplifier chip, which is electrically connected to the main control module 23. The main control module 23 adjusts the output power of the power amplifier chip, with an adjustment range of 0.5W-30W, to achieve independent volume adjustment for each speaker. A preset distance-power correspondence table is provided: when the real-time distance is <0.5m, the overall speaker power is adjusted to 0.5W-5W; when 0.5m ≤ real-time distance ≤ 3m, the overall speaker power is adjusted to 5W-20W. Furthermore, the power is linearly adjusted according to distance changes, increasing by 3W for every 0.5m increase in distance. When the real-time distance is greater than 3m, the overall speaker power is adjusted to 20W-30W. Simultaneously, the main control module 23 adjusts the power of the speaker in the corresponding direction to be 20%-40% higher than that in other directions based on the directional coordinate data output by the position recognition unit 32, thus creating a directional sound effect. For example, when the dancer is on the left side of the smart speaker, the power of the left speaker is adjusted to 120% of the current overall power, while the power of the front and right speakers is adjusted to 80% of the current overall power, ensuring that the volume is more suitable for the location of the person, while also reducing power consumption waste and the probability of disturbing others.
[0034] The audio track correction module 25 uses a VS1053 audio decoding chip and is electrically connected to the main control module 23; the storage module 26 uses a 16GB SD card to pre-store an audio track database and a motion feature sample library, containing standard motion feature data for various dance types such as street dance, jazz dance, square dance, and classical dance; the correction logic of the audio track correction module 25 is as follows: Extract the rhythm, melody, and intensity parameters of the currently playing audio track; The real-time motion rhythm output by the AI recognition module 24 is compared with the rhythm parameters of the current audio track to calculate the rhythm deviation value. When the deviation value is greater than 10 times / minute, the playback speed of the audio track is adjusted in real time, with an adjustment range of 0.8 times to 1.2 times, so that the rhythm of the audio track is consistent with the motion rhythm. Based on the motion amplitude in the motion feature data, adjust the intensity parameter of the audio track: when the motion amplitude is greater than the preset amplitude threshold, the audio track intensity is increased by 10%-15%; when the motion amplitude is less than the preset amplitude threshold, the audio track intensity is decreased by 10%-15%. By combining the action type in the action feature data, the corresponding sound effect segment is matched from the audio track database of the storage module 26. For example, jumping action is matched with drum sound effect, and rotating action is matched with melody transition sound effect. The sound effect segment is then inserted into the current audio track to achieve precise matching between the audio track and the dance action.
[0035] The wireless communication module uses an ESP8266 Wi-Fi module and an HC-05 Bluetooth module, which are electrically connected to the main control module 23. It supports Wi-Fi (802.11b / g / n) and Bluetooth 4.0 connections, enabling wireless connection between the smart speaker and mobile terminals such as mobile phones and tablets. Users can remotely control the speaker to start and stop, select music, set preset distance thresholds, distance-power correspondence, and motion recognition sensitivity through a dedicated APP installed on the mobile terminal, and view real-time distance, position, and motion recognition results.
[0036] The display panel 9 uses a 0.96-inch OLED display screen, which is installed on the front of the smart speaker and electrically connected to the main control module 23. It displays the distance data, position information, movement rhythm, movement type and current audio track correction parameters of the dancer in real time, so that users can intuitively view the operating status of the device.
[0037] The work process is as follows: Turn on the smart speaker, select dance music through the mobile terminal APP, and the speaker will start playing music; at the same time, the optical lens module 27 is activated, and the Kinect depth lens and wide-angle high-definition lens work together to collect image information of the dancers and transmit it to the AI recognition module 24. The AI recognition module 24 preprocesses the image information, calculates the real-time distance through the distance calculation unit 31, identifies the real-time position through the position recognition unit 32, extracts real-time action feature data through the action feature extraction unit 33, and transmits all recognition results to the main control module 23. The main control module 23 queries the preset distance-power correspondence table based on the real-time distance and adjusts the overall power of the multi-speaker module 6 to achieve adaptive adjustment of the overall volume of the speaker; at the same time, it adjusts the power of the speakers in the corresponding directions based on the real-time position to form a directional sound effect. The main control module 23 transmits real-time motion feature data to the audio track correction module 25. The audio track correction module 25, in conjunction with the audio track database in the storage module 26, performs real-time correction of the rhythm, intensity, and sound effects of the currently playing audio track, so that the audio track matches the movements of the dancers. Display panel 9 displays the device's operating parameters in real time, and users can view or adjust the device settings through a mobile terminal APP; when the distance, position or movement of the dancer changes, the AI recognition module 24 updates the recognition results in real time, and the main control module 23 and the audio track correction module 25 adjust synchronously to ensure that the volume and audio track are always adapted to the dance movements; After the dance ends, the smart speaker is turned off, all modules stop working, and the storage module 26 records the data from this operation for quick adaptation next time.
[0038] Example 2 differs from Example 1 in that the optical lens module 27 uses only one binocular depth lens to simultaneously acquire distance data, position coordinates, and motion posture, simplifying the device structure and reducing costs. The multi-speaker module 6 includes two independently controlled speakers, installed on the front and back of the smart speaker respectively, suitable for small dance scenarios. The audio track correction module 25 only adjusts the rhythm and intensity without adding additional sound effects, meeting the requirements for minimalist use. The remaining structure and operation are consistent with Example 1.
[0039] Example 3 differs from Example 1 in that the motion feature extraction unit 33 of the AI recognition module 24 uses the OpenPose skeletal keypoint recognition algorithm, which can extract 25 limb joint keypoints, improving the accuracy of motion feature recognition; the storage module 26 uses a 32GB SD card to expand the storage capacity of the audio track database and motion feature sample library, supporting users to import music tracks and motion feature data; the display panel 9 uses a 2.4-inch LCD screen, which can display more detailed device operation information and motion recognition screen. The remaining structure and working process are the same as in Example 1.
[0040] Based on the above-described preferred embodiments of the present invention, and through the foregoing description, those skilled in the art can make various changes and modifications without departing from the inventive concept. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.
Claims
1. A sound system with AI recognition and control function, comprising a main enclosure (1) and a bottom support plate (2), characterized in that: The main housing (1) is divided into a sound transmission chamber (4) and an energy chamber (5) by an internal partition (3). Multi-speaker modules (6) are installed on the inner walls of both sides and the bottom of the sound transmission chamber (4). A removable lithium battery (7) is installed on the side wall of the energy chamber (5) by the internal partition (3). A control component (8) and a display panel (9) are installed on the front of the main housing (1). An electronically controlled flip control module (10) is set on the front of the main housing (1) directly above the control component (8). The electronically controlled flip control module (10) includes a main control module (23), an AI recognition module (24), an audio track correction module (25), and an optical lens module (27) installed inside the electronically controlled flip cover (22); The AI recognition module (24) is used to receive image information collected by the optical lens module (27), process the image information through a preset AI algorithm, and identify and output the real-time distance, real-time position and real-time motion feature data of the dancer; The audio track correction module (25) is used to receive real-time motion feature data transmitted by the main control module (23), and combine it with the audio track database stored in the storage module (26) to correct the currently playing music track in real time so that the corrected audio track matches the real-time movements of the dancers.
2. The audio system with AI recognition and control function according to claim 1, characterized in that: The upper back of the main housing (1) is provided with a rear loading and unloading port (11) for easy loading and unloading of a detachable lithium battery (7), and the upper end of the main housing (1) is hinged with a flip cover (12) for closing the rear loading and unloading port (11).
3. The audio system with AI recognition and control function according to claim 1, characterized in that: The upper end of the main housing (1) is provided with a flip groove (13), and a top flip carrying frame (14) is hinged inside the flip groove (13).
4. The audio system with AI recognition and control function according to claim 1, characterized in that: The main housing (1) is mounted on the upper end of the bottom support plate (2) via a bottom bracket (15). The upper surface of the bottom support plate (2) is provided with a downwardly recessed bottom groove (16). The bottom of the main housing (1) is provided with a bottom opening (17) corresponding to the position of the multi-speaker module (6). A downwardly protruding bottom filter (18) is installed at the lower end of the bottom opening (17).
5. A sound system with AI recognition and control function according to claim 1, characterized in that: An LED light group (19) for lighting is fixedly installed on the front of the main housing (1) around the control component (8).
6. The audio system with AI recognition and control function according to claim 1, characterized in that: The front of the main housing (1) is provided with a front inclined surface (20) at the assembly end of the electronically controlled flip control module (10), and a flip storage slot (21) for storing the electronically controlled flip control module (10) is provided on the front inclined surface (20).
7. A sound system with AI recognition and control function according to claim 6, characterized in that: The electronically controlled flip control module (10) also includes an electronically controlled flip cover (22) and a storage module (26) hinged inside the flip storage slot (21); The optical lens module (27) is used to collect image information of the dancers, including the outline of the dancers, position coordinates, movement postures and distance data between the dancers and the smart speaker; The main control module (23) is used to receive real-time distance, real-time position and real-time motion feature data output by the AI recognition module (24), adjust the overall speaker power of the multi-speaker module (6) based on the real-time distance, and thereby adjust the overall volume of the speaker; adjust the independent power of the speakers in different positions in the multi-speaker module (6) based on the real-time position, and thereby adjust the volume of the speakers in different positions. The storage module (26) is used to store AI recognition algorithms, audio track databases, motion feature sample libraries, and device operating parameters.
8. A sound system with AI recognition and control function according to claim 7, characterized in that: The optical lens module (27) includes at least one depth optical lens (28) and at least one wide-angle optical lens (29). The depth optical lens (28) is used to collect distance data and three-dimensional position coordinates between the dancer and the smart speaker. The wide-angle optical lens (29) is used to collect panoramic movement posture and contour information of the dancer. The depth optical lens (28) and the wide-angle optical lens (29) work together to achieve comprehensive acquisition of image information of the dancer.
9. A sound system with AI recognition and control function according to claim 7, characterized in that: The AI recognition module (24) includes an image preprocessing unit (30), a distance calculation unit (31), a location recognition unit (32), and an action feature extraction unit (33).
10. A sound system with AI recognition and control function according to claim 7, characterized in that: The logic of the main control module (23) adjusting the overall power of the multi-speaker module (6) is as follows: a preset distance-power correspondence table is used. When the real-time distance is less than the preset near distance threshold, the overall speaker power is reduced and the overall speaker volume is decreased. When the real-time distance is greater than the preset far distance threshold, the overall speaker power is increased and the overall speaker volume is increased. When the real-time distance is between the preset near distance threshold and the preset far distance threshold, the current overall speaker power is kept unchanged, or linear adjustment is made according to the distance change.
Citation Information
Patent Citations
Audio system
CN102469402A
Method and device for controlling intelligent sound box
CN110062309A
Intelligent control system and intelligent sound box
CN110913307A
Dance music matching method and device and entertainment equipment
CN114419734A
Intelligent speaker and playing control method
WO2019153382A1