Intelligent loudspeaker box system based on gesture recognition and control system and control method thereof

By using a gesture recognition-based smart speaker system, which automatically controls the microphone to rise using an infrared detection module and a drive module, the problem of portable speaker systems needing to carry an additional microphone is solved, resulting in simpler operation and extended device lifespan.

CN120972735APending Publication Date: 2025-11-18广东台德智联科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511331525.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing portable speaker systems require an additional microphone when singing karaoke outdoors, which is inconvenient to operate and prone to damage. The existing press-operation method is not intuitive and suffers from severe mechanical wear.

Method used

The smart speaker system, which uses gesture recognition, defines a sensing area through an infrared detection module to detect user gestures and trajectories, and drives the module to automatically control the microphone to rise, replacing the traditional pressing operation.

Benefits of technology

No need to carry an extra microphone, easy and effortless operation, extended equipment lifespan, and avoidance of mechanical wear.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120972735A_ABST
    Figure CN120972735A_ABST
Patent Text Reader

Abstract

The invention relates to the field of sound box systems, and discloses an intelligent sound box system control method based on gesture recognition, which is applied to an intelligent sound box system provided with a sound box main body, a screen, a microphone, a plurality of infrared detection modules and a driving module. The microphone is at least partially accommodated in the sound box main body, the plurality of infrared detection modules define an induction area, the screen can be detected to be overturned to a turned-up state or a closed state relative to the sound box main body, and the driving module is used for driving the microphone to rise. The control method comprises the steps that it is detected that the screen is in a turned-up state, it is detected that the movement track in the induction area is consistent with a preset track, and the microphone is controlled to ascend. Therefore, a user does not need to additionally carry a microphone or use an external mobile phone to watch contents, the problem of short service life caused by repeated friction between the microphone and the locking mechanism is avoided through non-contact track recognition, and the operation is relatively simple, convenient and labor-saving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speaker systems, and more particularly to a smart speaker system with gesture-controlled microphone lifting function. Background Technology

[0002] With the rise of outdoor entertainment activities, the demand for portable karaoke devices is increasing. Current mainstream solutions typically rely on smartphones paired with portable speakers: the smartphone handles song selection, lyrics display, and audio processing, while the portable speaker amplifies the sound. This architecture has significant limitations: firstly, smartphone screens are limited in size, resulting in a poor user experience when selecting songs or viewing lyrics in bright outdoor light or shared environments with multiple users. More importantly, the issue of microphone portability and storage is particularly prominent: conventional portable speakers usually do not come with a microphone, requiring users to carry a separate microphone for outdoor karaoke. This undoubtedly increases the number of devices and the burden of carrying them, and separate microphones are easily lost or damaged during storage and transportation. To address this problem, existing technologies, such as patent CN116389967B, propose a solution that stores the microphone within the speaker body. This solution uses a recessed slot on the speaker to store the microphone and incorporates a press-to-release mechanism: the user needs to apply two different presses to the microphone—the first press lowers and locks it in place, and the second press raises it to release it—in order to access or store the microphone.

[0003] However, this pressing operation method still has obvious shortcomings: First, users need to apply precise force and memorize the operation logic, making the process less intuitive and convenient; second, as a component that is frequently picked up and put down, the microphone is prone to fatigue wear due to long-term repeated mechanical pressing operations, which can affect the reliability and service life of the device.

[0004] Application content

[0005] This application provides a control method for a smart speaker system based on gesture recognition, which aims to solve existing technical problems.

[0006] This application provides a control method for a smart speaker system based on gesture recognition, applicable to a smart speaker system comprising a speaker body, a screen, a microphone, multiple infrared detection modules, and a driving module. The microphone is at least partially housed within the speaker body. The multiple infrared detection modules define a sensing area. The screen can be detected to flip relative to the speaker body to a flipped-up or closed state. The driving module is used to drive the microphone to rise. The control method includes: detecting that the screen is in a flipped-up state, detecting that the movement trajectory within the sensing area matches a preset trajectory, and controlling the microphone to rise.

[0007] In one embodiment, before detecting the activity trajectory within the sensing area, the method further includes: detecting that a gesture within the sensing area matches a preset gesture.

[0008] In one embodiment, after detecting that the activity trajectory within the sensing area matches a preset trajectory, the method further includes: detecting that the object generating the activity trajectory hovers within the sensing area for more than a first preset duration S1.

[0009] In one embodiment, after detecting that the activity trajectory within the sensing area matches a preset trajectory, the method further includes: detecting that the gesture within the sensing area matches a preset gesture.

[0010] In one embodiment, before controlling the microphone to rise, the method further includes: detecting that an object generating an activity trajectory hovers within the sensing area for a period of more than a first preset duration S1.

[0011] In one embodiment, the first preset duration S1 is in the range of 0.5 seconds ≤ S1 ≤ 3 seconds. In another embodiment, the screen has a touch function, and before controlling the microphone to rise, the method further includes detecting that the screen has not been operated within a second preset duration S2.

[0012] In one embodiment, the starting time of the second preset duration S2 is when the previous step of the step for detecting the second preset duration ends.

[0013] In one embodiment, the range of the second preset duration S2 is: 0.3 seconds ≤ S2 ≤ 2 seconds.

[0014] In one embodiment, the screen has a touch function, and the minimum distance between the sensing area and the screen is greater than a preset distance L.

[0015] In one embodiment, the preset distance L is greater than 2 centimeters.

[0016] In one embodiment, the preset trajectory, with the user's position when normally operating the smart speaker system as the forward coordinate, is any one or more combinations of left to right, right to left, front to back, and top to bottom.

[0017] In one embodiment, the smart speaker system further includes a human body detection module for detecting whether there is a person within a certain range. After detecting that the screen is in a flipped-up state, the control method further includes: when the detection of no one for a third preset time S3, turning off the plurality of infrared detection modules.

[0018] This application also provides a gesture recognition-based smart speaker system, including a speaker body and a microphone. The speaker body includes an accommodating space for accommodating at least a portion of the microphone. The smart speaker system further includes: a screen that can be flipped relative to the speaker body to a flipped-up or closed state; multiple infrared detection modules defining a sensing area; a driving module for driving the microphone to rise; a screen detection module for detecting the state of the screen; and an MCU microcontroller module including a memory and a controller. The memory stores program instructions, and the controller calls and executes the program instructions in the memory to implement the aforementioned gesture recognition-based smart speaker system control method.

[0019] This application further provides an MCU microcontroller module applied to an intelligent speaker system comprising a speaker body, a screen, a microphone, multiple infrared detection modules, and a drive module; the microphone is at least partially housed within the speaker body, the multiple infrared detection modules define a sensing area, the screen can be detected to flip relative to the speaker body to a flipped-up or closed state, and the drive module is used to drive the microphone to rise; the MCU microcontroller module includes: a screen detection module for detecting the state of the screen; a trajectory detection module for detecting whether the movement trajectory within the sensing area matches a preset trajectory after the screen detection module detects a flipped-up state; a signal generation module for generating a control signal after the trajectory detection module detects a match; and a drive control module for controlling the drive module to drive the microphone to rise after reading the control signal.

[0020] With the control method and smart speaker system described above, users no longer need to carry a separate microphone or use an external mobile phone to view content, making it more convenient for users. Since non-contact trajectory recognition replaces traditional physical pressing operations, it fundamentally avoids repeated mechanical friction between the microphone and the locking mechanism, extending the device's lifespan. Furthermore, users only need to naturally wave their hand to trigger the microphone to automatically rise, making operation simpler and less strenuous. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of a smart speaker system provided in one embodiment of this application.

[0023] Figure 2 yes Figure 1A top-down view of the smart speaker system.

[0024] Figure 3 yes Figure 1 A schematic diagram showing the connection of some functional modules in a smart speaker system.

[0025] Figure 4 This is a flowchart of a smart speaker system control method based on gesture recognition provided in one embodiment of this application.

[0026] Figure 5 This is a flowchart of a control method provided in another embodiment of this application.

[0027] Figure 6 This is a flowchart of a control method provided in another embodiment of this application.

[0028] Figure 7 This is a flowchart of a control method provided in another embodiment of this application.

[0029] Figure 8 This is a flowchart of a control method provided in another embodiment of this application.

[0030] Figure 9 This is a flowchart of a control method provided in another embodiment of this application.

[0031] Figure 10 yes Figure 3 A schematic diagram of some components of the MCU microcontroller module in a smart speaker system.

[0032] Explanation of icon numbers:

[0033] 100. Smart speaker system; 10. Main body; 11. Top surface; 12. Receiving cavity; 20. Screen; 22. Hinge; 30. Microphone; 40. Infrared detection module; 42. Sensing area; 50. Drive module; 60. Screen detection module; 70. MCU microcontroller module; 80. Human body detection module; 102. Screen detection module; 104. Trajectory detection module; 106. Signal generation module; 108. Drive control module.

[0034] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0035] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0036] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of this application are only used to explain the relative positional relationship and movement of each component in a certain specific posture. If the specific posture changes, the directional indication will also change accordingly.

[0037] It should also be noted that when a component is described as "fixed to" or "set on" another component, it can be directly on the other component or there may be an intervening component present. When a component is described as "connected to" another component, it can be directly connected to the other component or there may be an intervening component present.

[0038] Furthermore, the use of terms such as "first" and "second" in this application is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. Additionally, the technical solutions of the various embodiments can be combined with each other, but only on the basis of being achievable by those skilled in the art. When the combination of technical solutions is contradictory or impossible to implement, such a combination of technical solutions should be considered non-existent and not within the scope of protection claimed in this application.

[0039] Please combine Figure 1-3This application provides an intelligent speaker system 100, including a speaker body 10, a screen 20, a microphone 30, multiple infrared detection modules 40, and a driver module 50. The microphone 30 is at least partially housed within the speaker body 10. For example, a receiving cavity 12 is formed on the top surface 11 of the speaker body 10, and the microphone 30 is at least partially housed within the receiving cavity 12. Since the intelligent speaker system 100 is equipped with a microphone 30, the user does not need to carry it separately, making it convenient for use. The screen 20 can be detected to be flipped relative to the speaker body 10 to either a flipped-up or closed state. For example, the screen 20 is connected to one side of the speaker body 10 via a hinge 22. When the screen 20 is in the flipped-up state, the top surface 11 is exposed; when the screen 20 is in the closed state, the top surface 11 is covered by the screen 20. Thus, the user can directly view the lyrics of a karaoke song and / or the corresponding video content on the screen 20 without needing a mobile phone, providing convenience and a better experience with a larger screen. The state of screen 20 can be detected by a screen detection module 60, which may include a Hall sensor disposed on speaker body 10 and a magnet disposed at a corresponding position on screen 20, detecting the state of screen 20 by sensing magnetic force; alternatively, it may consist of two conductive contacts disposed at corresponding positions on speaker body 10 and screen 20, detecting the state of screen 20 by whether the contacts are in contact. It should be noted that the detection of the state of screen 20 is not limited to the aforementioned methods, as this is a technique known to those skilled in the art and will not be elaborated upon here.

[0040] Multiple infrared detection modules 40 define a sensing area 42. For example, the smart speaker system 100 includes six infrared detection modules 40, arranged in two rows of three on each side of the microphone 30, forming an array. The working principle of the multiple infrared detection modules 40 is as follows: each infrared detection module includes at least one infrared emitting diode and multiple infrared receiving photodiodes (or phototransistors). The emitting diode actively emits an infrared beam of a specific wavelength (usually invisible to the human eye, such as 850nm or 940nm). When the sensed object, such as a user's hand, enters the sensing area 42, it reflects some infrared light. Because the distances between the multiple infrared receiving photodiodes and the infrared emitting diodes are different, the output signal intensity (current or voltage) of each receiver changes sequentially depending on the position of the user's hand. By analyzing the sequence of signal changes of each infrared receiving photodiode, the direction of the user's hand movement can be determined. It should be noted that the number and arrangement of the multiple infrared detection modules 40 are not limited to... Figure 2As shown, this is a technique well-known to those skilled in the art and will not be elaborated upon here. It should be understood that since each infrared detection module 40 has its own three-dimensional sensing range, the sensing area 42 is actually a three-dimensional space defined by the combination of multiple sensing ranges of multiple infrared detection modules 40. In addition, it should be pointed out that there need not be 6 infrared detection modules 40, nor need they be arranged on both sides of the microphone 30, as long as their spacing is sufficient to detect activity trajectories or trajectories and gestures.

[0041] The drive module 50 is used to drive the microphone 30 to rise. For example, the drive module 50 may include a motor, a lead screw, a nut, and a guide device; the motor drives the lead screw to rotate through a coupling, the nut engages with the lead screw but is restricted from rotating by the guide device and moves linearly along the lead screw axis, the nut is directly or indirectly connected to the microphone 30, thereby driving the microphone 30 to rise. Alternatively, the drive module 50 may include a motor, a pinion, a rack, and a fixed bracket; the motor drives the pinion to rotate, the pinion engages with a vertically mounted rack, the rotation of the pinion pushes the rack to move linearly up and down along a guide rail, the rack is directly or indirectly connected to the microphone 30, achieving linear rising. It should be noted that the specific implementation of the drive module 50 is not limited to the two methods mentioned above, and its implementation is a technology known to those skilled in the art, and will not be elaborated here. As for the microphone 30 being lowered after use and to be stored away, the microphone 30 can be lowered to its original storage position by the drive module 50, or it can be lowered manually.

[0042] Please combine Figure 4 This application provides a gesture recognition-based control method for a smart speaker system, used to control the aforementioned smart speaker system 100. In one embodiment, the control method includes the following steps:

[0043] The screen 20 is detected to be in a flipped-up state. Specifically, the screen detection module 60 can periodically detect the state of the screen 20. If the screen 20 is detected to be in a closed state, the detection step continues; if the screen is detected to be in a flipped-up state, the next step is performed.

[0044] The detected activity trajectory within the sensing area 42 matches a preset trajectory. Specifically, the smart speaker system 100 stores one or more preset trajectories. For example, in a coordinate system with the user's position when normally operating the smart speaker system 100 as the forward coordinate, the preset trajectory can be one or more from left to right, right to left, front to back, top to bottom, etc., or a combination of multiple trajectories (e.g., first from left to right and then back). In multiple cases (e.g., from left to right and from right to left), either from left to right or from right to left will be recognized as a preset trajectory. When the user operates within the sensing area 42, for example, when the user's hand sweeps across the sensing area 42 from left to right, the activity trajectory of the user's hand can be obtained through multiple infrared detection modules 40. If the detected activity trajectory does not match the preset trajectory, the detection step continues; if the detected activity trajectory matches the preset trajectory, the next step is initiated.

[0045] The microphone 30 is controlled to rise. Specifically, if the detected movement trajectory matches a preset trajectory, it can be determined that the user intends to pick up the microphone 30. At this time, the drive module 50 drives the microphone 30 to rise, allowing the user to easily obtain the microphone without pressing or other operations. Under the control method described above, the user does not need to carry a separate microphone or use an external mobile phone to view content, thus simplifying the user experience. Since non-contact trajectory recognition replaces traditional physical pressing operations, repeated mechanical friction between the microphone and the locking mechanism is fundamentally avoided, extending the device's lifespan. Furthermore, the user only needs to naturally wave their hand to trigger the microphone to rise automatically, making operation simpler and less strenuous.

[0046] It should be noted that the aforementioned control method is implemented under the control of the intelligent speaker system 100's control system, such as the built-in MCU microcontroller module 70. The MCU microcontroller module 70 can be a control circuit built around a controller with control functions, such as a microcontroller / microprocessor / central processing unit, and peripheral functional modules such as sensing user operation / external environment, equipped with necessary memory for storing program instructions, and can run based on languages ​​such as C / C++ / JAVA. In other words, the gesture recognition-based smart speaker system 100 provided in this application includes a speaker body 10, a microphone 30, a screen 20, multiple infrared detection modules 40, a driving module 50, a screen detection module 60, and an MCU microcontroller module 70. The speaker body 10 includes an accommodating space 12 for accommodating at least a portion of the microphone 30. The screen 20 can be flipped relative to the speaker body 10 to a flipped-up or closed state. The multiple infrared detection modules 40 define a sensing area 42. The driving module 50 is used to drive the microphone 30 to rise. The screen detection module 60 is used to detect the state of the screen 20. The MCU microcontroller module includes a memory and a controller. The memory is used to store program instructions, and the controller is used to call and execute the program instructions in the memory to realize the gesture recognition-based smart speaker system control method described above.

[0047] Please combine Figure 5 In another embodiment, before detecting the activity trajectory within the sensing area, the following step is included: detecting a gesture within the sensing area that matches a preset gesture. In addition to detecting the aforementioned activity trajectory, the multiple infrared detection modules 40 can also be used to detect user gestures. The principle is as follows: after the external emitting diode actively emits an infrared beam of a specific wavelength, multiple infrared receiving photodiodes output reflection intensity values ​​in real time. Because different gestures, such as a clenched fist or an open palm, have different reflective surfaces and efficiencies towards infrared light, the real-time output reflection intensity values ​​from multiple infrared receiving photodiodes will form a set of intensity distribution vectors. Key features such as total reflection intensity, intensity ratio of each channel, and maximum value position are extracted and compared with the pre-stored gesture "feature template." The gestures are classified using a threshold method or simple machine learning (such as the KNN algorithm), and the recognition result is output. Between detecting that the screen 20 is in a flipped-up state and initiating the detection of the activity trajectory, the system checks whether the gesture within the sensing area matches a preset gesture. If they do not match, the detection of the activity trajectory stops. This reduces misjudgments when the user unintentionally brushes their hand across the sensing area 42, improving the accuracy of the judgment. If the matching gesture matches the preset gesture, the system proceeds to the next step, which is to detect the activity trajectory within the sensing area 42. This embodiment reduces the possibility of misjudgments based on user operations and improves the accuracy of the judgment.

[0048] Furthermore, please combine Figure 6In another embodiment, after detecting that the activity trajectory within the sensing area matches the preset trajectory, the following step is further included: detecting that the object generating the activity trajectory hovers within the sensing area 42 for more than a first preset duration S1. Specifically, after detecting that the activity trajectory within the sensing area matches the preset trajectory, if the hovering time of the object generating the activity trajectory, such as the user's hand, within the sensing area 42 does not exceed the first preset duration S1, it is determined that the user does not intend to take the microphone, and the process returns to the initial detection step; if the hovering time exceeds the first preset duration S1, it is determined that the user intends to take the microphone, and the next step is initiated, i.e., the microphone is controlled to rise. Based on the aforementioned judgment of gesture and trajectory, this embodiment adds a waiting time judgment, adding another verification element and further improving the accuracy of the judgment. In addition, since the user's habit of taking the microphone 30 is to reach towards the microphone 30 located in the sensing area 42 with a grasping gesture, and then stop near the microphone 30 to wait for it to rise, the above control method also conforms to the user's actual operating habits. The entire detection process does not require the user to deliberately perform any additional operating steps, which is very convenient for the user.

[0049] It should be noted that, in other embodiments, the aforementioned step of detecting that a gesture within the sensing area matches a preset gesture can also be placed after the step of detecting that an activity trajectory within the sensing area matches a preset trajectory, such as... Figure 7 As shown.

[0050] It should be noted that, in other embodiments, the aforementioned step of detecting that the object generating the movement trajectory hovers in the sensing area for more than a first preset duration S1 is not limited to being located between the aforementioned step of detecting that the movement trajectory in the sensing area matches the preset trajectory and the step of controlling the microphone to rise, but can also be located between the step of detecting that the gesture in the sensing area matches the preset gesture and the step of controlling the microphone 30 to rise, such as... Figure 8 As shown.

[0051] Specifically, the first preset duration S1 is within the range of 0.5 seconds ≤ S1 ≤ 3 seconds. Research revealed that a hover time S1 less than 0.5 seconds is prone to accidental triggering, while a duration exceeding 3 seconds negatively impacts the user experience. Therefore, the aforementioned time setting ensures a certain level of accuracy in judgment without causing users to wait excessively.

[0052] In another embodiment, the screen 20 has a touch function. Before the control microphone 30 rises, the following step is included: detecting that the screen has not been operated within a second preset duration S2. Since the screen 20 has a touch function, when the user's hand moves within the sensing area 42, it may be operating the screen 20. Therefore, after meeting the various detection conditions before the control microphone 30 rises, detecting whether the screen 20 has been operated before controlling the microphone 30 rises can improve the accuracy of the judgment. It should be noted that in the aforementioned embodiment, the step before the control microphone 30 rises may be the step of detecting that the activity trajectory within the sensing area matches a preset trajectory, the step of detecting that the gesture within the sensing area matches a preset gesture, or the step of detecting that the object generating the activity trajectory hovers within the sensing area for more than a first preset duration S1. The aforementioned step of detecting that the screen has not been operated within the second preset duration S2 can be placed after these steps. Figure 9 One of the methods is shown.

[0053] Specifically, the second preset duration S2 begins at the end of the step preceding the step that performs the detection for the aforementioned second preset duration. That is, the second preset duration S2 can begin at the end of the step where the detected movement trajectory within the sensing area matches a preset trajectory, the step where the detected gesture within the sensing area matches a preset gesture, or the step where the detected object generating the movement trajectory hovers within the sensing area for more than the first preset duration S1. The range of the second preset duration S2 is: 0.3 seconds ≤ S2 ≤ 2 seconds. This ensures a certain level of accuracy in the judgment without causing the user to wait too long.

[0054] Please combine Figure 2 In another embodiment, the minimum distance between the sensing area 42 and the screen 20 is greater than a preset distance L. Setting the preset distance L between the sensing area 42 and the screen 20 reduces the likelihood of a user entering the sensing area 42 while operating the screen 20, thus reducing the probability of false alarms. Specifically, the preset distance L can be greater than 2 centimeters.

[0055] In another embodiment, the smart speaker system also includes a human body detection module 80 for detecting whether there is a person within a certain range; after detecting that the screen 20 is in a flipped-up state, the control method further includes the following steps: when the detection of no one for a third preset duration S3, multiple infrared detection modules 40 are turned off. In actual use, users may flip up the screen 20 and then temporarily leave or not want to remove the microphone 30. If multiple infrared detection modules 40 are kept in working state, power consumption will continue. Therefore, by detecting the human body, the human body detection module 80 can turn off multiple infrared detection modules 40 when the no-person state lasts for a third preset duration S3, which will help reduce power consumption. Specifically, the third preset duration S3 can be set to 3 minutes. When the human body detection module 80 detects someone, it activates multiple infrared detection modules 40. The human body detection module 80 can be an infrared detection module located at the edge of the screen 20.

[0056] Correspondingly, please combine Figure 10 This application also provides an MCU microcontroller module 70, applied to a smart speaker system 100 having a speaker body 10, a screen 20, a microphone 30, multiple infrared detection modules 40, and a drive module 50; the microphone 30 is at least partially housed within the speaker body 10, the multiple infrared detection modules 40 define a sensing area 42, the screen 20 can be detected to flip relative to the speaker body 10 to a flipped or closed state, and the drive module 50 is used to drive the microphone 30 to rise; the MCU microcontroller module 70 includes a screen detection module 102, a trajectory detection module 104, a signal generation module 106, and a drive control module 108; the screen detection module 102 is used to detect the state of the screen 20; the trajectory detection module 104 is used to detect whether the movement trajectory within the sensing area 42 matches a preset trajectory after the screen detection module 102 detects a flipped state; the signal generation module 106 is used to generate a control signal after the trajectory detection module 104 detects a match; the drive control module 108 is used to control the drive module 50 to drive the microphone 30 to rise after reading the control signal.

[0057] The above description is merely a preferred embodiment of this application and does not limit the patent scope of this application. Any equivalent structural transformations made based on the content of this application's specification and drawings, or direct / indirect applications in other related technical fields, are included within the patent protection scope of this application.

Claims

1. A control method for a smart speaker system based on gesture recognition, characterized in that, An intelligent speaker system comprising a speaker body, a screen, a microphone, multiple infrared detection modules, and a driving module; wherein the microphone is at least partially housed within the speaker body, the multiple infrared detection modules define a sensing area, the screen can be detected to flip relative to the speaker body to a flipped-up or closed state, and the driving module is used to drive the microphone to rise; the control method includes: The screen was detected to be in a flipped-up state; The detected activity trajectory within the sensing area matches a preset trajectory; Control the microphone to rise.

2. The control method for a smart speaker system based on gesture recognition as described in claim 1, characterized in that, Before detecting the activity trajectory within the sensing area, the following steps are also included: The gesture detected within the sensing area matches the preset gesture.

3. The control method for a smart speaker system based on gesture recognition as described in claim 2, characterized in that, After detecting that the activity trajectory within the sensing area matches the preset trajectory, the following steps are also included: An object that has generated an active trajectory is detected to hover within the sensing area for a period of more than a first preset duration S1.

4. The control method for a smart speaker system based on gesture recognition as described in claim 1, characterized in that, After detecting that the activity trajectory within the sensing area matches the preset trajectory, the process also includes: The gesture detected within the sensing area matches the preset gesture.

5. The control method for a smart speaker system based on gesture recognition as described in claim 1, characterized in that, Before controlling the microphone to rise, it also includes: An object that has generated a motion trajectory is detected to hover within the sensing area for a period of more than a first preset duration S1.

6. The control method for a smart speaker system based on gesture recognition as described in claim 3 or 5, characterized in that, The first preset duration S1 is in the range of: 0.5 seconds ≤ S1 ≤ 3 seconds.

7. The control method for a smart speaker system based on gesture recognition as described in any one of claims 1 to 5, characterized in that, The screen has a touch function, and before controlling the microphone to rise, it also includes: detecting that the screen has not been operated within a second preset time period S2.

8. The control method for a smart speaker system based on gesture recognition as described in claim 7, characterized in that, The starting time of the second preset duration S2 is when the previous step of the step that detects the second preset duration ends.

9. The control method for a smart speaker system based on gesture recognition as described in claim 8, characterized in that, The range of the second preset duration S2 is: 0.3 seconds ≤ S2 ≤ 2 seconds.

10. The control method for a smart speaker system based on gesture recognition as described in claim 1, characterized in that, The screen has a touch function, and the minimum distance between the sensing area and the screen is greater than a preset distance L.

11. The control method for a smart speaker system based on gesture recognition as described in claim 10, characterized in that, The preset distance L is greater than 2 centimeters.

12. The control method for a smart speaker system based on gesture recognition as described in claim 1, characterized in that, In the coordinate system with the user's position when operating the smart speaker system normally as the forward coordinate, the preset trajectory is any one or more combinations of left to right, right to left, front to back, and top to bottom.

13. The control method for a smart speaker system based on gesture recognition as described in claim 1, characterized in that, The smart speaker system also includes a human detection module for detecting whether there is a person within a certain range. After detecting that the screen is in a flip-up state, the control method further includes: When the unmanned state is detected for a third preset time S3, the multiple infrared detection modules are turned off.

14. A gesture recognition-based smart speaker system, comprising a speaker body and a microphone, wherein the speaker body includes an accommodating space for housing at least a portion of the microphone, characterized in that, The smart speaker system also includes: The screen can be flipped relative to the speaker body to a flipped-up or closed state; Multiple infrared detection modules define the sensing area; A driver module for driving the microphone to rise; A screen detection module is used to detect the state of the screen; MCU microcontroller module, including memory and controller; The memory is used to store program instructions; The controller is used to call and execute program instructions in the memory to implement the gesture recognition-based smart speaker system control method as described in any one of claims 1-13.

15. An MCU microcontroller module, characterized in that, This invention is applied to a smart speaker system comprising a speaker body, a screen, a microphone, multiple infrared detection modules, and a driving module; the microphone is at least partially housed within the speaker body, the multiple infrared detection modules define a sensing area, the screen can be detected to flip relative to the speaker body to a flipped-up or closed state, and the driving module is used to drive the microphone to rise. The MCU microcontroller module includes: The screen detection module is used to detect the state of the screen. The trajectory detection module is used to detect whether the activity trajectory within the sensing area matches a preset trajectory after the screen detection module detects the flipped-up state. The signal generation module is used to generate a control signal after the trajectory detection module detects a match; The drive control module is used to control the drive module to raise the microphone after reading the control signal.