Device control method and apparatus, electronic device, and medium

By combining voice recognition with continuous dynamic gestures, the problem of continuous operation control in vehicle-mounted equipment control is solved, enabling flexible, precise, and accurate adjustment of equipment status, and improving control efficiency and convenience.

CN114613362BActive Publication Date: 2026-01-06SHENZHEN HORIZON ROBOTICS TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210242711.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-11
Publication Date
2026-01-06
Estimated Expiration
2042-03-11

AI Technical Summary

Technical Problem

In existing technologies, controlling vehicle-mounted equipment via voice commands cannot achieve continuous operation control, resulting in low control efficiency and an inability to accurately adjust the equipment status.

Method used

The system identifies the target device through voice recognition and continuously adjusts it by combining preset dynamic gestures, thereby enabling the adjustment of the status of the in-vehicle equipment.

Benefits of technology

It improves the efficiency, convenience, and safety of selecting and operating vehicle-mounted equipment, and enables flexible, precise, and accurate adjustment of equipment status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114613362B_ABST
    Figure CN114613362B_ABST
Patent Text Reader

Abstract

The device control method and apparatus, the electronic device and the medium disclosed by the embodiments of the present disclosure, wherein the device control method comprises: in response to receiving a voice control instruction, performing voice recognition on the voice control instruction to obtain a first voice recognition result; determining a target device corresponding to the voice control instruction based on the first voice recognition result; and in response to detecting a preset dynamic gesture, continuously adjusting the state of the target device based on the continuous action of the dynamic gesture. The embodiments of the present disclosure can improve the efficiency and convenience of target device selection, and achieve continuous operation control of the target device, making the adjustment of the state of the target device more flexible, fine and accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to artificial intelligence technology, and in particular to a device control method and apparatus, electronic devices and media. Background Technology

[0002] Human-computer interaction (HCI) refers to the process of information exchange between humans and machines using a certain dialogue language and interactive methods to complete a specific task. Traditional HCI was mainly achieved through input / output devices such as keyboards, mice, and monitors. However, with the development of technologies such as artificial intelligence, humans and machines can now interact in a way similar to natural language.

[0003] With the increasing prevalence of intelligent vehicles, the number of onboard devices is gradually increasing, and the range of auxiliary functions they can perform is also expanding. For drivers while driving, manually operating and controlling these onboard devices to achieve their functions presents numerous inconveniences and safety risks. Summary of the Invention

[0004] To address the aforementioned technical problems, this disclosure is proposed. Embodiments of this disclosure provide a device control method and apparatus, an electronic device, and a medium.

[0005] According to one aspect of the present disclosure, a device control method is provided, comprising:

[0006] In response to receiving a voice control command, the voice control command is subjected to voice recognition to obtain a first voice recognition result;

[0007] Based on the first speech recognition result, the target device corresponding to the voice control command is determined;

[0008] In response to the detection of a preset dynamic gesture, the state of the target device is continuously adjusted based on the continuous action of the dynamic gesture.

[0009] According to another aspect of the present disclosure, a device control apparatus is provided, comprising:

[0010] The speech recognition module is used to respond to a received speech control command, perform speech recognition on the speech control command, and obtain a first speech recognition result;

[0011] The determining module is used to determine the target device corresponding to the voice control command based on the first voice recognition result obtained by the voice recognition module;

[0012] The detection module is used to detect preset dynamic gestures;

[0013] An adjustment module is used to continuously adjust the state of the target device based on the continuous action of the dynamic gesture in response to the detection module detecting the preset dynamic gesture.

[0014] According to another aspect of the present disclosure, a computer-readable storage medium is provided, the storage medium storing a computer program for performing the device control method described in any of the above embodiments of the present disclosure.

[0015] According to another aspect of the present disclosure, an electronic device is provided, the electronic device comprising:

[0016] processor;

[0017] Memory used to store the processor's executable instructions;

[0018] The processor is configured to read the executable instructions from the memory and execute the instructions to perform the device control method according to any of the above embodiments of this disclosure.

[0019] Based on the device control method, apparatus, electronic device, and medium provided in the above embodiments of this disclosure, upon receiving a voice control command, a first voice recognition result is obtained by performing voice recognition on the voice control command. Then, the target device corresponding to the voice control command is determined based on the first voice recognition result. Furthermore, when a preset dynamic gesture is detected, the state of the corresponding target device is continuously adjusted based on the continuous movement of the dynamic gesture. Therefore, the embodiments of this disclosure can determine the target device to be adjusted based on the voice control command without manually selecting the target device, thus improving the efficiency and convenience of target device selection and effectively avoiding the inconvenience of manually selecting the target device. In addition, the continuous adjustment of the state of the target device based on the continuous movement of the dynamic gesture realizes continuous operation control of the target device, making the adjustment of the target device's state more flexible, precise, and accurate, thereby improving the control effect of the target device.

[0020] This disclosure can be used to adjust the status of any device, such as home appliances, in-vehicle devices, and terminal devices. When applied to vehicles, this disclosure can improve the efficiency, convenience, and safety of selecting and operating in-vehicle devices, effectively avoiding the inconvenience and safety issues of manual operation by the driver while driving. Furthermore, continuous operation control of in-vehicle devices is achieved based on dynamic gestures, making the adjustment of the status of in-vehicle devices more flexible, precise, and accurate, thereby improving the control effect of in-vehicle devices.

[0021] The technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0022] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to offer a further understanding of the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0023] Figure 1 This is the system diagram to which this disclosure applies.

[0024] Figure 2 This is a schematic flowchart of a device control method provided in an exemplary embodiment of this disclosure.

[0025] Figure 3 This is a schematic diagram of drawing circles with individual fingers in an embodiment of this disclosure.

[0026] Figure 4 This is a schematic flowchart of a device control method provided in another exemplary embodiment of this disclosure.

[0027] Figure 5 This is a schematic flowchart of a device control method provided in yet another exemplary embodiment of this disclosure.

[0028] Figure 6 This is a schematic flowchart of a device control method provided in another exemplary embodiment of the present disclosure.

[0029] Figure 7 This is a schematic flowchart of a device control method provided in an exemplary embodiment of this disclosure.

[0030] Figure 8 This is a schematic flowchart of a device control method provided in yet another exemplary embodiment of this disclosure.

[0031] Figure 9 This is a schematic diagram of the structure of a device control apparatus provided in an exemplary embodiment of the present disclosure.

[0032] Figure 10 This is a schematic diagram of the structure of a device control apparatus provided in another exemplary embodiment of the present disclosure.

[0033] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an exemplary embodiment of this disclosure. Detailed Implementation

[0034] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.

[0035] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of this disclosure.

[0036] Those skilled in the art will understand that the terms "first," "second," etc., in the embodiments of this disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they indicate a necessary logical order between them.

[0037] It should also be understood that in the embodiments disclosed herein, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.

[0038] It should also be understood that any component, data or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless expressly defined or given to the contrary in the context.

[0039] Furthermore, the term "and / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this disclosure generally indicates that the preceding and following related objects have an "or" relationship.

[0040] It should also be understood that the description of the various embodiments in this disclosure emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.

[0041] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.

[0042] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.

[0043] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0044] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0045] The embodiments disclosed herein can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate together with a wide range of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems, etc.

[0046] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. Typically, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are executed by remote processing devices linked through communication networks. In distributed cloud computing environments, program modules can reside on local or remote computing system storage media, including storage devices.

[0047] Application Overview

[0048] Artificial intelligence (AI) enables machines to perform complex tasks that typically require human intelligence. Efficient and accurate human-machine interaction is essential for executing human instructions. In recent years, with the continuous development of AI technology, the application of voice recognition technology in in-vehicle devices has received increasing attention within the industry.

[0049] To avoid the inconvenience and safety issues of drivers manually operating and controlling on-board equipment while driving, related technologies utilize voice commands to operate and control on-board equipment.

[0050] However, through research, the inventors discovered that the method of controlling in-vehicle equipment via voice commands cannot achieve continuous operation and control of the equipment, resulting in poor control performance. For example, when controlling the opening of a vehicle window via the voice command "open window," the opening range can only be controlled according to the default settings, without precise control. If the opening range does not meet the user's expectations, multiple voice commands to "open window" are required to increase the opening range, leading to low control efficiency. Conversely, if the opening range exceeds the user's expectations, it is impossible to precisely reduce the opening range, thus failing to meet user needs.

[0051] In view of this, embodiments of the present disclosure propose a device control method and apparatus, electronic device and medium to improve the efficiency, convenience and safety of selecting and operating vehicle-mounted devices, while realizing continuous operation control of target devices.

[0052] This embodiment of the disclosure determines the target device to be adjusted through voice control commands and continuously adjusts the state of the target device through dynamic gestures. This eliminates the need for manual selection of the target device, improving the efficiency and convenience of target device selection and effectively avoiding the inconvenience of manual selection. It also enables continuous operation control of the target device, making the adjustment of the target device's state more flexible, precise, and accurate, thereby improving the control effect of the target device.

[0053] This disclosure can be used to adjust the status of any device, such as home appliances, in-vehicle devices, and terminal devices. When applied to vehicles, this disclosure can improve the efficiency, convenience, and safety of selecting and operating in-vehicle devices, effectively avoiding the inconvenience and safety issues of manual operation by the driver while driving. Furthermore, continuous operation control of in-vehicle devices is achieved based on dynamic gestures, making the adjustment of the status of in-vehicle devices more flexible, precise, and accurate, thereby improving the control effect of in-vehicle devices.

[0054] Exemplary System

[0055] Figure 1 This is the system diagram to which this disclosure applies. For example... Figure 1 As shown, voice control commands are acquired by the audio acquisition module 102 (e.g., a microphone). These voice control commands, or the voice control commands processed by the front-end signal, are input to the device control device 104 of this embodiment. The device control device 104 performs voice recognition on the received voice control commands, obtains a first voice recognition result, determines the target device 106 corresponding to the voice control command based on the voice recognition result, calls the image acquisition module 108 (e.g., a camera) to acquire a video stream, and performs preset dynamic gesture detection on the video stream acquired by the image acquisition module 108. When a preset dynamic gesture is detected, the state of the target device 106 is continuously adjusted based on the continuous movement of the dynamic gesture.

[0056] This disclosure can be used to adjust the state of any device, such as home appliances, in-vehicle devices, or terminal devices. Specifically, the target device 106 can be any device, including home appliances, in-vehicle devices, and terminal devices. When the target device 106 is an in-vehicle device, this disclosure addresses various interactive scenarios within the cockpit by using a combination of voice and dynamic gestures for human-computer interaction. It obtains control of the device by recognizing voice control commands, and then uses dynamic gestures to perform various possible continuous operations on the device. During continuous operation control, the speed of the dynamic gestures can also control the adjustment speed of the device, improving the efficiency, convenience, and safety of selecting and operating in-vehicle devices. This effectively avoids the inconvenience and safety issues associated with manual operation of in-vehicle devices by the driver while driving. Furthermore, the continuous movement of dynamic gestures enables continuous operation control of the in-vehicle device, making the adjustment of its state more flexible, precise, and accurate, thereby improving the control effect. The embodiments disclosed herein fully utilize the excellent permission interface capabilities of voice control and the fine adjustment capabilities of dynamic gestures, and are characterized by simple operation, good robustness, fine adjustment, high interaction efficiency, and wide range of functions.

[0057] Exemplary methods

[0058] Figure 2 This is a schematic flowchart of a device control method provided in an exemplary embodiment of this disclosure. This embodiment can be applied to electronic devices, such as... Figure 2 As shown, the device control method of this embodiment includes the following steps:

[0059] Step 202: In response to receiving a voice control command, perform voice recognition on the voice control command to obtain a first voice recognition result.

[0060] The voice control commands in this embodiment are either directly acquired by an audio acquisition module (such as a microphone) or obtained by front-end signal processing of the original voice control commands acquired by the audio acquisition module. This embodiment does not limit the scope of the voice control commands.

[0061] Front-end signal processing may include, but is not limited to: Voice Activity Detection (VAD), noise reduction, Acoustic Echo Cancellation (AEC), dereverberation processing, device control, beamforming (BF), etc.

[0062] Speech activity detection, also known as speech endpoint detection or speech boundary detection, refers to detecting the presence or absence of speech in an audio signal within a noisy environment. It accurately detects the start position of speech segments within the audio signal and is commonly used in speech processing systems such as speech coding and speech enhancement. This reduces speech coding speed, saves communication bandwidth, reduces mobile device power consumption, and improves recognition rates. The starting point of a VAD (Voice Activity Detection) is from silence to speech, and the ending point is from speech to silence. Determining the ending point of a VAD requires a period of silence. The speech obtained from the original audio signal after front-end signal processing includes the speech from the start to the end point of the VAD. Therefore, the voice control commands in this embodiment may also include a period of silence after the speech segment.

[0063] Step 204: Based on the first speech recognition result, determine the target device corresponding to the voice control command.

[0064] The target device corresponding to the voice control command is the device whose status needs to be adjusted. This target device can be any device such as a home appliance, in-vehicle device, or terminal device. In-vehicle devices refer to devices located in a vehicle, including, but not limited to, the following: left rearview mirror, right rearview mirror, interior rearview mirror, windows, air conditioners, seats, audio system, lights, etc. This disclosure does not limit the scope of the target device or the specific scope of the in-vehicle device.

[0065] Step 206: In response to detecting a preset dynamic gesture, continuously adjust the state of the target device based on the continuous action of the dynamic gesture.

[0066] Through step 206, the user can continuously adjust the state of the target device by making dynamic gestures until the state of the target device reaches the user's expected state effect, such as lowering the car window to the user's expected height. Stopping the dynamic gesture action will stop the adjustment of the target device's state.

[0067] Based on this embodiment, the target device to be adjusted can be determined based on voice control commands without the need for manual selection, which can improve the efficiency and convenience of target device selection and effectively avoid the inconvenience of manual selection. In addition, the continuous adjustment of the state of the target device based on the continuous action of dynamic gestures realizes continuous operation control of the target device, making the adjustment of the state of the target device more flexible, precise and accurate, thereby improving the control effect of the target device.

[0068] The preset dynamic gestures in this embodiment can be designed with the following features: (1) conforming to natural habits and easy to perform, so as to improve the convenience of the action; (2) dynamic gestures, which are more robust than static hand movements in a single frame image; (3) different from daily habitual movements, with a low probability of being falsely reported as other movements; (4) having different movement directions and being reusable.

[0069] Based on the above characteristics, in some implementation methods, the preset dynamic gesture can be, for example, drawing a circle, that is, drawing a circle in the air, such as including but not limited to drawing a circle with the left hand, the right hand, both hands, any one or more fingers, making a fist, bending fingers, etc. Figure 3 The diagram shown illustrates a circular motion drawn with individual fingers. The preset dynamic gestures in this embodiment are not limited to this; they can be any dynamic gesture exhibiting the above characteristics.

[0070] The preset dynamic gestures in this embodiment can simultaneously meet the above characteristics, possessing high robustness, few but precise gestures, conforming to natural habits, high recognition accuracy, and easy reusability, thereby improving recognition stability and accuracy and facilitating continuous operation control of the device.

[0071] Figure 4 This is a flowchart illustrating a device control method provided in another exemplary embodiment of this disclosure. Figure 4 As shown above, in the above Figure 2 Based on the illustrated embodiment, the device control method of this embodiment may further include the following steps:

[0072] Step 205: Determine the target dimension parameters to be adjusted for the target device.

[0073] The target dimension parameter refers to the dimension parameter that needs to be adjusted in the state of the target device. For example, when the target device is a window on a vehicle, the target dimension parameter could be the window's height adjustment; when the target device is a seat on a vehicle, the target dimension parameter could be the seat's fore-aft position, height, and backrest tilt angle; when the target device is a light on a vehicle, the target dimension parameter could be the light's brightness and color; when the target device is the left and right rearview mirrors on a vehicle, the target dimension parameter could be the left and right rearview mirrors' pitch and yaw angles. Similarly, when the target device is a home appliance such as a television, the target dimension parameter could be the television's channel, volume, and brightness. In this embodiment, the target dimension parameter to be adjusted can be any adjustable dimension parameter of the target device, and this embodiment does not limit the adjustable dimension parameters.

[0074] Accordingly, in this embodiment, step 206 may include:

[0075] Step 2062: In response to detecting a preset dynamic gesture, determine the direction of motion of the dynamic gesture.

[0076] Step 2064: Based on the movement direction of the dynamic gesture, determine the target adjustment direction of the target device in the target dimension parameters.

[0077] Optionally, in some implementations, the correspondence between the movement direction of the dynamic gesture and the device, the device's dimensional parameters, and the adjustment direction can be pre-defined. After determining the movement direction of the dynamic gesture, the target adjustment direction can be obtained by querying the correspondence based on the movement direction of the dynamic gesture, the target device, and the target dimensional parameters. Table 1 below shows a partial example of the correspondence between the circle-drawing direction and the device, the device's dimensional parameters, and the adjustment direction when the dynamic gesture is drawing a circle, in an embodiment of this disclosure. This is not intended to limit the specific content of the correspondence between the movement direction of the dynamic gesture and the device, the device's dimensional parameters, and the adjustment direction in the embodiments of this disclosure.

[0078] Table 1

[0079]

[0080] Step 2066: Based on the continuous movement of the dynamic gesture in the direction of the dynamic gesture, continuously adjust the target device in the target dimension parameter in the target adjustment direction.

[0081] Based on this embodiment, after determining the target dimension parameter to be adjusted of the target device, the target adjustment direction of the target device on the target dimension parameter can be determined by the movement direction of the dynamic gesture. Thus, the target dimension parameter to be adjusted and the target adjustment direction of the target device can be determined. Then, based on the continuous movement of the dynamic gesture in this movement direction, the continuous adjustment of the target device on the target dimension parameter towards the target adjustment direction can be realized, thereby realizing the continuous operation control of the target device on the target dimension parameter towards the target adjustment direction.

[0082] The target device in this embodiment can be a device whose state is determined based on a single dimension parameter. That is, the state of the device is determined by a single dimension parameter, and the device has only one adjustable dimension parameter. Each parameter value in this dimension parameter corresponds to a state of the device. For example, a window on a vehicle is a device whose state is determined based on a single dimension parameter: the height dimension. Different height values ​​of the window in the height dimension correspond to different states of the window.

[0083] Alternatively, the target device in this embodiment can also be a target device whose state is determined based on multiple dimensional parameters. That is, the state of the device is determined by these multiple dimensional parameters, and the device has multiple adjustable dimensional parameters. Each set of parameter values ​​on these multiple dimensional parameters corresponds to a state of the device. When the parameter value on any one of these dimensional parameters changes, the state of the device changes. For example, the left and right rearview mirrors on a vehicle are devices whose state is determined by two dimensional parameters: pitch angle and yaw angle. Each set (angle value on the pitch angle dimension and angle value on the yaw angle dimension) corresponds to a state of the left and right rearview mirrors, respectively. When the angle value of any or all of the parameter dimensions on the pitch angle and yaw angle dimensions changes, the state of the left and right rearview mirrors also changes.

[0084] In some implementations, in this embodiment of the disclosure, when the state of the target device is determined based on a dimension parameter, in step 205, the dimension parameter of the target device can be directly determined as the target dimension parameter.

[0085] Based on this embodiment, when the state of the target device is determined based on a single dimension parameter, and the target device has only one adjustable dimension parameter, then this single dimension parameter of the target device can be directly determined as the target dimension parameter without the user specifying the target dimension parameter that needs to be adjusted. This helps to improve the efficiency of determining the target dimension parameter, thereby improving the control efficiency of the target device.

[0086] In some implementations, in this embodiment of the disclosure, when the state of the target device is determined based on multiple dimensional parameters, in step 205, the target dimensional parameters can be determined based on the first speech recognition result.

[0087] In this embodiment, the user can directly carry relevant information about the target dimension parameter to be adjusted via voice control commands. For example, the voice control commands could be phrases like "I want to adjust the driver's seat forward and backward," "I want to adjust the driver's seat forward," "I want to adjust the driver's seat forward," or "I want to adjust the left rearview mirror tilt." This embodiment does not limit the content, form, or format of the relevant information about the target dimension parameter carried in the voice control commands. The first speech recognition result in text form obtained by performing speech recognition on the voice control command includes the relevant information about the target dimension parameter, and the target dimension parameter can be determined based on this relevant information.

[0088] For example, in a practical implementation, dimensional parameters for each device can be pre-set. After obtaining the first voice recognition result by performing voice recognition on the voice control command, for the target device, the dimensional parameter associated with or closest to the target dimensional parameter in the first voice recognition result is determined as the target dimensional parameter to be adjusted. For example, for the target device of the driver's seat, which has three dimensions: front-back dimension, height dimension, and backrest tilt angle dimension, based on the relevant information "front-back" of the target dimensional parameter in the first voice recognition result "I want to adjust the front-back of the driver's seat", the relevant information "front-back adjustment" of the target dimensional parameter in the first voice recognition result "I want to adjust the driver's seat, front-back adjustment", and the relevant information "adjust forward" of the target dimensional parameter in the first voice recognition result "I want to adjust the driver's seat forward", the dimensional parameter associated with or closest to the relevant information "front-back", "front-back adjustment", and "adjust forward" can be determined as the front-back dimension, which is then used as the target dimensional parameter to be adjusted for the driver's seat.

[0089] Furthermore, in specific implementations, a preset determination method can be adopted to determine, for the target device, the dimension parameter associated with or closest to the relevant information of the target dimension parameter in the first speech recognition result. For example, the dimension parameter whose name in the target device has the most identical characters to the relevant information of the target dimension parameter in the first speech recognition result can be identified as the associated or closest dimension parameter. Alternatively, an information list can be pre-defined, containing relevant information that may correspond to each dimension parameter of each device. Based on the relevant information of the target dimension parameter in the first speech recognition result, the information list can be queried for the target device to obtain the matching dimension parameter, which can then be used as the associated or closest dimension parameter. Additionally, embodiments of this disclosure can also employ other methods to determine the associated or closest dimension parameter of the relevant information of the target dimension parameter in the first speech recognition result, and this disclosure does not impose any limitations on these methods.

[0090] Based on this embodiment, when the state of the target device is determined based on multiple dimension parameters, the user can directly specify the target dimension parameter to be adjusted through voice control commands, without having to specify the target dimension parameter to be adjusted separately. This helps to improve the efficiency of determining the target dimension parameter, thereby improving the control efficiency of the target device.

[0091] In other implementations, when the state of the target device is determined based on multiple dimensional parameters, in step 205, in response to receiving a voice command for the dimensional parameters, the voice command for the dimensional parameters can be speech-recognized to obtain a second voice recognition result, and then the target dimensional parameters are determined based on the second voice recognition result.

[0092] In this embodiment, after sending a voice control command, the user can directly send a dimension parameter voice command. For example, after sending the voice control command "I want to adjust the driver's seat," the user can directly send the dimension parameter voice command "adjust forward and backward." Alternatively, the device implementing this embodiment can, after receiving the voice control command sent by the user, output a dimension parameter inquiry voice and receive the dimension parameter voice command sent by the user in response to the inquiry voice. For example, if the user sends the voice control command "I want to adjust the driver's seat," the device implementing this embodiment, after receiving the command, outputs the dimension parameter inquiry voice "Okay, how would you like to adjust it?" and receives the dimension parameter voice command "adjust forward and backward." Then, after performing speech recognition on the dimension parameter voice command and obtaining a second speech recognition result, the target dimension parameter can be determined based on the second speech recognition result.

[0093] In practical implementation, the dimension parameters of each device can be preset. After the second speech recognition result is obtained by performing speech recognition on the speech command of the dimension parameter, the dimension parameter associated with or closest to the second speech recognition result is determined for the target device as the target dimension parameter to be adjusted.

[0094] In practical implementation, a preset determination method can be adopted to determine the dimension parameter associated with or closest to the target device in the second speech recognition result. The specific determination method can refer to the implementation method described in the above embodiment for determining the related information of the target dimension parameter in the first speech recognition result and the associated or closest dimension parameter, which will not be elaborated here.

[0095] Based on this embodiment, when the state of the target device is determined based on multiple dimension parameters, the user can specify the dimension parameter to be adjusted through a voice command for a single dimension parameter, thereby determining the target dimension parameter that needs to be adjusted for the target device.

[0096] In some other implementations, when the state of the target device is determined based on multiple dimensional parameters, in step 205, the hand shape information corresponding to the dynamic gesture can also be obtained, and then the target dimensional parameters can be determined based on the hand shape information.

[0097] The hand shape information may include, but is not limited to, any of the following: finger extension form, number of fingers, and single / double hand information. Finger extension form may be, for example, straight or bent; the number of fingers may be, for example, one or two; and single / double hand information may be, for example, left hand, right hand, or both hands.

[0098] Specifically, a pre-defined correspondence can be established between hand shape information, the device, and the device's dimensional parameters. After obtaining the hand shape information corresponding to the dynamic gesture in step 205, the dimensional parameters corresponding to the target device and the obtained hand shape information are obtained from the above correspondence based on the target device and the obtained hand shape information, and are used as the target dimensional parameters.

[0099] Table 2 below shows a partial example of the correspondence between hand shape information, vehicle device, and dimensional parameters of vehicle device when the hand shape information is the number of fingers and the device is an in-vehicle device in this embodiment of the present disclosure.

[0100] Table 2

[0101]

[0102] Table 3 below shows a partial example of the correspondence between hand shape information, vehicle equipment, and dimensional parameters of the vehicle equipment when the hand shape information is single or double hand information and the device is an in-vehicle device in this embodiment of the present disclosure.

[0103] Table 3

[0104]

[0105]

[0106] Tables 2 and 3 above only exemplify a portion of the correspondence between hand shape information, device, and device dimensional parameters. For cases where the hand shape information is other than that in Tables 2 and 3, or the device is other than that in Tables 2 and 3 (e.g., other vehicle-mounted devices, home appliances, terminal devices, etc.), the content structure can be referenced in Tables 2 and 3, and will not be elaborated further in this embodiment.

[0107] Based on this embodiment, when the state of the target device is determined based on multiple dimensional parameters, the target dimensional parameters that need to be adjusted for the target device can be determined by the user's hand shape information.

[0108] Figure 5 This is a schematic flowchart of a device control method provided in yet another exemplary embodiment of this disclosure. For example... Figure 5 As shown above, in the above Figure 4 Based on the illustrated embodiment, step 2066 may include the following steps:

[0109] Step 20662: During the continuous movement of the dynamic gesture, the movement speed of the dynamic gesture is acquired in real time or according to a preset adjustment cycle.

[0110] In order to achieve real-time and dynamic adjustment of the state of the target device, the preset adjustment period can be set to a small value, such as 0.01s. The embodiments of this disclosure can be preset according to the specific device being operated and the adjustment effect, and can be updated as needed.

[0111] Step 20664: Based on the motion speed of the dynamic gesture, determine the target adjustment speed of the target device in the target dimension parameters.

[0112] Step 20666: Adjust the target device in the target dimension parameters at the target adjustment speed in the target adjustment direction.

[0113] In step 20666, the target device can be adjusted in the target adjustment direction at the target adjustment speed within the state limit range of the target dimension parameter. When the target device reaches the boundary of the state limit range in the target dimension parameter, such as when the window on a vehicle is lowered to the lowest or raised to the highest, the adjustment of the target device in the target dimension parameter in the target adjustment direction will no longer be carried out to avoid damaging the target device.

[0114] Based on this embodiment, the target adjustment speed of the target device in the target dimension parameter can be determined based on the movement speed of the dynamic gesture, and the target device can be adjusted in the target adjustment direction at the target adjustment speed. In this way, the faster the movement speed of the dynamic gesture, the faster the device adjustment speed, and vice versa. This realizes dynamic control of the device adjustment speed based on the movement speed of the dynamic gesture, realizes visualized control of the device adjustment speed, and improves the adjustment efficiency of the target device and the user's operating experience.

[0115] Optionally, in some implementations, adjustment speed configuration information of the target device on the target dimension parameters can be obtained. This adjustment speed configuration information is used to determine the relationship between gesture movement speed and device adjustment speed on each dimension parameter of the target device. For example, for a window on a vehicle, the relationship between gesture movement speed and window lifting speed on the lifting dimension can be linearly correlated with the device adjustment speed. Accordingly, in step 20664, based on the obtained adjustment speed configuration information, the device adjustment speed corresponding to the movement speed of the dynamic gesture obtained in step 20662 on the target dimension parameter can be determined as the target adjustment speed.

[0116] Based on this embodiment, the target adjustment speed of the target device in the target dimension parameters can be objectively and accurately determined based on the movement speed of the dynamic gesture according to the pre-set adjustment speed configuration information, so as to achieve accurate control of the adjustment speed of the target device state.

[0117] In some specific implementations, the adjustment speed configuration information of the target device in the target dimension parameters can be obtained from the first speech recognition result.

[0118] In this embodiment, the user can directly carry adjustment speed configuration information through voice control commands. For example, the voice control command could be the voice "I want to adjust the driver's side window; turning it three times will raise the entire window," which includes the adjustment speed configuration information "turning it three times will raise the entire window." This embodiment does not limit the content, form, or format of the adjustment speed configuration information carried in the voice control command. After performing voice recognition on the voice control command to obtain the first voice recognition result, the adjustment speed configuration information of the target device on the target dimension parameter can be obtained from the first voice recognition result, thereby determining the relationship between the gesture movement speed and the device adjustment speed on the target dimension parameter of the target device.

[0119] Based on this embodiment, users can directly set the adjustment speed configuration information of the target device on the target dimension parameters through voice control commands during the operation of the device, thereby realizing the real-time and dynamic configuration of the adjustment speed configuration information in specific scenarios and realizing the personalized configuration of the device adjustment effect.

[0120] Alternatively, in some other specific implementations, the adjustment speed configuration information of the target device on the target dimension parameter can be obtained in the following way: in response to receiving the adjustment speed configuration voice command, the adjustment speed configuration voice command is subjected to speech recognition to obtain a third speech recognition result. Then, the adjustment speed configuration information of the target device on the target dimension parameter is obtained from the third speech recognition result, thereby determining the relationship between the gesture movement speed and the device adjustment speed on the target dimension parameter of the target device. The speed adjustment configuration voice command can be a user-initiated voice command, such as sending the voice control command "I want to adjust the driver's seat forward" followed by the voice command "Turn three times to raise the entire window"; or it can be a user-initiated voice command based on the speed adjustment prompt voice output by the device used to implement the embodiments of this disclosure, such as sending the voice control command "I want to adjust the driver's seat forward" followed by the speed adjustment prompt voice output by the device used to implement the embodiments of this disclosure, "Okay, what speed would you like to adjust at?", followed by the voice command "Turn three times to raise the entire window". The embodiments of this disclosure do not limit the method and specific content of the user sending the speed adjustment configuration voice command.

[0121] Based on this embodiment, users can set the adjustment speed configuration information of the target device on the target dimension parameters through a single command during the operation of the device, thereby realizing the real-time and dynamic configuration of the adjustment speed configuration information in specific scenarios and realizing the personalized configuration of the device adjustment effect.

[0122] Alternatively, in some specific implementations, the adjustment speed configuration information of the target device on the target dimension parameter can be obtained from the pre-configured adjustment speed configuration information, thereby determining the relationship between the gesture movement speed and the device adjustment speed on the target dimension parameter of the target device.

[0123] The pre-configured speed adjustment information can be pre-configured by the user. Taking in-vehicle devices as an example, users can set or update the speed adjustment configuration information of each in-vehicle device through the speed adjustment configuration page provided by the vehicle's central control system, for example, through the configuration options for each in-vehicle device on the speed adjustment configuration page, or through human-machine voice interaction on the speed adjustment configuration page. Alternatively, users can also access the speed adjustment configuration permissions provided by the central control system through human-machine voice interaction and set the speed adjustment configuration information of each in-vehicle device through human-machine voice interaction. For other devices (such as home appliances, terminal devices, etc.), the speed adjustment configuration information of each device can be set or updated through the speed adjustment configuration page provided by the control device that provides unified control over these devices, using a similar method to that for in-vehicle devices.

[0124] When the user has not pre-configured the speed adjustment settings, the factory-preset information from the central control system (for in-vehicle devices), control devices (for home appliances, terminal devices, and other devices) can be used as the pre-configured speed adjustment settings.

[0125] Based on this embodiment, when the user has not set adjustment speed configuration information for the current scenario, the adjustment speed configuration information of the target device on the target dimension parameter can be obtained from the pre-configured adjustment speed configuration information, so as to determine the target adjustment speed of the target device in the current scenario.

[0126] For example, in specific applications, the speed adjustment configuration information can be pre-configured in the following way:

[0127] By setting up interfaces, such as the interfaces on the speed adjustment configuration page provided by the central control system (for in-vehicle devices) and control devices (for home appliances, terminal devices, and other devices), the system receives speed adjustment configuration requests sent by users. These requests include device identifier (ID), dimension parameter ID, gesture movement amplitude (e.g., one circle), and device adjustment amplitude (e.g., 0.5cm). The device ID is used to uniquely identify a device, and the dimension parameter ID is used to uniquely identify a dimension parameter.

[0128] Based on the gesture amplitude and device adjustment amplitude information in the speed adjustment configuration request, determine the relationship between the gesture speed and the device adjustment speed;

[0129] Based on the relationship between the device ID, dimension parameter ID, gesture speed, and device adjustment speed in the adjustment speed configuration request, configure the adjustment speed configuration information of the device identified by the device ID on the dimension parameter identified by the dimension parameter ID; or, based on the relationship between the device ID, dimension parameter ID, gesture speed, and device adjustment speed in the adjustment speed configuration request, update the adjustment speed configuration information corresponding to the device ID and the dimension parameter ID in the pre-configured adjustment speed configuration information.

[0130] Based on this embodiment, the configuration or update of the speed adjustment configuration information of the device in terms of dimensional parameters is realized.

[0131] In addition, in the above embodiments, during the execution of step 206 or 2066, in response to receiving the adjustment speed update voice command, the adjustment speed update voice command is subjected to speech recognition to obtain a fourth speech recognition result, and adjustment speed update configuration information is obtained from the fourth speech recognition result. The adjustment speed update configuration information is used to represent the relationship between the updated gesture movement speed and the device adjustment speed on each dimension parameter of the target device. Then, during the subsequent continuous action of the dynamic gesture, the movement speed of the dynamic gesture is obtained in real time or according to a preset adjustment cycle, and based on the above adjustment speed update configuration information, the updated device adjustment speed corresponding to the movement speed of the dynamic gesture on the target dimension parameter is determined. Then, the target device is adjusted in the target dimension parameter with the updated adjustment speed in the target adjustment direction.

[0132] During continuous adjustment of the target device, the user may find that the adjustment speed is too fast or too slow. Based on this embodiment, the user can send an adjustment speed update voice command to update the adjustment speed configuration information according to the adjustment effect requirements during the adjustment of the target device, thereby realizing real-time update of the adjustment speed of the target device, further improving the adjustment efficiency, adjustment effect and user operation experience of the target device.

[0133] In addition, the above embodiments of this disclosure may also include a step of preset dynamic gesture detection.

[0134] Figure 6 This is a schematic flowchart of a device control method provided in yet another exemplary embodiment of this disclosure. For example... Figure 6 As shown, in some implementations, preset dynamic gestures can be detected in the following ways:

[0135] Step 302: Determine the location of the sound source object that sends the voice control command.

[0136] For example, the location of the sound source object sending voice control commands can be determined by using the sound zone localization method.

[0137] Step 304: Based on the location of the sound source object, obtain an image sequence including the hand of the sound source object.

[0138] The image sequence includes multiple frames of images that have a temporal relationship.

[0139] Once the location of the sound source object is determined, an image acquisition module (such as a camera) can be invoked to acquire images of the sound source object. The acquired images are then used for hand detection and tracking to obtain a video stream that includes the hand of the sound source object. Multiple frames with temporal relationships are selected from the video stream according to a preset method (such as continuous selection or frame-by-frame selection) to form an image sequence of the sound source object's hand. Alternatively, images of uniform size containing the hand can be extracted from the selected multiple frames to obtain an image sequence of the sound source object's hand.

[0140] The method of extracting hand image sequences from selected multi-frame images, compared to image sequences of the sound source object, can improve the accuracy of gesture detection results because the images contain less background information and have less interference.

[0141] In a practical implementation, a first neural network, such as a convolutional neural network (CNN), can be used to detect and track the hand in the acquired image, resulting in a video stream that includes the hand of the sound source object. This first neural network can be trained in advance using sample images that include the hand.

[0142] Step 306: Perform hand keypoint detection on each frame of the image sequence in sequence to obtain the hand keypoint sequence.

[0143] The hand keypoint sequence is formed by the hand keypoints in each frame of the image based on temporal relationships.

[0144] In a practical implementation, a second neural network, such as a CNN, can be used to detect hand keypoints in each frame of the image. This second neural network can be trained in advance using sample images labeled with hand keypoint information.

[0145] Step 308: Based on the hand key point sequence, perform preset dynamic gesture detection.

[0146] In a practical implementation, the sequence of hand key points can be input into a third neural network, such as a CNN, which then outputs a preset gesture detection result indicating whether a preset dynamic gesture is performed. This third neural network can be pre-trained using sample videos of people performing the preset dynamic gesture.

[0147] Based on this embodiment, by acquiring an image sequence including the hand of the sound source object, a preset dynamic gesture is detected using a vision-based method, so as to trigger the adjustment of the state of the target device when the preset dynamic gesture is detected.

[0148] Accordingly, in Figure 6 Based on the illustrated embodiment, the direction of motion of the dynamic gesture can be determined based on the hand key point sequence obtained in step 306. For example, the direction of motion of the dynamic gesture can be determined based on the direction corresponding to the trajectory of the hand key point sequence.

[0149] Based on this embodiment, the direction of dynamic gesture movement is determined by using the hand key point sequence corresponding to the image sequence, based on visual technology.

[0150] Additionally, the movement speed of the dynamic gesture can be obtained based on the hand key points in the last frame and the previous frame of the image sequence obtained in step 304, as well as the acquisition time of the last frame and the acquisition time of the previous frame. The previous frame can be any frame in the image sequence preceding the last frame, such as the frame immediately preceding the last frame, or an image separated from the last frame by several frames. This embodiment does not impose any limitations on this.

[0151] For example, the movement speed of a dynamic gesture can be calculated based on the distance between the hand key points in the last frame of an image sequence and the hand key points in the previous frame, as well as the time between the acquisition time of the last frame and the acquisition time of the previous frame. The distance between the hand key points in the last frame and the hand key points in the previous frame can be the average distance between corresponding hand key points in the last and previous frames, or it can be the distance between preset hand key points (e.g., fingertip key points) in the last and previous frames, etc., and this disclosure does not limit this.

[0152] In a specific implementation, the sequence of hand key points can be input into the aforementioned third neural network. The third neural network outputs the direction and amplitude of the dynamic gesture corresponding to the hand key point sequence (e.g., the rotation angle). Then, based on the amplitude of the movement and the time corresponding to the image sequence, the speed of the dynamic gesture can be calculated. Alternatively, an image sequence carrying acquisition time information and labeled hand key points can be input into the aforementioned third neural network. The third neural network outputs the direction and speed of the dynamic gesture corresponding to the hand key point sequence, and so on. This disclosure does not limit the scope of the embodiments.

[0153] Based on this embodiment, the movement speed of a dynamic gesture can be accurately determined by using the key hand points corresponding to two frames in the image sequence and the image acquisition time.

[0154] Figure 7 This is a schematic flowchart of a device control method provided in another exemplary embodiment of this disclosure. Figure 7 As shown, in some other implementations, preset dynamic gestures can also be detected in the following way:

[0155] Step 402: Determine the location of the sound source object that sends the voice control command.

[0156] For example, the location of the sound source object sending voice control commands can be determined by using the sound zone localization method.

[0157] Step 404: Based on the position of the sound source object, use an optical time-of-flight (ToF) sensor to measure the distance information between each point on the hand of the sound source object and the ToF sensor, and obtain a set of distance information.

[0158] Once the location of the sound source object is determined, the distance between the ToF sensor and various points on the hand of the sound source object can be measured using a ToF sensor. A set of distance information is obtained at each measurement time, including the distance information between various points on the hand of the sound source object and the ToF sensor at that measurement time.

[0159] Step 406: Based on multiple sets of distance information with temporal relationships, a distance information sequence is obtained.

[0160] Step 408: Based on the distance information sequence, perform preset dynamic gesture detection.

[0161] Optionally, in some implementations, based on the distance information sequence, the distance between each point of the sound source object's hand and the ToF sensor can be known over time. Thus, it can be determined whether the sound source object's hand has made a preset dynamic gesture based on whether the distance change conforms to the distance change pattern corresponding to the preset dynamic gesture.

[0162] Alternatively, in other implementations, three-dimensional (3D) modeling can be performed on each set of distance information in the distance information sequence to obtain the corresponding hand pose. The hand pose corresponding to the distance information sequence can be used to determine whether the sound source object's hand makes a preset dynamic gesture.

[0163] Based on this embodiment, a preset dynamic gesture is detected using a ToF sensor, so that the state of the target device is adjusted when the preset dynamic gesture is detected.

[0164] Accordingly, in Figure 7 Based on the illustrated embodiment, the direction of motion of the dynamic gesture can be determined based on the distance information sequence obtained in step 406. For example, the direction of motion of the dynamic gesture corresponding to the distance information sequence obtained in step 406 can be determined according to the variation law of the distance corresponding to the preset dynamic gesture in different directions of motion.

[0165] Based on this embodiment, the distance changes between various points on the hand of the sound source object are detected by the ToF sensor, thereby realizing the determination of the direction of dynamic gesture movement.

[0166] Additionally, the movement speed of the dynamic gesture can be obtained based on the last set of distance information and the previous set of distance information in the distance information sequence obtained in step 406, as well as the measurement time corresponding to the last set of distance information and the measurement time corresponding to the previous set of distance information. The previous set of distance information can be any set of distance information in the distance information sequence preceding the last set of distance information. For example, it can be the previous set of distance information adjacent to the last set of distance information, or it can be a set of distance information separated from the last set of distance information by several sets of distance information. This embodiment of the present disclosure does not impose any limitations on this.

[0167] For example, the movement speed of a dynamic gesture can be calculated based on the distance change between the last set of distance information and the previous set of distance information in the distance information sequence, as well as the time between the measurement time corresponding to the last set of distance information and the measurement time corresponding to the previous set of distance information. The distance change between the last set of distance information and the previous set of distance information can be the average distance change between corresponding points on the hand in the last set of distance information and the previous set of distance information, or it can be the distance change between preset hand points (e.g., fingertips) in the last set of distance information and the previous set of distance information, etc. This disclosure does not limit this.

[0168] Based on this embodiment, the movement speed of a dynamic gesture can be accurately determined by using two sets of distance information and the measurement time in the distance information sequence.

[0169] Figure 8 This is a schematic flowchart of a device control method provided in yet another exemplary embodiment of this disclosure. For example... Figure 8 As shown, in some other implementations, preset dynamic gestures can also be detected in the following ways:

[0170] Step 502: Determine the location of the sound source object that sends the voice control command.

[0171] For example, the location of the sound source object sending voice control commands can be determined by using the sound zone localization method.

[0172] Step 504: Based on the location of the sound source object, use a wearable device to obtain the location of each point on the hand of the sound source object, and obtain the hand position information.

[0173] The hand position information includes the position information of various points on the hand.

[0174] The wearable device in this embodiment can be, for example, a smart glove, smart glasses, or other smart device. The smart glove can directly locate the position of each point on the hand at any time, and the smart glasses can obtain the position of each point on the hand through visual means. This embodiment does not limit the specific wearable device used or the method by which it obtains the position of each point on the hand of the sound source object.

[0175] Step 506: Determine the hand posture based on the hand position information.

[0176] Step 508: Determine the hand movement based on the hand posture at multiple moments.

[0177] Step 510: Confirm whether the hand movement is a preset dynamic gesture.

[0178] Step 512: In response to the hand movement, the action is a preset dynamic gesture, and the preset dynamic gesture is confirmed to be detected.

[0179] Otherwise, if the hand gesture is not a preset dynamic gesture, it is confirmed that no preset dynamic gesture has been detected.

[0180] Based on this embodiment, wearable devices can directly acquire the positions of various points on the hand of the sound source object, thereby determining the hand posture. Based on the hand posture at multiple moments, the hand movement can be determined, thereby confirming whether it is a preset dynamic gesture, so as to trigger the adjustment of the target device's state when a preset dynamic gesture is detected.

[0181] Accordingly, in Figure 8 Based on the illustrated embodiment, the direction of motion of the dynamic gesture can be determined based on the hand postures at multiple moments determined in step 506. For example, the direction of motion of the dynamic gesture can be determined based on the changes in hand postures at multiple moments. Alternatively, the direction of motion of the dynamic gesture can be directly determined based on the hand movements determined in step 508.

[0182] Based on this embodiment, by acquiring hand position information through a wearable device, the direction of dynamic gesture movement can be determined.

[0183] Additionally, the movement speed of the dynamic gesture can be obtained based on the last and previous moments among the multiple moments obtained in step 504, as well as the hand position information at the last moment and the hand position information at the previous moment. The moments can be the times when the wearable device acquires information about the positions of various points on the sound source object's hand. The wearable device can acquire the positions of various points on the sound source object's hand according to a preset information acquisition cycle (e.g., 0.01s), so the time interval between two information acquisition moments is 0.01s. The previous moment can be a moment before the last moment, or a moment located before the last moment and spaced a preset number of moments (e.g., 2) apart from the last moment; this embodiment of the present disclosure does not impose any limitations on this.

[0184] For example, the movement speed of a dynamic gesture can be calculated based on the change between the hand position information at the last moment and the hand position information at the previous moment, as well as the time between the measurement moments at the last moment and the previous moment. The change between the hand position information at the last moment and the hand position information at the previous moment can be the average of the distance changes between corresponding points on the hand in the hand position information at the last moment and the previous moment, or it can be the distance changes between preset hand points (e.g., fingertips) in the hand position information at the last moment and the previous moment, etc. This embodiment does not limit the scope of the invention.

[0185] Based on this embodiment, by utilizing the hand position information of the sound source object at different times obtained by the wearable device, the movement speed of dynamic gestures can be accurately determined.

[0186] The following are some exemplary application scenarios of the embodiments of this disclosure:

[0187] Scenario 1: Adjusting vehicle windows (car windows):

[0188] The user sends a voice control command, "I want to adjust the driver's side window with gestures." Upon receiving this command, the device in this embodiment performs voice recognition, determines the target device as the driver's side window based on the first voice recognition result, and gains control of the driver's side window. The user draws a circle clockwise, causing the driver's side window to rise continuously. During the continuous descent of the driver's side window, the user sends a voice command to update the adjustment speed, "Too slow, three circles will raise the entire window." The device in this embodiment then determines the update speed corresponding to the user's circling action and controls the driver's side window to rise at that updated speed. The user continues circling until the driver's side window is adjusted to the user's desired height.

[0189] Scenario 2: Adjusting the fore-and-aft position of the vehicle seats:

[0190] The user sends a voice control command, "I want to adjust the driver's seat forward, turn it one full circle and move it forward one centimeter." Upon receiving the voice control command, the device in this embodiment performs voice recognition. Based on the first voice recognition result, it determines the target device as the driver's seat, the target dimension parameter as forward / backward, and the adjustment speed configuration information as "turn one full circle and move it forward one centimeter," and then accesses the driver's seat control permissions. The user then draws a clockwise circle, and the driver's seat moves forward continuously. The user continues drawing circles until the driver's seat is adjusted to the desired position.

[0191] Scenario 3: Adjusting the left rearview mirror on the vehicle based on hand shape information:

[0192] The user sends a voice control command, "I want to adjust the left rearview mirror with gestures." Upon receiving this command, the device in this embodiment performs voice recognition, determines the target device as the left rearview mirror based on the first voice recognition result, and accesses control of the left rearview mirror. The user draws a circle counter-clockwise with their right hand, causing the left rearview mirror to continuously tilt downwards; the user draws a circle clockwise, causing the left rearview mirror to continuously tilt upwards; the user draws a circle counter-clockwise with their left hand, causing the left rearview mirror to continuously tilt outwards; the user draws a circle clockwise, causing the left rearview mirror to continuously tilt inwards. Alternatively, the user extends their right index finger and draws a circle counter-clockwise, causing the left rearview mirror to continuously tilt downwards; the user draws a circle clockwise, causing the left rearview mirror to continuously tilt upwards; the user simultaneously extends their right index and middle fingers and draws a circle counter-clockwise, causing the left rearview mirror to continuously tilt outwards; the user draws a circle clockwise, causing the left rearview mirror to continuously tilt inwards. The specific adjustment speed can be determined by obtaining pre-configured adjustment speed information, or, referring to the above scenarios one and two, the adjustment speed configuration information can be configured via user voice commands. The user continues drawing circles until the left rearview mirror is adjusted to the user's desired direction.

[0193] Scenario 4: Adjusting the air conditioner's fan speed:

[0194] The user sends a voice control command, "I want to adjust the airflow of the air conditioner using gestures." Upon receiving this command, the device in this embodiment performs voice recognition, determines the target device as an air conditioner based on the first voice recognition result, identifies the target dimension parameter as airflow, and accesses the air conditioner's control permissions. The user draws a circle clockwise to increase the airflow, and a circle counter-clockwise to decrease it. The specific adjustment speed can be determined from pre-configured adjustment speed settings. During the airflow adjustment process, the user sends a voice command to update the adjustment speed. The user continues drawing circles until the airflow reaches the user's desired level.

[0195] Users can adjust the temperature, direction, etc. of the air conditioner in a similar way, which will not be elaborated here.

[0196] Any of the device control methods provided in this disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices and servers. Alternatively, any of the device control methods provided in this disclosure can be executed by a processor, such as by a processor executing any of the device control methods mentioned in this disclosure by calling corresponding instructions stored in memory. Further details will not be elaborated below.

[0197] Exemplary device

[0198] The device control apparatus of this disclosure can be used to implement the device control methods of the above embodiments of this disclosure.

[0199] Figure 9This is a schematic diagram of the structure of a device control apparatus provided in an exemplary embodiment of this disclosure. For example... Figure 9 As shown, the device control apparatus of this embodiment includes: a voice recognition module 602, a first determination module 604, a detection module 606, and an adjustment module 608. Wherein:

[0200] The speech recognition module 602 is used to respond to a received speech control command, perform speech recognition on the speech control command, and obtain a first speech recognition result.

[0201] The first determining module 604 is used to determine the target device corresponding to the voice control command based on the first voice recognition result obtained by the voice recognition module 602.

[0202] The target device corresponding to the voice control command is the device whose status needs to be adjusted. This target device can be any device such as a home appliance, in-vehicle device, or terminal device. In-vehicle devices refer to devices located in a vehicle, including, but not limited to, the following: left rearview mirror, right rearview mirror, interior rearview mirror, windows, air conditioners, seats, audio system, lights, etc. This disclosure does not limit the scope of the target device or the specific scope of the in-vehicle device.

[0203] The detection module 606 is used to detect preset dynamic gestures.

[0204] The detection of preset dynamic gestures in this embodiment may include, but is not limited to, drawing circles.

[0205] The adjustment module 608 is used to continuously adjust the state of the target device determined by the first determination module 604 based on the continuous action of the detection module 606 detecting a preset dynamic gesture.

[0206] Based on this embodiment, the target device to be adjusted can be determined based on voice control commands without the need for manual selection, which can improve the efficiency and convenience of target device selection and effectively avoid the inconvenience of manual selection. In addition, the continuous adjustment of the state of the target device based on the continuous action of dynamic gestures realizes continuous operation control of the target device, making the adjustment of the state of the target device more flexible, precise and accurate, thereby improving the control effect of the target device.

[0207] Figure 10 This is a schematic diagram of the structure of a device control apparatus provided in another exemplary embodiment of this disclosure. For example... Figure 10 As shown, in Figure 9 Based on the illustrated embodiment, the device control device of this embodiment may further include: a second determining module 702, used to determine the target dimension parameters to be adjusted of the target device.

[0208] Accordingly, the adjustment module 608 may include: a first determining unit 6082, used to determine the direction of motion of the dynamic gesture; a second determining unit 6084, used to determine the target adjustment direction of the target device in the target dimension parameter based on the direction of motion of the dynamic gesture; and an adjustment unit 6086, used to continuously adjust the target device in the target dimension parameter in the target adjustment direction based on the continuous action of the dynamic gesture in the direction of motion.

[0209] Optionally, in some implementations, the state of the target device is determined based on a dimensional parameter. Accordingly, in this embodiment, the second determining module 702 is specifically used to determine that a dimensional parameter of the target device is a target dimensional parameter.

[0210] Optionally, in other implementations, the state of the target device is determined based on multiple dimensional parameters. Accordingly, in this embodiment, the second determining module 702 is specifically used to determine the target dimensional parameters based on the first speech recognition result.

[0211] Optionally, in some implementations, the state of the target device is determined based on multiple dimensional parameters. Accordingly, in this embodiment, the speech recognition module 602 is further configured to, in response to receiving a dimensional parameter speech command, perform speech recognition on the dimensional parameter speech command to obtain a second speech recognition result. The second determining module 702 is specifically configured to determine the target dimensional parameters based on the second speech recognition result.

[0212] Alternatively, in some implementations, the state of the target device is determined based on multiple dimensional parameters. See also... Figure 10 The device control apparatus of this embodiment may further include: a first acquisition module 704, specifically used to acquire hand shape information corresponding to dynamic gestures. The hand shape information may include, but is not limited to, any of the following: finger extension form, number of fingers, single / double hand information, etc. Specifically, the finger extension form may be, for example, straight or bent; the number of fingers may be, for example, one or two; the single / double hand information may be, for example, left hand, right hand, or both hands. Correspondingly, a second determination module 702 is specifically used to determine target dimension parameters based on the hand shape information acquired by the first acquisition module 704.

[0213] See also Figure 10In another embodiment of the device control apparatus, it may further include a second acquisition module 706 and a third determination module 708. The second acquisition module 706 is used to acquire the movement speed of the dynamic gesture in real time or according to a preset adjustment cycle during the continuous action of the dynamic gesture. The third determination module 708 is used to determine the target adjustment speed of the target device in the target dimension parameter based on the movement speed of the dynamic gesture. Accordingly, in this embodiment, the adjustment unit 6086 is specifically used to adjust the target device in the target dimension parameter in the target adjustment direction at the target adjustment speed.

[0214] See also Figure 10 In another embodiment of the device control apparatus, a third acquisition module 710 may be included, configured to acquire adjustment speed configuration information of the target device on target dimension parameters, wherein the adjustment speed configuration information represents the relationship between gesture movement speed and device adjustment speed on each dimension parameter of the target device. Accordingly, in this embodiment, the third determination module 708 is specifically configured to determine, based on the adjustment speed configuration information, the device adjustment speed corresponding to the movement speed of the dynamic gesture on the target dimension parameters as the target adjustment speed.

[0215] Optionally, in some implementations, the third acquisition module 710 is specifically used to acquire the adjustment speed configuration information of the target device on the target dimension parameters from the first speech recognition result.

[0216] Alternatively, in some other implementations, the third acquisition module 710 is specifically used to acquire the target device's adjustment speed configuration information on the target dimension parameters from the pre-configured adjustment speed configuration information.

[0217] Or, see again Figure 10 In some implementations, the speech recognition module 602 can also be used to respond to a received speech command for adjusting speed configuration, perform speech recognition on the speech command for adjusting speed configuration, and obtain a third speech recognition result. Accordingly, in this embodiment, the third acquisition module 710 is specifically used to acquire the adjustment speed configuration information of the target device on the target dimension parameter from the third speech recognition result;

[0218] See also Figure 10In another embodiment of the device control apparatus, it may further include: a configuration module 712, configured to receive an adjustment speed configuration request through a setting interface, the adjustment speed configuration request including a device identifier, a dimension parameter identifier, a gesture movement amplitude, and device adjustment amplitude information, wherein the device identifier is used to uniquely identify a device, and the dimension parameter identifier is used to uniquely identify a dimension parameter; determine the relationship between the gesture movement amplitude and the device adjustment amplitude information based on the gesture movement amplitude and the device adjustment amplitude information; configure the adjustment speed configuration information of the device identified by the device identifier on the dimension parameter identified by the dimension parameter identifier based on the relationship between the device identifier, the dimension parameter identifier, the gesture movement speed, and the device adjustment speed; or update the adjustment speed configuration information corresponding to the device identifier and the dimension parameter identifier in the pre-configured adjustment speed configuration information based on the relationship between the device identifier, the dimension parameter identifier, the gesture movement speed, and the device adjustment speed.

[0219] Optionally, in some implementations, the speech recognition module 602 is further configured to, in response to receiving an adjustment speed update speech command, perform speech recognition on the adjustment speed update speech command to obtain a fourth speech recognition result during the continuous adjustment of the target device in the target dimension parameters in the target adjustment direction based on the continuous action of the dynamic gesture in the motion direction. Correspondingly, in this embodiment, the third acquisition module 710 is further configured to acquire adjustment speed update configuration information from the fourth speech recognition result. The adjustment speed update configuration information is used to represent the relationship between the updated gesture movement speed and the device adjustment speed in each dimension parameter of the target device. The second acquisition module 706 is further configured to acquire the movement speed of the dynamic gesture in real time or according to a preset adjustment cycle during the subsequent continuous action of the dynamic gesture. The third determination module 708 is further configured to determine the updated device adjustment speed corresponding to the movement speed of the dynamic gesture in the target dimension parameters based on the adjustment speed update configuration information. The adjustment unit 6086 is further configured to adjust the target device in the target dimension parameters in the target adjustment direction with the updated adjustment speed.

[0220] See also Figure 10 In another embodiment of the device control apparatus, it may further include: a fourth determining module 714, used to determine the location of the sound source object that sends the voice control command.

[0221] Accordingly, in some implementations, the detection module 606 is specifically used to: acquire an image sequence including the hand of the sound source object based on the position of the sound source object, the image sequence including multiple frames of images with temporal relationship; sequentially perform hand key point detection on each frame of the image sequence to obtain a hand key point sequence, the hand key point sequence being formed by the hand key points in each frame of the image based on temporal relationship; and perform preset dynamic gesture detection based on the hand key point sequence.

[0222] Accordingly, in this embodiment, the first determining unit 6082 is specifically used to determine the direction of motion of a dynamic gesture based on the sequence of key points of the hand.

[0223] Accordingly, in this embodiment, the second acquisition module 706 is specifically used to acquire the motion speed of the dynamic gesture based on the hand key points in the last frame image and the hand key points in the previous frame image in the image sequence, as well as the acquisition time of the last frame image and the acquisition time of the previous frame image.

[0224] Correspondingly, in some other implementations, the detection module 606 is specifically used to: measure the distance information between each point of the sound source object's hand and the ToF sensor based on the position of the sound source object, and obtain a set of distance information; obtain a distance information sequence based on multiple sets of distance information with temporal relationships; and perform preset dynamic gesture detection based on the distance information sequence.

[0225] Accordingly, in this embodiment, the first determining unit 6082 is specifically used to determine the movement direction of the dynamic gesture based on the distance information sequence.

[0226] Accordingly, in this embodiment, the second acquisition module 706 is specifically used to acquire the motion speed of the dynamic gesture based on the last set of distance information and the previous set of distance information in the distance information sequence, as well as the measurement time corresponding to the last set of distance information and the measurement time corresponding to the previous set of distance information.

[0227] Accordingly, in some other implementations, the detection module 606 is specifically used to: based on the position of the sound source object, use a wearable device to obtain the position of each point of the sound source object's hand, and obtain hand position information, which includes the position information of each point of the hand; determine the hand posture based on the obtained hand position information; determine the hand action based on the hand posture at multiple moments; confirm whether the hand action is a preset dynamic gesture action, and confirm that a preset dynamic gesture has been detected in response to whether the hand action is a preset dynamic gesture action.

[0228] Accordingly, in this embodiment, the first determining unit 6082 is specifically used to determine the direction of motion of the dynamic gesture based on the posture of the hand at the multiple moments.

[0229] Accordingly, in this embodiment, the second acquisition module 706 is specifically used to acquire the movement speed of the dynamic gesture based on the last moment and the previous moment among the plurality of moments, as well as the hand position information of the last moment and the hand position information of the previous moment.

[0230] Exemplary electronic devices

[0231] Figure 11This is a schematic diagram of the structure of an electronic device provided in an exemplary embodiment of the present disclosure. Hereinafter, reference is made to… Figure 11 This describes an electronic device according to embodiments of the present disclosure. The electronic device may be either or both of a first device 100 and a second device 200, or a standalone device independent of them, which may communicate with the first and second devices to receive acquired input signals from them.

[0232] like Figure 11 As shown, the electronic device includes one or more processors 802 and memory 804.

[0233] The processor 802 may be a central processing unit (CPU) or other form of processing unit with data processing and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0234] The memory 804 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 802 may execute the program instructions to implement the device control methods of the various embodiments of this disclosure described above and / or other desired functions. Various contents such as input signals, signal components, and noise components may also be stored in the computer-readable storage medium.

[0235] In one example, the electronic device may also include an input device 806 and an output device 808, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0236] For example, when the electronic device is a first device 100 or a second device 200, the input device 806 can be the aforementioned microphone or microphone array for capturing input signals from a sound source. When the electronic device is a standalone device, the input device 806 can be a communication network connector for receiving the acquired input signals from the first device 100 and the second device 200.

[0237] In addition, the input device 806 may also include, for example, a keyboard, a mouse, etc.

[0238] The output device 808 can output various information to the outside, including determined distance information, direction information, etc. The output device 808 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0239] Of course, for the sake of simplicity, Figure 11 Only some of the components of the electronic device relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device may include any other suitable components depending on the specific application.

[0240] Exemplary computer program products and computer-readable storage media

[0241] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products comprising computer program instructions that, when executed by a processor, cause the processor to perform the steps of the device control methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.

[0242] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0243] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the steps of the device control methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.

[0244] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0245] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0246] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0247] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0248] The methods and apparatus of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the methods according to this disclosure.

[0249] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.

[0250] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0251] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A device control method, comprising: in response to receiving a voice control instruction, performing voice recognition on the voice control instruction to obtain a first voice recognition result; based on the first voice recognition result, determining a target device corresponding to the voice control instruction; determining a target dimension parameter of the target device to be adjusted, when a state of the target device is determined based on a plurality of dimension parameters, a group of parameter values on the plurality of dimension parameters corresponding to one state of the target device, the target dimension parameter being determined based on the first voice recognition result, a second voice recognition result of a dimension parameter voice instruction, or hand posture information of a preset dynamic gesture; in response to detecting the preset dynamic gesture, continuously adjusting the state of the target device based on a continuous action of the dynamic gesture in a direction, comprising: determining a motion direction of the dynamic gesture; based on the motion direction of the dynamic gesture, determining a target adjustment direction of the target device on the target dimension parameter; based on the continuous action of the dynamic gesture in the motion direction, continuously adjusting the target device on the target dimension parameter in the target adjustment direction.

2. The method of claim 1, wherein, The state of the target device is determined based on one dimension parameter. The determining of the target dimension parameter of the target device to be adjusted comprises: determining that the one dimension parameter of the target device is the target dimension parameter.

3. The method of claim 1, wherein, The state of the target device is determined based on a plurality of dimension parameters. The determining of the target dimension parameter of the target device to be adjusted comprises: based on the first voice recognition result, determining the target dimension parameter.

4. The method of claim 1, wherein, The state of the target device is determined based on a plurality of dimension parameters. The determining of the target dimension parameter of the target device to be adjusted comprises: in response to receiving the dimension parameter voice instruction, performing voice recognition on the dimension parameter voice instruction to obtain the second voice recognition result; based on the second voice recognition result, determining the target dimension parameter.

5. The method of claim 1, wherein, The state of the target device is determined based on a plurality of dimension parameters. The determining of the target dimension parameter of the target device to be adjusted comprises: obtaining the hand posture information corresponding to the dynamic gesture, the hand posture information comprising any one of the following: finger extension form, number of fingers, single or double hand information; based on the hand posture information, determining the target dimension parameter.

6. The method of any one of claims 1-5, wherein, The continuously adjusting of the target device on the target dimension parameter in the target adjustment direction based on the continuous action of the dynamic gesture in the motion direction comprises: during the continuous action of the dynamic gesture, obtaining a motion speed of the dynamic gesture in real time or according to a preset adjustment period; based on the motion speed of the dynamic gesture, determining a target adjustment speed of the target device on the target dimension parameter; adjusting the target device on the target dimension parameter in the target adjustment direction at the target adjustment speed.

7. The method of claim 6, further comprising: obtaining adjustment speed configuration information of the target device on the target dimension parameter, the adjustment speed configuration information being used to represent a relationship between a gesture motion speed and a device adjustment speed on each dimension parameter of the target device; determining the target adjustment speed of the target device on the target dimension parameter based on the motion speed of the dynamic gesture, including: determining, based on the adjustment speed configuration information, that the device adjustment speed corresponding to the motion speed of the dynamic gesture on the target dimension parameter is the target adjustment speed.

8. The method of claim 7, wherein, The obtaining of the adjustment speed configuration information of the target device on the target dimension parameter includes: obtaining the adjustment speed configuration information of the target device on the target dimension parameter from the first speech recognition result; or, in response to receiving the adjustment speed configuration speech instruction, performing speech recognition on the adjustment speed configuration speech instruction to obtain a third speech recognition result; obtaining the adjustment speed configuration information of the target device on the target dimension parameter from the third speech recognition result; or, obtaining the adjustment speed configuration information of the target device on the target dimension parameter from the preconfigured adjustment speed configuration information.

9. The method of claim 8, wherein, The preconfiguration of the adjustment speed configuration information includes: receiving an adjustment speed configuration request through a setting interface, the adjustment speed configuration request including device identification, dimension parameter identification, gesture motion amplitude and device adjustment amplitude information, the device identification being used to uniquely identify a device, and the dimension parameter identification being used to uniquely identify a dimension parameter; determining a relationship between a gesture motion speed and a device adjustment speed based on the gesture motion amplitude and the device adjustment amplitude information; configuring adjustment speed configuration information of a device identified by the device identification on a dimension parameter identified by the dimension parameter identification based on the relationship between the gesture motion speed and the device adjustment speed, or updating adjustment speed configuration information corresponding to the device identification and the dimension parameter identification in preconfigured adjustment speed configuration information based on the relationship between the gesture motion speed and the device adjustment speed.

10. The method of any one of claim 6, further comprising: during the continuous adjustment of the target device on the target dimension parameter to the target adjustment direction based on the continuous action of the dynamic gesture in the motion direction, in response to receiving an adjustment speed update speech instruction, performing speech recognition on the adjustment speed update speech instruction to obtain a fourth speech recognition result; obtaining adjustment speed update configuration information from the fourth speech recognition result, the adjustment speed update configuration information being used to represent a relationship between an updated gesture motion speed and a device adjustment speed on each dimension parameter of the target device; during a subsequent continuous action of the dynamic gesture, obtaining the motion speed of the dynamic gesture in real time or according to a preset adjustment period. determine, based on the adjusted speed update configuration information, an updated device adjustment speed corresponding to a motion speed of the dynamic gesture on the target dimension parameter; adjust, at the updated adjustment speed, the target device on the target dimension parameter in the target adjustment direction.

11. The method of any one of claims 6-10, further comprising: determining a position of a sound source object that sends the voice control instruction; based on the position of the sound source object, obtaining an image sequence including a hand of the sound source object, the image sequence including a plurality of images having a time sequence relationship; sequentially performing hand key point detection on each image in the image sequence to obtain a hand key point sequence, the hand key point sequence being formed by hand key points in the images based on the time sequence relationship; based on the hand key point sequence, performing preset dynamic gesture detection.

12. The method of claim 11, wherein, The determination of the motion direction of the dynamic gesture comprises: based on the hand key point sequence, determining the motion direction of the dynamic gesture.

13. The method of claim 11, wherein, The determination of the motion speed of the dynamic gesture comprises: based on the hand key point in the last image in the image sequence and the hand key point in the previous image, and the capture time of the last image and the capture time of the previous image, determining the motion speed of the dynamic gesture.

14. A device control apparatus, comprising: a voice recognition module configured to, in response to receiving a voice control instruction, perform voice recognition on the voice control instruction to obtain a first voice recognition result; a determination module configured to, based on the first voice recognition result obtained by the voice recognition module, determine a target device corresponding to the voice control instruction; a detection module configured to detect a preset dynamic gesture; an adjustment module configured to, in response to the detection module detecting the preset dynamic gesture, continuously adjust a state of the target device based on a continuous action of the dynamic gesture; a second determination module configured to determine a target dimension parameter of the target device to be adjusted, when the state of the target device is determined based on a plurality of dimension parameters, a group of parameter values of the plurality of dimension parameters corresponding to one state of the target device, the target dimension parameter being determined based on the first voice recognition result, a second voice recognition result of a dimension parameter voice instruction, or hand posture information of the preset dynamic gesture; the adjustment module comprises: a first determination unit configured to determine a motion direction of the dynamic gesture; a second determination unit configured to, based on the motion direction of the dynamic gesture, determine a target adjustment direction of the target device on the target dimension parameter; and an adjustment unit configured to, based on a continuous action of the dynamic gesture in the motion direction, continuously adjust the target device on the target dimension parameter in the target adjustment direction.

15. A computer-readable storage medium, the storage medium storing a computer program, the computer program being configured to execute the device control method of any one of claims 1-13.

16. An electronic device, comprising: a processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the device control method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Equipment control method and device, storage medium and device

    CN109886070A

  • Automobile skylight control method and electronic equipment

    CN110936797A

  • Control system and control method for vehicle

    TW201636233A