A device control method, apparatus and conference machine

By combining audio and image data to recognize the user's voice and behavioral information, control commands are generated, solving the problem of smart devices needing wake words and enabling convenient and accurate device control.

CN122135701APending Publication Date: 2026-06-02LENOVO (BEIJING) LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LENOVO (BEIJING) LTD
Filing Date
2026-01-30
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing technologies, smart device control requires a fixed wake word, which leads to inconvenient user experience.

Method used

By combining audio and image data, the system identifies the user's voice control and behavioral information, generates control commands, and realizes a control method that combines language and behavior, thus overcoming the limitations of wake words.

Benefits of technology

It improves the convenience and accuracy of users controlling smart devices, reduces accidental triggering, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135701A_ABST
    Figure CN122135701A_ABST
Patent Text Reader

Abstract

This application provides a device control method, apparatus, and conference machine, relating to the field of information processing technology. The method includes: obtaining audio data and image data; determining control information contained in the audio data and target user behavior information contained in the image data; generating a control command if the control information and the target user behavior information satisfy a target condition; the control command is used to control a target device within a space; the target device is determined based on at least one of the control information and the behavior information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing technology, and in particular to a device control method, apparatus, and conference machine. Background Technology

[0002] Smart devices are often installed in various spaces such as conference rooms and hotel rooms, and users can control them with commands. For example, users can control smart speakers with their voice.

[0003] Existing methods typically require a fixed wake word to trigger control of smart devices, which is inconvenient and affects the user experience. Summary of the Invention

[0004] This application is made in view of at least one of the above-mentioned technical problems existing in the prior art, and the present application can simplify the control method of the conference machine.

[0005] In a first aspect, embodiments of this application provide a device control method, including:

[0006] Obtain audio and image data; Determine the control information contained in the audio data and the target user's behavioral information contained in the image data; If the control information and the target user's behavior information meet the target conditions, a control command is generated; the control command is used to control the target device within the space; the target device is determined based on at least one of the control information and the behavior information.

[0007] Secondly, embodiments of this application provide a device control apparatus, including: The acquisition module is configured to acquire audio and image data; The determination module is configured to determine the control information contained in the audio data and the target user's behavioral information contained in the image data; The control module is configured to generate a control command if the control information and the target user's behavior information meet target conditions; the control command is used to control a target device within the space; the target device is determined based on at least one of the control information and the behavior information.

[0008] Thirdly, embodiments of this application provide a conference machine, including: a receiver, the receiver being used to receive audio data and image data; A processor, the processor being configured to determine control information contained in the audio data and behavioral information of a target user contained in the image data; If the control information and the target user's behavior information meet the target conditions, a control command is generated; the control command is used to control the target device within the space; the target device is determined based on at least one of the control information and the behavior information.

[0009] This application provides a device control method, apparatus, and conference machine that overcomes the limitations of existing wake-up words, allowing users to issue control commands through a combination of language and actions. For example, a user can say "It's a bit hot" while pointing to the air conditioner, eliminating the need to memorize specific control commands. Users do not need to remember wake-up words, making control convenient and improving the user experience. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart of a device control method provided in one embodiment of this application; Figure 2 This is a flowchart of a device control method provided in another embodiment of this application; Figure 3 This is a flowchart of a device control method provided in another embodiment of this application; Figure 4 This is a flowchart of a device control method provided in another embodiment of this application; Figure 5 This is a schematic diagram of the structure of a conference machine provided in one embodiment of this application. Detailed Implementation

[0012] To enable those skilled in the art to better understand the technical solutions of the embodiments of this application, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0013] like Figure 1 As shown in the figure, this application provides a device control method, including: Step 101: Obtain audio and image data.

[0014] Specifically, audio data from a space, such as a conference room or classroom, is collected using microphones or microphone arrays. Image data from the space is collected using cameras, such as images of people in the conference room, device positions, user actions, gestures, and gaze.

[0015] Step 102: Determine the control information contained in the audio data and the target user's behavioral information contained in the image data.

[0016] Control information can include text obtained from speech recognition, which contains only control statements for the device and not a pre-defined wake word, such as "turn on the light." It can also include voiceprint features, tone of voice, and emotional state. Target users can be specific individuals identified through facial recognition, speakers identified through sound source localization, or individuals making specific gestures. Behavioral information includes, but is not limited to: gestures (such as pointing or raising a hand), posture / actions (such as standing up, sitting down, or nodding), gaze / focus (looking at a device), facial expressions (fatigue, focus), and identity information (permissions associated with facial recognition).

[0017] Step 103: If the control information and the target user's behavior information meet the target conditions, generate a control command; the control command is used to control the target device in the space; the target device is determined based on at least one of the control information and the behavior information.

[0018] Target conditions are used to measure whether control information and behavioral information are directed at the same target device. For example, if the voice says "turn down the temperature," and the behavioral information shows the user looking at an air conditioner in a certain area, then the control information and behavioral information meet the target conditions.

[0019] The target device can be determined based on control information: if the voice explicitly mentions "air conditioner", the target device is directly identified as the air conditioner; it can be determined based on behavioral information: if the voice only says "turn this on", but the user points to the display screen, the target device is identified as the display screen through gaze / gesture information; it can also be determined based on a combination of both: if the voice says "lower the temperature", and the behavioral information shows that the user is looking at the air conditioner in a certain area, then the air conditioner in that area is identified as the target device.

[0020] This application's embodiments overcome the limitations of existing wake words, allowing users to issue control commands through a combination of language and actions. For example, a user can say "It's a bit hot" while pointing to the air conditioner, eliminating the need to memorize specific control commands. Users do not need to remember wake words, making control convenient and improving the user experience.

[0021] In some embodiments of this application, the method further includes: The audio and image data are aligned to generate multimodal data; If, based on multimodal data, it is determined that the generation time of control information is consistent with the generation time of behavioral information, and that the control information and behavioral information match, then it is determined that the control information and the behavioral information of the target user satisfy the target conditions.

[0022] Based on timestamps, audio segments generated at the same time are aligned with image frames, and the aligned audio data and corresponding image data are integrated into structured multimodal data, such as data formats including fields such as timestamps, control information, and behavioral information.

[0023] Based on the timestamps of control information and behavior information in multimodal data, it is determined whether the time difference between the two is within a preset threshold (e.g., ≤1 second). If the time difference is within the threshold, it is considered that the time is consistent; otherwise, it is considered that the time is inconsistent (e.g., if the user says "turn on the lights" 5 seconds ago, and then makes a gesture pointing to the lights 5 seconds later, it is considered that the time is inconsistent). By judging whether the time is consistent, scenarios where the voice and behavior are unrelated are excluded, ensuring that the control information and behavior information are generated synchronously by the target user in the same interaction scenario, rather than independent actions at different times.

[0024] Determining whether control information matches behavioral information can be assessed from two dimensions: instruction intent and instruction object. Specifically, this includes: Intent matching: The core intent of the control information is consistent with the action meaning of the behavioral information. For example, if the control information is a "turn on / off" command, the behavioral information is an action such as "point to the device" or "raise hand to confirm"; if the control information is a "adjust parameters" command, the behavioral information is an action such as "focus your gaze on the device" or "swipe gesture".

[0025] Object matching: The device object mentioned in the control information is consistent with the target pointed to by the behavior information. For example, if the user's gaze is on the air conditioner and the gesture is pointing to the air conditioner, it matches the object "air conditioner" in the control information "lower the temperature". If the control information is "turn up the lights", but the behavior information is pointing to the curtains, it is considered an object mismatch.

[0026] Only when both time consistency and information matching verifications pass will the control information and the target user's behavior information be deemed to meet the target conditions, thereby triggering the generation of subsequent control instructions. If either verification fails, the target conditions are deemed not met, and no control instructions are generated.

[0027] Compared to existing solutions that only recognize wake words to trigger commands, this application's embodiments use dual verification to eliminate false triggering scenarios from the source. For example, scenario 1: During a meeting discussion, the control keyword "Should we turn on the air conditioner?" is mentioned, but there is no corresponding action; the command is not triggered due to information mismatch. Scenario 2: The user unintentionally makes a control-like gesture, such as pointing at a light, but there is no corresponding control voice, such as not saying "Turn on the lights"; the command is not triggered due to information mismatch.

[0028] For ambiguous control information in voice, such as "this" or "that" without a clear object, behavioral information matching and verification are used. The specific device is determined by combining the pointing target and gaze focus in the behavioral information, so as to avoid control errors caused by ambiguous command objects.

[0029] In multi-person meeting room scenarios, the system can accurately distinguish between the user who initiates the instruction and other irrelevant users, avoiding confusion caused by multiple people speaking or making actions at the same time. For example, if user A says "turn down the temperature" and points to the air conditioner, and user B makes an irrelevant gesture at the same time, the system will only recognize user A's instruction and behavior and will not be affected by user B.

[0030] In some embodiments of this application, such as Figure 2 As shown, the method also includes: Step 201: Determine the first device that matches the control intent of the target user based on the control information.

[0031] Step 202: Determine a second device that matches the target user's behavioral intent based on the target user's behavioral information.

[0032] Step 203: If it is determined that the first device and the second device are the same device, determine that the control information and the target user's behavior information meet the target conditions.

[0033] The audio data is converted into text, and then target instructions, such as "turn on," "turn down," and "turn off," are extracted from the text based on a preset set of target instructions using a Trie tree or regular expressions. The context of the target instructions in the text is extracted and input into a large language model to obtain the target user's control intent. For example, the control intent of "It's a bit hot, turn this on" is "turn on the cooling device," and the control intent of "turn up that" is "increase the brightness of the lighting device."

[0034] The image data is analyzed to extract the target user's behavioral information, which may include gestures, gaze focus, and body movements. Based on this behavioral information, the target user's behavioral intent is determined. For example, "pointing at the air conditioner and focusing gaze on the air conditioner" corresponds to "selecting the air conditioner device," and "looking at the ceiling light and raising one's hand upwards" corresponds to "selecting the ceiling light device." If the verification passes, the target condition is deemed met, and subsequent control commands are generated. If the verification fails, the target condition is deemed not met, and no control commands are generated.

[0035] For ambiguous commands in voice without a clear target, behavioral information is used to locate a second device, which is then verified against the first device to accurately pinpoint the final target device, avoiding control errors caused by ambiguous command targets. In conference room scenarios with multiple types of devices with the same function, a combination of control intent and behavioral direction can accurately distinguish the target device, avoiding confusion between devices with the same function.

[0036] In some embodiments of this application, determining the control information contained in the audio data includes: Sound source localization is performed based on audio data to determine the location information of the target user in space; The noise data in the audio data is processed based on location information to generate the target audio; If the target audio contains first data that matches the target instruction, control information is determined. The control information contains first data and second data, where the second data is the data in the target audio that is adjacent to the first data.

[0037] By employing sound source localization algorithms, such as time delay estimation and beamforming, the three-dimensional coordinates or regional location of the audio signal source (i.e., the target user) in space can be calculated by analyzing the time difference, phase difference, or intensity difference of the audio signals received by different microphone units.

[0038] By combining the target user's location information, the noise sources in the audio data are identified, such as ambient noise from other areas of the conference room, idle chatter from irrelevant people, and equipment operating noise, all of which are audio emanating from non-target locations. Directional noise reduction techniques, such as adaptive beamforming and spatial filtering, are employed to enhance the audio signal in the target location area while suppressing, filtering, or attenuating noise data from non-target locations. After noise reduction processing, the output audio is clear and has minimal noise interference.

[0039] A preset target command set is established, including commands such as "turn on," "turn off," "increase the temperature," and "increase the lights." Audio data is converted into text, and the text is scanned using a Trie tree or regular expressions to check for first data that matches a command in the target command set. If first data is detected, the speech segments adjacent to the first data in the target audio are extracted as second data, including the contextual speech before and after the first data. For example, if the target audio is "It's a bit hot, turn on the air conditioner," the first data is "turn on the air conditioner," and the second data is "It's a bit hot, turn on the air conditioner." The first and second data are combined to form complete control information.

[0040] This application embodiment focuses on the target user's speech through targeted noise reduction, filtering out environmental noise and irrelevant personnel's speech, significantly improving the detection accuracy of the first data. The control information not only includes the first data but also integrates contextual speech, i.e., the second data, providing complete speech evidence for contextual arbitration of the large language model. This effectively distinguishes between real instructions and dialogue mentions, further reducing the probability of false triggering.

[0041] In some embodiments of this application, determining the behavioral information of the target user contained in the image data includes: Feature extraction is performed on image data to identify target users contained within the image data; If a third data matching the target behavior is detected in the image data, behavioral information is determined. The behavioral information includes the third data and a fourth data, where the fourth data is the data in the image data that is adjacent to the third data.

[0042] Specifically, image data can be input into a target model to identify the target user in the image data. The target model can be a YOLO model, a Faster R-CNN model, etc. A preset set of target behaviors is provided, including various behaviors related to device control, such as gestures pointing at the device, gaze at the device, hand-raising confirmation actions, standing / sitting actions, fatigue state actions, etc.

[0043] Frame-level analysis is performed on the image data to detect whether it contains third data that matches the target behavior set. If third data is detected, visual information of the adjacent third data in the image data is extracted as fourth data. For example, consecutive frame images before and after the frame corresponding to the third data, such as image frames one second before and one second after the pointing action, capture the complete process of the behavior, such as the complete action chain of raising hand → pointing → lowering hand.

[0044] By combining the third and fourth data, complete behavioral information is formed, providing a complete visual basis for multimodal fusion decision-making.

[0045] Behavioral information not only includes third data but also integrates context (fourth data), providing a complete visual basis for multimodal information fusion decisions and improving control accuracy. This application's embodiments distinguish between effective control behavior and unintentional behavior. For example, if the third data is a "gesture pointing to a light" and the fourth data shows "the gesture lasts for 3 seconds, and the gaze is simultaneously focused on the light," it is determined to be effective control behavior; if the fourth data shows "the gesture only lasts for 0.5 seconds, and the gaze is not focused," it is determined to be unintentional behavior.

[0046] In some embodiments of this application, such as Figure 3 As shown, the method also includes Step 301: Determine the user's identity information based on the image data.

[0047] Step 302: Based on the user's identity information, determine the user's permission information, which includes the devices that the user can control.

[0048] Step 303: If the target device matches the permission information, send the control command to the target device.

[0049] Facial features are extracted from the image data and compared with a pre-set user identity database. The user identity database stores the facial features and corresponding identity identifiers of authorized users in the conference room, such as name, employee ID, and role.

[0050] User access levels are categorized by user identity, such as administrator, regular user, and visitor. The scope of devices and operation types that each user level can control are clearly defined. For example, administrators can control all devices, including meeting scheduling and delay, and camera adjustment. Regular users can only control basic devices such as air conditioning, lighting, and volume. Visitors have no control permissions.

[0051] The system compares the target device and operation with the established permission information. If the target device is on the user's controllable device list and the operation is within the permitted range (e.g., a regular user controlling the air conditioner temperature, which complies with the permission), the match is considered successful. If the target device is not on the controllable list (e.g., a regular user controlling a camera) or the operation exceeds the permission (e.g., a regular user adjusting a meeting schedule), the match is considered unsuccessful.

[0052] This application embodiment clarifies the control boundaries of different users through identity-permission binding, avoiding meeting disorder, equipment damage or information leakage caused by unauthorized operations, and greatly improving the security and management controllability of the system.

[0053] In some embodiments of this application, the method further includes: Based on image data, the behavioral information of multiple users was determined; these users were participants in the same meeting agenda. If, based on the meeting agenda and the behavioral information of multiple users, it is determined that the current meeting meets the target suspension conditions, a prompt message is generated to remind the user to pause the current meeting.

[0054] Identify all participants in the current meeting from the video stream, perform multi-dimensional behavioral analysis on each participant's video frames, and extract behavioral information, specifically including: State-related behaviors: yawning, rubbing eyes, looking down and daydreaming, leaning forward / backward, frequently checking the time, and other behaviors indicating fatigue or lack of concentration; Action-related behaviors: all or most people standing up, whispering to each other, prolonged silence without interaction, etc.

[0055] Statistical analysis includes the percentage of users exhibiting fatigue behavior, such as yawning if more than 80% of users do so, and the duration of fatigue behavior, such as users showing signs of fatigue for 5 consecutive minutes.

[0056] The meeting agenda may include the preset total meeting duration, the current duration, agenda nodes, and remaining agenda content.

[0057] Preset target pause conditions, and combine the meeting agenda with multi-user behavior characteristics to set rules for triggering pause prompts. Specifically, these may include: Duration-related conditions: The current meeting has lasted for a preset time threshold, and the agenda does not include a break. User status conditions: The percentage of users exhibiting fatigue behavior reaches a preset proportion (e.g., ≥60%), and the duration of this status exceeds a set time (e.g., ≥3 minutes).

[0058] This application embodiment captures multi-user behavior information through image data and judges the meeting status by combining the meeting agenda. It can generate prompt information without the user actively requesting a break. The prompt information can remind the meeting organizer to pause the meeting so that the participants can get a break in time.

[0059] In some embodiments of this application, the method further includes: Based on image data, behavioral information of multiple users was determined; these users were participants in the same meeting agenda. If, based on the behavioral information of multiple users, it is determined that the current meeting meets the target termination conditions, the devices within the control space are shut down.

[0060] Continuous, multi-dimensional behavioral analysis was performed on the video frames of the participating users to extract behavioral information, including: Leaving behaviors include: users getting up, walking towards the meeting room door, taking belongings away from their seats, and turning off their personal devices. Scenario-related behaviors: the number of people remaining in the meeting room, whether there is continuous interaction (such as communication, operation of meeting equipment), and the status of tables and chairs in their original positions.

[0061] Preset target termination conditions: Based on the actual needs of the meeting room usage scenario, set rules to trigger the automatic shutdown of the equipment. Specifically, these may include: Conditions for personnel leaving the meeting: All users participating in the same meeting agenda leave the meeting (or the percentage of users leaving the meeting is ≥95%), and the duration of this departure exceeds a preset threshold (e.g., 3 minutes). Scene status conditions: All participating users have stopped interacting (no communication, no device operation), and most users have stood up and are preparing to leave (e.g., ≥80% of users are standing and walking towards the door).

[0062] The embodiments of this application can avoid energy waste or equipment damage caused by forgetting to turn off the equipment, while reducing the workload of administrative staff in inspection and shutdown, and reducing the management cost of the conference room.

[0063] like Figure 4 As shown, this application provides a device control apparatus, including: The acquisition module 401 is configured to acquire audio data and image data; The determination module 402 is configured to determine the control information contained in the audio data and the target user's behavioral information contained in the image data; The control module 403 is configured to generate a control command if the control information and the target user's behavior information meet the target conditions; the control command is used to control the target device within the space; the target device is determined based on at least one of the control information and the behavior information.

[0064] In some embodiments of this application, the determining module 402 is configured to align audio data and image data to generate multimodal data; if, based on the multimodal data, the generation time of the control information is determined to be consistent with the generation time of the behavior information, and the control information matches the behavior information, it is determined that the control information and the behavior information of the target user satisfy the target conditions.

[0065] In some embodiments of this application, the determining module 402 is configured to determine a first device that matches the control intention of the target user based on control information; determine a second device that matches the behavior intention of the target user based on the behavior information of the target user; and if the first device and the second device are determined to be the same device, determine that the control information and the behavior information of the target user satisfy the target conditions.

[0066] In some embodiments of this application, the determining module 402 is configured to perform sound source localization based on audio data to determine the location information of the target user in space; process the noise data in the audio data based on the location information to generate target audio; if the target audio is detected to contain first data that matches the target instruction, determine control information, the control information containing the first data and the second data, the second data being the data in the target audio that is adjacent to the first data.

[0067] In some embodiments of this application, the determining module 402 is configured to perform feature extraction on image data to determine the target user contained in the image data; if the target user is detected to contain third data matching the target behavior in the image data, behavior information is determined, the behavior information containing the third data and fourth data, the fourth data being the data in the image data adjacent to the third data.

[0068] In some embodiments of this application, the control module 403 is configured to determine the user's identity information based on image data; determine the user's permission information based on the user's identity information, the permission information including devices that the user can control; and send control instructions to the target device if the target device matches the permission information.

[0069] In some embodiments of this application, the control module 403 is configured to determine the behavioral information of multiple users based on image data; the multiple users are personnel participating in the same meeting agenda; if it is determined that the current meeting meets the target suspension condition based on the meeting agenda and the behavioral information of multiple users, a prompt message is generated to prompt the user to suspend the current meeting.

[0070] In some embodiments of this application, the control module 403 is configured to determine the behavioral information of multiple users based on image data; the multiple users are people participating in the same meeting agenda; if it is determined based on the behavioral information of multiple users that the current meeting meets the target termination condition, the device in the control space is turned off.

[0071] like Figure 5 As shown in the figure, this application embodiment provides a conference machine, including: Receiver, used to receive audio and image data; The processor determines control information contained in the audio data and behavioral information of the target user contained in the image data; if the control information and the behavioral information of the target user satisfy the target conditions, it generates control instructions; the control instructions are used to control the target device within the space; the target device is determined based on at least one of the control information and the behavioral information.

[0072] The conference machine can connect to microphones, cameras, and other devices in the conference room via Bluetooth, WIFI, etc., to receive audio data collected by the microphone and image data collected by the camera, and transmit control commands to the target device.

[0073] This application provides a computer program product that, when executed by a processor, implements the method described in any of the above embodiments.

[0074] It should be understood that in the embodiments of this application, the processor may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0075] It should also be understood that the memory mentioned in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (Read-Only Memory). Only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus RAM (DR RAM).

[0076] It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, the memory (storage module) is integrated into the processor.

[0077] It should be noted that the memories described herein are intended to include, but are not limited to, these and any other suitable types of memories.

[0078] In addition to the data bus, this bus may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled "bus" in the diagram.

[0079] It should also be understood that the first, second, third, fourth and various numerical designations used herein are merely for descriptive convenience and are not intended to limit the scope of this application.

[0080] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0081] In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software. The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware processor, or by a combination of hardware and software modules in the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are omitted here.

[0082] In the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0083] Those skilled in the art will recognize that the various illustrative logical blocks (ILBs) and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0084] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0085] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0086] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0087] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0088] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A device control method, comprising: Obtain audio and image data; Determine the control information contained in the audio data and the target user's behavioral information contained in the image data; If the control information and the target user's behavior information meet the target conditions, a control command is generated; the control command is used to control the target device within the space. The target device is determined based on at least one of the control information and the behavior information.

2. The method as described in claim 1, The method further includes: The audio and image data are aligned to generate multimodal data; If, based on the multimodal data, it is determined that the generation time of the control information is consistent with the generation time of the behavior information, and the control information matches the behavior information, then it is determined that the control information and the target user's behavior information satisfy the target conditions.

3. The method of claim 1, further comprising: Based on the control information, a first device matching the control intent of the target user is determined; Based on the target user's behavioral information, a second device matching the target user's behavioral intent is determined; If it is determined that the first device and the second device are the same device, then the control information and the target user's behavior information satisfy the target conditions.

4. The method of claim 1, wherein determining the control information contained in the audio data includes: Based on the audio data, sound source localization is performed to determine the location information of the target user in the space; Based on the location information, the noise data in the audio data is processed to generate the target audio; If the target audio contains first data that matches the target instruction, control information is determined. The control information includes the first data and second data, where the second data is data in the target audio that is adjacent to the first data.

5. The method of claim 1, wherein determining the target user's behavioral information contained in the image data includes: Feature extraction is performed on the image data to determine the target user contained in the image data; If a third data matching the target behavior is detected in the image data, behavioral information is determined. The behavioral information includes the third data and a fourth data, wherein the fourth data is data in the image data that is adjacent to the third data.

6. The method of claim 1, further comprising: Based on the image data, the user's identity information is determined; Based on the user's identity information, the user's permission information is determined, and the permission information includes the devices that the user can control; If the target device matches the permission information, the control command is sent to the target device.

7. The method of claim 1, further comprising: Based on the image data, behavioral information of multiple users is determined; The multiple users mentioned are individuals participating in the same meeting agenda; If, based on the meeting agenda and the behavioral information of the multiple users, it is determined that the current meeting meets the target pause conditions, a prompt message is generated to prompt the user to pause the current meeting.

8. The method of claim 1, further comprising: Based on the image data, behavioral information of multiple users is determined; The multiple users are people participating in the same meeting agenda; If, based on the behavioral information of the multiple users, it is determined that the current meeting meets the target termination condition, the devices within the space are controlled to shut down.

9. A device for controlling equipment, comprising: The acquisition module is configured to acquire audio and image data; The determination module is configured to determine the control information contained in the audio data and the target user's behavioral information contained in the image data; The control module is configured to generate a control command if the control information and the target user's behavior information meet the target conditions; the control command is used to control the target device within the space. The target device is determined based on at least one of the control information and the behavior information.

10. A conference machine, comprising: Receiver, the receiver being used to receive audio data and image data; A processor, the processor being configured to determine control information contained in the audio data and behavioral information of a target user contained in the image data; If the control information and the target user's behavior information meet the target conditions, a control command is generated. The control commands are used to control target devices within the space; The target device is determined based on at least one of the control information and the behavior information.