Device control method, apparatus and electronic device

By combining sound source location information and image data, candidate devices are screened from multiple acquisition devices, solving the problems of subjectivity and resource waste in device selection in existing technologies. This achieves efficient and flexible target device determination, improving the accuracy and real-time performance of device selection.

CN122420508APending Publication Date: 2026-07-17LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LENOVO (BEIJING) LTD
Filing Date
2026-04-30
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

When identifying a target device with high-quality images from multiple acquisition devices, existing technologies suffer from issues such as strong subjectivity in human selection, waste of computational resources, and poor flexibility.

Method used

By combining sound source location information and image data, candidate devices are screened from multiple acquisition devices. Initial screening is performed using audio information, reducing computational resource requirements. Then, target devices are determined through image analysis, improving the accuracy and flexibility of device selection.

Benefits of technology

It improves the accuracy and flexibility of target device selection, reduces computing resource requirements, reduces resource consumption, and enhances the efficiency and real-time performance of device selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122420508A_ABST
    Figure CN122420508A_ABST
Patent Text Reader

Abstract

This disclosure provides a device control method, apparatus, and electronic device, which can be applied in the field of computer technology. The method includes: obtaining azimuth information corresponding to each acquisition device, wherein the azimuth information represents the relative azimuth between the location of the sound source and the corresponding acquisition device; determining candidate devices from multiple acquisition devices based on the azimuth information corresponding to each acquisition device; and determining a target device from the candidate devices based on image data acquired by the candidate devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more specifically to a device control method, apparatus, and electronic device. Background Technology

[0002] In some scenarios, to expand the capture range of video and audio, multiple capture devices can be used simultaneously to capture video and sound from a target area. How to determine the target device from among these multiple capture devices that can capture high-quality video is a problem that needs to be solved. Summary of the Invention

[0003] According to a first aspect of this disclosure, a device control method is provided, comprising: obtaining azimuth information corresponding to each acquisition device, wherein the azimuth information represents the relative azimuth between the location of a sound source and the corresponding acquisition device; determining candidate devices from a plurality of acquisition devices based on the azimuth information corresponding to each acquisition device; and determining a target device from the candidate devices based on image data acquired by the candidate devices.

[0004] According to embodiments of this disclosure, candidate devices are determined from multiple acquisition devices based on the azimuth information corresponding to each acquisition device, including: selecting relative azimuths that meet the first target conditions from multiple azimuth information as target azimuths; and determining the acquisition device corresponding to the target azimuth as a candidate device.

[0005] According to an embodiment of this disclosure, the azimuth information includes the angle value between the location of the sound source and the corresponding acquisition device. Selecting the azimuth that meets the first preset condition from multiple relative azimuths as the target azimuth includes: determining the relative azimuths whose angle values ​​meet the angle condition from multiple azimuth information as the target azimuths.

[0006] According to an embodiment of this disclosure, the azimuth information includes the angle value between the location of the sound source and the corresponding acquisition device. Selecting the azimuth that meets the first preset condition from multiple relative azimuths as the target azimuth includes: obtaining a sorting result corresponding to multiple azimuth information based on the angle value in the azimuth information; and selecting a target number of relative azimuths as the target azimuths based on the sorting result.

[0007] According to embodiments of this disclosure, determining a target device from candidate devices based on image data acquired by candidate devices includes: determining the target region where the sound source is located in the image data corresponding to each candidate device; performing target detection on the target region and determining the orientation information of the sound source in the image data relative to the acquisition device based on the detection result; and determining the candidate device whose orientation information satisfies the second target condition as the target device.

[0008] According to embodiments of this disclosure, determining a target device from candidate devices based on image data collected by candidate devices includes: selecting at least one target image from multiple image data that satisfies a second target condition based on the image data corresponding to the candidate devices; determining the candidate device corresponding to the target image as the target device; the image performance is determined based on at least one of the speaker's posture, image integrity, and image brightness.

[0009] According to embodiments of this disclosure, the method further includes: obtaining target image data corresponding to the target device; and generating a display screen based on the target image data.

[0010] According to embodiments of this disclosure, generating a display screen based on target image data includes at least one of the following: adjusting the target image data according to the display parameters of the display device, and displaying the adjusted target image data as a target screen on the display device; processing the target image data, and displaying the processed target image data as a target screen on the display device; the processing includes at least one of adjusting the speaker image in the target image data and adjusting the image as a whole.

[0011] According to a second aspect of this disclosure, a device control apparatus is provided, comprising: an acquisition module for acquiring azimuth information corresponding to each acquisition device, the azimuth information representing the relative azimuth between the location of a sound source and the corresponding acquisition device; a first determination module for determining candidate devices from a plurality of acquisition devices based on the azimuth information corresponding to each acquisition device; and a second determination module for determining a target device from the candidate devices based on image data acquired by the candidate devices.

[0012] According to a third aspect of this disclosure, an electronic device is provided, comprising: a first communication unit for data communication with a plurality of acquisition devices, each acquisition device being configured to determine azimuth information, the azimuth information representing the relative azimuth between the location of a sound source and the corresponding acquisition device; and at least one processor configured to: obtain azimuth information corresponding to each acquisition device; determine candidate devices from the plurality of acquisition devices based on the azimuth information corresponding to each acquisition device; and determine a target device from at least two candidate devices based on image data of the candidate devices. Attached Figure Description

[0013] The foregoing contents, as well as other objects, features, and advantages of this disclosure, will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0014] Figure 1 This diagram schematically illustrates an application scenario of the device control method according to an embodiment of the present disclosure.

[0015] Figure 2A flowchart illustrating a device control method according to an embodiment of the present disclosure is shown schematically.

[0016] Figure 3 A schematic diagram illustrating a device control method according to an embodiment of the present disclosure is provided.

[0017] Figure 4 This schematically illustrates one of the schematic diagrams for determining a candidate device from a plurality of acquisition devices based on the azimuth information corresponding to each acquisition device, according to an embodiment of the present disclosure.

[0018] Figure 5 This schematically illustrates a second diagram of a process for determining candidate devices from multiple acquisition devices based on the azimuth information corresponding to each acquisition device, according to an embodiment of the present disclosure.

[0019] Figure 6 This schematically illustrates one of the schematic diagrams for determining a target device from candidate devices based on image data acquired by candidate devices according to an embodiment of the present disclosure;

[0020] Figure 7 This schematically illustrates a second schematic diagram of determining a target device from candidate devices based on image data acquired by candidate devices according to an embodiment of the present disclosure.

[0021] Figure 8 A schematic diagram of a device control method according to an embodiment of the present disclosure is shown for reference.

[0022] Figure 9 A schematic block diagram of a device control apparatus according to an embodiment of the present disclosure is shown.

[0023] Figure 10 A block diagram schematically illustrates an electronic device suitable for implementing a device control method according to an embodiment of the present disclosure. Detailed Implementation

[0024] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0025] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0026] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0027] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0028] In the technical solution disclosed herein, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation entry points are provided for users to choose to authorize or refuse.

[0029] Before introducing the technical solutions provided in the embodiments of this disclosure, the relevant technologies involved in this disclosure will be explained first.

[0030] Because a single acquisition device has a limited field of view and sound pickup range, in situations where the target area is large, such as a conference room or classroom, multiple acquisition devices can be deployed at different locations within the target area to capture sound and images from multiple directions as comprehensively as possible without blind spots. However, when displaying the image, the output will typically show one main image or a limited number of images. Therefore, it is necessary to select at least one target acquisition device from among the multiple acquisition devices to capture the main image.

[0031] Common methods for determining target acquisition devices include those based on fixed rules or manual control, and those based on image detection. On the one hand, methods based on fixed rules or manual control may require personnel to manually select the target acquisition devices. Human selection is highly subjective and may lead to inconsistent output image quality and / or image delays. For example, when multiple people take turns speaking, image lag may occur due to switching delays. Methods based on fixed rules lack flexibility and are difficult to adapt to the current scenario. On the other hand, methods based on image detection require image analysis of all acquired images, resulting in high computational costs and unnecessary resource waste.

[0032] Based on the above, this disclosure provides a device control method that determines the target device by using both sound source and image dimensions. This effectively improves the accuracy and flexibility of target device selection, avoids misjudgments caused by a single dimension, and enhances the robustness of target device determination. Before performing image analysis on the acquired footage, using audio information for filtering effectively reduces the demand for computing resources and unnecessary resource consumption, thereby improving the efficiency and real-time performance of target device selection.

[0033] Figure 1 The diagram illustrates an application scenario of the device control method according to an embodiment of the present disclosure.

[0034] like Figure 1 As shown, application scenario 100 according to an embodiment of this disclosure may include a first acquisition device 101, a second acquisition device 102, a third acquisition device 103, a network 104, a server 105, and at least one terminal device 106. The network 104 serves as a medium for providing communication links between the first acquisition device 101, the second acquisition device 102, the third acquisition device 103, and the server 105, as well as between the terminal device 106 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables. For example, a user can use the terminal device 106 to interact with the server 105 to receive and display screen information sent by the server.

[0035] Terminal device 106 can be a display device such as a screen or projector, or an electronic device such as a smartphone, wearable device, personal computer, extended reality device, or conference machine. Extended reality devices can include virtual reality devices, augmented reality devices, and mixed reality devices.

[0036] Terminal device 106 can connect to the server via wired connection (such as HDMI, DVI, VGA, etc.) or wireless connection (such as Wi-Fi, Bluetooth, etc.) to receive and display screen information pushed by the server. A client application for the target application can be installed and run on the terminal device, allowing users to view the screen information pushed by the server through the target application. Furthermore, this embodiment of the disclosure does not limit the form of the target application, including but not limited to applications, mini-programs, etc., installed on the terminal device, and may also be in the form of a webpage.

[0037] The acquisition devices can be, for example, electronic devices with camera functions distributed at different locations in the target area. Each acquisition device has a corresponding field of view and is used to acquire video data of the target area.

[0038] Server 105 can be a server providing various services, used to receive audio or video data collected from various acquisition devices, and analyze and process the received data using specific algorithms and rules. The processed results are then distributed to terminal devices to generate display screens. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services such as cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data. The server can also be a backend server for the aforementioned target application, providing backend services to the client of the target application.

[0039] It should be noted that the device control method provided in this embodiment can generally be applied to either server 105 or terminal device 106. When applied to a terminal device, if multiple terminal devices exist in the target area, one terminal device can be selected as the master device, and the master device executes the device control method. For example, a terminal device can be selected as the master device based on its priority, and the master device executes the device control method. The priority of the terminal device can be determined, for example, based on preset device type, device performance indicators, etc. Accordingly, the device control device provided in this embodiment can generally be located in server 105 or terminal device 106.

[0040] It should be understood that Figure 1 The number and type of data acquisition devices, terminal devices, networks, and servers shown are merely illustrative. Depending on implementation needs, any number and type of data acquisition devices, terminal devices, networks, and servers can be included.

[0041] Figure 2A flowchart illustrating a device control method according to an embodiment of the present disclosure is shown schematically. Figure 3 A schematic diagram of a device control method according to an embodiment of the present disclosure is shown.

[0042] like Figure 2 , Figure 3 As shown, the device control method according to the embodiments of this disclosure may include operations S210 to S230.

[0043] In operation S210, the azimuth information corresponding to each acquisition device is obtained. The azimuth information represents the relative azimuth between the location of the sound source and the corresponding acquisition device.

[0044] In some embodiments, the acquisition device may be a device with sound acquisition function and image acquisition function. The acquisition device is evenly or unevenly distributed in the area to be detected in order to achieve coverage of sound sources and images from different directions.

[0045] When sound is detected in the area to be detected, each acquisition device collects the sound information. Based on a sound source localization algorithm, the collected sound signals are analyzed to determine the location of the sound source relative to the corresponding acquisition device, thus obtaining the location information of each acquisition device. The system also receives the location information sent by each acquisition device.

[0046] In operation S220, candidate devices are determined from multiple acquisition devices based on the azimuth information corresponding to each acquisition device.

[0047] In some embodiments, the selection rules for candidate devices can be set according to actual needs. For example, a directional angle range can be set, and the acquisition devices whose directional information display sound sources are located within this directional angle range can be identified as candidate devices.

[0048] The location information corresponding to each acquisition device is compared with the preset screening conditions, and the acquisition devices that meet the screening conditions are identified as candidate devices.

[0049] In operation S230, the target device is determined from the candidate devices based on the image data acquired by the candidate devices.

[0050] In some embodiments, image data acquired by the candidate device is obtained, and the image data is analyzed to identify specific targets in the image and related information about the targets. Specific targets in the image may include, for example, tasks, objects, etc., and related information about the targets may include their position, size, and orientation.

[0051] Based on the results of image analysis, the target device can be determined from the candidate devices. For example, if the task is to track a specific person, then the acquisition device corresponding to the image of the person in the image acquired by the candidate devices can be identified by facial recognition technology, which is the target device. Alternatively, based on a comprehensive judgment of factors such as the position and clarity of the target in the image, the acquisition device that can most accurately acquire the target information can be selected as the target device.

[0052] The device control method provided in this disclosure pre-screens candidate devices based on azimuth information and then determines the target device based on image data, which can more accurately select the acquisition device suitable for completing a specific task and improve the accuracy of device selection.

[0053] Pre-screening candidate devices using location information can quickly narrow down the selection range, avoiding image analysis of all acquisition devices. This effectively improves resource utilization, reduces the computing power requirements of individual devices, and increases response speed, thereby quickly identifying target devices capable of responding to and processing specific tasks.

[0054] On the other hand, by locating and identifying the target device through both orientation information and image data, the reliability of the target device determination can be effectively improved. When one type of information is interfered with or erroneous, the other type of information can supplement and correct it, thus avoiding affecting the accuracy of the target device determination.

[0055] According to an embodiment of this disclosure, in operation S220, determining candidate devices from multiple acquisition devices based on the azimuth information corresponding to each acquisition device may include: filtering relative azimuths that meet the first target condition from multiple azimuth information as target azimuths; and determining the acquisition device corresponding to the target azimuth as a candidate device.

[0056] In some embodiments, the first target condition can be set according to the actual needs of the application scenario, and the first target condition is different for different application scenarios. By traversing and judging the azimuth information obtained from multiple acquisition devices, the relative azimuths that meet the first target condition are filtered out as target azimuths. Based on the filtered target azimuths and the correspondence between the azimuth information and the acquisition devices, the acquisition devices corresponding to the target azimuths are determined as candidate devices.

[0057] In some embodiments, the acquisition device may be a camera device that includes an audio acquisition component (i.e., the camera device and the audio acquisition component are integrated into one unit), or it may be a camera device associated with an audio acquisition component (i.e., the camera device and the audio acquisition component are independent devices, and the two can be associated with each other via wired or wireless means).

[0058] When the acquisition device is a camera device that includes an audio acquisition component, the location information can be directly sent from the camera device to the server or terminal device. When the acquisition device is a camera device associated with an audio acquisition component, the audio acquisition component is placed close to the camera device to ensure spatial consistency between audio and image. In this case, the location information can be sent from the audio acquisition component to the server or terminal device.

[0059] For example, the audio acquisition component can be a microphone matrix, which consists of multiple microphones arranged around the camera device, with each microphone outputting one audio channel. The relative position between the sound source and the camera device can be determined by calculating the time difference between the microphone audio data.

[0060] For each acquisition device, each microphone independently collects audio signals from the environment, converts them into electrical signals, and after processing such as analog-to-digital conversion, outputs a single digital audio signal. This audio data can be transmitted in real time to the storage or processing module of the acquisition device. After receiving the audio data from each channel, the location information of the sound source relative to the acquisition device can be determined by calculating the time difference between the audio data output by different microphones.

[0061] For example, the azimuth information of a sound source relative to a data acquisition device can be determined by calculating the direction of arrival (DOA) data. This calculation process can be performed independently by each data acquisition device, which then sends the result (i.e., the azimuth information) to a server or terminal device. The server or terminal device can then use the received azimuth information to identify candidate devices from among the multiple data acquisition devices. For instance, after acquiring audio data, each data acquisition device can calculate the azimuth information and send it to the server or terminal device.

[0062] Alternatively, the calculation of location information can be performed uniformly by a processing unit, which can be deployed in a server or terminal device. For example, the acquisition devices send the acquired audio data to the processing unit, and the processing unit calculates the location information corresponding to each acquisition device based on the received audio data.

[0063] This embodiment of the disclosure filters directional information by setting a first target condition, which can quickly narrow down the range of devices to be selected and improve the efficiency of candidate device selection. Filtering candidate devices based on directional information can effectively reduce the number of devices that need to be analyzed in the subsequent image analysis, thereby reducing the demand for computing resources, reducing the time overhead of subsequent image processing, and achieving a reduction in computing power requirements and an improvement in response speed.

[0064] Figure 4The schematic diagram illustrates one of the schematic diagrams for determining candidate devices from a plurality of acquisition devices based on the azimuth information corresponding to each acquisition device, according to an embodiment of the present disclosure.

[0065] See Figure 4 According to embodiments of this disclosure, selecting relative azimuths that satisfy a first target condition from multiple azimuth information as target azimuths may include: determining relative azimuths whose angle values ​​satisfy angle conditions from multiple azimuth information as target azimuths.

[0066] In some embodiments, the orientation information may include, for example, an angle value and the number of the acquisition device. The angle value is determined based on the coordinate system of the acquisition device itself and is used to characterize the deviation angle of the sound source relative to the centerline of the acquisition device. The sign of the angle value is used to reflect the left-right direction between the sound source and the acquisition device. For example, if the sound source comes from the left 30° direction of acquisition device IP1, then the angle value corresponding to acquisition device IP1 is 30°; if the sound source comes from the right 15° direction of acquisition device IP2, then the angle value corresponding to acquisition device IP2 is -15°.

[0067] The angle condition in the first target condition can be set according to the actual application scenario and requirements. For example, the angle condition can be an angle range or a specific angle value. For instance, in a scenario where it is necessary to locate a sound source in a specific direction, the angle condition can be a horizontal angle between -30° and 30°.

[0068] From the location information of multiple acquisition devices, the system filters out the relative locations whose angle values ​​meet the set angle conditions and determines these relative locations as target locations. For example, if the system receives location information from 8 acquisition devices, and the angle values ​​of 3 acquisition devices meet the angle conditions, the relative locations of these 3 acquisition devices are determined as target locations. Based on the acquisition device numbers in the location information, these 3 acquisition devices are also identified as candidate devices.

[0069] The embodiments disclosed herein filter target locations by angle conditions, which can quickly focus on candidate devices closely related to the sound source, eliminate interference from other irrelevant devices, and meet specific business needs.

[0070] Figure 5 The schematic diagram illustrates a second schematic of determining a candidate device from multiple acquisition devices based on the azimuth information corresponding to each acquisition device, according to an embodiment of the present disclosure.

[0071] See Figure 5 According to embodiments of this disclosure, selecting relative azimuths that meet the first target condition from multiple azimuth information as target azimuths may include: obtaining a sorting result corresponding to multiple azimuth information based on the angle values ​​in the azimuth information; and selecting a target number of relative azimuths as target azimuths from the multiple azimuth information based on the sorting result.

[0072] In some embodiments, the angle values ​​may include a plus or minus sign, which reflects the left-right relationship between the sound source and the acquisition device. The angle values ​​can be sorted according to their absolute values, and the relative orientations of a number of targets are selected from multiple orientation information sources based on the sorting results. The number of targets can be determined according to the actual application scenario and requirements; it can be a fixed value or calculated based on a fixed percentage. For example, if there are 20 orientation information sources and the fixed percentage is 10%, then the number of targets is 2, meaning two relative orientations are selected from multiple orientation information sources as target orientations based on the sorting results.

[0073] For example, the angle values ​​can be sorted in ascending or descending order. For instance, the angle values ​​can be sorted in ascending order of absolute value from smallest to largest, and the azimuth information of the highest-ranking target quantity can be selected. If the target quantity is 3, then the relative azimuths corresponding to the first 3 angle values ​​in the sorting result are selected as the target azimuths.

[0074] The selected target locations are associated with the corresponding acquisition devices, and the acquisition devices corresponding to the target locations are identified as candidate devices.

[0075] The method for determining candidate devices based on angle value sorting in this disclosure allows for flexible selection of the number of targets according to actual needs, meeting the requirements for the number of candidate devices in different scenarios. For example, in dynamically changing scenarios, angle value sorting can dynamically adjust the target orientation in real time, always focusing on the acquisition device most closely related to the current position of the sound source, thereby improving the accuracy and real-time performance of candidate device selection.

[0076] In some embodiments, orientation information can indicate the direction from which the sound source originates, but it cannot determine the orientation of the sound source facing the acquisition device. For example, the orientation information of the acquisition device is the same whether the sound source is facing or away from the acquisition device. Based on this, this disclosure further proposes to determine the target device from candidate devices using image data.

[0077] Figure 6 The schematic diagram illustrates one of the schematic diagrams for determining a target device from candidate devices based on image data acquired by candidate devices according to an embodiment of the present disclosure.

[0078] See Figure 6 According to embodiments of this disclosure, determining a target device from candidate devices based on image data acquired by candidate devices may include: determining the target region where the sound source is located in the image data corresponding to each candidate device; performing target detection on the target region and determining the orientation information of the sound source in the image data relative to the acquisition device based on the detection result; and determining the candidate device whose orientation information satisfies the second target condition as the target device.

[0079] In some embodiments, image data acquired by each device is obtained. The image data may be, for example, a frame from a captured real-time video stream or a separately captured still image.

[0080] The acquired image data can be preprocessed, and the preprocessed image data can be analyzed to determine the target area where the sound source is located. First, based on some sound source localization methods or prior knowledge, a rough estimate can be made of the approximate area where the sound source might appear in the image, determining a possible search range as a candidate target area. Alternatively, a simple method based on image features can be used for rough localization. For example, if the sound source is an object emitting a specific color of light, a color thresholding method can be used to find areas in the image with that color feature as candidate target areas for the sound source.

[0081] Furthermore, image segmentation algorithms can be used to divide the candidate target region into multiple regions, and the region that best matches the characteristics of the sound source can be selected as the target region based on the characteristics of the sound source.

[0082] In some embodiments, the target area is scanned and analyzed to detect sound source objects present in the target area, and the location and category information of the sound source are output. The orientation information of the sound source relative to the acquisition device can be determined based on the sound source location information obtained from target detection and the geometric relationship of the image. For example, the angle of the sound source relative to the acquisition device can be determined by analyzing the position and posture of the sound source in the image, combined with the viewpoint and coordinates of the acquisition device. For some sound sources with obvious directional features (such as faces), their orientation can be further determined using feature extraction algorithms. For example, for face sound sources, a facial landmark detection algorithm can be used to detect key points such as the eyes, nose, and mouth of the face, and then the orientation of the face can be determined based on the positional relationship of these key points.

[0083] The second target condition can be a condition related to a specific range of the sound source orientation, set according to the specific application scenario and requirements.

[0084] The orientation information of the sound source corresponding to each candidate device is compared with the second target condition. For candidate devices whose orientation information meets the second target condition, they are identified as target devices. There can be one or more target devices. For example, if only one candidate device meets the second target condition, it can be directly identified as the target device. If multiple devices meet the second target condition, further filtering can be performed based on other conditions (such as device reliability, image quality, etc.), or all multiple devices can be used as target devices for subsequent processing.

[0085] This embodiment of the disclosure, by determining the target region where the sound source is located in each candidate device image and performing target detection on the target region, can further reduce the computational load while obtaining the sound source orientation information. Furthermore, by performing secondary screening of the acquisition device through image data, the accuracy of target device selection can be further improved, thereby enhancing the quality of the acquired image. For example, even if the angle between the sound source and a certain acquisition device is very small, if image detection reveals that the sound source is facing away from the acquisition device, that acquisition device will not be selected as a target device.

[0086] Figure 7 The schematic diagram illustrates a second schematic of determining a target device from candidate devices based on image data acquired by candidate devices according to an embodiment of the present disclosure.

[0087] See Figure 7 According to an embodiment of this disclosure, determining a target device from candidate devices based on image data collected by candidate devices includes: selecting at least one target image from multiple image data that satisfies a second target condition based on the image data corresponding to the candidate devices; determining the candidate device corresponding to the target image as the target device; the image performance is determined based on at least one of the speaker's posture, image integrity, and image brightness.

[0088] In some embodiments, for image data acquired by a candidate device, the image performance of the acquired device can be determined by evaluating at least one of the speaker's posture, image integrity, and image brightness.

[0089] For example, the speaker's posture in an image can be detected using corresponding algorithms in computer vision (such as pose estimation algorithms). For instance, these algorithms can identify key points on the human body (such as joints, head position, etc.) and determine the speaker's posture, such as orientation, standing position, and gestures, based on the positional relationships of these key points. The speaker's posture is then evaluated according to preset posture standards. For example, if the desired posture is a speaker facing forward, standing naturally, and making clear gestures, then image data conforming to these standard postures will be assigned a higher image performance value.

[0090] Object detection algorithms can be used to detect the completeness of the speaker and other related targets (such as the podium, background signage, etc.) in an image, and the completeness of each image can be determined based on the object detection results. For example, if the speaker and related targets are fully displayed in the image without obvious occlusion or obstruction, the image completeness is high; if some targets are missing, the image completeness is low.

[0091] The average brightness of each image can be determined by converting it to a grayscale image and calculating the average value of all pixels in the grayscale image. The brightness of an image can be scored based on a preset brightness range. If the image brightness is within a suitable range (neither too bright nor too dark), the brightness score is higher; if the brightness is too high or too low, affecting image quality, the brightness score is lower.

[0092] The system can comprehensively calculate a score for each image by assigning weights to the speaker's posture, image integrity, and image brightness. The weights can be adjusted based on the specific application scenario. For example, in a meeting setting, where speaker posture is more important, a higher weight can be assigned to the posture score. In outdoor scenarios, where image brightness and integrity are more crucial, higher weights can be assigned to brightness and integrity scores.

[0093] Based on the comprehensive score, at least one target image that meets the second objective condition can be selected from multiple image datasets. The second objective condition can be that the comprehensive score of the image reaches a certain threshold, or it can be that the top N images are selected from high to low comprehensive scores, where N is a positive integer greater than 0.

[0094] The candidate device corresponding to the selected target image is determined as the target device.

[0095] This disclosure embodiment filters out high-quality image data by comprehensively considering factors such as speaker posture, image integrity, and image brightness, thereby achieving target device selection. This helps improve the accuracy of target device selection and ensures that the target device can provide image data with good image performance to meet subsequent display requirements. By determining image performance through multiple dimensions, the flexibility of target device selection can be effectively improved to adapt to the device selection needs of different application scenarios.

[0096] Figure 8 A schematic diagram of a device control method according to an embodiment of the present disclosure is shown for reference.

[0097] See Figure 8 According to embodiments of this disclosure, the device control method further includes: obtaining target image data corresponding to the target device; and generating a display screen based on the target image data.

[0098] In some embodiments, after determining the target device, image data acquired by the target device is obtained as target image data. The target image data is then processed according to task requirements (e.g., cropping, scaling, color adjustment), generating a display screen, and outputting the display screen to a display device for display. The display device may be, for example, a monitor, a projector, etc.

[0099] Adapting the target image based on display parameters ensures that the target image can be fully displayed on different display devices, avoiding display problems caused by differences in display devices.

[0100] According to embodiments of this disclosure, generating a display screen based on target image data includes at least one of the following: adjusting the target image data according to the display parameters of the display device, and displaying the adjusted target image data as a target screen on the display device; processing the target image data, and displaying the processed target image data as a target screen on the display device; the processing includes at least one of adjusting the speaker image in the target image data and adjusting the image as a whole.

[0101] In some embodiments, the target image data can be adjusted according to the display parameters of the display device to obtain a target image for display on the display device. For example, display parameters may include resolution, refresh rate, color space, color depth, etc. The attributes of the target image data (such as resolution, color mode, etc.) can be compared with the display parameters of the display device, and the target image data can be adjusted according to the comparison result to eliminate the differences between the target image data and the characteristics of the display device, ensuring that the image can be presented with optimal effect on the display device and improving display quality.

[0102] In some embodiments, the target image data can be processed, and the processed target image data can be used as the display screen. The processing may include, for example, adjusting the sound source target (such as a speaker) in the target image data, adjusting the image frame, etc., to improve the display screen quality and enhance the user's visual experience.

[0103] Adjustments to the speaker's image in the target image data can include speaker image enhancement and speaker pose optimization. Speaker image enhancement can include, for example, strengthening the details of the speaker's facial features, enhancing facial expression dynamics, and adjusting lighting (such as eliminating unsightly shadows on the face and making facial lighting more uniform). In cases where the speaker's pose is unsatisfactory (such as tilting, looking down, or turning their head to the side), image warping and pose estimation algorithms can be used to correct the speaker's pose. For example, by detecting key points on the speaker's head (such as the center points of the eyes, nose, and mouth), estimating head pose parameters (such as rotation angle and tilt angle), and using image warping techniques (such as affine transformation or perspective transformation) to adjust the head to a frontal or appropriate pose.

[0104] Adjusting the target image data can include color and tone optimization, sharpness enhancement, noise reduction, and removal of interfering factors. Color balance can be achieved by adjusting the gain values ​​of the three primary color channels in the target image data. For example, if the target image data has a yellowish tint, the color deviation can be corrected by appropriately reducing the gain of the red and green channels and increasing the gain of the blue channel, making the image more natural and realistic. When occlusions are present in the image, image inpainting algorithms can automatically remove the occlusions and fill in the pixels in the occluded areas to improve image quality.

[0105] For example, when there is only one target image data point, the displayed image can be obtained by optimizing that target image data. When there is more than one target image data point, the optimized displayed image can be obtained by fusing information from multiple target image data points. For example, images from different angles may contain different information about the target object, and image processing techniques can be used to fuse this information together to obtain a more complete and clearer displayed image. For instance, when optimizing a speaker's image, images from different angles may show different details of the speaker's face, and fusing this image information can yield a more comprehensive and clearer image of the speaker.

[0106] By optimizing the target image data, the visual quality of the target image can be effectively improved, further enhancing the user's visual experience.

[0107] Based on the above-described equipment control method, embodiments of this disclosure also provide an equipment control device. The following will be combined with... Figure 9 The device is described in detail.

[0108] Figure 9 A schematic block diagram of a device control apparatus according to an embodiment of the present disclosure is shown.

[0109] like Figure 9 As shown, the device control device 900 of this embodiment includes an acquisition module 910, a first determination module 920, and a second determination module 930.

[0110] The acquisition module 910 is used to acquire azimuth information corresponding to each acquisition device. The azimuth information represents the relative azimuth between the location of the sound source and the corresponding acquisition device. In one embodiment, the acquisition module 910 can be used to perform the operation S210 described above, which will not be repeated here.

[0111] The first determining module 920 is used to determine candidate devices from multiple acquisition devices based on the azimuth information corresponding to each acquisition device. In one embodiment, the first determining module 920 can be used to perform the operation S220 described above, which will not be repeated here.

[0112] The second determining module 930 is used to determine the target device from the candidate devices based on the image data acquired by the candidate devices. In one embodiment, the second determining module 930 can be used to perform the operation S230 described above, which will not be repeated here.

[0113] According to embodiments of this disclosure, any plurality of modules among the obtaining module 910, the first determining module 920, and the second determining module 930 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the obtaining module 910, the first determining module 920, and the second determining module 930 can be at least partially implemented as a hardware circuit, such as a field-programmable gate array, a programmable logic array, a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit, or implemented by any other reasonable means of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in a suitable combination of any of these. Alternatively, at least one of the obtaining module 910, the first determining module 920, and the second determining module 930 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0114] Figure 10 A block diagram schematically illustrates an electronic device suitable for implementing a device control method according to an embodiment of the present disclosure.

[0115] like Figure 10 As shown, an electronic device 1000 according to an embodiment of the present disclosure includes a processor 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage portion 1008 into a random access memory 1003. The processor 1001 may include, for example, a general-purpose microprocessor, an instruction set processor and / or an associated chipset and / or a dedicated microprocessor. The processor 1001 may also include onboard memory for caching purposes. The processor 1001 may include a single processing unit or multiple processing units for performing different steps of the method flow according to an embodiment of the present disclosure.

[0116] Random access memory 1003 stores various programs and data required for the operation of electronic device 1000. Processor 1001, read-only memory 1002, and random access memory 1003 are interconnected via bus 1004. Processor 1001 executes various steps of the method flow according to embodiments of the present disclosure by executing programs in read-only memory 1002 and / or random access memory 1003. It should be noted that the programs may also be stored in one or more memories other than read-only memory 1002 and random access memory 1003. Processor 1001 may also execute various steps of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.

[0117] According to embodiments of this disclosure, the electronic device 1000 may further include an input / output interface 1005, which is also connected to a bus 1004. The electronic device 1000 may also include one or more of the following components connected to the input / output interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube, liquid crystal display, etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card, such as a local area network card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the input / output interface 1005 as needed. A removable medium 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1010 as needed so that computer programs read from it can be installed into the storage section 1008 as needed.

[0118] In some embodiments, the electronic device 1000 further includes a first communication unit. The first communication unit is used to communicate data with multiple acquisition devices, each acquisition device being used to determine directional information, the directional information representing the relative directional position of the sound source to the corresponding acquisition device.

[0119] Embodiments of this disclosure also provide a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0120] According to embodiments of this disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In embodiments of this disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include the read-only memory 1002, and / or random access memory 1003, and / or one or more memories other than read-only memory 1002 and random access memory 1003 described above.

[0121] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the methods provided in the embodiments of this disclosure.

[0122] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 709, and / or installed from a removable medium 711. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0123] In embodiments of this disclosure, the computer program can be downloaded and installed from a network via communication section 709, and / or installed from removable medium 711. When the computer program is executed by processor 701, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0124] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The program code can execute entirely on a user computing device, partially on a user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0126] Those skilled in the art will understand that the features described in the various embodiments of this disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments of this disclosure can be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

Claims

1. A device control method, comprising: Obtain the azimuth information corresponding to each acquisition device, wherein the azimuth information represents the relative azimuth between the location of the sound source and the corresponding acquisition device; Based on the location information corresponding to each acquisition device, candidate devices are determined from the multiple acquisition devices; Based on the image data collected by the candidate devices, the target device is determined from the candidate devices.

2. The method according to claim 1, wherein determining candidate devices from the plurality of acquisition devices based on the azimuth information corresponding to each acquisition device includes: Select the relative orientation that meets the first target condition from the multiple orientation information as the target orientation; The acquisition device corresponding to the target location is identified as a candidate device.

3. The method according to claim 2, wherein the azimuth information includes the angle value between the location of the sound source and the corresponding acquisition device, and the step of filtering the relative azimuth information that satisfies the first target condition as the target azimuth includes: The relative orientation of the multiple orientation information whose angle values ​​meet the angle conditions is determined as the target orientation.

4. The method of claim 2, wherein the azimuth information includes the angle value between the location of the sound source and the corresponding acquisition device, and the step of selecting the azimuth that meets the first preset condition from multiple relative azimuths as the target azimuth includes: Based on the angle value in the azimuth information, a sorting result corresponding to multiple azimuth information is obtained; Based on the sorting results, the relative orientation of the target quantity is selected from multiple orientation information as the target orientation.

5. The method according to claim 1, wherein determining the target device from the candidate devices based on the image data acquired by the candidate devices comprises: Determine the target region where the sound source is located in the image data corresponding to each candidate device; Target detection is performed on the target area, and the orientation information of the sound source in the image data relative to the acquisition device is determined based on the detection results. Candidate devices whose orientation information satisfies the second objective condition are identified as target devices.

6. The method according to claim 1, wherein determining the target device from the candidate devices based on the image data acquired by the candidate devices comprises: Based on the image data corresponding to the candidate devices, at least one target image whose image performance satisfies the second target condition is selected from multiple image data; The candidate device corresponding to the target image is identified as the target device; The image representation is determined based on at least one of the speaker's posture, image integrity, and image brightness.

7. The method according to claim 1, further comprising: Obtain the target image data corresponding to the target device; Based on the target image data, a display screen is generated.

8. The method according to claim 7, wherein generating the display screen based on the target image data comprises at least one of the following: The target image data is adjusted according to the display parameters of the display device, and the adjusted target image data is displayed on the display device as the target image. The target image data is processed, and the processed target image data is displayed as the target screen on the display device; the processing includes at least one of adjusting the speaker image in the target image data and adjusting the image as a whole.

9. A device for controlling equipment, comprising: The acquisition module is used to acquire the azimuth information corresponding to each acquisition device, wherein the azimuth information represents the relative azimuth between the location of the sound source and the corresponding acquisition device; The first determining module is used to determine candidate devices from a plurality of acquisition devices based on the azimuth information corresponding to each acquisition device. The second determining module is used to determine the target device from the candidate devices based on the image data collected by the candidate devices.

10. An electronic device, comprising: The first communication unit is used to communicate data with multiple acquisition devices. Each acquisition device is used to determine directional information, which represents the relative directional position between the sound source and the corresponding acquisition device. At least one processor, said processor being used for: Obtain the location information corresponding to each acquisition device; Based on the location information corresponding to each acquisition device, candidate devices are determined from the multiple acquisition devices; Based on the image data of the candidate devices, the target device is determined from the at least two candidate devices.