Portrait processing method and system, processor, microphone and camera all-in-one machine and storage medium
By using a camera all-in-one unit to collect and process images in collaboration with a primary camera, the problems of inaccurate image capture and system latency in video conferencing were solved. This enabled optimal image display in dynamic environments, enhancing the interactive experience and visual realism of video conferencing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-05-15
AI Technical Summary
Existing video conferencing systems struggle to accurately capture optimal facial information when participants change position or perspective, resulting in a poor interactive experience. Furthermore, the system latency and computational load issues caused by fixed camera layouts and centralized processing have not been effectively resolved.
The system employs an architecture that integrates a microphone and a camera, working in tandem with a first camera. By acquiring the speaker's location data, the system controls the microphone and camera to capture and process data from different perspectives. This achieves complementary dual-perspective data of close-up details and the overall scene, reducing data transmission and processor load, and enabling rapid selection of the best portrait.
It enables accurate and efficient selection of the best image that is clear and meets the needs of real-world scenarios when the position or perspective of the participants changes, significantly improving the interactive experience and visual realism of video conferencing, while reducing latency and computing power consumption.
Smart Images

Figure CN122053779A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a portrait processing method and system, a processor, a microphone / camcorder, and a storage medium. Background Technology
[0002] With the increasing prevalence of video conferencing applications, users have ever-growing demands for the quality of close-up portraits in video conferences. Simply achieving a clear close-up portrait is no longer sufficient to meet actual meeting needs. Users expect close-up portraits to present the most realistic viewing angle to optimize the meeting experience. Related technologies employ multiple cameras fixed at different positions in front to capture images from multiple angles to obtain the best possible perspective. However, this solution requires the portrait to be facing directly forward. When the position or viewing angle of participants changes during the meeting, it becomes difficult to obtain accurate portrait information, thus affecting the interactive experience of the meeting. Summary of the Invention
[0003] This application provides a portrait processing method and system, a processor, a microphone / camcorder, and a storage medium to at least solve the above-mentioned technical problems.
[0004] According to a first aspect of this application, a portrait processing method is provided, applied to a processor. The method includes: acquiring location data of a speaker, the location data being spatial location data of a meeting; controlling a microphone-camera integrated unit to acquire first portrait data of the speaker according to the location data, and controlling the microphone-camera integrated unit to perform first portrait processing on the first portrait data to obtain a first processing result; controlling a first camera to acquire second portrait data of the speaker according to the location data; performing second portrait processing on the second portrait data to obtain a second processing result; wherein the second portrait data includes an image from a global perspective, and the first portrait data includes an image from a directional coverage perspective pointing towards the speaker; determining at least one candidate image containing the speaker based on the first processing result and the second processing result; and determining a target image for display based on the at least one candidate image.
[0005] Optionally, obtaining the speaker's location data includes: determining the speaker's location data based on the speaker's orientation data; wherein the orientation data originates from the microphone and / or the processor.
[0006] Optionally, determining the speaker's location data based on the speaker's directional data includes: when there are at least two directional data points; constructing a virtual ray pointing towards the speaker's direction based on each directional data point, with the device that acquired the directional data as the origin; and determining the speaker's location data based on the intersection of at least two virtual rays.
[0007] Optionally, determining at least one candidate image containing the speaker based on the first processing result and the second processing result includes: obtaining first portrait feature data and a first portrait recognition result based on the first processing result; obtaining second portrait feature data and a second portrait recognition result based on the second processing result; performing cross-device portrait deduplication processing on the first portrait feature data and the second portrait feature data to obtain a comprehensive portrait recognition result; and determining at least one candidate image containing the speaker from the first portrait data and the second portrait data based on the comprehensive portrait recognition result.
[0008] Optionally, the step of performing cross-device deduplication processing on the first image recognition result and the second image recognition result based on the first image feature data and the second image feature data to obtain a comprehensive image recognition result includes: performing face feature matching on the first face feature data and the second face feature data to obtain a face matching result; wherein, the face matching result includes the face similarity between the object corresponding to each first image identifier and the object corresponding to each second image identifier; performing human shape feature matching on the first human shape feature data and the second human shape feature data to obtain a human shape matching result; wherein, the human shape matching result includes the human shape similarity between the object corresponding to each first image identifier and the object corresponding to each second image identifier; based on the... Based on the face matching results and / or the human figure matching results, calculate the comprehensive similarity between the object corresponding to each of the first human figure identifiers and the object corresponding to each of the second human figure identifiers; assign the first human figure identifier and the second human figure identifier whose comprehensive similarity satisfies the matching condition to the same comprehensive human figure identifier; wherein, the first human figure feature data includes the first face feature data and / or the first human figure feature data, and the first human figure recognition result includes at least one first human figure identifier; the second human figure feature data includes the second face feature data and / or the second human figure feature data, and the second human figure recognition result includes at least one second human figure identifier; the comprehensive human figure recognition result includes at least one of the comprehensive human figure identifiers.
[0009] Optionally, determining the target image for display based on at least one of the candidate images includes: selecting the target image from at least one of the candidate images based on a first angle between the orientation of the speaker's feature point and the center line of the field of view of the microphone-camera integrated device, and a second angle between the orientation of the speaker's feature point and the center line of the field of view of the first camera.
[0010] Optionally, selecting the target image from at least one candidate image based on a first angle between the orientation of the speaker's feature point and the center line of the field of view of the microphone-camera unit, and a second angle between the orientation of the speaker's feature point and the center line of the field of view of the first camera, includes: if the second angle is smaller than the first angle, then the candidate image captured by the first camera from the at least one candidate image is taken as the target image; if the second angle is larger than the first angle, then the target image is selected from the candidate images captured by the microphone-camera unit from the at least one candidate image.
[0011] Optionally, the microphone-camera system has multiple cameras. When the candidate images originate solely from the multiple cameras of the microphone-camera system: the microphone-camera system acquires the location data and collects at least two candidate image data for the speaker; the candidate image with the highest image quality score among the candidate images collected by the microphone-camera system is selected as the target image; wherein, the image quality score is used to indicate the comprehensive score of the candidate image under multiple image indicators, and the image indicators include at least one of the following: portrait integrity, portrait centering, image sharpness, and image distortion.
[0012] Optionally, after determining the target image for display based on at least one of the candidate images, the method further includes: if the currently determined target image and the previously determined target image meet a difference condition, performing a gradual transition processing on the currently determined target image to obtain the processed target image for display; wherein the difference condition includes at least one of the following: the currently determined target image and the previously determined target image were acquired by different devices, and the image difference between the currently determined target image and the previously determined target image is greater than a difference threshold.
[0013] According to a second aspect of this application, a portrait processing method is also provided, applied to a microphone camera, the method comprising: in response to a first shooting instruction sent by a processor based on speaker location data, acquiring first portrait data of the speaker; performing first portrait processing on the first portrait data to obtain a first processing result; wherein the first processing result is used to determine a target image containing the speaker.
[0014] Optionally, the first image processing of the first image data to obtain a first processing result includes: performing image recognition on each image in the first image data to obtain an image recognition result corresponding to each image; wherein the image recognition result includes at least one first image identifier; performing in-device image deduplication processing on the image recognition results corresponding to multiple images in the first image data to obtain a first image recognition result; wherein the first processing result includes the first image recognition result.
[0015] Optionally, the step of performing in-device deduplication processing on the image recognition results corresponding to multiple images in the first portrait data to obtain the first portrait recognition result includes: mapping the image recognition result corresponding to the first camera image to the spatial coordinate system of the second camera to obtain the projection recognition result corresponding to the first camera image; wherein, the first portrait data includes a first camera image captured by the first camera and a second camera image captured by the second camera; performing object matching on the projection recognition result corresponding to the first camera image and the image recognition result corresponding to the second camera image; for the same object, retaining one first portrait identifier of the object from the projection recognition result corresponding to the first camera image and the image recognition result corresponding to the second camera image; and determining the first portrait recognition result based on the first portrait identifiers of multiple objects.
[0016] Optionally, the method further includes: acquiring the speaker's location data; wherein the location data is used to determine the speaker's position data.
[0017] According to a third aspect of this application, a portrait processing system is also provided, the system comprising: a microphone-camera integrated unit, a first camera, and a processor connected to the microphone-camera integrated unit and the first camera; wherein the processor is configured to: acquire location data of a speaker; send a first shooting command to the microphone-camera integrated unit and a second shooting command to the first camera according to the location data; the microphone-camera integrated unit is configured to: in response to the first shooting command, acquire first portrait data of the speaker; perform first portrait processing on the first portrait data to obtain a first processing result; and send the first processing result to the processor; the first camera is configured to: in response to the second shooting command, acquire second portrait data of the speaker; perform second portrait processing on the second portrait data to obtain a second processing result; and send the second processing result to the processor; wherein the second portrait data includes an image from a global perspective; the processor is further configured to: determine at least one candidate image containing the speaker based on the first processing result and the second processing result; and determine a target image for display based on the at least one candidate image.
[0018] According to a fourth aspect of this application, a processor is also provided, including a first memory and a first processor, wherein the first memory stores a computer program or instructions, and when the computer program or instructions are executed by the first processor, the first processor causes the first processor to perform the portrait processing method as described in any of the above embodiments.
[0019] According to a fifth aspect of this application, a microphone-camera integrated device is also provided, including a second memory, a second processor, a microphone array, and a camera; the microphone array is used to acquire sound data, and the camera is used to acquire images; the second memory stores a computer program or instructions, which, when executed by the second processor, cause the second processor to perform the portrait processing method as described in any of the above embodiments.
[0020] Optionally, the field of view of the integrated camera is greater than or equal to 90 degrees and less than or equal to 360 degrees.
[0021] Optionally, the camera-camera system includes at least two cameras, and the at least two cameras have different shooting angles.
[0022] Optionally, the camera in the camcorder is a fisheye camera.
[0023] According to a sixth aspect of this application, a computer-readable storage medium is also provided, having stored thereon a computer program or instructions that, when executed by a processor, implement the steps in the portrait processing method as described in any of the above embodiments.
[0024] This application provides a technical solution that firstly acquires the speaker's location data; then, based on the location data, controls a microphone-camera integrated unit to collect first image data of the speaker, and controls the microphone-camera integrated unit to perform first image processing on the first image data to obtain a first processing result; next, based on the location data, controls a first camera to collect second image data of the speaker; performs second image processing on the second image data to obtain a second processing result; then, based on the first and second processing results, determines at least one candidate image containing the speaker; and based on the at least one candidate image, determines a target image for display. In this solution, by acquiring the speaker's real-time location data and dynamically controlling the microphone-camera integrated unit and the first camera to perform collaborative acquisition and processing, the inaccurate image capture problem caused by relying on fixed camera orientation in traditional solutions is effectively overcome. Specifically, the system directionally controls the microphone-camera integrated unit to collect close-up images of the speaker based on the location data, and performs first image processing on the device, which not only reduces data transmission and the computing load on the central processing unit, but also significantly shortens processing latency, achieving rapid selection of the best image. Meanwhile, the second portrait data includes images from a global perspective, while the first portrait data includes images from a directional overlay perspective pointing towards the speaker. By fusing close-up details with global scene information, the system can more comprehensively evaluate portrait quality and scene adaptability. This mechanism enables the system to accurately and efficiently select the best image that is both clear and meets the requirements of real-world scene representation even when the speaker's position or perspective changes, thereby significantly improving the interactive experience and visual realism in video conferencing. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 A flowchart of an optional portrait processing method provided in this application embodiment; Figure 2 A schematic diagram of an optional conference room scenario provided for an embodiment of this application; Figure 3 A flowchart illustrating an optional cross-device deduplication method provided in this application embodiment; Figure 4 A flowchart of an optional in-device deduplication method provided in this application embodiment; Figure 5 This is a schematic diagram of an optional in-device deduplication scenario provided in an embodiment of this application; Figure 6A flowchart illustrating an optional portrait selection method provided in this application embodiment; Figure 7 Flowchart of another optional portrait selection method provided in the embodiments of this application; Figure 8 Flowchart of another optional portrait processing method provided in the embodiments of this application; Figure 9 This is a structural block diagram of an optional portrait processing system provided in an embodiment of this application. Detailed Implementation
[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0028] In the following description, specific embodiments of this application will be illustrated with reference to steps and symbols performed by one or more computers, unless otherwise stated. Therefore, these steps and operations will be referred to several times as being performed by a computer. Computer performance as referred to in this application includes operations performed by a computer processing unit on electronic signals represented by data in a structured format. This operation transforms the data or maintains it at a location in the computer's memory system, which can be reconfigured or otherwise alter the operation of the computer in a manner well known to those skilled in the art. The data structure maintained by the data is the physical location of the memory, which has specific characteristics defined by the data format. However, the principles of this application are illustrated with specific embodiments and are not intended to be limiting. Those skilled in the art will understand that many of the steps and operations described below can also be implemented in hardware.
[0029] The terms "module" or "unit" as used in this application can be considered as software objects executing on the computing system. The different components, modules, engines, and services described in this application can be considered as implementation objects on the computing system. While the apparatus and methods described in this application are preferably implemented in software, they can also be implemented in hardware, both of which are within the scope of protection of this invention.
[0030] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used in the embodiments of this application may also include the plural forms. It should be further understood that the term “comprising” as used in the specification of this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. It should be understood that when an element is “connected” or “coupled” to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein may include wireless connection or wireless coupling. The term “and / or” as used herein includes all or any unit and all combinations of one or more associated listed items.
[0031] In the long-term development of video conferencing systems, the industry has generally regarded improving the clarity and stability of close-up portraits as the core direction of technological evolution, forming a conventional path that relies on fixed-position multi-camera shooting and centralized image processing. This path implies a widely accepted premise: to obtain the best viewpoint, participants should face the direction of a fixed camera, and system latency and data processing burden are the necessary costs to bear in pursuit of high-definition image quality. However, the inventors of this application recognize that this technological paradigm actually forcibly constrains real, dynamic meeting interactions into a static equipment layout and processing flow. Its deeper problem lies in the fact that existing solutions have failed to recognize that the definition of "best portrait" has evolved from a single aspect of image clarity to a comprehensive requirement that simultaneously includes realistic scene relevance, natural interactive expressiveness, and real-time presentation synchronization.
[0032] Specifically, through in-depth analysis of real-world meeting scenarios, the inventors discovered a fundamental collaborative flaw in existing technological approaches: the fixed camera layout and centralized processing architecture inherently contradict the freely changing postures and line of sight of attendees, as well as the flexible physical layout of meeting rooms. This is not a simple matter of insufficient equipment performance or inadequate algorithm optimization, but rather a system-level design blind spot. The industry is accustomed to localized optimization by increasing the number of cameras, improving individual machine performance, or enhancing central computing power, neglecting to redefine the problem from an architectural perspective of dynamic perception, distributed collaboration, and multi-view fusion. For example, when an attendee naturally turns to the side or behind the speaker, their facial micro-expressions are extremely valuable to distant observers, but a fixed-forward camera interprets this as an "invalid back view." Similarly, high-performance global cameras deployed to ensure clear images of people in edge seats are mostly idle, while the massive amounts of data they collect inevitably introduce display latency when fed to the central processor. These phenomena have long been considered "normal" due to technological limitations, rather than a design problem that can be systematically solved.
[0033] Therefore, the core issue addressed in this application is: how to construct a system method that can dynamically adapt to changes in real meeting spaces and human interaction, and, while ensuring real-time performance, collaboratively utilize acquisition and processing resources with different characteristics to ultimately output the best human image that is both clear and naturally integrated into the scene.
[0034] In view of this, this application provides a portrait processing method and system, a processor, a microphone-camera all-in-one device, and a storage medium. This solution adopts an architecture where the microphone-camera all-in-one device and a first camera work together: first, the speaker's spatial location data in the meeting is acquired; then, based on the location data, differentiated perspective acquisition by the two devices is driven; the first camera simultaneously acquires second portrait data covering the entire meeting scene from a global perspective and completes the second portrait processing, achieving complementary dual-perspective data of "close-up details + global scene". The microphone-camera all-in-one device and the processor adopt a synchronous initial processing mode, and finally, the candidate images containing the speaker are determined by combining the two types of processing results, without needing to transmit the entire original data to a remote processor. This aims to solve the problems of inaccurate portrait capture, high computational load, and processing latency in traditional fixed layouts, improving the real-time performance and visual experience of meeting displays.
[0035] To at least partially solve the above-mentioned technical problems, embodiments of this application provide a portrait processing method. Figure 1 A flowchart of an optional portrait processing method provided in this application embodiment is shown below. Figure 1 As shown, the portrait processing method is applied to the processor and includes the following steps S110 to S160.
[0036] Step S110: Obtain the speaker's location data; wherein, the location data is the meeting space location data; Step S120: According to the location data, control the microphone camera to collect the first image data of the speaker, and control the microphone camera to perform first image processing on the first image data to obtain the first processing result; Step S130: Based on the location data, control the first camera to acquire second image data of the speaker; Step S140: Perform second portrait processing on the second portrait data to obtain a second processing result; wherein, the second portrait data includes an image from a global perspective, and the first portrait data includes an image from a directional overlay perspective pointing towards the speaker; Step S150: Based on the first processing result and the second processing result, determine at least one candidate image containing the speaker; Step S160: Determine the target image for display based on at least one candidate image.
[0037] In the portrait processing method provided in this application embodiment, the microphone-camera all-in-one machine can complete the first portrait processing, which can reduce the amount of data transmission and processor computing power consumption, shorten the processing time, avoid display delay, and achieve rapid selection of the best portrait. By controlling the microphone-camera all-in-one machine and the first camera to collect portrait data according to the speaker's position data, interference from irrelevant areas is reduced, providing a high-quality data foundation for subsequent image selection. The second portrait data includes images from a global perspective, and the first portrait data includes images from a directional coverage perspective pointing to the speaker. By fusing the first processing result for close-up details with the second processing result for global perspective information to select the target image, the clarity of the portrait and the relevance of the scene can be taken into account, improving the image quality of the best portrait and thus enhancing the visual experience.
[0038] The microphone-camera all-in-one device can be a device that integrates a microphone and a camera to capture images from a directional coverage angle pointing at the speaker. Each microphone-camera all-in-one device is equipped with at least two cameras to capture images from different directions, with a field of view between 100 and 360 degrees, enabling multi-angle portrait capture. The microphone-camera all-in-one device may include one or more microphones to acquire the speaker's voice information, and the speaker's location data can be determined from the speaker's voice data. In some embodiments, the microphone-camera all-in-one device is also equipped with a local processor to perform local processing on the image data acquired by the camera and the voice information acquired by the microphone, reducing data transmission links and improving data processing efficiency.
[0039] The first camera is used to capture images from a global perspective of the meeting space and may include one or more cameras. For example, the first camera can be a wide-angle camera, or a combination of a wide-angle camera and a telephoto camera. When a single camera is used, it can be positioned in the front area of the meeting space; when multiple cameras are used, they can be placed at different locations within the meeting space, such as in front, to the left, or to the right, to obtain more comprehensive global image data from different perspectives.
[0040] The processor receives data from the camcorder and the first camera, and processes the received data to perform image processing operations such as speaker location, image deduplication, and optimal view decision-making. Furthermore, the processor may include at least one auxiliary microphone array to acquire the speaker's voice information, thereby determining the speaker's location data based on the acquired voice information.
[0041] It should be noted that the processor, the first camera, and the auxiliary microphone array can be set up independently, or they can be integrated into the same device in pairs, or all three can be integrated into the same device. They can be flexibly set according to the application scenario, and this application embodiment does not limit them.
[0042] Figure 2 This is a schematic diagram of an optional conference room scenario provided in an embodiment of this application, with reference to... Figure 2 As shown in the example, the above-mentioned image processing method is applied in a small to medium-sized video conferencing scenario. The conference room is equipped with a microphone and camera all-in-one device 210, a first camera 220, and a processor 230. The microphone and camera all-in-one device is placed in the middle of the conference table, and the first camera is installed in the middle of the front wall of the conference room. The processor establishes a communication connection with the two devices through a wired network. The participants are distributed in different positions around the conference table. During the conference, different participants will speak in turn. It is necessary to obtain the spatial position data of the speakers in real time to provide a basis for subsequent image acquisition and processing.
[0043] In some embodiments, step S110 can be implemented in the following ways: S111, determine the speaker's location data based on the speaker's orientation data; wherein, the orientation data comes from the microphone and / or processor.
[0044] In this embodiment, the acquisition of azimuth data can be based on the microphone array built into the all-in-one microphone camera, the auxiliary microphone array associated with the processor, or both. Specifically, the microphone array of the all-in-one microphone camera captures sound signals in the conference room in real time. The speaker's azimuth data relative to the all-in-one microphone camera is calculated using information such as the intensity and phase of the sound signals. This azimuth data includes horizontal and vertical azimuth angles, with the geometric center of the all-in-one microphone camera as the reference point. The horizontal azimuth angle indicates the speaker's deflection direction relative to the all-in-one microphone camera on the horizontal plane, and the vertical azimuth angle indicates the speaker's deflection direction relative to the all-in-one microphone camera on the vertical plane. Simultaneously, the auxiliary microphone array associated with the processor is installed on the side wall of the conference room. It also captures sound signals in real time and calculates the speaker's azimuth data relative to the auxiliary microphone array. This azimuth data also includes horizontal and vertical azimuth angles, with the geometric center of the auxiliary microphone array as the reference point.
[0045] In some embodiments, step S111 above can be implemented in the following ways: S1111: When there are at least two azimuth data points; based on each azimuth data point, construct a virtual ray pointing in the direction of the speaker, with the device that acquired the azimuth data as the origin; S1112, Determine the speaker's position data based on the intersection of at least two virtual rays.
[0046] By using multiple devices to collaboratively capture the speaker's location information, the speaker's position can be accurately determined. The specific steps are as follows: S1 to S4.
[0047] S1, Location Data Acquisition: The microphone array built into the microphone camera captures the speaker's voice and determines the speaker's approximate location data A; at the same time, the auxiliary microphone array associated with the processor also captures the speaker's location and obtains location data B.
[0048] S2, Position Calculation: The processor receives azimuth data A and azimuth data B, treats the two azimuth data as two virtual rays, and determines the current speaker's specific position by calculating the intersection of the two virtual rays.
[0049] S3, Camera Scheduling: Based on the determined speaker's position, the processor sends instructions to all or some cameras (the built-in camera of the camera and the first camera) to guide them to adjust parameters such as angle and focal length in order to obtain a close-up view of the speaker. Specifically, the processor analyzes and selects the most suitable camera for output based on the speaker's angle, image clarity, and centering.
[0050] S4, Multi-Device Positioning Optimization: When multiple microphone-camera units exist in the system, their microphone arrays capture the speaker's location data separately and send it to the processor to generate multiple virtual rays. Ideally, these virtual rays will converge at the same intersection point (i.e., the speaker's position). However, due to factors such as environmental noise and device accuracy, deviations may occur. In this case, the processor selects the two virtual rays with the highest confidence level and uses their intersection point as the final speaker's position, further improving positioning accuracy.
[0051] Continue to refer to Figure 2 The following example illustrates the process using two sets of directional data. Assume we have acquired two sets of directional data: directional data A from the integrated microphone camera and directional data B from the auxiliary microphone array. For directional data A, construct a first virtual ray L1 pointing towards the speaker, with the geometric center of the integrated microphone camera as the origin O1. For directional data B, construct a first virtual ray L2 pointing towards the speaker, with the geometric center of the auxiliary microphone array as the origin O2. Solve for the intersection point P of the first virtual ray L1 and the second virtual ray L2. The coordinates (Xp, Yp, Zp) of this intersection point P are the speaker's coordinates in the XYZ coordinate system. Here, the X-axis points horizontally towards the front of the conference room, the Y-axis points horizontally towards the right side of the conference room, and the Z-axis points vertically upwards.
[0052] In one optional example, the microphone arrays of multiple microphone cameras capture the speaker's location data separately and send them together to the processor to generate multiple virtual rays, such as... Figure 2 The L1 and L3 are shown. At this point, the processor will filter the two virtual rays with the highest confidence (e.g., L1 and L3). Figure 2The intersection of L1 and L2 (as shown) is used as the final speaker's position.
[0053] Compared to separate camera and microphone solutions, the integrated camera and microphone unit in this application embodiment has significant advantages in speaker positioning, specifically in the following three aspects.
[0054] Firstly, it improves positioning efficiency. By shortening data links and reducing processing latency, the microphone array (sound capture) and camera (image verification) are integrated into the same all-in-one device and equipped with a local processor. This enables local collaborative processing of "sound-image" data, eliminating the need to transmit sound data separately to a remote processor or separate camera before matching the image, significantly shortening the data link. After the microphone array captures the speaker's location data, the local processor can directly coordinate with the camera on the same device to adjust its angle / focus, simultaneously completing "sound positioning-image capture," reducing data transmission steps and aligning with the design goal of "fast processing efficiency" (such as completing initial optimization locally, reducing processor load). In the "camera scheduling" step of this process, the integrated design improves the response speed of camera parameter adjustment commands, avoiding delays in speaker close-up capture caused by command transmission latency.
[0055] Secondly, it improves collaborative stability. Avoiding synchronization deviations across multiple devices, accurate speaker capture relies on strong collaboration between "voice location data" and "image target matching," an advantage inherent in integrated solutions. Hardware synchronization advantages: Microphones and cameras integrated into the same device, sharing a clock and hardware control module, avoids time synchronization errors caused by "microphones and cameras belonging to different devices" in separate solutions (e.g., sound data has been transmitted, but the camera is not yet ready to capture the image). In the connection between "location data capture" and "camera scheduling" in this process, data synchronization errors can be effectively reduced. More direct data association: This application embodiment designs "speaker positioning combined with microphone location data and camera image verification." In the integrated solution, the "location data" captured by the microphone can be directly associated with the image captured by the camera of the same device, eliminating the need for cross-device data source matching (separate solutions require additional confirmation of "which camera's image corresponds to a certain sound data"), reducing association errors and improving positioning stability. In the "multi-device positioning optimization" step, the "location data-image data" of multiple microphone-camera integrated devices can form independent data units and be transmitted to the processor, avoiding positioning deviations caused by data source confusion in separate solutions.
[0056] Thirdly, it enhances anti-interference capabilities. It reduces environmental and link interference. Speaker capture is susceptible to environmental noise (such as conference room echoes) and data link interference (such as packet loss in split-type transmissions). The integrated solution can be specifically optimized for these issues. Environmental noise suppression is more precise: In the integrated solution, the microphone and camera are closer together, allowing the microphone to focus on specific areas through "image recognition of speaker position" (e.g., amplifying only the sound around the person captured by the camera), reducing noise interference from irrelevant areas. In the "location data capture" step of this process, the camera can output the person's coordinates in real time, guiding the microphone array to dynamically adjust its pickup range, improving noise suppression accuracy compared to split-type solutions (in split-type solutions, the microphone cannot directly correlate with the camera's person's position, resulting in lower noise suppression accuracy). Less link interference: In a split-type solution, audio data needs to be transmitted to the processor or camera via the network, which may result in packet loss and delay, leading to distortion of the location data; in an integrated solution, audio and image data are transmitted within the local device without relying on external network links, reducing interference during transmission. In the "location calculation" step, the integrity of the location data received by the processor is increased, which aligns with the goal of "improving the accuracy of speaker positioning" (such as the precise location data required for multi-microphone array collaboration).
[0057] The process of obtaining the first processing result and the second processing result in steps S120 to S140 above can be understood through the following exemplary description.
[0058] After acquiring the speaker's location data (Xp, Yp, Zp), the processor sends a control command to the microphone all-in-one camera via a pre-established communication link. This control command contains the speaker's location data and acquisition parameter configuration information. The microphone all-in-one camera can have multiple built-in cameras with a field of view covering 100 to 360 degrees, enabling multi-angle portrait capture. Upon receiving the control command, the microphone all-in-one camera first analyzes the speaker's location data and, combined with its own position in the meeting space, calculates the speaker's positional relationship relative to each camera. It then adjusts the angle, focal length, and other parameters of each camera to ensure that the camera's field of view accurately covers the speaker's area, thus acquiring the first portrait data of the speaker.
[0059] The camera performs initial image processing on the acquired portrait data. This processing is completed by the camera's built-in local processor, eliminating the need to transmit raw data to a remote processor, thus reducing data transmission volume and processing latency. Initial image processing includes image preprocessing, portrait detection, and preliminary feature extraction. The image preprocessing stage includes noise reduction, white balance adjustment, and contrast optimization. Pre-set image processing algorithms remove environmental noise and correct color balance, improving overall image quality. The portrait detection stage employs a deep learning-based portrait detection model trained on numerous conference scene portrait samples. This model can quickly and accurately identify portrait regions from image frames and output portrait frame coordinates. The portrait frame coordinates are based on the top-left corner of the image frame and include four parameters: the top-left horizontal coordinate, the top-left vertical coordinate, width, and height. The preliminary feature extraction stage extracts basic facial features (such as facial contour features and key organ position features) and basic human shape features (such as body contour features and posture features) from the detected portrait regions, forming a preliminary feature data set. This feature data set, along with the corresponding portrait frame coordinates, constitutes the first processing result. The camera transmits this first processing result to the processor, providing data support for subsequent cross-device processing.
[0060] The first camera can be a combination of a wide-angle camera and multiple telephoto cameras. The wide-angle camera is used to capture a global view of the conference room, while the multiple telephoto cameras are used to assist in capturing close-up images of specific areas. The first camera is installed in the middle of the front wall of the conference room. After the processor obtains the speaker's position data (Xp, Yp, Zp), it sends a capture control command to the first camera. This command contains the speaker's position data and capture mode configuration information.
[0061] After receiving the control command, the first camera first analyzes the speaker's location data and, combined with its own coordinates, calculates the speaker's position within its field of view. The wide-angle camera adjusts its shooting parameters to ensure its global view image fully encompasses the speaker's area and the entire conference room scene, acquiring a sequence of image frames that presents the overall layout of the conference room and the general state of all participants. Simultaneously, based on the speaker's location data, the first camera selects the telephoto camera with the optimal shooting angle from among several telephoto cameras, adjusting its focal length and angle to precisely focus on the speaker, acquiring a sequence of close-up image frames that supplement the presentation with detailed information about the speaker.
[0062] The global perspective image captured by the wide-angle camera and the close-up image captured by the telephoto camera together constitute the second portrait data. The image format of the second portrait data is consistent with that of the first portrait data, which facilitates subsequent unified processing. The first camera transmits the captured second portrait data to the processor in real time, providing raw data for subsequent second portrait processing.
[0063] After receiving the second portrait data transmitted from the first camera, the processor initiates the second portrait processing flow. The second portrait processing is similar to the first portrait processing flow of the camcorder, but it is specifically optimized for the global perspective characteristics of the second portrait data. First, image preprocessing is performed, including noise reduction, white balance adjustment, and edge enhancement. Edge enhancement improves the clarity of boundaries in different areas of the overall image, facilitating accurate identification of the location and contours of each participant. Next, portrait detection is performed using the same portrait detection model as the microphone camera. The system identifies the portrait regions of all participants from each image frame of the second portrait data, outputting the frame coordinates for each region. These frame coordinates, also originating from the top left corner of the image frame, include four parameters: top left horizontal coordinate, top left vertical coordinate, width, and height. Then, feature extraction is performed. For each detected portrait region, corresponding facial feature data and human shape feature data are extracted. Facial feature data includes multiple dimensions such as facial key point coordinates and facial texture features, while human shape feature data includes multiple dimensions such as height proportion features, body contour features, and posture features, forming a feature data set corresponding to each portrait region.
[0064] During feature extraction, the processor standardizes the extracted feature data to ensure consistency in dimensions and data format with the feature data in the first processing result of the camera, facilitating subsequent cross-device feature matching. The processor then associates the frame coordinates of each portrait region in the second portrait data with the standardized feature data set to form the second processing result. This second processing result contains portrait recognition information for all participants from a global perspective, providing complete data support for subsequent cross-device portrait deduplication and candidate image selection.
[0065] In some embodiments, step S150 above can be implemented in the following manner: S151, Based on the first processing result, obtain the first portrait feature data and the first portrait recognition result; S152, Based on the second processing result, obtain the second facial feature data and the second facial recognition result; S153, Based on the first portrait feature data and the second portrait feature data, perform cross-device portrait deduplication processing on the first portrait recognition result and the second portrait recognition result to obtain a comprehensive portrait recognition result; S154, Based on the comprehensive facial recognition results, determine at least one candidate image containing the speaker from the first facial data and the second facial data.
[0066] In this embodiment, the first processing result includes the frame coordinates of the first portrait data acquired by the camera-microphone combo unit, the preliminarily extracted facial feature data, and the human shape feature data. The second processing result includes the frame coordinates of the second portrait data acquired by the first camera, the standardized facial feature data, and the human shape feature data. The processor integrates and analyzes these two processing results to filter out candidate images containing the speaker, providing a basis for determining the subsequent target image.
[0067] The first portrait feature data extracted from the first processing result includes first face feature data and first human figure feature data. The first face feature data is multi-dimensional facial feature information extracted by the camera from the speaker's close-up image, including the coordinate sequence of facial key points, the grayscale value distribution sequence of facial texture, and the curvature sequence of facial contours. Each dimension of feature information is presented as a vector composed of multiple values. The first human figure feature data includes the coordinate sequence of the speaker's body contour, the numerical sequence of posture angles, and the numerical sequence of height proportions, also presented as a multi-value vector. The first portrait recognition result includes at least one first portrait identifier, which is a unique identifier assigned by the camera to each detected portrait to distinguish different portrait objects. In this scenario, since the camera specifically collects images of the speaker, the first portrait recognition result mainly includes the first portrait identifier corresponding to the speaker and possibly the first portrait identifiers of a few surrounding attendees.
[0068] The second portrait feature data extracted from the second processing result includes second face feature data and second human figure feature data. The second face feature data consists of multi-dimensional facial feature information extracted by the processor for each portrait region in the second portrait data. Its feature dimensions are consistent with the first face feature data, including the coordinate sequence of facial key points, the gray value distribution sequence of facial texture, and the curvature sequence of facial contours, etc. Each feature dimension is presented in the form of a standardized multi-value vector. The second human figure feature data includes the coordinate sequence of the body contour of each portrait, the numerical sequence of posture angles, the numerical sequence of height proportions, etc., and its feature dimensions are consistent with the first human figure feature data, also in the form of a standardized multi-value vector. The second portrait recognition result includes at least one second portrait identifier, which is a unique identifier assigned by the processor to each portrait detected in the second portrait data, used to distinguish different participants.
[0069] Because duplicate human images may appear in the images acquired by the camera of the all-in-one camera and the first camera, a cross-device (all-in-one camera and first camera) deduplication operation needs to be performed. In an optional embodiment, the above step S153 can be implemented in the following way: S1531, perform face feature matching on the first face feature data and the second face feature data to obtain face matching results; wherein, the face matching results include the face similarity between the objects corresponding to each first image identifier and the objects corresponding to each second image identifier; S1532, Perform human shape feature matching on the first human shape feature data and the second human shape feature data to obtain human shape matching results; wherein, the human shape matching results include the human shape similarity between the object corresponding to each first human shape identifier and the object corresponding to each second human shape identifier; S1533, Based on the face matching result and / or the human figure matching result, calculate the comprehensive similarity between the object corresponding to each first human figure identifier and the object corresponding to each second human figure identifier; S1534, assign the same comprehensive image identifier to the first image identifier and the second image identifier whose comprehensive similarity meets the matching condition; wherein, the first image feature data includes first face feature data and / or first human shape feature data, and the first image recognition result includes at least one first image identifier; the second image feature data includes second face feature data and / or second human shape feature data, and the second image recognition result includes at least one second image identifier; the comprehensive image recognition result includes at least one comprehensive image identifier.
[0070] Figure 3 A flowchart of an optional cross-device deduplication method provided in this application embodiment is shown below. Figure 3 As shown, the cross-device deduplication process includes the following steps S1 to S3.
[0071] S1 extracts features from the target portrait (speaker), including facial features (such as facial key points and texture information) and human features (such as height and body shape), while also acquiring the portrait cutout data.
[0072] S2 performs cross-device matching across different cameras, calculating the similarity of the trajectory of the person captured by different cameras based on facial feature similarity, human shape feature similarity, and the relative positional relationship of the people. Figure 3 Camera 1 in the text can correspond to the aforementioned first camera, and camera 2 can correspond to the aforementioned camera of the microphone-camera all-in-one machine.
[0073] S3 assigns the same person_id to trajectories with similarity reaching a set threshold, identifies them as the same target person image, and completes cross-device deduplication.
[0074] Face matching and human figure matching include distance matrix calculation and Hungarian matching (IoU algorithm). During cross-device deduplication, the matching results of the previous frame are comprehensively considered to reduce mismatches caused by single-frame image errors. At the same time, for fast-moving participants, the calculation weight of trajectory similarity is dynamically adjusted, prioritizing the human motion trend to assist in the judgment, ensuring the real-time performance and accuracy of the deduplication results.
[0075] The cross-device deduplication operation follows these steps: First, the optimal match for the current frame is obtained by comprehensively considering face matching, human figure matching, target position constraints, and the matching result of the previous frame; then, pairwise camera matching is performed to obtain the optimal pairwise camera matching result; finally, the pairwise matching is merged using a greedy method to obtain the final cross-camera matching result.
[0076] In this embodiment, facial feature matching employs a feature vector similarity calculation method, such as using cosine similarity to form a facial similarity matrix, which represents the complete facial matching result. Human shape feature matching uses the same calculation method as facial feature matching and is also presented in the form of a similarity matrix.
[0077] The overall similarity is calculated using a weighted concatenation method. Weight coefficients are pre-assigned to face similarity and human shape similarity, with w1 for face similarity and w2 for human shape similarity. Both w1 and w2 range from 0 to 1, and their sum is 1. For each combination of a first and second human image identifier, the corresponding face similarity is multiplied by w1, and the human shape similarity is multiplied by w2. The two results are then concatenated to obtain the overall similarity vector. This vector comprehensively reflects the degree of matching between the two identifiers in terms of face and human shape features. For example, if the face similarity between a first and second human image identifier is 's' and the human shape similarity is 't', then the overall similarity vector is (s×w1, t×w2).
[0078] The matching condition is that the overall matching degree corresponding to the comprehensive similarity vector reaches a preset threshold. This preset threshold is calibrated through a large amount of experimental data to ensure the accuracy of the matching results. For first and second image identifiers whose comprehensive similarity vectors meet the matching condition, they are determined to correspond to the same image object and are assigned the same comprehensive image identifier. This comprehensive image identifier is a unique string identifier used to uniquely identify the image object in subsequent processing. For first and second image identifiers that do not meet the matching condition, they are determined to correspond to different image objects and are assigned different comprehensive image identifiers.
[0079] Through the aforementioned cross-device deduplication process, a comprehensive facial recognition result is obtained. This result includes comprehensive facial identifiers corresponding to all participants, as well as image frames from the first and second facial data associated with each comprehensive facial identifier. Based on the comprehensive facial recognition result, all image frames associated with the comprehensive facial identifier corresponding to the speaker are selected. These image frames are the candidate images containing the speaker. The candidate images include close-up images from the directional coverage perspective captured by the microphone-camera all-in-one device, as well as images from the global perspective and telephoto close-up images captured by the first camera.
[0080] In some embodiments, the camera module can have multiple cameras. When candidate images originate only from the multiple cameras of the camera module, it is necessary to perform in-device deduplication for the same portrait in different images. The specific deduplication method can be achieved through the following steps: the camera module acquires location data and collects at least two candidate image data for the speaker; the candidate image with the highest image quality score among the candidate images collected by the camera module is selected as the target image; wherein, the image quality score is used to indicate the comprehensive score of the candidate image under multiple image indicators, which include at least one of the following: portrait integrity, portrait centering, image sharpness, and image distortion.
[0081] Figure 4 This is a flowchart of an optional in-device deduplication method provided in an embodiment of this application. Figure 5 This is a schematic diagram of an optional in-device deduplication scenario provided in an embodiment of this application, with reference to... Figure 4 and Figure 5 As shown, taking a camera all-in-one unit containing three cameras as an example, the deduplication process within the device includes the following steps: S1: Input the intrinsic parameters of the three cameras and the extrinsic parameters of the left and right cameras and the middle camera, and obtain the image frames captured by each camera.
[0082] S2 identifies the heads in each image frame and outputs a preliminary list of head bounding boxes.
[0083] S3 maps the head frame identified by the central camera of the camera all-in-one machine to the image space of the left and right cameras through intrinsic and extrinsic parameters.
[0084] S4 employs a Hungarian algorithm based on IoU (Intersection over Union) to match the mapped head bounding boxes with the head bounding boxes directly detected by the left and right cameras. It determines which belong to the same target, removes duplicate head targets in overlapping areas in real time, and outputs a deduplicated list of head bounding boxes within the device. (Reference) Figure 5 As shown, the same target in different scenes is outlined with boxes and connected by lines.
[0085] In this embodiment, there are scenarios where candidate images originate solely from multiple cameras of the camera all-in-one device. For example, the first camera may malfunction and fail to capture images properly, or the speaker in a meeting may be positioned extremely close to the camera all-in-one device, where the multiple cameras of the camera all-in-one device are sufficient to meet the acquisition requirements. In this case, the camera all-in-one device obtains the speaker's position data (Xp, Yp, Zp) through a communication link with the processor. Combining this with the layout and field of view of its multiple cameras, it adjusts the shooting angle, focal length, and other parameters of each camera to ensure that the shooting field of view of each camera covers the speaker, thereby acquiring at least two candidate image data. These candidate image data come from different lenses of the camera all-in-one device, presenting close-up views of the speaker from different angles.
[0086] The image quality score is calculated based on multiple image metrics, each with its own quantitative evaluation method. Quantitative evaluation of portrait integrity involves first identifying key areas that a speaker's complete portrait should include, such as the face, neck, and upper body. Then, the pixel coverage of these key areas in candidate images is detected, and the coverage ratio of each key area is calculated. The coverage ratios of all key areas are then weighted and stitched together to obtain a portrait integrity evaluation vector. A higher overall value for this vector indicates better portrait integrity.
[0087] Quantitative evaluation of portrait centering: Establish an image coordinate system with the center of the candidate image as the origin, obtain the center coordinates of the speaker's portrait frame, calculate the distance between the center coordinates and the origin of the image coordinate system. The smaller the distance, the higher the portrait centering. After normalizing the distance, obtain the portrait centering evaluation value. The normalized evaluation value ranges from 0 to 1, and the closer it is to 1, the higher the centering.
[0088] Quantitative evaluation of image sharpness: An image sharpness evaluation algorithm is used to detect the gray-level changes in the edge regions of the candidate image and calculate the average gray-level gradient of the edge regions. The larger the average gray-level gradient, the sharper the image edges and the higher the overall image sharpness. After standardizing the average value, the image sharpness evaluation value is obtained. The standardized evaluation value ranges from 0 to 1.
[0089] Quantitative assessment of image distortion: During the camera calibration process of the camcorder, the distortion parameter model of each camera is obtained in advance. Based on the model, the deformation of the preset standard shape (such as rectangle, circle) in the candidate image is calculated. The smaller the deformation, the lower the degree of image distortion. After normalizing the deformation, the image distortion evaluation value is obtained. The normalized evaluation value ranges from 0 to 1. The closer it is to 1, the lower the degree of distortion.
[0090] The image quality score is calculated using a weighted concatenation method. Each image indicator is assigned a corresponding weight coefficient: w3 for portrait integrity, w4 for portrait centering, w5 for image sharpness, and w6 for image distortion. Each weight coefficient ranges from 0 to 1, and the sum of all weight coefficients is 1. For each candidate image, the evaluation value of each image indicator is multiplied by its corresponding weight coefficient, and the results are then concatenated to obtain an image quality score vector. This score vector comprehensively reflects the overall image quality of the candidate image.
[0091] The local processor of the camera calculates the image quality score vector for each of the acquired candidate images. Then, it compares the score vectors of all candidate images and selects the candidate image with the highest overall score as the target image. If multiple candidate images have the same or similar overall scores, it further compares the evaluation values of the image sharpness index and selects the candidate image with higher image sharpness as the target image. This ensures that the final output target image has the best image quality, providing remote users with a clear, complete, centered, and distortion-free close-up image of the speaker.
[0092] In some embodiments, step S160 above can be implemented in the following ways: S161, select a target image from at least one candidate image based on the first angle between the orientation of the speaker's feature point and the center line of the field of view of the camera and the second angle between the orientation of the speaker's feature point and the center line of the field of view of the first camera.
[0093] In this embodiment, the speaker's feature points are selected from key facial organ feature points, such as the face, eyes, nose, mouth, center of the pupils, and tip of the nose. The coordinates of these feature points are accurately obtained from candidate images using an image feature extraction algorithm. The center line of the field of view of the integrated camera is the central axis of the field of view of each camera in the integrated camera, and its direction is determined by the camera's mounting angle and shooting parameters. The center line of the field of view of the first camera is the central axis of the field of view of the telephoto camera used to capture close-ups of the speaker, and its direction is also determined by the camera's mounting angle and shooting parameters.
[0094] The calculation process for the first included angle is as follows: First, extract the coordinates of the speaker's feature points from the candidate image and determine the orientation vector of the feature points, which indicates the orientation direction of the speaker's face; then, obtain the direction vector of the center line of the field of view of the camera; finally, calculate the angle between the orientation vector of the feature points and the direction vector of the center line of the field of view of the camera. This angle is the first included angle. The value of the first included angle is between 0 degrees and 90 degrees. The smaller the first included angle, the closer the speaker's face is to the shooting direction of the camera, and the better the angle of the captured image.
[0095] The calculation process for the second included angle is similar to that for the first included angle: extract the speaker's feature point orientation vector, obtain the direction vector of the center line of the first camera's field of view, calculate the included angle between the two vectors, and obtain the second included angle. The value range of the second included angle is also between 0 degrees and 90 degrees. The smaller the second included angle, the closer the speaker's face is to the shooting direction of the first camera, and the better the image perspective.
[0096] In some embodiments, step S161 above can be implemented in the following manner: S1611, if the second included angle is smaller than the first included angle, then the candidate image captured by the first camera in at least one candidate image shall be used as the target image; S1612, if the second included angle is greater than the first included angle, then select the target image from the candidate images acquired by the camera in at least one of the candidate images.
[0097] In this embodiment, the preferred source of candidate images is determined by comparing the sizes of the first angle and the second angle. If the second angle is smaller than the first angle, it indicates that the shooting angle of the first camera is more in line with the speaker's facial orientation, and the target image is selected from the candidate images captured by the first camera. During the selection process, further screening is required based on image quality indicators, including portrait integrity, portrait centering, image sharpness, and image distortion.
[0098] Image integrity is assessed by detecting the pixel ratio of the speaker's image and the completeness of key parts in the image. If the speaker's face, upper body, and other key parts are not missing and the pixel ratio of the image reaches the preset ratio, the image integrity is considered good. Image centering is assessed by calculating the offset between the center of the speaker's image frame and the center of the image frame. The smaller the offset, the higher the image centering. Image sharpness is assessed by calculating the image sharpness value, which is obtained by detecting the grayscale change rate of the edge areas in the image. The higher the grayscale change rate, the sharper the image. Image distortion is assessed by comparing the degree of deformation of standard geometric shapes in the image. The smaller the degree of deformation, the lower the degree of image distortion.
[0099] For the candidate images captured by the first camera, each of the above-mentioned image quality indicators is quantitatively evaluated to obtain an evaluation value for each indicator. Then, the evaluation values of each indicator are weighted and stitched together to obtain an image quality score vector. Based on this score vector, the image with the highest score is selected as the target image.
[0100] If the second included angle is greater than the first included angle, it indicates that the camcorder has a better shooting angle. In this case, the target image is selected from the candidate images captured by the camcorder. Since the camcorder has multiple cameras, the candidate images it captures come from different lenses. The selection process also incorporates image quality metrics. For each candidate image captured by a lens, metrics such as portrait integrity, portrait centering, image sharpness, and image distortion are evaluated to obtain an image quality score vector for each image. The image with the highest score is selected as the target image.
[0101] Figure 6 A flowchart of an optional portrait selection method provided in this application embodiment is shown below. Figure 6 As shown, taking a system containing two microphone-camera all-in-one devices as an example, the master device is the first camera, and the slave devices are the cameras of the two microphone-camera all-in-one devices. After completing the speaker location and deduplication of the portrait, the processor needs to filter the best portrait view from the image data of multiple cameras. The decision-making process mainly revolves around three dimensions: "device selection", "image quality judgment", and "display stability assurance".
[0102] Device selection logic: Prioritize the output of images from the microphone and camera unit and the first camera: By calculating the angle difference between the same speaker and different cameras, the angle difference is used as the main weight (weight > 50%). If the speaker is facing the microphone and camera unit and the angle difference between them is smaller, the microphone and camera unit will be selected for output; if the speaker is facing the first camera and the angle difference between them is smaller, the first camera will be selected for output.
[0103] If the first camera is selected for output, the image captured by the first camera is directly used as the output view; if the camcorder is selected for output, it is necessary to further filter from its multiple lenses (taking 3 as an example): combining parameters such as portrait integrity (whether there is cropping), portrait centering degree, image clarity, and image distortion degree, the optimal lens is determined through weighted calculation, and the image of that lens is finally output.
[0104] Image quality assessment: The core assessment indicators are image sharpness, distortion, and image consistency (weight > 40%). Sharpness is quantified and evaluated using an image sharpness algorithm; distortion is calculated based on camera lens parameters and shooting distance correction; and image consistency is determined by referring to the light balance and color parameters of lenses within the same device (because the light balance and color of multiple lenses in the camcorder are consistent, the cost of consistency adjustment can be reduced).
[0105] Display stability calculation: To ensure smooth screen switching, the main processor records the historical output device and view parameters. When it is necessary to switch the output device or lens, it judges the difference between the screen before and after the switch (such as the position of the portrait, color deviation, etc.). If the difference exceeds the set threshold, it will achieve a smooth switch through a gradual transition algorithm to avoid visual stuttering caused by sudden switching.
[0106] Figure 7 A flowchart illustrating another optional portrait selection method provided in this application embodiment is shown below. Figure 7 As shown, combining the main view and the deduplication results, a schematic diagram of each target portrait is output, showing the optimal view selection logic of targets 1-5 under different perspectives of the main device (first camera), slave device 1 (camera of the microphone all-in-one camera), and slave device 2 (camera of the microphone all-in-one camera). For example, target 1 displays the optimal view selected from the three devices, target 2 displays the optimal view selected from the main device and slave device 1, etc., and if the main device cannot see the target, it will not be displayed even if the slave device can see it (such as target 5).
[0107] In some embodiments, after step S160 above, the portrait processing method may further include the following steps: Step S170: If the currently determined target image and the previously determined target image meet the difference conditions, perform a gradient transition process on the currently determined target image to obtain the processed target image for display; wherein, the difference conditions include at least one of the following: the currently determined target image and the previously determined target image are acquired by different devices, and the image difference between the currently determined target image and the previously determined target image is greater than the difference threshold.
[0108] In this embodiment, the determination of the difference condition is used to ensure the smoothness of screen switching and avoid visual discomfort to users due to sudden image changes. First, it is determined whether the acquisition device of the currently determined target image is the same as that of the previously determined target image. If the acquisition device is different, it means that the image source has changed and the difference condition is met; if the acquisition device is the same, the screen difference between the two is further calculated.
[0109] The calculation of image difference involves multiple dimensions, including differences in portrait position, color, and brightness. Portrait position difference is calculated by determining the coordinate offset of the speaker's image frame center in the two images; color difference is calculated by determining the average difference in the RGB color values of corresponding pixels in the two images; brightness difference is calculated by determining the difference in the average brightness values of the two images. The difference values from each dimension are weighted and concatenated to obtain an image difference vector. If the overall difference level corresponding to this vector is greater than a preset difference threshold, the difference condition is considered met.
[0110] When the difference condition is met, a gradual transition is applied to the currently determined target image. This transition uses a frame interpolation algorithm to generate multiple transition frames between the current and previous target images. The number of transition frames is dynamically adjusted based on the degree of difference between the images; the greater the difference, the more transition frames are used. The image parameters in the transition frames gradually transition from those of the previous target image to those of the current target image, including smooth movement of the person's position, gradual color adjustments, and gradual changes in brightness. Through the seamless connection of these transition frames, a smooth switch between the current and previous target images is achieved, avoiding visual stuttering or abrupt changes and ensuring the smoothness of the displayed image.
[0111] The current target image after the gradient transition processing is the final target image used for display. This target image has the best viewing angle, good image quality and smooth display effect, and can provide remote users with a clear, realistic and comfortable visual experience for the meeting.
[0112] Figure 8 A flowchart illustrating another optional portrait processing method provided in this application embodiment is shown below. Figure 8 As shown in the embodiment of this application, a portrait processing method is also provided, applied to a microphone-camera integrated camera, the method comprising: Step S810: In response to the first shooting command sent by the processor based on the speaker's location data, first portrait data is collected for the speaker; Step S820: Perform first portrait processing on the first portrait data to obtain a first processing result; wherein, the first processing result is used to determine the target image containing the speaker.
[0113] For specific processing methods and beneficial effects, please refer to the portrait processing method in any of the foregoing embodiments; repeated content will not be repeated here.
[0114] Optionally, step S820 above can be implemented in the following way: S821, Perform facial recognition on each image in the first facial image data to obtain the image recognition result corresponding to each image; wherein, the image recognition result includes at least one first facial image identifier; S822, perform in-device deduplication processing on the image recognition results corresponding to multiple images in the first portrait data to obtain the first portrait recognition result; wherein, the first processing result includes the first portrait recognition result.
[0115] Optionally, step S822 above can be implemented in the following way: The image recognition result corresponding to the first camera image is mapped to the spatial coordinate system of the second camera to obtain the projection recognition result corresponding to the first camera image; wherein, the first portrait data includes the first camera image captured by the first camera and the second camera image captured by the second camera; object matching is performed on the projection recognition result corresponding to the first camera image and the image recognition result corresponding to the second camera image; for the same object, a first portrait identifier of the object is retained from the projection recognition result corresponding to the first camera image and the image recognition result corresponding to the second camera image; the first portrait recognition result is determined based on the first portrait identifiers of multiple objects.
[0116] Optionally, the method further includes: acquiring the speaker's directional data; wherein the directional data is used to determine the speaker's location data.
[0117] Figure 9 This is a structural block diagram of an optional portrait processing system provided in an embodiment of this application, with reference to... Figure 9 As shown in the illustration, this application also provides a portrait processing system 900 for implementing the portrait processing method described in any of the foregoing embodiments. The system includes: a microphone / camera integrated unit 910, a first camera 920, and a processor 930 connected to the microphone / camera integrated unit 910 and the first camera 920.
[0118] The processor is used to: acquire the speaker's location data; and, based on the location data, send a first shooting command to the microphone and camera, and a second shooting command to the first camera.
[0119] The camera is used to: in response to a first shooting command, collect first image data of the speaker; perform first image processing on the first image data to obtain a first processing result; and send the first processing result to the processor.
[0120] The first camera is used to: in response to a second shooting command, acquire second portrait data of the speaker; perform second portrait processing on the second portrait data to obtain a second processing result; and send the second processing result to the processor; wherein the second portrait data includes an image from a global perspective; The processor is also configured to: determine at least one candidate image containing a speaker based on the first processing result and the second processing result; and determine a target image for display based on the at least one candidate image.
[0121] This application also provides a processor, including a first memory and a first processor. The first memory stores a computer program or instructions. When the computer program or instructions are executed by the first processor, the first processor performs the portrait processing method as described in any of the above embodiments.
[0122] This application also provides a microphone-camera integrated device, including a second memory, a second processor, a microphone array, and a camera; the microphone array is used to collect sound data, and the camera is used to collect images; the second memory stores computer programs or instructions, and when the computer programs or instructions are executed by the second processor, the second processor performs the portrait processing method described in any of the above embodiments.
[0123] Optionally, the field of view of the integrated camera is greater than or equal to 90 degrees and less than or equal to 360 degrees.
[0124] Optionally, the camera-camera system includes at least two cameras, and the at least two cameras have different shooting angles.
[0125] Optionally, the camera in the camcorder is a fisheye camera.
[0126] Fisheye cameras, with their ultra-wide field of view (typically 180 degrees or more), can capture a wider range of spaces within the coverage area of a single camera, reducing the number of cameras required for the entire camera kit and lowering hardware costs and installation complexity. Simultaneously, fisheye cameras effectively cover blind spots in meeting room corners and edges that are difficult for traditional cameras to reach, avoiding missed images due to dispersed participants or those in off-center locations. They are particularly suitable for small to medium-sized roundtable meetings and multi-person seating scenarios. Combined with fisheye image distortion correction algorithms (such as distortion correction technology based on spherical projection), they can output images without significant distortion while retaining the advantage of a wide field of view, further improving the completeness and effectiveness of image capture in various scenarios.
[0127] This application also provides a computer-readable storage medium storing a computer program or instructions thereon, which, when executed by a processor, implements the steps in the portrait processing method as described in any of the above embodiments.
[0128] In the embodiments of this application, the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0129] The portrait processing method provided in the above embodiments can achieve the following beneficial effects: High deployment flexibility: With the manual / automatic positioning solution of the Maicam all-in-one machine, there is no need to fix the equipment in a specific location, and it supports cascading of multiple devices. The equipment layout can be flexibly adjusted according to the size of the conference room and the number of participants. At the same time, during deployment, it is only necessary to ensure that the host (first camera) and the Maicam all-in-one machine are on the same straight line, reducing layout constraints.
[0130] High processing efficiency: The camcorder has a local processor that can perform deduplication and preliminary image optimization of portraits from local lenses without transferring all data to the main processor, reducing data transmission and computational pressure on the main processor; in addition, the light balance and color of multiple lenses in the same camcorder are consistent, which can reduce the number of image consistency adjustment steps, further shorten the processing time, and avoid display delay.
[0131] Excellent positioning accuracy: By capturing the speaker's directional data through a multi-microphone array, multiple directional lines are generated, and the optimal intersection point is determined through confidence level filtering. This effectively reduces the impact of environmental noise and equipment accuracy deviation on positioning, thereby improving the accuracy of speaker positioning.
[0132] Comprehensive field of view: Combining the multi-lens camera (100-360 degree field of view) and the front camera (wide-angle / wide-angle + telephoto), it can cover the front, middle and surrounding areas of the conference room and the participants. It can not only output close-ups of the speaker, but also use RoomView panoramic output (displayed in a small window at the bottom of the screen) to help present the status of other participants, solving the problem that traditional solutions only output the speaker and remote users cannot see the micro-expressions of other participants.
[0133] Stable image output quality: The optimal portrait image output decision comprehensively considers the viewing angle, image quality, and display stability. Through weight allocation and algorithm optimization, it ensures that the output portrait view fits the real scene, has high clarity, no obvious distortion, and smooth screen transitions, thus improving the visual experience of the meeting.
[0134] Equipment and speaker positioning accuracy control: The automatic positioning of the microphone camera relies on image feature matching and coordinate system transformation. It is necessary to solve the problem of the difficulty in recognizing image feature points under different shooting angles. In speaker positioning, the intersection of multiple directional lines is easily affected by environmental interference and deviations occur. It is necessary to optimize the confidence screening algorithm to ensure that the intersection point is consistent with the actual speaker position.
[0135] Optimal balance of portrait decision weights: Image output decisions involve multiple dimensions of parameters such as viewing angle (weight > 50%), output quality (weight > 40%), and display stability. Extensive experiments are needed to verify the rationality of weight allocation in different scenarios to avoid performance degradation in other dimensions due to excessive weight in one dimension (such as over-pursuing viewing angle while ignoring sharpness).
[0136] Cross-device collaboration and data synchronization: Data needs to be transmitted in real time between the multi-microphone camera, the front camera and the main processor. Data transmission delay and time synchronization issues between multiple devices need to be resolved to avoid portrait deduplication errors and positioning deviations caused by data asynchrony.
[0137] In this embodiment, all collection and processing of privacy-sensitive data employs data encryption technology. Image data collected by the integrated camera and the first camera are transmitted using an encrypted transmission protocol to prevent data theft or tampering during transmission. When processing facial feature data, the processor de-identifies the feature data, removing sensitive information directly linked to personal identity. All processed image and feature data are stored in an encrypted storage module, accessible only to authorized devices and processes, ensuring the privacy and data security of participants. Furthermore, the operation of all devices and data processing procedures comply with relevant laws, regulations, and industry standards, and do not violate laws, social ethics, or harm public interests.
[0138] In the above embodiments, the descriptions of each embodiment have different focuses. Parts not described in detail in a particular embodiment can be found in the relevant descriptions of other embodiments. Furthermore, those skilled in the art will recognize that, based on the spirit of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for processing human images, characterized in that, Applied to a processor, the method includes: Obtain the speaker's location data, which is the spatial location data of the meeting; According to the location data, the microphone and camera are controlled to collect first image data of the speaker, and the microphone and camera are controlled to perform first image processing on the first image data to obtain a first processing result; Based on the location data, control the first camera to acquire second image data of the speaker; The second portrait data is subjected to a second portrait processing to obtain a second processing result; The second portrait data includes images from a global perspective, and the first portrait data includes images from a directional overlay perspective pointing towards the speaker. Based on the first processing result and the second processing result, at least one candidate image containing the speaker is determined; A target image for display is determined based on at least one of the candidate images.
2. The method according to claim 1, characterized in that, The acquisition of the speaker's location data includes: Based on the speaker's directional data, the speaker's location data is determined; wherein the directional data originates from the microphone and / or the processor.
3. The method according to claim 2, characterized in that, Determining the speaker's location data based on the speaker's location data includes: When there are at least two orientation data points; Based on each of the aforementioned directional data, a virtual ray pointing towards the speaker's direction is constructed with the device that acquired the directional data as the origin. The speaker's location data is determined based on the intersection of at least two of the virtual rays.
4. The method according to claim 1, characterized in that, The step of determining at least one candidate image containing the speaker based on the first processing result and the second processing result includes: Based on the first processing result, obtain the first portrait feature data and the first portrait recognition result; Based on the second processing result, obtain the second facial feature data and the second facial recognition result; Based on the first portrait feature data and the second portrait feature data, cross-device portrait deduplication processing is performed on the first portrait recognition result and the second portrait recognition result to obtain a comprehensive portrait recognition result; Based on the comprehensive facial recognition results, at least one candidate image containing the speaker is determined from the first facial data and the second facial data.
5. The method according to claim 4, characterized in that, The step of performing cross-device deduplication processing on the first and second image recognition results based on the first and second image feature data to obtain a comprehensive image recognition result includes: Face feature matching is performed on the first face feature data and the second face feature data to obtain face matching results; wherein, the face matching results include the face similarity between the objects corresponding to each first image identifier and the objects corresponding to each second image identifier; Human feature matching is performed on the first human feature data and the second human feature data to obtain human feature matching results; wherein, the human feature matching results include the human feature similarity between the objects corresponding to each of the first human image identifiers and the objects corresponding to each of the second human image identifiers; Based on the face matching results and / or the human figure matching results, calculate the comprehensive similarity between the object corresponding to each of the first human figure identifiers and the object corresponding to each of the second human figure identifiers; The first and second portrait icons that meet the matching conditions in terms of comprehensive similarity are assigned the same comprehensive portrait icon. Wherein, the first portrait feature data includes the first face feature data and / or the first human figure feature data, and the first portrait recognition result includes at least one first portrait identifier; the second portrait feature data includes the second face feature data and / or the second human figure feature data, and the second portrait recognition result includes at least one second portrait identifier; the comprehensive portrait recognition result includes at least one of the comprehensive portrait identifiers.
6. The method according to claim 1, characterized in that, The step of determining the target image for display based on at least one of the candidate images includes: The target image is selected from at least one candidate image based on the first angle between the orientation of the speaker's feature point and the center line of the field of view of the camera and the second angle between the orientation of the speaker's feature point and the center line of the field of view of the first camera.
7. The method according to claim 6, characterized in that, Based on a first angle between the orientation of the speaker's feature point and the center line of the camera's field of view, and a second angle between the orientation of the speaker's feature point and the center line of the first camera's field of view, the target image is selected from at least one candidate image, including: If the second included angle is smaller than the first included angle, then at least one of the candidate images captured by the first camera will be used as the target image; If the second included angle is greater than the first included angle, then the target image is selected from the candidate images acquired by the microphone camera from at least one candidate image.
8. The method according to any one of claims 1 to 7, characterized in that, The microphone-camera system has multiple cameras. When the candidate image originates solely from the multiple cameras of the microphone-camera system: The microphone and camera all-in-one device acquires the location data and collects at least two candidate image data for the speaker; The candidate image with the highest image quality score among at least one candidate image captured by the camera is taken as the target image; The image quality score is used to indicate the comprehensive score of the candidate image under multiple image indicators, which include at least one of the following: portrait integrity, portrait centering, image sharpness, and image distortion.
9. The method according to claim 1, characterized in that, After determining the target image for display based on at least one of the candidate images, the process further includes: If the currently determined target image meets the difference condition with the previously determined target image, a gradient transition processing is performed on the currently determined target image to obtain the processed target image for display. The difference conditions include at least one of the following: the currently determined target image and the previously determined target image were acquired by different devices; or the image difference between the currently determined target image and the previously determined target image is greater than a difference threshold.
10. A method for processing human images, characterized in that, Applied to a camcorder, the method includes: In response to a first shooting command sent by the processor based on the speaker's location data, first image data of the speaker is collected; The first portrait data is subjected to a first portrait processing to obtain a first processing result; wherein, the first processing result is used to determine a target image containing the speaker.
11. The method according to claim 10, characterized in that, The first image processing of the first image data to obtain a first processing result includes: Perform facial recognition on each image in the first facial image data to obtain an image recognition result corresponding to each image; wherein, the image recognition result includes at least one first facial image identifier; The image recognition results corresponding to multiple images in the first portrait data are subjected to in-device deduplication processing to obtain a first portrait recognition result; wherein, the first processing result includes the first portrait recognition result.
12. The method according to claim 11, characterized in that, The step of performing in-device deduplication processing on the image recognition results corresponding to multiple images in the first portrait data to obtain the first portrait recognition result includes: The image recognition result corresponding to the first camera image is mapped to the spatial coordinate system of the second camera to obtain the projection recognition result corresponding to the first camera image; wherein, the first portrait data includes the first camera image captured by the first camera and the second camera image captured by the second camera. Object matching is performed on the projection recognition result corresponding to the first camera image and the image recognition result corresponding to the second camera image; For the same object, a first human image identifier of the object is retained from the projection recognition result corresponding to the first camera image and the image recognition result corresponding to the second camera image; The first portrait recognition result is determined based on the first portrait identifier of the multiple objects.
13. The method according to claim 10, characterized in that, The method further includes: Obtain the speaker's location data; wherein the location data is used to determine the speaker's position data.
14. A human image processing system, characterized in that, The system includes: a microphone / camcorder combo unit, a first camera, and a processor connected to the microphone / camcorder combo unit and the first camera; wherein, The processor is used to: acquire the speaker's location data; send a first shooting command to the microphone and camera all-in-one device and a second shooting command to the first camera according to the location data; The microphone and camera are used to: in response to the first shooting command, collect first portrait data of the speaker; perform first portrait processing on the first portrait data to obtain a first processing result; and send the first processing result to the processor. The first camera is configured to: in response to the second shooting command, acquire second portrait data of the speaker; perform second portrait processing on the second portrait data to obtain a second processing result; and send the second processing result to the processor; wherein the second portrait data includes an image from a global perspective; The processor is further configured to: determine at least one candidate image containing the speaker based on the first processing result and the second processing result; and determine a target image for display based on the at least one candidate image.
15. A processor, characterized in that, It includes a first memory and a first processor, wherein the first memory stores a computer program or instructions, and when the computer program or instructions are executed by the first processor, the first processor performs the portrait processing method as described in any one of claims 1 to 9.
16. A microphone / camcorder combo camera, characterized in that, The device includes a second memory, a second processor, a microphone array, and a camera; the microphone array is used to acquire sound data, and the camera is used to acquire images; the second memory stores a computer program or instructions, which, when executed by the second processor, cause the second processor to perform the portrait processing method as described in any one of claims 10 to 13.
17. The all-in-one camera with microphone and camcorder according to claim 16, characterized in that, The field of view of the camera is greater than or equal to 90 degrees and less than or equal to 360 degrees.
18. The camera / video recorder according to claim 17, characterized in that, The camera and microphone combo unit includes at least two cameras, and the at least two cameras have different shooting angles.
19. The all-in-one camera with microphone and camcorder according to claim 17, characterized in that, The camera in the aforementioned camcorder is a fisheye camera.
20. A computer-readable storage medium, characterized in that, It stores a computer program or instructions thereon, which, when executed by a processor, implement the steps in the portrait processing method as described in any one of claims 1 to 9, or implement the steps in the portrait processing method as described in any one of claims 10 to 13.