Object display method, apparatus, system, device, medium, and product
By using face recognition and voice tracking detection, the target area is identified and the object is adaptively displayed in the video image, solving the problem of the inability to adaptively display objects in the video image and realizing the adaptive display of speakers and participants.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG DAHUA TECH CO LTD
- Filing Date
- 2022-07-22
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies cannot adaptively display objects in video images based on different meeting scenarios, especially they cannot effectively identify and display speakers and all participants.
By performing face recognition and voice tracking detection on a preset area, the positions of face objects and sound sources are obtained. The target area is determined based on the intersection position, and the corresponding face objects are highlighted in the video image to achieve adaptive display.
This invention enables adaptive display of objects in video images based on different meeting scenarios. It prioritizes displaying the speaker when a speaker is detected and displays all participants when no speaker is detected, thus solving the problem of adaptive display in existing technologies.
Smart Images

Figure CN115426474B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video image processing technology, and in particular to an object display method, apparatus, system, device, medium and product. Background Technology
[0002] Live video streaming and video conferencing offer efficient and convenient solutions for remote work, significantly improving work efficiency. Related technologies provide a method for automatically selecting people by performing face detection on camera images, outputting the face detection results, calculating the selection area based on the results, and processing the camera image according to the selected area, ensuring that all participants are displayed prominently in the camera image. However, sometimes meeting scenarios require focusing on the speaker, not all participants, and the speaker's identity is not fixed; the speaker may also be moving during the video conference.
[0003] There is currently no effective solution to the problem that related technologies cannot adaptively display objects based on different meeting scenarios. Summary of the Invention
[0004] Therefore, it is necessary to provide an object display method, device, system, equipment, medium, and product that can adaptively display objects in video images based on different meeting scenarios, in order to address the above-mentioned technical problems.
[0005] Firstly, this application provides a method for displaying objects, including:
[0006] Perform facial recognition detection and voice tracking detection on the preset area;
[0007] When a face object and a voice signal are detected, the positions of the face object and the sound source in the video image are obtained, and a target area is determined in the video image based on the intersection position of the face object and the sound source.
[0008] If the face object is detected but the voice signal is not detected, the position of the face object in the video image is obtained, and the target area is determined in the video image based on the position of the face object.
[0009] The corresponding face object is highlighted based on the target area.
[0010] In one embodiment, the recognition result of the face object includes a face detection box, and obtaining the position of the face object in the video image includes:
[0011] The vertex coordinates and size of the face detection box are adjusted based on a preset image resolution, wherein the aspect ratio of the preset image resolution is 1:1.
[0012] In one embodiment, upon detecting a face object and a voice signal, the method further includes obtaining the positions of the face object and the sound source in a video image, and determining a target region in the video image based on the intersection position of the face object and the sound source.
[0013] With the camera lens in its initial accelerator state, determine the relative positional relationship between the target area and the camera's field of view;
[0014] If the target area is not completely contained within the camera's field of view, rotate the camera lens horizontally until the target area is completely contained within the camera's field of view.
[0015] In one embodiment, after determining the target region in the video image, the method further includes:
[0016] Compare the current target area with the historical target area determined in the previous stage, and determine whether the deviation between the current target area and the historical target area is greater than a preset threshold.
[0017] If it is determined that the deviation between the current target region and the historical target region is greater than a preset threshold, digital image processing is performed on the image of the current target region, wherein the digital image processing includes cropping and scaling.
[0018] In one embodiment, digital image processing of the image of the current target region includes:
[0019] If the scaling factor of the current target region is less than the scaling factor of the historical target region, the image of the current target region is first reduced in size and then shifted; or,
[0020] If the scaling factor of the current target area is greater than that of the historical target area, the image of the current target area is first translated and then magnified.
[0021] In one embodiment, the method further includes:
[0022] In response to a first instruction, a first preset mode is activated, wherein the first preset mode is configured to, upon detecting the voice signal, acquire the positions of the face object and the sound source in the video image, and determine the target region in the video image based on the intersection position of the face object and the sound source; and / or,
[0023] In response to the second instruction, a second preset mode is activated, wherein the second preset mode is configured to obtain the position of the face object in the video image, and determine the target region in the video image based on the position of the face object.
[0024] In one embodiment, when both the first preset mode and the second preset mode are activated, the method further includes:
[0025] If the voice signal is not detected within a preset time period or after a preset number of detections, the display mode of the video image will be switched from the first preset mode to the second preset mode.
[0026] In one embodiment, highlighting a corresponding face object based on the target region includes:
[0027] The target area includes a geometric selection box, which is used to select the corresponding face object; or, the target area includes a geometric shape, which is used to mark the corresponding face object; or, the target area is centered in the video image, and the corresponding face object is displayed in the target area.
[0028] Secondly, this application provides a data processing device, including: a face recognition module, a voice tracking module, and a main control module, wherein the face recognition module and the voice tracking module are respectively connected to the main control module;
[0029] The face recognition module is configured to perform face recognition detection in a preset area, and the voice tracking module is configured to perform voice tracking detection in the preset area;
[0030] The main control module is configured to, when a face object and a voice signal are detected, acquire the positions of the face object and the voice source in a video image, and determine a target region in the video image based on the intersection position of the face object and the voice source; when a face object is detected but no voice signal is detected, acquire the position of the face object in the video image, and determine the target region in the video image based on the position of the face object; and highlight the corresponding face object based on the target region.
[0031] Thirdly, this application provides a system for determining a face object, comprising: a camera, a microphone, a playback device, and the data processing device described in the second aspect above, wherein the camera, the microphone, and the playback device are respectively connected to the data processing device; the camera is used to capture video of a preset area; the microphone is used to collect audio signals of the preset area; and the playback device is used to output the video image and audio signals processed by the data processing device.
[0032] Fourthly, this application provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the object display method described in the first aspect above.
[0033] Fifthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the object display method described in the first aspect above.
[0034] In a sixth aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the object display method described in the first aspect above.
[0035] The aforementioned object display method, apparatus, system, device, medium, and product perform face recognition detection and voice tracking detection on a preset area; when a face object and voice signal are detected, the positions of the face object and sound source in the video image are obtained, and a target area is determined in the video image based on the intersection position of the face object and the sound source; when a face object is detected but no voice signal is detected, the position of the face object in the video image is obtained, and a target area is determined in the video image based on the position of the face object; the corresponding face object is highlighted based on the target area. This solves the problem of not being able to adaptively display objects in video images based on different meeting scenarios, and achieves the beneficial effect of adaptively displaying objects in video images based on different meeting scenarios. Attached Figure Description
[0036] Figure 1 This is a diagram illustrating the application environment of an object display method in one embodiment;
[0037] Figure 2 This is a flowchart illustrating an object display method in one embodiment;
[0038] Figure 3 This is a schematic diagram showing the positions of a face object and a sound source in a video image in one embodiment;
[0039] Figure 4 This is a schematic diagram of the camera's field of view in one embodiment;
[0040] Figure 5 A flowchart illustrating the overall method for selecting a face object in one embodiment;
[0041] Figure 6 This is a flowchart illustrating the selection of a face object in a first preset mode in one embodiment.
[0042] Figure 7 This is a schematic diagram of the target region in a video image in one embodiment;
[0043] Figure 8 This is a flowchart illustrating the selection of a face object in a second preset mode in one embodiment.
[0044] Figure 9 This is a flowchart of adjusting the target area in one embodiment;
[0045] Figure 10 This is a schematic diagram of the structure of a data processing device in one embodiment;
[0046] Figure 11 This is a schematic diagram of the structure of an object display system in one embodiment;
[0047] Figure 12 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0049] The object display method provided in this application embodiment can be applied to, for example, Figure 1In the application environment shown, the application environment can be live video streaming or video conferencing. The terminal device 100 includes a camera 101, a microphone 102, a display screen 103, and a speaker 104. In addition, the terminal device 100 has a data processing device (not shown in the figure) installed inside. The camera 101, microphone 102, display screen 103, and speaker 104 are respectively connected to the data processing device. During the execution of the method, the camera 101 captures video of a preset area, the microphone 102 collects audio signals from the preset area, the data processing device processes the video and audio signals, and sends the processing results to the display screen 103 and speaker 104 for output. Specifically, the data processing device performs face recognition detection and voice tracking detection on a preset area based on video and voice signals; when a face object and voice signal are detected, the device obtains the positions of the face object and the sound source in the video image, and determines a target area 105 in the video image based on the intersection position of the face object and the sound source; when a face object is detected but no voice signal is detected, the device obtains the position of the face object in the video image, and determines a target area 105 in the video image based on the position of the face object; and highlights the corresponding face object based on the target area 105.
[0050] One embodiment provides an object display method that can be applied to Figure 1 In the application environment shown, Figure 2 Here is a flowchart of the method, which includes the following steps:
[0051] Step S201: Perform face recognition detection and voice tracking detection on the preset area.
[0052] The data processing device receives audio and video data (video and audio signals) captured by the camera and microphone, and performs face recognition detection and voice tracking detection, which can be performed in parallel. The voice tracking detection result includes the sound source localization angle, which refers to the angle of the sound source relative to the camera lens when the camera lens is in its initial aligned state. The location of the sound source in the video image can be obtained by acquiring the relative positional relationship between the camera and microphone, and then using this relative positional relationship and the sound source localization angle. The relative positional relationship between the camera and microphone can be pre-calibrated, and the calibration parameters are stored in the data processing device.
[0053] Step S202: When a face object and a voice signal are detected, the positions of the face object and the sound source in the video image are obtained, and the target area is determined in the video image based on the intersection position of the face object and the sound source.
[0054] Upon detecting a voice signal, indicating that someone has spoken at the meeting, the first preset mode will be activated first to highlight the speaker in the video image. There can be one or more faces, one or more sound sources, and one or more target areas. To avoid introducing external interference, an effective monitoring area can be set in the video image; only faces and sound sources located within the effective monitoring area are considered valid data. Please refer to [link / reference]. Figure 3 , Figure 3 This diagram illustrates the positions of the face and sound source in the video image in this embodiment. A Cartesian coordinate system is established in the video image, with the horizontal axis as the X-axis and the vertical axis as the Y-axis. The intersection of the face and sound source means that the face and sound source have at least the same horizontal coordinate. The effective monitoring area is Area1. S1 to S5 represent the face objects, presented as detection boxes in the diagram. X1 to X5 represent the horizontal coordinates of the sound source, presented as vertical lines for ease of understanding, and the horizontal coordinates of these lines are defined. Since X1 and S2 are outside Area1, they are invalid data; there is no face object on the vertical line corresponding to X3, so it is also invalid data; S1 is on the vertical line corresponding to X2, and S3 and S4 are on the vertical line corresponding to X4. S1, S3, and S4 are all within the effective monitoring area, so the target areas are the locations of S1, S3, and S4, respectively. It should be noted that the sound source localization result includes coordinates along the horizontal and / or vertical directions of the video image. That is, the sound source localization result may only have the horizontal coordinate, only the vertical coordinate, or both the horizontal and vertical coordinates.
[0055] Step S203: When a face object is detected but no voice signal is detected, the position of the face object in the video image is obtained, and the target area is determined in the video image based on the position of the face object.
[0056] If no voice signal is detected, it means that no one is speaking at the meeting. In this case, the first preset mode will be disabled and the second preset mode will be enabled to highlight all participants in the video image.
[0057] Step S204: Highlight the corresponding face object based on the target area.
[0058] Display the corresponding face object in the target area. Optionally, the target area includes a geometric selection box to select the corresponding face object, such as surrounding the face object with a box or ellipse; or, the target area includes geometric shapes to mark the corresponding face object, such as overlaying an indicator arrow above the face object; or, the target area is centered in the video image and the corresponding face object is displayed in the target area.
[0059] In steps S201 to S204 above, face recognition detection and voice tracking detection are performed on the preset area to identify two situations: someone speaking and no one speaking. In response to these two meeting situations, an adaptive switch is made between a first preset mode and a second preset mode. That is, when a voice signal is detected, the first preset mode is activated first to highlight the speaker in the video image. When no voice signal is detected, the first preset mode is blocked and the second preset mode is activated to highlight all participants in the video image. This solves the problem of not being able to adaptively display objects in the video image based on different meeting scenarios and achieves the beneficial effect of adaptively displaying objects in the video image based on different meeting scenarios.
[0060] In one embodiment, the first preset mode and the second preset mode can be enabled or disabled based on user instructions. Optionally, the data processing device, in response to the first instruction, activates the first preset mode, wherein the first preset mode is configured to, upon detecting a voice signal, acquire the positions of a face object and a sound source in a video image, and determine a target region in the video image based on the intersection position of the face object and the sound source. Optionally, the data processing device, in response to the second instruction, activates the second preset mode, wherein the second preset mode is configured to acquire the position of a face object in a video image, and determine a target region in the video image based on the position of the face object. Optionally, the data processing device, in response to the first and second instructions, activates the first and second preset modes, and adaptively switches between the first and second preset modes depending on whether someone speaks at the meeting; that is, when both the first and second preset modes are enabled, the first preset mode is executed first, and under certain conditions, it can automatically switch to the second preset mode.
[0061] In one embodiment, a method is provided for automatically switching from the first preset mode to the second preset mode when both the first preset mode and the second preset mode are activated: if no voice signal is detected within a preset duration or after a preset number of detections, the display mode of the video image is switched from the first preset mode to the second preset mode.
[0062] In one embodiment, the recognition result of the face object includes a face detection box, and obtaining the position of the face object in the video image includes: adjusting the vertex coordinates and size of the face detection box based on a preset image resolution, wherein the aspect ratio of the preset image resolution is 1:1.
[0063] refer to Figure 3S1 to S5 represent face objects, and the vertical lines corresponding to X1 to X5 represent the horizontal coordinates of the sound source. Assume the video image resolution is W1×H1, and the preset image resolution is M×M, where W1 is the width and H1 is the height. Assume the vertex coordinates of the face detection box are (P1, Q1), and the face detection box size is W2×H2. The following is the calculation formula for normalizing the vertex coordinates and size of the face detection box:
[0064] The vertex coordinates of the face detection bounding box are adjusted as follows: P1' = P1 × W1 / M, Q1' = P1 × H1 / M;
[0065] The results of the face detection bounding box size adjustment are: W2' = W2 × W1 / M, H2' = H2 × H1 / M.
[0066] By normalizing the face detection bounding box, it can adapt to video image input of any ratio. That is, in the face detection bounding box, the same coordinate in the original image will have the same normalized coordinate under the same aspect ratio at different resolutions. For the preset image resolution of this embodiment, it is sufficient to ensure that the aspect ratio is 1:1.
[0067] In one embodiment, upon detecting a face object and a voice signal, the positions of the face object and the sound source in the video image are obtained. After determining the target area in the video image based on the intersection position of the face object and the sound source, the camera's shooting angle is adjusted until the face object contained in the target area is displayed in the center of the video image.
[0068] Related technologies often employ cameras with limited fields of view, which cannot be adjusted once fixed. Returning to this application, the same problem exists after determining the target area: the target area may be located at the boundary of the video image, causing the face to not be centered. To solve this problem, the camera in this embodiment has an adjustable shooting angle. By adjusting the camera's shooting angle, the face contained within the target area is centered in the video image. Specifically, the camera lens can be adjusted left and right, thus ensuring that even when the target area is at the boundary, it can still be centered in the video image when the camera lens is initially centered.
[0069] Furthermore, in one embodiment, a method for adjusting the camera's shooting angle is provided. Adjusting the camera's shooting angle until the face object contained in the target area is centered in the video image includes:
[0070] With the camera lens initially centered, determine the relative position of the target area to the camera's field of view. If the target area is not entirely contained within the camera's field of view, rotate the camera lens horizontally until the target area is completely contained within the camera's field of view. Here, "the target area is not entirely contained within the camera's field of view" means that, with the camera lens initially centered, at least a portion of the target area is outside the field of view.
[0071] Figure 4 A diagram illustrating the camera's field of view is provided, such as... Figure 4 As shown, with the camera lens initially in its centered state, the field of view is Area 1. When the lens is adjusted to the far left, the field of view becomes Area 2, and when it's adjusted to the far right, the field of view becomes Area 3. Only the sound source localization result for Area 1 is valid. When the target area is in Area A or A+B, the lens will adjust slightly to the left to center the face in the video image. When the target area is entirely in Areas A, B, or C, the lens remains centered and does not adjust. When the target area is in Area C or B+C, the lens will adjust slightly to the right to center the face in the video image.
[0072] In one embodiment, after determining the target region in the video image, the method further includes: comparing the current target region with the historical target region determined in the previous stage, and determining whether the deviation between the current target region and the historical target region is greater than a preset threshold; if it is determined that the deviation between the current target region and the historical target region is greater than the preset threshold, performing digital image processing on the image of the current target region, wherein the digital image processing includes cropping and scaling.
[0073] Because face objects have inherent errors and fluctuations, the size of the face detection box changes when a person shakes their head left or right, or tilts their head up or down. This embodiment, through the above steps, achieves an anti-shaking effect, avoiding the shaking phenomenon caused by repeated scaling or translation. The deviation between the current target region and the historical target region includes any of the following: the center point of the current target region deviates from the center point of the historical target region, and the deviation value is greater than a first threshold; the area of the current target region deviates from the area of the historical target region, and the deviation is greater than a second threshold.
[0074] Furthermore, in one embodiment, a method for digital image processing of an image of a current target region is provided. The method includes: if the scaling factor of the current target region is less than the magnification factor of the historical target region, the image of the current target region is first reduced and then translated; or, if the scaling factor of the current target region is greater than the magnification factor of the historical target region, the image of the current target region is first translated and then enlarged.
[0075] When zooming out, if you pan first and then zoom out, the face object may become invisible during the panning process until it is zoomed out. Conversely, when zooming in, if you zoom out first and then pan, the face object may become invisible during the zooming process until it is panned. This setting ensures that the face object remains visible during both zooming and panning, and also provides some anti-shaking effect.
[0076] The object display method will be described below through preferred embodiments.
[0077] Figure 5 This is a flowchart illustrating the overall method for selecting face objects in one embodiment. In this embodiment, the first preset mode is set to be able to select speakers appearing within the camera's field of view, and the second preset mode is set to be able to select all participants appearing within the camera's field of view. The identified face objects are surrounded by detection boxes, and the target region is set as a rectangle, such as... Figure 5 As shown, the process includes the following steps:
[0078] Step S501: Perform face recognition detection on the video image, perform voice tracking detection on the voice signal, and normalize the face object and sound source localization results.
[0079] Step S502: Enable the first preset mode and the second preset mode;
[0080] Step S503: Determine if anyone is speaking; if yes, proceed to step S506; if no, proceed to step S504.
[0081] Step S504: Switch to the second preset mode;
[0082] Step S505: Determine the target region based on the face object;
[0083] Step S506: Switch to the first preset mode;
[0084] Step S507: Determine the target area based on the face object and sound source localization results, and adjust the left and right angles of the camera lens;
[0085] Step S508: Further process the target area;
[0086] Step S509: Determine the digital zoom ratio based on the target area, and adjust the video image according to the digital zoom ratio.
[0087] Combination Figure 5 In one embodiment, Figure 6 A flowchart for selecting a face object in the first preset mode is provided, such as... Figure 6 As shown, the process includes the following steps:
[0088] Step S601: Activate the first preset mode;
[0089] Step S602: Enable voice tracking detection and adjust the camera lens to the initial homing state;
[0090] Step S603: Determine whether a voice signal is detected; if yes, proceed to step S604; if no, proceed to step S619.
[0091] Step S604: Determine whether a face object has been detected; if yes, proceed to step S605; if no, proceed to step S615.
[0092] Step S605: Determine whether the camera lens is in the initial alignment state; if yes, proceed to step S606; if no, proceed to step S616.
[0093] Step S606: Determine the sound source localization angle, and output the X coordinate of the sound source in the video image based on the sound source localization angle;
[0094] Step S607: Compare the X coordinate of the sound source with the position of the face object to determine the target area;
[0095] Step S608: Determine whether the target area does not exist; if yes, proceed to step S621; if no, proceed to step S609.
[0096] Step S609: Expand the target area outwards by the width of the maximum face selection area on each of the four sides;
[0097] Step S610: Determine whether the target area exceeds any of the four boundaries; if yes, proceed to step S617; if no, proceed to step S611.
[0098] Step S611: Adjust the aspect ratio of the target area to 1:1;
[0099] Step S612: Determine whether the target area exceeds any of the four boundaries; if yes, proceed to step S618; if no, proceed to step S613.
[0100] Step S613: Determine whether the deviation of the center point of the target area is greater than the first threshold or whether the deviation of the area of the target area is greater than the second threshold; if yes, proceed to step S614; if no, proceed to step S622.
[0101] Step S614: Perform digital image processing on the target area and output the coordinates of the face object;
[0102] Step S615: The target area is the entire video image with no scaling ratio. The lens angle is corrected and the count is reset to zero.
[0103] Step S616: Obtain the current face object and sound source, and map the current face object and sound source to the original video image;
[0104] Step S617, adjust the target area;
[0105] Step S618: Adjust the target area and determine the left / right tilt angle of the lens;
[0106] Step S619: Determine if there is no continuous voice count timeout; if yes, proceed to step S620; if no, proceed to step S603.
[0107] Step S620: Switch to the second preset mode;
[0108] Step S621, return the lens angle to center;
[0109] Step S622: Clear the counter to zero.
[0110] In this embodiment, if the X-coordinate deviation of the current target area center compared to the historical target area center exceeds 30% of the width of the historical target area, or the Y-coordinate deviation exceeds 30% of the height of the historical target area, or the area deviation of the current target area compared to the historical target area differs by 20%, then digital image processing of the target area is required, including cropping, scaling, and translation. Outputting the new coordinates of the digitally scaled face object relative to the original video image provides effective information for subsequent OSD (On-Screen Display) information overlay.
[0111] Furthermore, Figure 7 A schematic diagram of the target region in the video image is given, such as... Figure 7As shown, M1 is the original video image size, M2 is the target region, and M3 and M4 are face objects. According to the above steps S604, S611, and S612, all face objects and the target region are obtained. That is, the coordinates (x, y) of the top left corner vertex of all face objects, as well as their width and height, are known. The coordinates (X, Y) of the top left corner vertex of the target region, as well as its width W and height H, are known. Then, the coordinate position of the scaled face object relative to the original video image is (x1, y1), width w1, and height h1. x1 = (xX) × width of the original video image / W, y1 = (yY) × height of the original image / H, w1 = width × width of the original image / W, and h1 = height × height of the original image / H.
[0112] Combination Figure 5 In one embodiment, Figure 8 A flowchart for selecting face objects in the second preset mode is provided, such as... Figure 8 As shown, the process includes the following steps:
[0113] Step S801: Switch to the second preset mode and adjust the camera lens to the initial alignment state;
[0114] Step S802: Determine whether a face object has been detected; if yes, proceed to S810; if no, proceed to S803.
[0115] Step S803: Calculate the target region containing all facial objects within the camera's field of view;
[0116] Step S804: Expand the target area outwards by the width of the maximum face selection area on each of the four sides;
[0117] Step S805: The target area exceeds any of the four boundaries; if so, proceed to S811; otherwise, proceed to S806.
[0118] Step S806: Adjust the aspect ratio of the target area to 1:1;
[0119] Step S807: The target area exceeds any of the four boundaries; if so, proceed to S812; otherwise, proceed to S808.
[0120] Step S808: If the deviation of the center point of the target area is greater than the first threshold or the deviation of the area of the selected region is greater than the second threshold; if yes, proceed to S809; otherwise, proceed to S802.
[0121] Step S809: Perform digital image processing on the target area and output the coordinates of the face object;
[0122] Step S810: Set the target area to the entire video image without scaling.
[0123] Step S811, adjust the target area;
[0124] Step S812: Adjust the target area and determine the left / right tilt angle of the lens.
[0125] The output of the new coordinates of the digitally scaled face object relative to the original video image can provide effective information for the subsequent overlay of OSD (On-Screen Display) information.
[0126] Combination Figure 6 and Figure 8 In one embodiment, Figure 9 A flowchart for adjusting the target area is provided, such as... Figure 9 As shown, the process includes the following steps:
[0127] Step S901: Expand the target area outwards by the width of the maximum face selection area on each of the four sides;
[0128] Step S902: Determine whether the width / height of the target area exceeds the original video image; if yes, proceed to step S908; if no, proceed to step S903.
[0129] Step S903: Determine whether the upper boundary of the target area exceeds the original video image; if yes, proceed to step S909; if no, proceed to step S904.
[0130] Step S904: Determine whether the lower boundary of the target area exceeds the original video image; if yes, proceed to step S911; if no, proceed to step S905.
[0131] Step S905: Determine whether the left boundary of the target area exceeds the original video image; if yes, proceed to step S910; if no, proceed to step S906.
[0132] Step S906: Determine whether the right boundary of the target area exceeds the original video image; if yes, proceed to step S912; if no, proceed to step S907.
[0133] Step S907: Adjust the aspect ratio of the target area to 1:1;
[0134] Step S908: Set the target area to the entire video image;
[0135] Step S909: Move the target area down by an amount exceeding the allowable amount;
[0136] Step S910: Move the target area to the right by the amount of excess.
[0137] Step S911: Move the target area upwards by an excess amount;
[0138] Step S912: Shift the target area to the left by an amount exceeding the allowable limit.
[0139] Based on the same inventive concept, this application also provides a data processing apparatus for implementing the object display method described above. Figure 10 A schematic diagram of the structure of a data processing device in one embodiment, such as Figure 10 As shown, it includes: a face recognition module, a voice tracking module, and a main control module. The face recognition module and the voice tracking module are respectively connected to the main control module. The face recognition module is configured to perform face recognition detection in a preset area, and the voice tracking module is configured to perform voice tracking detection in a preset area. The main control module is configured to, when a face object and a voice signal are detected, obtain the positions of the face object and the sound source in the video image, and determine a target area in the video image based on the intersection position of the face object and the sound source; when a face object is detected but no voice signal is detected, obtain the position of the face object in the video image, and determine a target area in the video image based on the position of the face object; and highlight the corresponding face object based on the target area.
[0140] The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations in one or more data processing device embodiments provided below can be found in the limitations of the object display method above, and will not be repeated here.
[0141] In one embodiment, the main control module is further configured to adjust the vertex coordinates and size of the face detection box based on a normalized coordinate system, wherein the aspect ratio of the normalized coordinate system is 1:1.
[0142] In one embodiment, the main control module is further configured to adjust the camera's shooting angle until the face object contained in the current target area is displayed in the center of the video image.
[0143] In one embodiment, the main control module is further configured to: determine the relative positional relationship between the target area and the camera's field of view when the camera lens is in the initial aligning state; and rotate the camera lens horizontally until the target area is completely contained within the camera's field of view when the target area is not fully contained within the camera's field of view.
[0144] In one embodiment, the main control module is further configured to: compare the current target region with the historical target region obtained when determining the face object in the previous stage, and determine whether the deviation between the current target region and the historical target region is greater than a preset threshold; if it is determined that the deviation between the current target region and the historical target region is greater than the preset threshold, perform digital image processing on the image of the current target region, wherein the digital image processing includes cropping and scaling.
[0145] In one embodiment, the main control module is further configured to: shrink and then pan the image of the current target area when the scaling factor of the current target area is less than the magnification factor of the historical target area; or pan and then enlarge the image of the current target area when the scaling factor of the current target area is greater than the magnification factor of the historical target area.
[0146] In one embodiment, the main control module is further configured to: in response to a first instruction, activate a first preset mode, wherein the first preset mode is configured to, upon detecting a voice signal, acquire the positions of a face object and a sound source in a video image, and determine a target region in the video image based on the intersection position of the face object and the sound source; and / or, in response to a second instruction, activate a second preset mode, wherein the second preset mode is configured to acquire the position of a face object in a video image, and determine a target region in the video image based on the position of the face object.
[0147] In one embodiment, the main control module is further configured to: when both the first preset mode and the second preset mode are activated, the method further includes switching the display mode of the video image from the first preset mode to the second preset mode if no voice signal is detected within a preset duration or after a preset number of detections.
[0148] Each module in the aforementioned data processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0149] This application also provides an object display system. Figure 11 This is a schematic diagram of the architecture of an object display system in one embodiment, such as... Figure 11As shown, the system includes: a camera, a microphone, a playback device, and a data processing device as described in the above embodiment. The camera, microphone, and playback device are respectively connected to the data processing device. The camera is used to capture video images of a preset area; the microphone is used to collect audio signals from the preset area; and the playback device is used to output the video images and audio signals processed by the data processing device. In this embodiment, the components of the object display system are independent of each other and connected by cables. The playback device includes at least a display screen and a speaker. The playback device can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart in-vehicle devices, etc., and portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc.
[0150] In one embodiment, the relative positional relationship between the camera and multiple microphones is preset, calibration parameters of the relative positional relationship are obtained, and the calibration parameters are written into the data processing device.
[0151] In one embodiment, the camera lens can tilt left and right, and the data processing device can control the camera lens angle.
[0152] In one embodiment, the components of the object display system are integrated into one unit, such as... Figure 1 As shown, the object display system includes Figure 1 The terminal device 100 shown integrates a camera 101, a microphone 102, a playback device (display screen 103 and speaker 104), and a data processing device into one unit.
[0153] This application also provides a computer device, the internal structure of which can be shown in the following embodiment: Figure 12 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an object display method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.
[0154] Those skilled in the art will understand that Figure 12 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0155] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, performs the following steps:
[0156] Step S201: Perform face recognition detection and voice tracking detection on the preset area;
[0157] Step S202: When a face object and a voice signal are detected, the positions of the face object and the sound source in the video image are obtained, and the target area is determined in the video image based on the intersection position of the face object and the sound source.
[0158] Step S203: When a face object is detected but no voice signal is detected, the position of the face object in the video image is obtained, and the target area is determined in the video image based on the position of the face object.
[0159] Step S204: Highlight the corresponding face object based on the target area.
[0160] This application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0161] Step S201: Perform face recognition detection and voice tracking detection on the preset area;
[0162] Step S202: When a face object and a voice signal are detected, the positions of the face object and the sound source in the video image are obtained, and the target area is determined in the video image based on the intersection position of the face object and the sound source.
[0163] Step S203: When a face object is detected but no voice signal is detected, the position of the face object in the video image is obtained, and the target area is determined in the video image based on the position of the face object.
[0164] Step S204: Highlight the corresponding face object based on the target area.
[0165] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0166] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0167] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for displaying an object, characterized in that, The method is applied to a data processing device connected to a single camera, which is used to capture images of a preset area, including: Perform facial recognition detection and voice tracking detection on the preset area; When a face object and a voice signal are detected, the positions of the face object and the sound source in the video image are obtained, and a target area is determined in the video image based on the intersection position of the face object and the sound source. If the face object is detected but the voice signal is not detected, the position of the face object in the video image is obtained, and the target area is determined in the video image based on the position of the face object. Highlight the corresponding face object based on the target area; After determining the target region in the video image, the method further includes: Compare the current target area with the historical target area determined in the previous stage, and determine whether the deviation between the current target area and the historical target area is greater than a preset threshold. If it is determined that the deviation between the current target region and the historical target region is greater than a preset threshold, digital image processing is performed on the image of the current target region, wherein the digital image processing includes cropping and scaling; The method further includes: In response to a first instruction, a first preset mode is activated, wherein the first preset mode is configured to, upon detecting the voice signal, acquire the positions of the face object and the sound source in the video image, and determine the target region in the video image based on the intersection position of the face object and the sound source; and / or, In response to the second instruction, a second preset mode is activated, wherein the second preset mode is configured to obtain the position of the face object in the video image, and determine the target region in the video image based on the position of the face object; the first instruction and the second instruction are user instructions; When both the first preset mode and the second preset mode are activated, the method further includes: If the voice signal is not detected within a preset time period or after a preset number of detections, the display mode of the video image will be switched from the first preset mode to the second preset mode. The step of highlighting the corresponding face object based on the target region includes: When the first preset mode is activated, the speaker's face is highlighted in the target area; When the second preset mode is activated, the faces of all participants are highlighted in the target area.
2. The object display method according to claim 1, characterized in that, The recognition result of the face object includes a face detection bounding box, and obtaining the position of the face object in the video image includes: The vertex coordinates and size of the face detection box are adjusted based on a preset image resolution, wherein the aspect ratio of the preset image resolution is 1:
1.
3. The object display method according to claim 1, characterized in that, Upon detecting a face object and a voice signal, the method further includes obtaining the positions of the face object and the sound source in a video image, and determining a target region in the video image based on the intersection position of the face object and the sound source. With the camera lens in its initial accelerator state, determine the relative positional relationship between the target area and the camera's field of view; If the target area is not completely contained within the camera's field of view, rotate the camera lens horizontally until the target area is completely contained within the camera's field of view.
4. The object display method according to claim 1, characterized in that, Digital image processing of the image of the current target region includes: If the scaling factor of the current target region is less than the scaling factor of the historical target region, the image of the current target region is first reduced in size and then shifted; or, If the scaling factor of the current target area is greater than that of the historical target area, the image of the current target area is first translated and then magnified.
5. The object display method according to any one of claims 1 to 4, characterized in that, The corresponding facial objects highlighted based on the target region include: The target area includes a geometric selection box, which is used to select the corresponding face object; or, the target area includes a geometric shape, which is used to mark the corresponding face object; or, the target area is centered in the video image, and the corresponding face object is displayed in the target area.
6. A data processing apparatus, characterized in that, include: The system includes a face recognition module, a voice tracking module, and a main control module, with the face recognition module and the voice tracking module respectively connected to the main control module; the data processing device is connected to a single camera, which is used to capture images of a preset area; The face recognition module is configured to perform face recognition detection in a preset area, and the voice tracking module is configured to perform voice tracking detection in the preset area; The main control module is configured to, when a face object and a voice signal are detected, acquire the positions of the face object and the sound source in the video image, and determine a target region in the video image based on the intersection position of the face object and the sound source; when a face object is detected but no voice signal is detected, acquire the position of the face object in the video image, and determine the target region in the video image based on the position of the face object; and highlight the corresponding face object based on the target region. The main control module is configured to compare the current target area with the historical target area determined in the previous stage, and determine whether the deviation between the current target area and the historical target area is greater than a preset threshold. If it is determined that the deviation between the current target region and the historical target region is greater than a preset threshold, digital image processing is performed on the image of the current target region, wherein the digital image processing includes cropping and scaling; The main control module is configured to, in response to a first instruction, activate a first preset mode, wherein the first preset mode is configured to, upon detecting the voice signal, acquire the positions of the face object and the sound source in the video image, and determine the target region in the video image based on the intersection position of the face object and the sound source; and / or, in response to a second instruction, activate a second preset mode, wherein the second preset mode is configured to, acquire the position of the face object in the video image, and determine the target region in the video image based on the position of the face object; the first instruction and the second instruction are user instructions; The main control module is configured to switch the display mode of the video image from the first preset mode to the second preset mode when both the first preset mode and the second preset mode are activated, and if the voice signal is not detected within a preset time or after a preset number of detections. The main control module is configured to highlight the speaker's face in the target area when the first preset mode is activated; and to highlight the faces of all participants in the target area when the second preset mode is activated.
7. An object display system, characterized in that, include: The system includes a camera, a microphone, a playback device, and the data processing apparatus as described in claim 6, wherein the camera, the microphone, and the playback device are respectively connected to the data processing apparatus; the camera is used to capture video of a preset area. The microphone is used to collect voice signals from the preset area; the playback device is used to output video images and voice signals processed by the data processing device.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the object display method according to any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the object display method according to any one of claims 1 to 5.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the object display method according to any one of claims 1 to 5.