Video stream display method, device, computer equipment and storage medium

By obtaining the location of multiple video streams and sound source, determining the target video stream and identifying the spokesperson, the problem that the picture cannot be focused on the spokesperson in a multi-person conference scene is solved, and a better conference screen display effect is achieved.

CN114422743BActive Publication Date: 2025-05-06HUIZHOU VISION NEW TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111583153.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-22
Publication Date
2025-05-06
Estimated Expiration
2041-12-22

AI Technical Summary

Technical Problem

In scenes where multiple people participate, the pictures captured by the camera cannot effectively highlight the focus in the current scene, especially in multi-person meeting scenes. How to focus the displayed pictures on the spokesperson and present a better meeting picture is a problem that needs to be solved urgently.

Method used

By obtaining the multi-channel video stream and sound source location of the current scene, each video stream corresponds to an image acquisition area. According to the sound source location and image acquisition area, the target video stream is determined from the multi-channel video stream, the target object (object with lipstick action) is identified, and the video stream to be displayed is determined based on the recognition results, and the picture corresponding to the video stream to be displayed is finally displayed.

Benefits of technology

The target video stream used to identify the speaker is determined by determining the location of the sound source, and the efficiency of identifying the speaker is improved. The video stream to be displayed is determined based on the recognition results, so that the displayed picture is focused on the speaker, and a better meeting picture is presented.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114422743B_ABST
    Figure CN114422743B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose a video stream display method, apparatus, computer equipment and storage medium; the embodiments of the present application obtain multiple video streams and sound source positions of the current scene, and an image acquisition area corresponding to each of the video streams; according to the sound source positions and the image acquisition area, a target video stream is determined from the multiple video streams; according to the target video stream, a target object is identified, and the target object is an object with lip movements; according to the identification result of the target object, a video stream to be displayed is determined from the multiple video streams; and the picture corresponding to the video stream to be displayed is displayed. In the embodiments of the present application, the target video stream used to identify the speaker is determined by the sound source position, so that the efficiency of identifying the speaker can be improved. At the same time, the video stream to be displayed is determined according to the identification result, so that the displayed picture can be focused on the speaker, presenting a better conference picture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a video stream display method, apparatus, computer equipment and storage medium. Background Art

[0002] With the development of video technology, more and more occasions use cameras to collect and play live images in real time. However, in scenes with multiple participants, the images captured by the camera usually cannot highlight the key points of the current scene.

[0003] Especially in the scenario of multi-person meetings, there are often multiple different speakers during the meeting. How to focus the displayed image on the speakers and present a better meeting picture is a problem that needs to be solved urgently. Summary of the invention

[0004] The embodiments of the present application provide a video stream display method, apparatus, computer equipment, and storage medium, which can focus the displayed image on the speaker and present a better conference image.

[0005] An embodiment of the present application provides a video stream display method, including: obtaining multiple video streams and a sound source position of a current scene, and an image acquisition area corresponding to each of the video streams; determining a target video stream from the multiple video streams according to the sound source position and the image acquisition area; identifying a target object according to the target video stream, wherein the target object is an object with lip movements; determining a video stream to be displayed from the multiple video streams according to the recognition result of the target object; and displaying a picture corresponding to the video stream to be displayed.

[0006] An embodiment of the present application also provides a video stream display device, including: an acquisition unit, used to acquire multiple video streams and a sound source position of a current scene, and an image acquisition area corresponding to each of the video streams; a first determination unit, used to determine a target video stream from the multiple video streams based on the sound source position and the image acquisition area; an identification unit, used to identify a target object based on the target video stream, wherein the target object is an object with lip movements; a second determination unit, used to determine a video stream to be displayed from the multiple video streams based on the identification result of the target object; and a display unit, used to display the screen corresponding to the video stream to be displayed.

[0007] An embodiment of the present application also provides a computer device, including a memory storing multiple instructions; the processor loads instructions from the memory to execute the steps in any one of the video stream display methods provided in the embodiments of the present application.

[0008] An embodiment of the present application further provides a computer-readable storage medium, which stores a plurality of instructions, and the instructions are suitable for a processor to load to execute the steps in any one of the video stream display methods provided in the embodiments of the present application.

[0009] The embodiment of the present application can obtain multiple video streams and sound source positions of the current scene, and an image acquisition area corresponding to each of the video streams; determine a target video stream from the multiple video streams according to the sound source position and the image acquisition area; identify a target object according to the target video stream, wherein the target object is an object with lip movements; determine a video stream to be displayed from the multiple video streams according to the recognition result of the target object; and display the picture corresponding to the video stream to be displayed. In the present application, the target video stream used to identify the speaker is determined by the sound source position, so that the efficiency of identifying the speaker can be improved. At the same time, the video stream to be displayed is determined according to the recognition result, so that the displayed picture can be focused on the speaker, presenting a better conference picture. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0011] Figure 1 is a schematic diagram of a scene of a video stream display system provided in an embodiment of the present application;

[0012] Figure 2 It is a flowchart of a video stream display method provided in an embodiment of the present application;

[0013] Figure 3 is a structural diagram of a video stream display system provided in an embodiment of the present application;

[0014] Figure 4 It is a flowchart of a data processing module provided in an embodiment of the present application;

[0015] Figure 5 is a flowchart of a video stream display method provided by another embodiment of the present application;

[0016] Figure 6 is a structural schematic diagram of a video stream display device provided in an embodiment of the present application;

[0017] Figure 7 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0018] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0019] Embodiments of the present application provide a video stream display method, apparatus, computer equipment, and storage medium.

[0020] The video stream display device can be integrated into an electronic device, which can be a terminal, a server, etc. The terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop, or a personal computer (PC); the server can be a single server or a server cluster composed of multiple servers.

[0021] In some embodiments, the video stream display device may also be integrated into multiple electronic devices. For example, the video stream display device may be integrated into multiple servers, and the video stream display method of the present application may be implemented by multiple servers.

[0022] In some embodiments, the server may also be implemented in the form of a terminal.

[0023] For example, refer to Figure 1 In some implementations, a scene schematic diagram of a video stream display system is provided. The image rendering system may include a display data acquisition module 1000, a server 2000, and a terminal 3000.

[0024] Among them, the data acquisition module can obtain multiple video streams and sound source locations of the current scene, and an image acquisition area corresponding to each video stream.

[0025] Among them, the server can determine the target video stream from multiple video streams based on the sound source location and the image acquisition area; identify the target object based on the target video stream, and the target object is an object with lip movements; based on the recognition result of the target object, determine the video stream to be displayed from multiple video streams.

[0026] Among them, the terminal can display the picture corresponding to the video stream to be displayed.

[0027] It should be noted that the serial numbers of the following embodiments are not intended to limit the preferred order of the embodiments.

[0028] In this embodiment, a video stream display method is provided, such as Figure 2As shown, the specific process of the video stream display method can be as follows:

[0029] 110. Obtain multiple video streams and sound source positions of the current scene, and an image acquisition area corresponding to each video stream.

[0030] The sound source position refers to the position where the sound is emitted in the current scene, for example, the position where the speech sound is emitted in a conference scene. The sound can be collected by setting a microphone array in the current scene, and the sound source position can be calculated according to the sound source localization algorithm.

[0031] In some embodiments, the multiple video streams include a panoramic video stream and at least one close-up video stream. The image acquisition area refers to the area range of the image that can be acquired by the image acquisition device corresponding to the video stream corresponding to the current scene. Among them, the panoramic video stream refers to a video stream containing a panoramic picture of the current scene, and its corresponding image acquisition area is the current scene panorama, which can be acquired by a camera with a wide-angle lens. The close-up video stream refers to a video stream containing a partial scene of the current scene, and its corresponding image acquisition area is the partial scene of the current scene, which can be acquired by a camera with a telephoto lens.

[0032] In some embodiments, the method for obtaining the sound source position may include steps 1.1 to 1.2, as follows:

[0033] 1.1. Collect the sound information of the current scene;

[0034] 1.2. Process the collected sound information through the sound source localization algorithm to obtain the sound source location.

[0035] Among them, the sound source localization algorithm can adopt TDOA (Time Difference of Arrival), GCC-PHAT (Generalized Cross Correlation PHAse Transformation), and the like.

[0036] 120. Determine a target video stream from multiple video streams based on a sound source location and an image acquisition area.

[0037] The target video stream is a video stream determined based on the correlation between the sound source position and the image acquisition area. The correlation may be that the sound source position is within the image acquisition area, or that the distance between the sound source position and the center of the image acquisition area is less than a preset distance.

[0038] In some embodiments, step 120 may include the steps of: determining a target image acquisition area from acquisition areas corresponding to multiple multi-channel video streams based on the sound source location and the association relationship between the image acquisition areas, and determining the video stream corresponding to the target image acquisition area as the target video stream.

[0039] In some implementations, due to interference such as reflection and noise, the sound source position determined by sound source localization may have errors. Therefore, the possible area where the sound source exists is determined by the sound source position, so as to determine the target video stream corresponding to the area, thereby increasing the accuracy of the acquired image information. Specifically, step 120 may include steps 2.1 to 2.4, as follows:

[0040] 2.1. Determine the sound source area according to the sound source location;

[0041] 2.2. For each video stream, determine the overlapping area of ​​the sound source area and the image acquisition area;

[0042] 2.3. Determine the overlapping area that meets the preset first area size as the target area;

[0043] 2.4. Determine the video stream corresponding to the target area as the target video stream.

[0044] The sound source area refers to the area where the sound source is located, and the area is located in the same plane as the image acquisition area. The sound source area can be determined based on the sound source location and a preset area parameter value, and the preset area parameter value can be set based on the current scene or experience, for example, a circular area with the sound source location as the center and a preset radius value as the radius is used as the sound source area, and so on.

[0045] In some embodiments, step 2.1 may include the steps of: obtaining a reference point; using the reference point as a vertex and the line connecting the reference point and the sound source position as an angle bisector to determine an angle that satisfies a preset first angle; determining the area corresponding to the angle in the current scene as the sound source area. Among them, the reference point can be any boundary point in the current scene. In some embodiments, the reference point can be a position point for collecting sound information of the current scene, for example, a point determined according to a microphone array for measuring the sound source position, which can be any position point on the microphone array or a midpoint. It should be noted that the reference point, the sound source position, the image acquisition area, the sound source area, and the target area are all on the same plane, which can be a horizontal plane. For example, the sound source position is a position point where the real sound source position calculated by the sound source localization algorithm is projected onto the horizontal plane.

[0046] The preset first area size is a size condition of the area set according to the current scene or experience. It can be a specific value, such as being greater than or equal to one third of the size of the image acquisition area corresponding to any video stream, or being greater than or equal to the area size determined according to the sound source area, such as being greater than or equal to one half of the size of the sound source area.

[0047] In some implementations, step 2.3 may include the step of determining an overlapping area with the same size of sound source areas as a target area.

[0048] 130. According to the target video stream, identify a target object, where the target object is an object with lip movements.

[0049] The target object refers to an object with lip movements identified from the image information of the target video stream. Generally speaking, a person with lip movements is speaking and can therefore be used as a speaker in the current scene. The lip movements can be lip movements of a person when speaking determined according to the prior art.

[0050] Since different video streams correspond to different image acquisition areas, determining the target video stream used to identify the speaker by the sound source position can reduce the amount of identification data and improve the efficiency of identifying the speaker.

[0051] In some implementations, due to interference such as reflection and noise, the sound source position determined by sound source localization may have errors. Therefore, the area used to identify the speaker is determined by the sound source position, so as to determine the target video stream corresponding to the area, thereby increasing the accuracy of the acquired image information. Step 130 may include steps 3.1 to 3.3, as follows:

[0052] 3.1. Determine the recognition area based on the location of the sound source;

[0053] 3.2. According to the recognition area, the target image information is obtained from the target video stream, and the target image information is the image information corresponding to the recognition area;

[0054] 3.3. Identify the target object based on the target image information.

[0055] The recognition area refers to the area used to identify the target object determined according to the sound source position, and the area is located in the same plane as the image acquisition area. The sound source position is located in the recognition area, and the recognition area can be determined according to the sound source position and the preset area parameter value. The preset area parameter value can be set according to the current scene or experience. For example, a circular area with the sound source position as the center and the preset radius value as the radius is used as the recognition area, etc. The recognition area can also be the sound source area.

[0056] In some embodiments, step 3.1 may include the steps of: obtaining a reference point; using the reference point as a vertex and a line connecting the reference point and the sound source position as an angle bisector to determine an angle that satisfies a preset second angle; and determining the area corresponding to the angle in the current scene as an identification area.

[0057] The target image information refers to the image information in the area where the recognition area is projected onto the screen captured by the target video stream. Specifically, the coordinate position of the recognition area can be obtained, and the coordinate position can be projected into the coordinate system where the screen captured by the target video stream is located to obtain the projected area, and the image information in the area can be used as the target image information.

[0058] The identification area that may contain the target object is determined by the sound source position, and the image information corresponding to the identification area is obtained from the target video stream through the identification area, and then it is determined whether there is a target object from the image information.

[0059] In some implementations, to improve recognition efficiency, step 3.3 may include steps 3.3.1 to 3.3.2, as follows:

[0060] 3.3.1. When an object with lip movements is identified from the target image information, the object with lip movements is regarded as the target object;

[0061] 3.3.2. When no object with lip movements is identified from the target image information, the identification area is expanded to a preset second area size to identify the target object.

[0062] For example, first set the recognition area to a sector-shaped area with an angle of 30°. When the target object is not recognized in this area, the recognition area is expanded to a sector-shaped area with an angle of 40°, and recognition is performed again. When the target object is not recognized in this area, the recognition area is expanded to a sector-shaped area with an angle of 50°, and so on, until the target object is recognized or the recognition area is expanded to the upper limit value.

[0063] Since the sound source position determined by sound source localization may have errors, the pre-set recognition area may not be able to recognize the target object when performing lip movement recognition. At this time, by gradually expanding the size of the recognition area, the recognition range can be expanded to correct the recognition result. At the same time, gradually expanding the size of the recognition area can also make the area to be recognized each time smaller than the next recognition, so as to obtain the recognition result in the smallest area possible and improve the recognition efficiency.

[0064] In some embodiments, in order to further improve the recognition efficiency, step 3.3.2 may include the following steps: when no object with lip movements is identified from the target image information, expanding the recognition area to a preset second area size to obtain an enlarged area; using the non-overlapping area between the recognition area and the enlarged area as the target recognition area; based on the target recognition area, obtaining target image information from the target video stream, the target image information being image information corresponding to the recognition area; and identifying the target object based on the target image information.

[0065] 140. Determine a video stream to be displayed from multiple video streams according to the recognition result of the target object.

[0066] The video stream to be displayed refers to a video stream used to display the current scene. The target object can be focused on and displayed through the video stream to be displayed.

[0067] In some implementations, in order to provide a better display effect of the current scene, a display strategy determined according to the target object recognition result is provided. Step 140 may include steps 4.1 to 4.4, as follows:

[0068] 4.1. When a target object is identified, the area to be displayed is determined according to the target object;

[0069] 4.2. When the target object is not recognized, the area to be displayed is determined based on all objects in the target image information;

[0070] 4.3. Obtain the image acquisition area corresponding to each video stream;

[0071] 4.4. Determine the video stream to be displayed according to the area to be displayed and the image acquisition area.

[0072] The area to be displayed refers to the area to be displayed by the video stream to be displayed. When the target object is identified, the area where the target object is located, such as the sound source area or the identification area, can be used as the area to be displayed. When the target object is not identified, the area where all objects in the target image information are located is used as the area to be displayed. The area to be displayed can be in the same plane as the image acquisition area, or in the same plane as the image corresponding to the target video stream. When comparing the area to be displayed with different plane areas such as the image acquisition area, the area to be displayed can be projected onto the plane where the image acquisition area is located before comparison.

[0073] In some implementations, when multiple target objects are identified, the area to be displayed is determined according to the multiple target objects. In this case, the area to be displayed is the area where the multiple target objects are located.

[0074] The area to be displayed is determined by the recognition result of the target object, and the area to be displayed is compared with the image acquisition area to determine the video stream to be displayed. For example, the overlapping area between the area to be displayed and each image acquisition area can be determined, and the video stream corresponding to the image acquisition area with the largest overlapping area can be used as the video stream to be displayed.

[0075] In some embodiments, in order to focus on the speaker and provide a better display effect of the current scene, step 4.4 may include the steps of: determining the area size ratio of the area to be displayed to each image acquisition area, and taking the video stream corresponding to the image acquisition area with the highest ratio of the area size to be displayed / the image acquisition area size as the video stream to be displayed. In some embodiments, in order to avoid an incomplete display of the speaker's picture, the ratio of the area size to be displayed / the image acquisition area size is less than a preset value, which may be 1.

[0076] 150. Display a picture corresponding to the video stream to be displayed.

[0077] In some implementations, by cropping the display image and focusing on the speaker to provide a better display effect of the current scene, step 150 may include steps 5.1 to 5.3, as follows:

[0078] 5.1. Obtain the display screen of the video stream to be displayed;

[0079] 5.2. According to the area to be displayed, crop the display screen of the video stream to be displayed to obtain a cropped display screen;

[0080] 5.3. Display the cropped display screen.

[0081] The cropped display screen is a screen corresponding to the to-be-displayed area in the display screen of the to-be-displayed video stream.

[0082] By cropping the display screen of the video stream to be displayed into the screen corresponding to the area to be displayed, the speaker can be further focused to provide a better display effect of the current scene.

[0083] The video stream display method provided by the embodiment of the present application can be applied to various scenarios involving multiple people. For example, taking a multi-person conference as an example, obtain multiple video streams and the sound source position of the current scene, and an image acquisition area corresponding to each video stream; determine the target video stream from the multiple video streams according to the sound source position and the image acquisition area; identify the target object according to the target video stream, and the target object is an object with lip movements; determine the video stream to be displayed from the multiple video streams according to the recognition result of the target object; and display the picture corresponding to the video stream to be displayed. The solution provided by the embodiment of the present application is used to determine the target video stream used to identify the speaker through the sound source position, so as to improve the efficiency of identifying the speaker. At the same time, the video stream to be displayed is determined according to the recognition result, so that the displayed picture can be focused on the speaker, presenting a better meeting picture.

[0084] The method described in the above embodiment will be further described in detail below.

[0085] In this embodiment, a multi-person conference scenario is taken as an example to describe the method of the embodiment of the present application in detail.

[0086] like Figure 3 As shown, a structural schematic diagram of a video stream display system is provided, which includes a data acquisition module, a data processing module and a terminal.

[0087] The data acquisition module consists of an infrared thermal imager, an ultrasonic module, a dual camera module, and an array microphone. The camera module collects information and sends it to the data processing module. The details are as follows:

[0088] The data acquisition module includes two cameras, which are a wide-angle lens and a telephoto lens. The wide-angle lens has a large field of view and a wide visible range, but the distant view is blurred. The telephoto lens has a small field of view and a narrow visible range, but the distant view is clear. When the field of view overlaps, the camera of the telephoto lens is switched to the camera of the wide-angle lens. When the field of view is outside the range of the telephoto lens, it is switched to the camera of the wide-angle lens. The dual-camera switching method includes: 1. The dual-camera module includes two types of cameras, a wide-angle lens camera and a telephoto lens camera. The wide-angle lens camera has a short focal length and a wide field of view, and the captured images are more, and the objects in the images account for a smaller proportion. On the contrary, the telephoto lens camera has a long focal length and a narrow field of view, and the captured images are fewer, and the objects in the images account for a larger proportion. The wide-angle lens camera and the telephoto lens camera can each produce two video streams, one video stream is used for actual image presentation, which can be called a preview stream, and the other video stream is used to provide AI for lip movement detection and face recognition, which can be called an AI image stream. 2. The terminal screen can only be presented from one of the preview streams of the two cameras, but the AI ​​image streams of the two cameras can be provided to the image AI thread for lip movement detection and face recognition. 3. The image AI thread decides to perform lip movement recognition and face recognition on one of the two AI image streams based on the angle information of the sound source positioning, and then outputs it to the UVC thread to decide which preview stream to switch to and crop, and finally presents the face focus effect.

[0089] Infrared thermal imager, used to measure the temperature of target objects.

[0090] The ultrasonic module is used to detect the distance of the target object in combination with the infrared thermal imager. Since the infrared thermal imager is essentially a camera, there are requirements for the minimum imaging distance of its lens. For example, the distance between the object to be measured and the lens must be greater than 25cm to ensure a clear thermal image. Therefore, the ultrasonic module can be used to detect the target distance and indicate the distance requirement of the target.

[0091] The matrix microphone module is used to locate the sound source and determine the speaker's location.

[0092] The data processing module includes a UVC thread, a UAC thread, an image AI thread, and an audio AI thread. The data processing module obtains information collected by the camera module and performs data processing. Figure 4 As shown, the workflow of the threads in the data processing module is as follows:

[0093] The UVC thread is used to collect video stream information from dual cameras. Each camera will output two video streams, one for output to the terminal to present real-time images, and the other for the image AI thread to analyze lip movement analysis and face recognition.

[0094] The UAC thread is used to collect audio stream information from the array microphone. There are two types of audio information output. One type of audio information is to directly output the PCM format audio stream data of one microphone to the terminal for audio playback. The other type is to combine the PCM format audio stream data collected by all microphones and give it to the audio AI thread for sound source positioning.

[0095] The image AI thread is used to analyze and process the image information of the two cameras output by the UVC thread, and output a decision to the UVC thread. The decision includes feedback on which camera's video stream to display, and zooming in and cropping the image information of the video stream to focus on the speaker. Specifically, the image AI thread obtains two types of information, one is the video stream information of the two cameras provided by the UVC thread, and the other is the sound source angle information provided by the audio AI thread. After the image AI thread obtains the sound source angle information, it determines the sound source angle of the current speaker, and determines which camera's video stream to obtain to analyze the lip movements based on the field of view angle range of the two cameras, determines the recognition area corresponding to the lip movements, and recognizes the face information. Finally, the UVC thread is fed back to switch the camera for display, and zoom in and crop to focus on the speaker.

[0096] The audio AI thread is used to analyze and process the audio stream data in PCM format output by the array MIC given by the UAC, locate the sound source, and output the sound source angle information to the image AI thread for decision-making.

[0097] Among them, the data processing module also includes a policy management module, which is used to obtain the data processed by the data processing module, make scenario decisions, and realize speaker tracking, speech subtitle display and participant sign-in.

[0098] The terminal is used to display images, and the terminal may be a TV (television).

[0099] like Figure 5 As shown, a specific process of a video stream display method is as follows:

[0100] 210. The array microphone collects environmental sounds in real time.

[0101] 220. The audio AI thread uses the sound source localization algorithm to determine and output the sound source angle information based on the collected ambient sound.

[0102] Before collecting ambient sound through the display microphone, the following steps may be included: the policy management module controls the infrared thermal imager and the ultrasonic module to detect the body temperature of the participants. The ultrasonic module starts the distance detection function. When the distance of the target participant reaches the imaging requirement of the infrared thermal imager, the infrared thermal imager starts to detect the body temperature of the target participant. When the temperature exceeds the requirement, the participant is not allowed to attend the meeting.

[0103] The sound source angle refers to the angle between the sound source position and the array microphone. The midpoint of the line segment formed by the array microphone can be taken as the vertex, and the angle formed by the sound source position, the midpoint of the line segment formed by the array microphone and any vertex of the line segment formed by the array microphone can be taken as the sound source angle.

[0104] The array microphone collects ambient sound in real time and sends it to the UAC thread. After being processed by the UAC thread, it is sent to the terminal for playback and to the image AI thread for sound source positioning.

[0105] 231. When the audio AI thread does not output the sound source angle information, the image AI thread controls the terminal to display the image captured by the wide-angle camera.

[0106] When there is no sound source angle output, it will enter the listening mode. In this mode, the UVC thread will output the image presentation of the wide-angle camera by default. When the image AI thread analyzes the face recognition of one of the AI ​​image streams of the two cameras, it will inform the UVC thread to switch to the image presentation of the corresponding camera. At the same time, step 210 is executed to collect environmental sounds in real time. If both AI image streams have face recognition, the image of the wide-angle camera is output first. If neither of the two AI image streams has face recognition, it will not focus, and the image of the wide-angle camera is output first.

[0107] 232. When the audio AI thread outputs the sound source angle information, the image AI thread determines the sound source area based on the sound source angle information.

[0108] When there is a sound source angle output, the image AI thread will divide the sound source angle into sectors within the range of ±15° to ±30°, which is the sound source area.

[0109] 240. The image AI thread determines the target video stream from the two video streams according to the sound source area.

[0110] The image AI thread determines which camera to collect the image for recognition based on the sound source area. If the sound source area is completely covered by both cameras, the image AI thread will give priority to processing the image information of the telephoto camera. If the sound source area is within the wide-angle camera, the image AI thread will process the image information of the wide-angle camera.

[0111] 250. The image AI thread identifies a target object according to a target video stream, where the target object is an object with lip movements.

[0112] The image AI thread performs lip movement analysis based on the camera used for identification determined in step 340 and the image information captured by the camera to identify the target object.

[0113] 261. When a target object is identified, the image AI thread determines the area to be displayed based on the target object.

[0114] At the beginning, the area corresponding to ±15° of the sound source angle is used as the sector for face recognition. If no person with lip movement is recognized, the area corresponding to ±20° of the sound source angle is used as the sector for face recognition. The sector size is increased by 5° each time until a person with lip movement is recognized, and the sector at this time is used as the area to be displayed. If there are multiple speakers, the area to be displayed should cover all speakers.

[0115] Finally, the UVC thread controls the image information of the output camera and crops the output image information so that users can see the final face focusing effect.

[0116] 262. When the target object is not identified, the image AI thread determines the area to be displayed based on all objects in the target image information.

[0117] 270. The image AI thread determines the video stream to be displayed according to the area to be displayed.

[0118] 280. The image AI thread crops the display screen of the video stream to be displayed according to the area to be displayed to obtain a cropped display screen.

[0119] 290. The terminal displays the cropped display image.

[0120] When no person with lip movement is recognized, the area corresponding to ±15° of the sound source angle is used as the sector for face recognition. If a face is recognized, the area corresponding to ±20° of the sound source angle is used as the sector for face recognition. The sector size is increased by 5° each time until a face is recognized, and the sector at this time is used as the area to be displayed. If there are multiple people in the area to be displayed, the area to be displayed must cover all of them.

[0121] Finally, the UVC thread controls the image information of the output camera and crops the output image information so that users can see the final face focusing effect.

[0122] When no face is recognized, the listening mode is entered and step 210 is executed to collect environmental sounds in real time.

[0123] As can be seen from the above, the embodiment of the present application obtains the sound source angle and switches between dual cameras to achieve focusing on the speaker, so that the displayed image can be focused on the speaker and present a better conference picture.

[0124] In order to better implement the above method, the embodiment of the present application also provides a video stream display device, which can be integrated in an electronic device, and the electronic device can be a terminal, a server, etc. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop, a personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers.

[0125] For example, in this embodiment, the method of the embodiment of the present application is described in detail by taking the video stream display device being specifically integrated in a terminal as an example.

[0126] For example, Figure 6 As shown, the video stream display device may include an acquisition unit 310, a first determination unit 320, an identification unit 330, a second determination unit 340, and a display unit 350, as follows:

[0127] (I) Acquisition Unit 310

[0128] Used to obtain multiple video streams and sound source locations of the current scene, and an image acquisition area corresponding to each video stream.

[0129] In some implementations, the method for obtaining the sound source position may include steps 6.1 to 6.2, as follows:

[0130] 6.1. Collect the sound information of the current scene;

[0131] 6.2. Process the collected sound information through the sound source localization algorithm to obtain the sound source location.

[0132] (II) First Determination Unit 320

[0133] Used to determine the target video stream from multiple video streams based on the sound source location and image acquisition area.

[0134] In some implementations, the first determining unit 320 may be specifically used in steps 7.1 to 7.4 as follows:

[0135] 7.1. Determine the sound source area according to the sound source location;

[0136] 7.2. For each video stream, determine the overlapping area of ​​the sound source area and the image acquisition area;

[0137] 7.3. Determine the overlapping area that meets the preset first area size as the target area;

[0138] 7.4. Determine the video stream corresponding to the target area as the target video stream.

[0139] (III) Identification Unit 330

[0140] It is used to identify the target object according to the target video stream, where the target object is an object with lip movements.

[0141] In some implementations, the identification unit 330 may have a method for including steps 8.1 to 8.3 as follows:

[0142] 8.1. Determine the recognition area based on the location of the sound source;

[0143] 8.2. According to the recognition area, the target image information is obtained from the target video stream, where the target image information is the image information corresponding to the recognition area;

[0144] 8.3. Identify the target object based on the target image information.

[0145] In some embodiments, step 8.3 may include steps 8.3.1 to 8.3.2, as follows:

[0146] 8.3.1. When an object with lip movements is identified from the target image information, the object with lip movements is regarded as the target object;

[0147] 8.3.2. When no object with lip movements is identified from the target image information, the identification area is expanded to a preset second area size to identify the target object.

[0148] (IV) Second Determination Unit 340

[0149] It is used to determine the video stream to be displayed from multiple video streams according to the recognition result of the target object.

[0150] In some implementations, the second determining unit 340 may be specifically used in steps 9.1 to 9.4 as follows:

[0151] 9.1. When a target object is identified, the area to be displayed is determined according to the target object;

[0152] 9.2. When the target object is not identified, the area to be displayed is determined based on all objects in the target image information;

[0153] 9.3. Obtain the image acquisition area corresponding to each video stream;

[0154] 9.4. Determine the video stream to be displayed based on the area to be displayed and the image acquisition area.

[0155] (V) Display unit 350

[0156] Used to display the screen corresponding to the video stream to be displayed.

[0157] In some implementations, the display unit 350 may be specifically used in steps 10.1 to 10.3 as follows:

[0158] 10.1. Obtain the display screen of the video stream to be displayed;

[0159] 10.2. Crop the display screen of the video stream to be displayed according to the area to be displayed to obtain a cropped display screen;

[0160] 10.3. Display the cropped display screen.

[0161] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can refer to the previous method embodiments, which will not be repeated here.

[0162] Therefore, the embodiment of the present application can determine the target video stream used to identify the speaker by the sound source position, which can improve the efficiency of identifying the speaker. At the same time, the video stream to be displayed is determined according to the recognition result, so that the displayed picture can be focused on the speaker and present a better conference picture.

[0163] Correspondingly, an embodiment of the present application also provides a computer device, which may be a terminal or a server, and the terminal may be a smart phone, a tablet computer, a laptop computer, a touch screen, a game console, a personal computer, a personal digital assistant (PDA), or other terminal devices.

[0164] like Figure 7 As shown, Figure 7 The schematic diagram of the structure of the computer device provided in the embodiment of the present application is that the computer device 400 includes a processor 410 having one or more processing cores, a memory 420 having one or more computer-readable storage media, and a computer program stored in the memory 420 and executable on the processor. The processor 410 is electrically connected to the memory 420. It can be understood by those skilled in the art that the computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0165] The processor 410 is the control center of the computer device 400. It uses various interfaces and lines to connect various parts of the entire computer device 400, executes various functions of the computer device 400 and processes data by running or loading software programs and / or modules stored in the memory 420, and calling data stored in the memory 420, thereby monitoring the computer device 400 as a whole.

[0166] In the embodiment of the present application, the processor 410 in the computer device 400 will load instructions corresponding to the processes of one or more application programs into the memory 420 according to the following steps, and the processor 410 will run the application programs stored in the memory 420 to implement various functions:

[0167] Obtain multiple video streams and sound source positions of the current scene, and an image acquisition area corresponding to each video stream; determine the target video stream from the multiple video streams based on the sound source position and the image acquisition area; identify the target object based on the target video stream, where the target object is an object with lip movements; determine the video stream to be displayed from the multiple video streams based on the recognition result of the target object; and display the picture corresponding to the video stream to be displayed.

[0168] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0169] Optional, such as Figure 7 As shown, the computer device 400 further includes: a touch screen 430, a radio frequency circuit 440, an audio circuit 450, an input unit 460, and a power supply 470. The processor 410 is electrically connected to the touch screen 430, the radio frequency circuit 440, the audio circuit 450, the input unit 460, and the power supply 470, respectively. Those skilled in the art will appreciate that Figure 7 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0170] The touch display screen 430 can be used to display a graphical user interface and receive operation instructions generated by the user acting on the graphical user interface. The touch display screen 430 may include a display panel and a touch panel. Among them, the display panel may be used to display information input by the user or information provided to the user and various graphical user interfaces of the computer device, which may be composed of graphics, text, icons, videos and any combination thereof. Optionally, the display panel may be configured in the form of a liquid crystal display (LCD, Liquid Crystal Display), an organic light-emitting diode (OLED, Organic Light-Emitting Diode) and the like. The touch panel may be used to collect the user's touch operation on or near it (such as the user using any suitable object or accessory such as a finger, stylus, etc. on the touch panel or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute corresponding programs. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the touch point coordinates, and then sends it to the processor 410, and can receive the command sent by the processor 410 and execute it. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 410 to determine the type of touch event, and then the processor 410 provides a corresponding visual output on the display panel according to the type of touch event. In an embodiment of the present application, the touch panel and the display panel can be integrated into the touch display screen 430 to realize the input and output functions. However, in some embodiments, the touch panel and the display panel can be used as two independent components to realize the input and output functions. That is, the touch display screen 430 can also be used as a part of the input unit 460 to realize the input function.

[0171] The radio frequency circuit 440 may be used to send and receive radio frequency signals, so as to establish wireless communication with a network device or other computer devices through wireless communication, and to send and receive signals between the network device or other computer devices.

[0172] The audio circuit 450 can be used to provide an audio interface between the user and the computer device through a speaker and a microphone. The audio circuit 450 can transmit the electrical signal converted from the received audio data to the speaker, which is converted into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 450 and converted into audio data, and then the audio data is output to the processor 410 for processing, and then sent to another computer device through the radio frequency circuit 440, or the audio data is output to the memory 420 for further processing. The audio circuit 450 may also include an earphone jack to provide communication between an external headset and the computer device.

[0173] The input unit 460 may be used to receive input numbers, character information or user feature information (such as fingerprint, iris, facial information, etc.), and to generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0174] The power supply 470 is used to supply power to various components of the computer device 400. Optionally, the power supply 470 can be logically connected to the processor 410 through a power management system, so that the power management system can manage charging, discharging, and power consumption. The power supply 470 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0175] although Figure 7 Not shown, the computer device 400 may also include a camera, a sensor, a wireless fidelity module, a Bluetooth module, etc., which will not be described in detail here.

[0176] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0177] From the above, it can be seen that the computer device provided in this embodiment can determine the target video stream used to identify the speaker through the sound source position, which can improve the efficiency of identifying the speaker. At the same time, the video stream to be displayed is determined according to the recognition result, so that the displayed picture can be focused on the speaker, presenting a better conference picture.

[0178] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0179] To this end, an embodiment of the present application provides a computer-readable storage medium, in which multiple computer programs are stored, and the computer program can be loaded by a processor to execute the steps in any video stream display method provided in the embodiment of the present application. For example, the computer program can execute the following steps:

[0180] Obtain multiple video streams and sound source positions of the current scene, and an image acquisition area corresponding to each video stream; determine the target video stream from the multiple video streams based on the sound source position and the image acquisition area; identify the target object based on the target video stream, where the target object is an object with lip movements; determine the video stream to be displayed from the multiple video streams based on the recognition result of the target object; and display the picture corresponding to the video stream to be displayed.

[0181] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0182] The storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0183] Since the computer program stored in the storage medium can execute the steps in any one of the video stream display methods provided in the embodiments of the present application, the beneficial effects that can be achieved by any one of the video stream display methods provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0184] The above is a detailed introduction to a video stream display method, device, storage medium and computer equipment provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A video stream display method, characterized in that: include: Obtain multiple video streams and sound source locations of the current scene, and an image acquisition area corresponding to each of the video streams; Determining a target video stream from the multiple video streams according to the sound source position and the image acquisition area; According to the target video stream, a target object is identified, wherein the target object is an object with lip movements; Determining a video stream to be displayed from the multiple video streams according to a recognition result of the target object; The step of determining a video stream to be displayed from the multiple video streams according to the recognition result of the target object comprises: When the target object is identified, determining a to-be-displayed area according to the target object; When the target object is not identified, determining the area to be displayed according to all objects in the target image information; the target image information is image information corresponding to the identification area determined according to the sound source position; Obtaining an image acquisition area corresponding to each of the video streams; Determining a video stream to be displayed according to the area to be displayed and the image acquisition area; The step of determining the video stream to be displayed according to the area to be displayed and the image acquisition area includes: Determine the overlapping area between the area to be displayed and each image acquisition area, and use the video stream corresponding to the image acquisition area with the largest overlapping area as the video stream to be displayed; Display the picture corresponding to the video stream to be displayed.

2. The video stream display method according to claim 1, characterized in that: Determining a target video stream from the multiple video streams according to the sound source position and the image acquisition area includes: Determine the sound source area according to the sound source position; For each of the video streams, determining an overlapping area between the sound source area and the image acquisition area; Determine the overlapping area that meets the preset first area size as the target area; The video stream corresponding to the target area is determined as the target video stream.

3. The video stream display method according to claim 1, characterized in that: The step of identifying a target object according to the target video stream includes: Determine the recognition area according to the sound source position; According to the identified area, target image information is acquired from the target video stream, where the target image information is image information corresponding to the identified area; The target object is identified according to the target image information.

4. The video stream display method according to claim 3, characterized in that: The step of identifying the target object according to the target image information includes: When an object with lip movements is identified from the target image information, the object with lip movements is used as a target object; When no object with lip movements is identified from the target image information, the identification area is expanded to a preset second area size to identify the target object.

5. The video stream display method according to claim 1, characterized in that: The displaying the picture corresponding to the video stream to be displayed includes: Acquire a display screen of the video stream to be displayed; According to the to-be-displayed area, cropping the display picture of the to-be-displayed video stream to obtain a cropped display picture; Shows the cropped display.

6. The video stream display method according to claim 1, characterized in that: The method for obtaining the sound source position comprises: Collect sound information of the current scene; The collected sound information is processed by the sound source localization algorithm to obtain the sound source position.

7. A video stream display device, characterized in that: include: An acquisition unit, used to acquire multiple video streams and sound source positions of the current scene, and an image acquisition area corresponding to each of the video streams; A first determining unit, configured to determine a target video stream from the multiple video streams according to the sound source position and the image acquisition area; an identification unit, configured to identify a target object according to the target video stream, wherein the target object is an object with lip movements; A second determining unit, configured to determine a video stream to be displayed from the multiple video streams according to a recognition result of the target object; The step of determining a video stream to be displayed from the multiple video streams according to the recognition result of the target object comprises: When the target object is identified, determining a to-be-displayed area according to the target object; When the target object is not identified, determining the area to be displayed according to all objects in the target image information; the target image information is image information corresponding to the identification area determined according to the sound source position; Obtaining an image acquisition area corresponding to each of the video streams; Determining a video stream to be displayed according to the area to be displayed and the image acquisition area; The step of determining the video stream to be displayed according to the area to be displayed and the image acquisition area includes: Determine the overlapping area between the area to be displayed and each image acquisition area, and use the video stream corresponding to the image acquisition area with the largest overlapping area as the video stream to be displayed; The display unit is used to display the picture corresponding to the video stream to be displayed.

8. A computer device, characterized in that: It comprises a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps in the video stream display method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the video stream display method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Sound and image linkage-based sound source orientation system and method

    CN111551921A

  • Video conferencing

    US20080218582A1

  • Automatic Switching Between Different Cameras at a Video Conference Endpoint Based on Audio

    US20160057385A1