Video processing method and apparatus, device, and medium
By identifying and correcting the relative position parameters of human figures in online conference video streams, the problem of failing to recreate the real scene after stitching together close-up images was solved, thus improving the immersive experience for remote users.
Patent Information
- Application Number
- PCT/CN2025/106012
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-13
- Filing Date
- 2025-06-30
- Publication Date
- 2026-02-19
AI Technical Summary
In online meetings, existing technologies cannot effectively reproduce the real-life situation after stitching together multiple close-up shots, resulting in a poor viewing immersion for remote users.
By acquiring the video stream to be processed, identifying the human image area and confirming the relative position parameters of the real person, and using the video stream captured by a monocular or multi-view camera, the relative position relationship of the human image area in the video image is corrected to ensure that the close-up image is closer to the relative position of the real person on site when displayed in a small window. Grouping and layout display methods are adopted to improve the immersive experience.
It enables a better reproduction of the relative positions of real people on site when viewed remotely, thus enhancing the immersive experience for remote users.
Smart Images

Figure CN2025106012_19022026_PF_FP_ABST
Abstract
Description
Video processing method, apparatus, device and medium
[0001] This application claims priority from the Chinese patent application No. CN2024111095672, filed on August 13, 2024, and entitled "Video processing method, apparatus, device and medium", the whole content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] Embodiments of the present application relate to the field of video technology, in particular to a video processing method, apparatus, device and medium. BACKGROUND
[0003] In an online meeting scenario, a global image is usually obtained by an image acquisition module (such as a camera hung on a wall or a camera integrated in a meeting device) for image acquisition of a meeting scene. However, during the online meeting, instead of directly displaying the global image, a specific layout combination of multiple close-up pictures of people can be displayed. The overall implementation process of this display mode is to extract multiple region images from the global image to obtain separate close-up pictures of each person, and then determine the relative layout relationship of the close-up pictures according to the relative position relationship of the people in the global image, so that the picture obtained by splicing the close-up pictures in the relative layout relationship is also close to the relative position relationship in the global picture. This way allows the remote participants to see the close-up pictures of multiple people in a micro way while roughly understanding the positions of the people in the picture from a macro perspective, and can more intuitively obtain the live picture of the meeting scene, such as intuitively confirming which participant the current speaker is facing, and obtaining a better meeting immersion.
[0004] The inventors analyzed the presentation effect of the existing extraction and display of close-up pictures in a specific layout combination, and found that although the close-up pictures are displayed in a layout according to the relative position relationship of the people in the global image, the close-up pictures are not spliced into the global image or part of the global image when displayed in a layout. In addition, due to the misplacement of real people, the angle and depth of field of the image acquisition module, and other reasons, the spliced close-up pictures cannot restore the real live state from the intuitive picture content and macro orientation feeling, and the viewing immersion effect of the remote user is poor. SUMMARY
[0005] The present application provides a video processing method, apparatus, device and medium to solve the technical problem that the existing extraction and display of close-up pictures in a combination can cause the spliced close-up pictures to not restore the real live state from the intuitive picture content and macro orientation feeling.
[0006] In a first aspect, the embodiments of the present application provide a video processing method, which comprises:
[0007] obtaining a to-be-processed video stream, the to-be-processed video stream being obtained by an image acquisition module from image acquisition of a multi-person scene;
[0008] confirming relative position parameters of the corresponding real persons according to the person regions in the video picture of the to-be-processed video stream;
[0009] confirming a to-be-displayed region from the person regions, confirming a display layout of each sub-picture window corresponding to the to-be-displayed region according to the relative position parameters of the to-be-displayed region, and displaying a corresponding region picture in each sub-picture window, the region picture being a partial picture including the corresponding to-be-displayed region in the video picture of the to-be-processed video stream.
[0010] According to the above, the video stream obtained by image acquisition of the multi-person scene, the relative position parameters of the corresponding real persons are first confirmed according to the person regions, and the corrected relative position relationship of the person regions in the video picture is used as the correction result, when the close-up pictures corresponding to each person in the video stream are displayed through the small windows, the corrected relative position relationship of the real persons on the scene is used as the reference for the display layout of the small windows displaying the close-up pictures, so that the relative position relationship of the real persons on the scene is as far as possible to be presented macroscopically when the multiple small windows are displayed, and a better scene immersion can be obtained when the audience remotely views the scene pictures through the multiple small windows.
[0011] The to-be-processed video stream is a video stream collected by a monocular camera.
[0012] Accordingly, the relative position parameters of the corresponding real persons are confirmed according to the person regions in the video picture of the to-be-processed video stream, including:
[0013] identifying the person regions in the video picture of the to-be-processed video stream, and confirming a first distance value and a first angle value corresponding to each person region according to a parameter value of a preset dimension in the person region and a pre-stored parameter value of the image acquisition module, the first distance value being used to represent the relative size relationship between the corresponding shooting distances of each person region, the shooting distance being the horizontal distance between the corresponding real person and the image acquisition module, and the first angle value being used to represent the included angle of the projection of the connecting line between the corresponding real person and the image acquisition module in the horizontal plane relative to a preset reference line;
[0014] confirming the relative position parameters of the corresponding real persons according to the first distance value and the first angle value.
[0015] According to the imaging principle and the final imaging effect of the image acquisition module, the relative position relationship of the real people in the real scene is confirmed based on the pre-stored device parameters and the parameters obtained by processing the collected video screen, and the relative position relationship of the portrait area in the video screen is corrected as the correction result, so that when the close-up screen corresponding to each portrait is displayed through the small window, the corrected relative position relationship of the real people in the field is used as the reference for the display layout of the small window displaying the close-up screen, so that the relative position relationship of the real people in the field is as close as possible when multiple small windows are displayed, and the audience can obtain a better sense of immersion when remotely viewing the field screen through multiple small windows.
[0016] The method comprises the steps of:
[0017] The method comprises the steps of:
[0018] The method comprises the steps of:
[0019] The method comprises the steps of:
[0020] According to the imaging principle and the final imaging effect of the image acquisition module, various parameters describing the conversion relationship between the real scene and the final imaging are extracted, and the first distance value and the first angle value used to represent the relative position relationship of the real people in the real scene are confirmed based on these parameters, so that the accurate relative position relationship between the real people can be obtained.
[0021] The method comprises the steps of:
[0022] The method comprises the steps of:
[0023] According to the imaging principle and the final imaging effect of the image acquisition module, various parameters describing the conversion relationship between the real scene and the final imaging are extracted, and the first distance value and the first angle value used to represent the relative position relationship of the real people in the real scene are confirmed based on these parameters, so that the accurate relative position relationship between the real people can be obtained.
[0024] The relative position parameters of the real person corresponding to each portrait area are determined according to the first distance value and the first angle value, and the relative position parameters of the real person corresponding to each portrait area include:
[0025] The relative position coordinates of the real person corresponding to each portrait area in the same rectangular coordinate system are determined according to the first distance value and the first angle value of each portrait area.
[0026] The relative position coordinates of the real person corresponding to each portrait area in the same rectangular coordinate system are determined according to the first distance value and the first angle value of each portrait area.
[0027] The relative position parameters of the real person corresponding to each portrait area are determined according to the first distance value and the first angle value of each portrait area.
[0028] The relative position parameters of the real person corresponding to each portrait area are determined according to the first distance value and the first angle value of each portrait area.
[0029] The relative position parameters of the real person corresponding to each portrait area are determined according to the first distance value and the first angle value of each portrait area.
[0030] The relative position parameters of the real person corresponding to each portrait area are determined according to the first distance value and the first angle value of each portrait area.
[0031] The relative position parameters of the real person corresponding to each portrait area are determined according to the first distance value and the first angle value of each portrait area.
[0032] The relative position parameters of the real person corresponding to each portrait area are determined according to the first distance value and the first angle value of each portrait area.
[0033] The relative position parameters of the real person corresponding to each portrait area are determined according to the first distance value and the first angle value of each portrait area.
[0034] The relative position parameters of the real person corresponding to each portrait area are determined according to the first distance value and the first angle value of each portrait area.
[0035] The method comprises the following steps:
[0036] The method comprises the following steps:
[0037] The method comprises the following steps:
[0038] The method comprises the following steps:
[0039] The method comprises the following steps:
[0040] The method comprises the following steps:
[0041] The method comprises the following steps:
[0042] The method comprises the following steps:
[0043] The method comprises the following steps:
[0044] The method comprises the following steps:
[0045] The method comprises the following steps:
[0046] The method comprises the following steps:
[0047] The method comprises the following steps:
[0048] Specifically, based on the relative positions of the display areas within the same group, the display positions of the sub-windows corresponding to the display areas within the same group are determined within the same row, including:
[0049] Based on the order of the horizontal position parameters in the relative position parameters of the display areas in the same group, the child screen windows corresponding to the display areas in the same group are arranged in the same row for display.
[0050] As described above, the display is arranged in order of size according to the horizontal position. This is equivalent to controlling the display area in the same group to be displayed in a way that roughly follows the distribution of the real people on site. Remote users can obtain an overall display effect that is closer to the real layout on site when multiple close-up images are displayed at the same time, thus achieving a better sense of immersion on site.
[0051] Secondly, embodiments of this application also provide a video processing apparatus, which includes:
[0052] The video acquisition unit is used to acquire the video stream to be processed, which is obtained by the image acquisition module from images of a multi-person scene.
[0053] The parameter confirmation unit is used to confirm the relative position parameters of the corresponding real person based on the human image area in the video frame of the video stream to be processed.
[0054] The layout display unit is used to identify the area to be displayed from the portrait area, determine the display layout of the sub-screen window corresponding to each area to be displayed according to the relative position parameters of the area to be displayed, and display the corresponding area screen in each sub-screen window. The area screen is a partial screen of the corresponding area to be displayed in the video screen of the video stream to be processed.
[0055] The video stream to be processed is the video stream captured by a monocular camera;
[0056] Correspondingly, the parameter confirmation unit includes:
[0057] The distance and angle confirmation module is used to identify the human figure region in the video frame of the video stream to be processed, and to confirm the first distance value and the first angle value corresponding to each human figure region based on the parameter values of the preset dimension in the human figure region and the parameter values of the pre-stored image acquisition module. The first distance value is used to characterize the relative size relationship between the shooting distances corresponding to each human figure region. The shooting distance is the horizontal distance between the corresponding real person and the image acquisition module. The first angle value is used to characterize the angle between the projection of the line connecting the corresponding real person and the image acquisition module on the horizontal plane and the preset reference line.
[0058] The position confirmation module is used to confirm the relative position parameters of the real person corresponding to each portrait area based on the first distance value and the first angle value.
[0059] The distance angle confirmation module comprises:
[0060] The region identification sub-module is configured to identify a portrait region in a video frame of the video stream to be processed and confirm a portrait size of each portrait region, wherein the portrait size is confirmed according to a distance between at least two preset portrait reference points in the portrait region.
[0061] The distance confirmation sub-module is configured to confirm a first distance value corresponding to each portrait region according to each portrait size.
[0062] The angle confirmation sub-module is configured to confirm a first angle value corresponding to each portrait region according to a pre-recorded field of view angle of the image acquisition module, a frame size of the video stream to be processed and a position of a central reference point, wherein the position of the central reference point is confirmed according to the distance between the at least two portrait reference points.
[0063] The distance confirmation sub-module comprises:
[0064] The normalization sub-module is configured to divide each portrait size by a minimum size in the portrait sizes to obtain a first value sequence, sort the values in the first value sequence from small to large and associate each value to a portrait region sorted from large to small according to the portrait size as a first distance value corresponding to the portrait region.
[0065] The position confirmation module comprises:
[0066] The coordinate conversion sub-module is configured to confirm a relative position coordinate of a real person corresponding to each portrait region in a same rectangular coordinate system according to the first distance value and the first angle value corresponding to each portrait region.
[0067] The video stream to be processed is a multi-path video stream captured by a multi-view camera.
[0068] The parameter confirmation unit comprises:
[0069] The three-dimensional mapping module is configured to map the portrait region in the video frame of the corresponding multi-path video stream to be processed to a three-dimensional space confirmed by the multi-view camera according to pre-stored calibration information of the multi-view camera.
[0070] The coordinate confirmation module is configured to confirm a relative position parameter of the real person corresponding to the portrait region in the three-dimensional space.
[0071] The layout display unit comprises:
[0072] The main picture display module is configured to confirm a to-be-displayed region from a portrait region of a video stream captured by a main camera of the multi-view camera, confirm a display layout of a sub-picture window corresponding to each to-be-displayed region according to a relative position parameter of the to-be-displayed region, and display a region picture corresponding to each sub-picture window in each sub-picture window.
[0073] The layout display unit includes:
[0074] The size confirmation module is configured to confirm a to-be-displayed region as a portrait region in which at least one portrait size meets a preset display condition.
[0075] The layout display unit includes:
[0076] The grouping confirmation module is configured to confirm to-be-displayed regions from the portrait regions, confirm a grouping result of the to-be-displayed regions according to a number of the to-be-displayed regions, and display sub-picture windows corresponding to each group of to-be-displayed regions in a same row according to the display layout.
[0077] The in-group confirmation module is configured to confirm display positions of the sub-picture windows corresponding to the to-be-displayed regions in the same group in the same row according to relative position parameters of the to-be-displayed regions in the same group.
[0078] The picture display module is configured to display the region pictures in the corresponding sub-picture windows according to the display positions.
[0079] The grouping confirmation module includes:
[0080] The to-be-displayed confirmation submodule is configured to confirm to-be-displayed regions from the portrait regions.
[0081] The first grouping module is configured to, in a case where the number of the to-be-displayed regions matches a target layout number in a preset array layout number, confirm a grouping manner of the to-be-displayed regions as an array manner corresponding to the target layout number.
[0082] The second grouping module is configured to, in a case where the number of the to-be-displayed regions does not match the target layout number in the preset array layout number, confirm that the grouping manner of the to-be-displayed regions includes grouping according to distances from a same reference point and grouping according to distances between the to-be-displayed regions.
[0083] The in-group confirmation module includes:
[0084] The horizontal ordering submodule is configured to, according to a size order of horizontal position parameters in the relative position parameters of the to-be-displayed regions in the same group, sequentially arrange display positions of the sub-picture windows corresponding to the to-be-displayed regions in the same group in the same row.
[0085] In a third aspect, an electronic device is also provided in the embodiments of the present application.
[0086] one or more processors;
[0087] a memory for storing one or more computer programs;
[0088] When the one or more computer programs are executed by the one or more processors, the electronic device implements the video processing method as in the first aspect.
[0089] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the video processing method as in the first aspect. BRIEF DESCRIPTION OF DRAWINGS
[0090] FIG. 1 is a method flowchart of a video processing method provided by the embodiments of the present application.
[0091] FIG. 2 is a method flowchart of another video processing method provided by the embodiments of the present application.
[0092] FIG. 3 is a schematic diagram of a video frame of a video stream collected by the related art.
[0093] FIG. 4 is a schematic diagram of displaying a close-up frame of the video stream shown in FIG. 3 according to the related art.
[0094] FIG. 5 is a schematic diagram of a first kind of real person distribution in a multi-person live scene.
[0095] FIG. 6 is a schematic diagram of a video frame of a video stream collected from the multi-person live scene shown in FIG. 5.
[0096] FIG. 7 is a schematic diagram of displaying a close-up frame of the video stream shown in FIG. 6 according to the related art.
[0097] FIG. 8 is a schematic diagram of a second kind of real person distribution in a multi-person live scene.
[0098] FIG. 9 is a schematic diagram of a video frame of a video stream collected from the multi-person live scene shown in FIG. 8.
[0099] FIG. 10 is a schematic diagram of displaying a close-up frame of the video stream shown in FIG. 9 according to the related art.
[0100] FIG. 11 is a schematic diagram of an imaging process when an image collection module collects an image.
[0101] FIG. 12 is a schematic diagram of a principle of confirming a first distance value based on FIG. 11.
[0102] FIGS. 13 and 14 are schematic diagrams of a principle of confirming a first angle value based on FIG. 11.
[0103] FIG. 15 is a schematic diagram of a relationship between a coordinate system of a three-dimensional space of a camera and a real space.
[0104] FIG. 16 is a schematic diagram of mapping a real person into the same Cartesian coordinate system based on a relative position parameter.
[0105] FIG. 17 is a schematic diagram of a real person's relative position confirmed by the video stream shown in FIG. 6 according to an embodiment of the present application.
[0106] FIG. 18 is a schematic diagram of displaying a close-up view of the video stream shown in FIG. 6 based on FIG. 17.
[0107] FIG. 19 is a schematic diagram of a structure of a video processing apparatus according to an embodiment of the present application.
[0108] FIG. 20 is a schematic diagram of a structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0109] The present application will be further described below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are intended to explain, not to limit, the present application. In addition, it should be noted that only the parts related to the present application are shown in the accompanying drawings for the purpose of description.
[0110] It should be noted that, due to the limitation of the length, all the optional embodiments are not enumerated in the present application. Those skilled in the art, after reading the present application, should be able to think of any combination of technical features as long as the technical features do not contradict each other, which can constitute an optional embodiment.
[0111] The embodiments of the present application will be described in detail below.
[0112] In the related art, when extracting and displaying multiple close-up pictures based on a global picture corresponding to a video stream, the close-up pictures are usually displayed based on the relative position relationship of the close-up picture corresponding to the portrait in the global picture. FIG. 3 and FIG. 4 show a schematic diagram of a video picture of a video stream and a corresponding generated and displayed close-up picture. In FIG. 3, when the image acquisition module acquires a global picture, it captures 6 real people P1-P6 in the multi-person live scene and records 6 portraits PA-PF in the video picture. According to the related art, the close-up pictures of the six portraits in the global picture are extracted, and the picture data of the six portrait areas in the global picture is filled into the corresponding small windows. The multiple small windows are displayed in rows and columns. When displaying the small windows, the layout of the small windows corresponding to different portrait areas is not random display, but is arranged according to the relative position of the portrait in the global picture as shown in FIG. 4. For example, in FIG. 3, the six portraits can be divided into three portraits PA, PB and PC from left to right in the upper part and three portraits PD, PE and PF from left to right in the lower part. When displaying the six portrait areas extracted from FIG. 3 in the small windows of 2 rows and 3 columns, the relative position relationship between the above portraits is taken as a reference in the horizontal and vertical directions, and the display layout shown in FIG. 4 is presented in the multiple small windows. In this way, the remote user can feel the corresponding live conversation state from a macroscopic perspective while viewing the close-up picture, that is, the close-up picture is highlighted and the live picture is felt at the same time, and better conference immersion is obtained.
[0113] According to the above layout method of the small windows, it is possible that the relative position relationship in the global picture finally acquired due to the misplacement of the real people, the angle and depth of field of the image acquisition module, etc. is not necessarily equivalent to the relative position relationship in the multi-person live scene, resulting in a deviation of the relative position relationship of the portraits displayed in the small windows compared to the relative position relationship in the multi-person live scene.
[0114] For example, the first display of the close-up picture from the multi-person live image collection shown in FIG. 5, FIG. 6 and FIG. 7, according to the relative position relationship of the first multi-person live human distribution schematic diagram of FIG. 5, it is assumed that the scene details at this time are that the human P5 is talking to the human P1, and the global picture presented in the collected video picture is as shown in FIG. 6. If the relative technology takes the up-down-left-right relationship of the portrait PA, the portrait PB, the portrait PC and the portrait PE in FIG. 6 as the relative position relationship of the human P1, the human P2, the human P3 and the human P5, and displays the close-up picture through the small window, the layout schematic diagram shown in FIG. 7 will be obtained. Among them, the human P5 stands at a position relatively close to the image collection module, and the position of the portrait (the portrait PE) imaged in the global picture is higher than that of the human P1 (corresponding to the portrait PA) and the human P3 (corresponding to the portrait PC), which may finally lead to that in the layout shown in FIG. 7, the human P5 and the human P1 who are originally face-to-face talking, present a state of portrait PE and portrait PA talking back-to-back in the layout after the arrangement according to the relative position relationship of the portraits in the video picture.
[0115] For example, the second display of the close-up picture from the multi-person live video collection shown in FIG. 8, FIG. 9 and FIG. 10, according to the relative position relationship of the second multi-person live human distribution schematic diagram of FIG. 8, it is assumed that the scene details at this time are that the human P5 is talking to the human P3, and the global picture presented in the collected video picture is as shown in FIG. 9. If the relative technology takes the up-down-left-right relationship of the portrait PA, the portrait PB, the portrait PC and the portrait PE in FIG. 9 as the relative position relationship of the human P1, the human P2, the human P3 and the human P5, and displays the close-up picture through the small window, the layout schematic diagram shown in FIG. 10 will be obtained. Among them, the human P5 stands at a position relatively far from the image collection module, and the position of the portrait (the portrait PE) imaged in the global picture is more left than that of the human P3 (corresponding to the portrait PC), which may finally lead to that in the layout shown in FIG. 10, the human P5 and the human P3 who are originally face-to-face talking, present a state of portrait PE and portrait PA talking back-to-back in the layout after the actual arrangement.
[0116] Based on the above analysis, it can be found that when the related technology extracts the close-up picture for combined display, the multi-close-up picture spliced therefrom cannot restore the real live state from the intuitive picture content and the macroscopic orientation feeling, and the viewing immersion effect of the remote user is poor.
[0117] To solve the above technical problems, the embodiment of the present application provides a video processing method. The video stream obtained by image acquisition of a multi-person scene is processed. The relative position parameters of the real persons corresponding to the portrait regions are first confirmed, i.e., the relative position parameters of the real persons in the actual real scene, and the relative position relationship of the portrait regions in the video picture is corrected based on the relative position parameters. When displaying the close-up pictures of the respective portraits through the small windows, the corrected relative position relationship of the real persons in the actual real scene is used as the reference for the display layout of the small windows displaying the close-up pictures, so that the relative position relationship of the real persons in the actual real scene is as much as possible to be presented macroscopically when the multiple small windows are displayed, and the audience can obtain a better sense of immersion when remotely viewing the live pictures through the multiple small windows.
[0118] FIG. 1 is a method flowchart of a video processing method provided by the embodiment of the present application. The video processing method is applied to various electronic devices that can be used for processing and remotely sending the video stream based on the live acquisition of a multi-person scene. The video processing method is realized by the electronic devices, such as various mobile terminals, personal computers, and the like.
[0119] As shown in FIG. 1, the video processing method includes but is not limited to steps S01-S03.
[0120] Step S01: obtaining a to-be-processed video stream. The to-be-processed video stream is obtained by image acquisition of a multi-person scene by an image acquisition module.
[0121] Step S02: confirming the relative position parameters of the real persons corresponding to the portrait regions in the video picture of the to-be-processed video stream.
[0122] Step S03: confirming the to-be-displayed regions from the portrait regions, confirming the display layout of the sub-picture window corresponding to each to-be-displayed region according to the relative position parameters of the to-be-displayed regions, and displaying the corresponding region picture in each sub-picture window. The region picture is a local picture including the corresponding to-be-displayed region in the video picture of the to-be-processed video stream.
[0123] In the embodiment of the present application, the to-be-processed video stream obtained by image acquisition of a multi-person scene is described as a whole. The relative position parameters of the real persons corresponding to the portrait regions in the actual real scene are first confirmed. Then, the to-be-displayed regions are confirmed from the portrait regions. When the to-be-displayed regions are displayed in the multiple sub-picture windows distributed in an array or close to the array, the positions of the sub-picture windows in which the to-be-displayed regions are displayed are confirmed according to the relative position parameters, so that the portraits displayed in the sub-picture windows are as much as possible to be close to the distribution of the real persons in the actual real scene. The relative position relationship of the real persons in the actual real scene is as much as possible to be presented macroscopically when the multiple small windows are displayed. The audience can obtain a better sense of immersion when remotely viewing the live pictures through the multiple small windows.
[0124] The to-be-processed video stream in the embodiments of the present application can be an image acquisition module such as a monocular camera or a multi-view camera, and a video stream obtained by image acquisition of a multi-person scene. The to-be-processed video stream obtained by the monocular camera or the multi-view camera has a corresponding manner to realize the corresponding relative position parameters of the real person. On the basis of having confirmed the relative position parameters, different types of to-be-processed video streams can confirm the specific display area displayed in the sub-picture window according to the relative position parameters in the same manner.
[0125] The entire video processing process will be described in detail below based on the complete processing process corresponding to the acquisition manner of the specific to-be-processed video stream.
[0126] As shown in FIG. 2, for the to-be-processed video stream acquired by the camera, the video processing method includes but is not limited to steps S110-S140:
[0127] Step S110: Obtain the to-be-processed video stream, which is obtained by image acquisition of a multi-person scene by an image acquisition module.
[0128] The to-be-processed video stream of the embodiments of the present application mainly refers to the need to simultaneously present multiple close-up pictures from the same scene to the user, and the layout of the multiple close-up pictures when simultaneously presented substantially meets the layout of the corresponding entities in the scene. For example, after obtaining a global picture of a multi-person scene by image acquisition, multiple area pictures corresponding to each person in the multi-person scene are extracted from the global picture as close-up pictures and simultaneously displayed to the user. At the same time, the multiple close-up pictures are displayed simultaneously, and the display layout of the multiple close-up pictures substantially maintains the relative position relationship of the corresponding real persons in the multi-person scene when the multiple close-up pictures are simultaneously displayed. Each close-up picture corresponds to a sub-picture window for display. The sub-picture window can also be referred to as a small window, a sub-window, a sub-area, etc., and specifically refers to a small area in the display interface for presenting a single close-up picture.
[0129] The to-be-processed video stream can be a real-time acquired video stream or a saved video stream, which is obtained by image acquisition of a multi-person scene by an image acquisition module, and the real person is taken as a close-up target accordingly. In the specific implementation process, considering the requirement of close-up on the picture quality and the need to extract multiple close-up pictures from a global video stream, the image acquisition module can be configured as a high-definition device, such as a 4K device, an 8K device, etc. The video stream can come from a single-view camera or a multi-view camera fusion, and accordingly, the close-up picture with excellent picture quality can be extracted from the high-definition to-be-processed video stream. Steps S120 and S130 in the embodiments of the present application are specifically used to describe the detailed processing process of the relative position parameters when a single-view camera is used, i.e., only one to-be-processed video stream.
[0130] Step S120: identifying a human region in a video frame of the video stream to be processed, and confirming a first distance value and a first angle value corresponding to each human region according to a parameter value of a preset dimension in the human region and a parameter value of the pre-stored image acquisition module, the first distance value being used to represent a relative size relationship between the corresponding shooting distances of each human region, the shooting distance being a horizontal distance between the corresponding real person and the image acquisition module, and the first angle value being used to represent an included angle of a projection of a line connecting the corresponding real person and the image acquisition module in a horizontal plane relative to a preset reference line.
[0131] In the existing image processing field, it is actually impossible to confirm the depth data of the object in the shooting target through a video, that is, it is impossible to confirm the position parameter of the object in the shooting target in the real space through a video. In the embodiment of the present application, through the depth analysis of the imaging principle of the image acquisition module and the characteristics of the object in the scene, a strategy of confirming the position parameter of the object in the real space through a video in the form of estimation is proposed. In the specific confirmation, considering that the video frame is generated inside the image acquisition module, the video frame is generated by the light emitted or reflected by the shooting target entering the photosensitive unit inside the image acquisition module through the lens of the image acquisition module, and there are some fixed geometric relationships in this process, which can be used as the basis for finally estimating the position of the real person in the real space. Finally, considering the installation characteristics of the image acquisition module and the characteristics of the real person, it is confirmed that the shooting distance of the real person and the image acquisition module and the included angle of the line connecting the real person and the image acquisition module relative to the preset reference line can be confirmed. The position parameter of the real person in the real space with the image acquisition module as the reference is obtained. It should be understood that the shooting distance confirmed based on the video stream to be processed collected by the monocular camera refers to the shooting distance of the corresponding real person and the image acquisition module, the relative size relationship between the corresponding shooting distances of each human region is estimated based on the human state in the image, and the numerical value of each human region for describing the relative size relationship is the corresponding first distance value, and the included angle of the line connecting the real person and the image acquisition module relative to the preset reference line is the first angle value.
[0132] The feasibility of describing the relative position relationship of the real person based on the first distance value and the first angle value can be further explained with reference to the combination of FIGS. 11-14.
[0133] First, please refer to FIG. 11, which is a schematic diagram of the imaging process when the image acquisition module is acquiring images. The light reflected by the real person 11 passes through the lens of the image acquisition module, enters the photosensitive unit 10 of the image acquisition module, and completes signal acquisition to generate images through the photosensitive unit 10. In this relatively short distance, the light reflected by the real person 11 can be regarded as straight-line propagation. In addition, based on the basic structure of the image acquisition module and the relatively uniform target when a person is the target for shooting, a large number of angle relationships (equal, complementary, and supplementary) and a large number of similar triangles can be confirmed based on FIG. 11. Based on these angle relationships and similar triangles, the corresponding edge and angle specific values can be calculated, and then the principle diagram for confirming the first distance value as shown in FIG. 12 and the principle diagram for confirming the first angle value based on FIG. 11 as shown in FIG. 13 and FIG. 14 can be obtained. The edge and angle relationships confirmed in FIG. 11 and the edge and angle calculation based thereon are basic mathematical implementations, such as the solution of equal top angles, equal homologous angles, and right triangles, which will not be expanded here.
[0134] In FIG. 12, taking the imaging of a head portrait as an example, d represents the real distance, h represents the height of the head, z represents the imaging height of the head, and f represents the lens distance of the image acquisition module from the center of the lens of the image acquisition module to the lens. According to the similarity of triangles, d = (f * h) / z can be obtained. Moreover, the head imaging is in the picture. For a panoramic picture with a height H, the head has a height hp in the picture, so z / Z = hp / H, and then d = f * H / (Z * hp) * h. To calculate d, five parameters need to be known. For the same image area of a person, f, H, and Z in the five parameters are the same, h is difficult to measure in real time for each user on the spot, and the height of the head is generally close to the same value and can be regarded as the same value in the processing process. In the case where the state of the image acquisition module does not change, only the corresponding hp (i.e., the parameter value of the preset dimension in the image area) of each image area needs to be extracted from the panoramic image, and the real distance corresponding to each image area can be obtained through the corresponding hp. However, the five parameters may change due to the focusing of the image acquisition module, for example, the height H changes when focusing, so in the embodiments of the present application, the real distance corresponding to each image area is further divided by one of them as a reference. In the case where f, H, and Z are the same and h is regarded as the same, the division eliminates the influence of all the same parameters, and only the height of the head in the picture can be obtained to obtain the relative size relationship between the shooting distances of each image area. The specific numerical value representing the relative size relationship is used as the first distance value of each image area.
[0135] In FIG. 13 and FIG. 14, two different calculation methods are adopted to confirm the first angle value in consideration of different position states. Based on the imaging process schematic diagram in FIG. 11, for the actual picture formed by the light rays entering the photosensitive unit 10 being collected by the photosensitive unit, there is the relationship shown in FIG. 13 or FIG. 14 in the actual picture. Wherein W is the pixel size of the width in the collected picture. According to the pre-stored horizontal field angle a of the image collection module, L can be calculated according to the following formula: According to the horizontal coordinate x of the center point of the detected head in the panoramic picture, the current θ is converted. In the position state of FIG. 13, tan θ = L / (x-W / 2), and in the position state shown in FIG. 14, tan(θ-π / 2)=(W / 2-x) / L. According to the above formula, θ, that is, the first angle value, can be calculated.
[0136] Based on the above argument, it can be found that the position of the portrait area is represented by the first distance value and the first angle value, which is equivalent to scaling the distribution of the live person in another coordinate system. Obviously, such scaling will only change the distance, but not the relative position relationship. For example, the distribution relationship of the live person is that A is in the southeast direction of B, and after scaling, A is still in the southeast direction of B. The distribution of the small window is only determined by the relative position relationship. Based on the relative position relationship obtained in the embodiment of the present application, which can accurately describe the relative position relationship of the live person, the relative position relationship in the embodiment of the present application can accurately provide a reference for the layout of the small window. Based on the overall design idea of confirming the first distance value and the first angle value described in the foregoing, the process of confirming the first distance value and the first angle value in step S120 can include: identifying the portrait area in the video picture of the video stream to be processed, and confirming the portrait size of each portrait area, the portrait size being confirmed according to the distance between at least two preset portrait reference points in the portrait area; confirming the first distance value corresponding to each portrait area according to each portrait size; and confirming the first angle value corresponding to each portrait area according to the pre-recorded field angle of the image collection module, the picture size of the video stream to be processed, and the position of the center reference point. When confirming the first angle value, for a portrait area with a certain size, a point located in the center of the face is used to represent the position of the portrait area, which is the center reference point. In the case where the positions of at least two portrait reference points have been determined, the position of the center reference point can be determined according to the distance between the at least two portrait reference points, and the position of the center reference point can also be determined according to the coverage position of the portrait area, with the centroid as the position of the center reference point.
[0137] In the implementation process, the portrait area in the video picture is identified by a portrait recognition or detection method. The portrait size can be represented by a body surface parameter with relatively small difference between people, such as the height of the head (i.e., the portrait reference points are the highest point on the head and the lowest point on the chin, and the distance between the two preset portrait reference points corresponds to the height of the head), the distance between the eye corners (i.e., the portrait reference points are the points on the outer corners of the eyes, and the distance between the two preset portrait reference points corresponds to the distance between the eye corners), etc. The first distance value and the first angle value can be calculated accordingly. The field of view angle and other basic parameters of the image acquisition module are pre-stored accordingly. In addition, it should be noted that the portrait area actually covers a region. In the embodiment of the present application, the center of the portrait reference points is selected as the coordinate position for calculating the first angle value. In addition, the coordinate position confirmed by the average of the coordinate components of all pixels in the entire portrait area can also be used.
[0138] It should be understood that the video stream to be processed can actually be regarded as a collection of multiple consecutive images. To cope with the changes in the relative position relationship between multiple people that may occur during image acquisition and to adjust the layout in real time when displaying the close-up picture in the display sub-picture window, the present solution can be implemented frame by frame or at intervals of a certain number of frames. In addition, the portrait areas in the same frame should be processed one by one or in parallel until the layout of the subsequent sub-picture window is confirmed after all the portrait areas in the same frame are processed. In terms of the existing video frame rate, the total time corresponding to a certain number of frames (e.g., 2 frames, 3 frames, 5 frames, etc.) is very short. Even if changes in the relative position relationship occur in the middle, the subsequent layout can be quickly switched to the corresponding reconfirmed layout, which basically does not affect the immersion of the user during viewing. Of course, processing at intervals of frames (or multiple frames) is only an exemplary representation. In actual processing, it can also be processed at intervals of a fixed time length (e.g., 0.5 seconds, 1 second, etc.).
[0139] As described above, according to the imaging principle and the final imaging effect of the image acquisition module, various parameters describing the conversion relationship between the real scene and the final imaging are extracted, and the first distance value and the first angle value used to represent the relative position relationship between real people in the real scene are confirmed based on these parameters, so that the accurate relative position relationship between real people can be obtained.
[0140] For obtaining the first distance value, the minimum size in the portrait size can be sequentially divided by each portrait size to obtain a first value sequence. The values in the first value sequence are sorted from small to large, and are associated with the portrait areas sorted from large to small by the portrait size one by one as the first distance value of the corresponding portrait area. For detecting various sizes of heads in the panoramic picture, the heads are sorted from large to small by the head height, and the sorting result is denoted as hp1, hp2, …, hpn. The first distance value of the head is denoted as d1, d2, …, dn, and the first angle value of the head is denoted as a1, a2, …, an. nSubstituting the head heights mentioned above into the formula d = f*H / (Z*hp)*h from the previous text, we can estimate the actual distance d corresponding to each portrait region. i `=f*H / (Z*hp i )*h, where 1≤i≤n. Divide the estimated true distances for each portrait region by the maximum distance, specifically by dividing each true distance by the maximum distance to obtain the corresponding first distance value di, i.e., d. i =d i ` / d n Simplifying this expression, we can obtain d. i =hp n / hp i That is, d1 = hp n / hp1,d2=hp n / hp2,d3=hp n / hp3,…,d n =1, which means dividing the smallest size in the portrait dimensions by each portrait dimension to obtain the first numerical sequence. The values in this first numerical sequence are then sorted from smallest to largest and associated with the portrait regions sorted from largest to smallest portrait dimensions. This yields the first distance value for each portrait region. This association is essentially equivalent to dividing the smallest size in the portrait dimensions by each portrait dimension to obtain the first distance value for the corresponding portrait region. The sorted and associated values are only used to more intuitively present the size relationship based on the normalized description during the demonstration. It should be understood that the sorting, substitution, and simplification based on the head height description described above are used to demonstrate the mathematical processing to obtain the final first distance value. For the electronic device that specifically implements this video processing method, directly based on d... i =hp n / hp i This means dividing the minimum head height by the head height of each portrait region to obtain the first distance value for the corresponding portrait region. This is equivalent to normalizing the head height that can be directly obtained from the image to obtain the corresponding first distance value.
[0141] The above-mentioned relative distance values obtained through normalization can present the distance status between each real person and the image acquisition module in a more intuitive way, resulting in higher processing efficiency in the future.
[0142] Step S130: Determine the relative position parameters of the real person corresponding to each portrait area based on the first distance value and the first angle value.
[0143] Specifically, the relative position parameters can be determined according to the first distance value and the first angle value of each portrait region, and the relative position coordinates of the real person corresponding to each portrait region in the same rectangular coordinate system can be determined. That is, for the relative distance and the angle of each portrait region, the relative position of the corresponding real person in the real space can be obtained. For example, for the horizontal coordinate x of a certain portrait region in the same rectangular coordinate system, x n = d n *cos θ n , for the vertical coordinate y of a certain portrait region in the same rectangular coordinate system, y n = d n *sin θ n , and the relative position coordinates (x n , y n ) of each object in the space are obtained, where d n represents the first distance value of the n-th portrait region corresponding to the real person, and θ n represents the first angle value of the n-th portrait region corresponding to the real person. Through the first distance value and the first angle value, the relative position coordinates of each real person in the same rectangular coordinate system are determined, which can be quickly grouped and laid out in the subsequent process. Based on the relative position parameters, a diagram of mapping the real person to the same rectangular coordinate system based on the relative position parameters can be obtained as shown in FIG. 16. It should be understood that FIG. 16 is only used to represent the relationship between the first distance value, the first angle value, and the relative position parameters in the same rectangular coordinate system, and it is not necessary to map the relative position parameters of each portrait region to the rectangular coordinate system in the specific processing process.
[0144] For the image acquisition module, in addition to the horizontal field of view angle a of the image acquisition module, the vertical field of view angle β of the image acquisition module can also be preset. Based on the same processing logic of the horizontal field of view angle, the offset angle κ of the portrait region in the vertical direction based on the optical axis in the camera coordinate system can be obtained according to the height H and the vertical coordinate y of the portrait region.
[0145] The direction of the optical axis of the image acquisition module is taken as the positive direction of the x-axis, the y-axis direction is the horizontal direction of the image acquisition module, and the z-axis direction is the vertical direction of the image acquisition module. The camera three-dimensional coordinate system is established, and the horizontal angle θ and the vertical offset angle κ of the real person based on the camera three-dimensional coordinate system are obtained. Based on the normalized distance d of each target obtained in the foregoing, the position relationship of each object is converted into a vector μ(x, y, z) in the camera three-dimensional coordinate system.
[0146] When installing the image acquisition module, it is also possible to make the image acquisition module not face the horizontal direction but have a certain change in pitch angle in order to obtain a better imaging effect. In order to more accurately restore the real position of each real person in space, the pitch angle γ of the image acquisition module is introduced, as shown in FIG. 15, that is, the image acquisition module has a downward inclination based on the horizontal line, that is, rotates γ around the y axis, and the corresponding rotation matrix is:
[0147] The vector corresponding to the real person in the camera three-dimensional coordinate system is converted into a vector in the coordinate system in the real space, and only needs to be multiplied by the inverse of the rotation matrix R to obtain the real vector in the space, that is:
[0148] Based on the vector v in the conference room space, x and y in the vector v are obtained, and the relative plane position of the real person object in the real space is obtained, that is, the relative position parameter is obtained.
[0149] Step S140: Confirm the to-be-displayed region from the portrait region, confirm the display layout of the sub-picture window corresponding to each to-be-displayed region according to the relative position parameter of the to-be-displayed region, and display the corresponding region picture in each sub-picture window. The region picture is a local picture including the corresponding to-be-displayed region in the video picture of the to-be-processed video stream.
[0150] For multiple portrait regions in a video picture, corresponding real persons are usually partially in a speaking state at the same time. The portrait regions corresponding to the real persons in the speaking state have a higher necessity to be displayed to a remote user in a multi-person communication scenario. In addition, a camera in the multi-person communication scenario can usually track a target in the speaking state, so that the target is closer to the center and has a larger portrait in the video picture. Therefore, the size of the portrait region can be directly sorted, and a portrait region with a larger size is considered as a portrait region corresponding to a real person in the speaking state or a real person adjacent to the real person in the speaking state. The portrait region is regarded as a to-be-displayed region, so that close-up presentation of the real person in the speaking state can be ensured. In an optional implementation, the to-be-displayed region can be confirmed from the multiple portrait regions, and at least a portrait region with a size reaching a preset display condition is regarded as the to-be-displayed region. According to the size of the portrait region, the to-be-displayed region is selected, and only the portrait region meeting the preset display condition is displayed in close-up, so that unnecessary portrait display can be reduced, and the pertinence of close-up display can be improved. The preset display condition is mainly a size-based condition, that is, only a portrait region reaching a certain size is regarded as the to-be-displayed region. The size of the portrait region can reach a preset absolute size value (for example, a pixel value of a length and a width), or the size of the portrait region can reach a relative size value (for example, the size is in a top few positions in the sorting or the size is in a certain percentage range in the top positions in the sorting). For example, in a panoramic picture, there are 30 portrait regions, but according to the size relationship of all the portrait regions, there are 5 portrait regions with sizes reaching a preset absolute size value, which indicates that there are 5 persons who have a higher communication interaction degree and need to be displayed in close-up. Therefore, the size of the portrait region can be used as a reference for whether the portrait region needs to be displayed in close-up. If the to-be-processed video stream is collected by a high-definition camera, the close-up effect of each portrait region can be ensured, and all the portrait regions can be regarded as to-be-displayed regions automatically or according to a user operation. It is considered that a speaking target is tracked in a conference or a teaching process, that is, an image collection module focuses on a current speaker and displays the current speaker in the center. Through an embodiment of the present application, close-up change can be implemented to adapt to a change of the tracking target.
[0151] When a close-up picture based on a portrait region is displayed based on a video stream to be processed, a local picture including a region to be displayed can be displayed through a sub-picture window, and a complete video stream to be processed can be simultaneously displayed or not displayed. For the region to be displayed, it is actually a face or head portrait region identified based on a portrait, and this part of the region is usually an irregular region, while the sub-picture window is usually a rectangular window with the same aspect ratio. When the region to be displayed is displayed to the sub-picture window, a rectangular region picture can be confirmed as the center of the region to be displayed, and the region picture as a local picture in a video picture includes a corresponding region to be displayed. The aspect ratio of the region picture is consistent with the aspect ratio of the sub-picture window. When the region picture is filled into the sub-picture window for display, the portrait region is enlarged or reduced as the center, which can ensure that the portrait does not appear stretched or compressed horizontally or vertically.
[0152] The display process of the region to be displayed can be to first confirm the region to be displayed from the portrait region, count the number of portrait regions (i.e., the region to be displayed) that meet the preset display condition, then confirm the grouping result of the region to be displayed according to the number of the region to be displayed, each group of the region to be displayed corresponds to a sub-picture window displayed in the same row in the display layout; then according to the relative position parameters of the region to be displayed in the same group, confirm the display position of the sub-picture window corresponding to the region to be displayed in the same group in the same row; and finally display the region picture in the corresponding sub-picture window according to the display position. The region to be displayed is displayed in the grouping and row display manner, which can effectively improve the display utilization of the display surface. It is equivalent to still take the structure shown in FIG. 4 as the basic display architecture, and the embodiment of the application mainly further optimizes which portrait region is displayed in each sub-picture window under the structure shown in FIG. 4, so that the whole is more consistent with the layout on site.
[0153] According to the number of the region to be displayed, the grouping result can be confirmed in different grouping manners. First, the region to be displayed is confirmed from the portrait region; in a case where the number of the region to be displayed matches a target layout number in the preset array layout number, the grouping manner of the region to be displayed is grouped in an array manner corresponding to the target layout number; and in a case where the number of the region to be displayed does not match a target layout number in the preset array layout number, the grouping manner of the region to be displayed includes grouping according to the distance from the same reference point and grouping according to the distance between the regions to be displayed.
[0154] The display of the sub-picture window can have multiple preset array layouts, which refers to the local arrangement of the sub-picture window in rows and columns, and each row has the same number of sub-picture windows. Each array layout has a corresponding array layout number, for example, the array layout number corresponding to the 3-row 4-column array layout is 12, and if the number of to-be-displayed regions is exactly 12, that is, in the case of exactly displaying the number of to-be-displayed regions in an array manner, the layout is directly displayed in an array manner (3-row 4-column array layout). In specific implementation, the number of regions that can be evenly divided into multiple groups has a corresponding array layout, for example, the array layout of 10 to-be-displayed regions can be 2 rows and 5 columns, and the array layout of 14 to-be-displayed regions can be 2 rows and 7 columns. In the case that a large number of close-up pictures are not usually confirmed, the array layout number corresponding to the array layout may be, for example, 4 (2 rows and 2 columns), 6 (2 rows and 3 columns), 8 (2 rows and 4 columns), 9 (3 rows and 3 columns), 10 (2 rows and 5 columns), 12 (3 rows and 4 columns), 14 (2 rows and 7 columns), and the like. In specific implementation, the effect presented when finally grouped can also be considered, that is, the number of regions that can be evenly divided into multiple groups is not the array layout number, for example, 10 to-be-displayed regions, if the corresponding sub-picture window adopts an array layout, a 2-row 5-column layout may appear, but based on the imaging result of the camera, the actual person in the live scene is less likely to be distributed in two rows, that is, the 2-row 5-column layout is more likely to deviate from the actual person layout in the live scene. At this time, a 3-row (3+3+4) layout can be used, and at this time, the 10 to-be-displayed regions are not the array layout number corresponding to the array layout.
[0155] In the case that the number of to-be-displayed regions cannot fill an array or the layout according to the array deviates greatly from the actual person layout, that is, in the case that the number of to-be-displayed regions does not have a matching target layout number in the preset array layout number, flexible display can be performed according to specific scenes or parameters. For example, grouping according to the distance from the same reference point and grouping according to the distance between to-be-displayed regions. In this case, the layout can be further adjusted on the basis of the complete array shown in FIG. 4, for example, on the basis of FIG. 4, 3 sub-picture windows of the same size are displayed in the first row, and 2 sub-picture windows of the same size are displayed in the second row, and in the case that the first row has more sub-picture windows than the second row, the width (and the height accordingly) of part of the sub-picture windows is adjusted, so that the total width of the sub-picture windows in different rows is the same, and then the effect of aligning the first row of sub-picture windows and the second row of sub-picture windows left and right as units is presented.
[0156] In the relative position parameters of each to-be-displayed region, based on the expression of the same rectangular coordinate system as the horizontal coordinate and the vertical coordinate, according to the up-down direction pointed by the vertical coordinate, in addition, the uppermost and the lowermost can also be grouped from the middle, that is, according to the vertical coordinate from large to small and from small to large, each time a fixed number of to-be-displayed regions are taken as a group, and the fixed number is the column number of the preset sub-picture window array, and the fixed number of to-be-displayed regions corresponds to a row of sub-picture windows. After completing the complete grouping from the uppermost and the lowermost to the middle, if there are less than a fixed number of to-be-displayed regions left, which are not enough to form a complete group to be displayed in a row of sub-picture windows, these to-be-displayed regions are allocated to the two adjacent complete groups, and the specific allocation standard is to confirm the distance of each to-be-displayed region in the adjacent two complete groups according to the relative position parameters confirmed in the foregoing, and the smaller the total distance to which to-be-displayed region is considered to be closer to the complete group, and the to-be-displayed region is added to the group accordingly. It is equivalent to each group allocating to-be-displayed regions as a standard array layout of a row of sub-picture window array based on the column number. Finally, if there are remaining to-be-displayed regions in the middle position, the remaining ungrouped to-be-displayed regions are divided into the groups that are closer according to the comprehensive distance between the remaining to-be-displayed regions and the to-be-displayed regions in the adjacent groups. On the basis of completing the basic standard array layout, the to-be-displayed regions that are not enough to be displayed in a row of sub-picture windows are flexibly adjusted and allocated to a row of standard array layout, that is, the number of sub-picture windows in one group or two groups may be more than the column number of the sub-picture window array in the final display.
[0157] The relative position parameters are used to represent the relative position relationship of real persons. In the case that the real persons are considered to be in the same plane, a rectangular coordinate system coinciding with the plane is constructed, and each real person can be represented by a horizontal position parameter (i.e., a horizontal coordinate) and a vertical position parameter (i.e., a vertical coordinate) in the plane. In the grouping based on the vertical coordinate, the sub-picture windows in the same group are displayed in the display layout of the same row. In the final display link after grouping, the relative position relationship of the portraits displayed in the sub-picture windows needs to be as close as possible to the relative position relationship of the real persons. At this time, the display positions of the sub-picture windows corresponding to the to-be-displayed regions in the same group in the same row can be arranged in order according to the size order of the horizontal position parameters (i.e., horizontal coordinates) in the relative position parameters of the to-be-displayed regions in the same group. The display strategy in the group is basically the same as that shown in FIG. 4, and the to-be-displayed regions in the same group are displayed in order according to the size order of the horizontal position. It is equivalent to controlling the to-be-displayed regions in the same group to be displayed in the distribution mode of the real persons on the scene. The remote user can obtain a more close-to-reality overall display effect of the on-site layout when multiple close-up pictures are displayed at the same time, and a better on-site immersion is achieved.
[0158] For a single-channel video stream to be processed collected by a single camera, the remote user can obtain a more realistic overall display effect closer to the actual layout of the scene while multiple close-up pictures are displayed simultaneously based on existing hardware without increasing hardware costs, achieving better immersion in the live scene.
[0159] In another optional embodiment, for a multi-channel video stream to be processed collected by a multi-camera, i.e., multiple video streams to be processed, in the case where the multi-camera has completed calibration, the human figure regions in the video pictures of the corresponding multiple video streams to be processed can be mapped to the three-dimensional space confirmed by the multi-camera according to the pre-stored calibration information; then the relative position parameters of the real people corresponding to the human figure regions in the three-dimensional space are confirmed. For a multi-camera that has completed calibration, the depth information of the target can be detected, the real world can be virtually restored to the three-dimensional space, and the spatial positions of the real people corresponding to each human figure can be obtained in the three-dimensional space, i.e., the relative position parameters are obtained, and accurate and real relative position parameters are obtained.
[0160] For a multi-channel video stream collected by a multi-camera, not every video stream can be sent to the remote end for display. The video stream collected by the main camera of the multi-camera can be taken as the display target, the display area to be displayed can be confirmed from the human figure region in the video stream collected by the main camera, the display layout of the sub-picture window corresponding to each display area to be displayed can be confirmed according to the relative position parameters of the display area to be displayed, and the corresponding area picture can be displayed in each sub-picture window. The main camera is a camera arranged in a communication space (such as a conference room or a classroom) and can capture as much as possible the front face or side face of all real people in a multi-person communication scene. In the case where the relative position parameters have been confirmed, the display area to be displayed can be confirmed according to the size of the human figure, and the sub-picture window corresponding to the display area to be displayed can be confirmed according to the relative position parameters. The process can be referred to the description in the foregoing specific implementation process. Through the video stream collected by the calibrated multi-camera, the live scene of multi-person communication can be proportionally restored in the three-dimensional space, and the accurate relative position relationship of the real people can be obtained. When the close-up display is performed through multiple sub-picture windows, the macroscopic distribution of the human figures is as close as possible to the communication live scene, and the immersion in the communication process is effectively improved.
[0161] Overall, the scheme confirms the relative position relationship of real people in the real scene based on the imaging principle of the image acquisition module and the final imaging effect by pre-stored device parameters and parameters obtained by processing the collected video pictures, and takes this as the correction result of the relative position relationship of the portrait area in the video picture. When displaying the close-up pictures corresponding to each portrait through the small window, the corrected relative position relationship closer to the real people on the scene is taken as the reference for the display layout of the small window displaying the close-up picture, so that the relative position relationship of the real people on the scene is as far as possible to be presented macroscopically when multiple small windows are displayed, and a better on-site immersion can be obtained when the audience remotely views the on-site picture through multiple small windows. The display effect changes of FIG. 7 and FIG. 18 can be compared in detail. FIG. 7 is based on FIG. 6, and the relative position relationship of the portrait area in the panoramic picture is directly used for the layout display of the sub-picture window, in which the effect of the portrait PE and the portrait PA talking back to back appears, which obviously affects the immersive experience of the remote user. Through the scheme, the relative position relationship of the real people (the positions corresponding to each real person are respectively denoted as PA', PB', PC' and PE') as shown in FIG. 17 can be obtained, and the relative position relationship of the real people is taken as the layout reference of the close-up picture displayed in the sub-picture window. Correspondingly, a macro display state more consistent with the on-site state is presented in the schematic diagram shown in FIG. 18, that is, more consistent with the conversation details shown in FIG. 6. The overall feeling of the remote user is that the portrait PE and the portrait PA are face-to-face.
[0162] FIG. 19 is a structural schematic diagram of a video processing device provided by an embodiment of the present application. As shown in FIG. 19, the video processing device includes a video acquisition unit 210, a parameter confirmation unit 220 and a layout display unit 230.
[0163] The video acquisition unit 210 is configured to acquire a to-be-processed video stream, and the to-be-processed video stream is obtained by an image acquisition module performing image acquisition on a multi-person scene. The parameter confirmation unit 220 is configured to confirm relative position parameters of corresponding real people according to portrait areas in a video picture of the to-be-processed video stream. The layout display unit 230 is configured to confirm to-be-displayed areas from the portrait areas, confirm a display layout of a sub-picture window corresponding to each to-be-displayed area according to relative position parameters of the to-be-displayed areas, and display a corresponding area picture in each sub-picture window. The area picture is a local picture including the corresponding to-be-displayed area in the video picture of the to-be-processed video stream.
[0164] On the basis of the above-mentioned embodiments, the to-be-processed video stream is a video stream collected by a monocular camera.
[0165] Correspondingly, the parameter confirmation unit 220 includes:
[0166] The distance-angle confirmation module is configured to identify a portrait area in a video picture of the video stream to be processed, and confirm a first distance value and a first angle value corresponding to each portrait area according to a parameter value of a preset dimension in the portrait area and a parameter value of the image acquisition module stored in advance. The first distance value is used to represent a relative size relationship between shooting distances corresponding to each portrait area. The shooting distance is a horizontal distance between a corresponding real person and the image acquisition module. The first angle value is used to represent an included angle of a projection of a line connecting the corresponding real person and the image acquisition module on a horizontal plane relative to a preset reference line.
[0167] The position confirmation module is configured to confirm a relative position parameter of a real person corresponding to each portrait area according to the first distance value and the first angle value.
[0168] On the basis of the above embodiment, the distance-angle confirmation module comprises:
[0169] The area identification sub-module is configured to identify a portrait area in a video picture of the video stream to be processed, and confirm a portrait size of each portrait area. The portrait size is confirmed according to a distance between at least two preset portrait reference points in the portrait area.
[0170] The distance confirmation sub-module is configured to confirm a first distance value corresponding to each portrait area according to each portrait size.
[0171] The angle confirmation sub-module is configured to confirm a first angle value corresponding to each portrait area according to a field of view angle of the image acquisition module recorded in advance, a picture size of the video stream to be processed and a position of a central reference point. The position of the central reference point is confirmed according to a distance between at least two portrait reference points.
[0172] On the basis of the above embodiment, the distance confirmation sub-module comprises:
[0173] The normalization sub-module is configured to sequentially divide a minimum size in the portrait size by each portrait size to obtain a first value sequence, sort the values in the first value sequence from small to large, and sequentially associate the values to the portrait areas sorted from large to small according to the portrait sizes as the first distance values of the corresponding portrait areas.
[0174] On the basis of the above embodiment, the position confirmation module comprises:
[0175] The coordinate conversion sub-module is configured to confirm a relative position coordinate of a real person corresponding to each portrait area in a same rectangular coordinate system according to the first distance value and the first angle value corresponding to each portrait area.
[0176] The video stream to be processed is a multi-path video stream captured by a multi-view camera.
[0177] Correspondingly, the parameter confirmation unit 220 comprises:
[0178] a three-dimensional mapping module, configured to map a human region in a video picture of a corresponding multi-path video stream to be processed to a three-dimensional space confirmed by the multi-view camera according to pre-stored calibration information of the multi-view camera;
[0179] a coordinate confirmation module, configured to confirm a relative position parameter of a real person corresponding to the human region in the three-dimensional space.
[0180] On the basis of the above embodiment, the layout display unit 240 comprises:
[0181] a main picture display module, configured to confirm a to-be-displayed region from the human region in the video stream collected by a main camera in the multi-view camera, confirm a display layout of a sub-picture window corresponding to each to-be-displayed region according to a relative position parameter of the to-be-displayed region, and display a corresponding region picture in each sub-picture window.
[0182] On the basis of the above embodiment, the layout display unit 240 comprises:
[0183] a size confirmation module, configured to confirm at least a human region in which a human size reaches a preset display condition as a to-be-displayed region.
[0184] On the basis of the above embodiment, the layout display unit 240 comprises:
[0185] a grouping confirmation module, configured to confirm a to-be-displayed region from the human region, confirm a grouping result of the to-be-displayed region according to a number of the to-be-displayed region, and display a sub-picture window corresponding to each group of to-be-displayed regions in a same row in the display layout.
[0186] an intra-group confirmation module, configured to confirm a display position of a sub-picture window corresponding to a to-be-displayed region in a same row according to a relative position parameter of the to-be-displayed region in the same group.
[0187] a picture display module, configured to display a region picture in a corresponding sub-picture window according to the display position.
[0188] On the basis of the above embodiment, the grouping confirmation module comprises:
[0189] a to-be-displayed confirmation submodule, configured to confirm a to-be-displayed region from the human region;
[0190] a first grouping module, configured to, in a case where a number of the to-be-displayed regions matches a target layout number in a preset array layout number, confirm a grouping manner of the to-be-displayed regions as an array manner corresponding to the target layout number.
[0191] The second grouping module is configured to, in a case where the number of the to-be-displayed areas does not match the target layout number in the preset array layout number, determine that the grouping manner of the to-be-displayed areas comprises grouping according to distances to the same reference point and grouping according to distances between the to-be-displayed areas.
[0192] On the basis of the above-described embodiments, the group determination module comprises:
[0193] The horizontal ordering sub-module is configured to, according to a size order of the horizontal position parameter in the relative position parameter of the to-be-displayed areas in the same group, sequentially arrange display positions of the sub-picture windows corresponding to the to-be-displayed areas in the same group in the same row.
[0194] The video processing apparatus provided by the embodiments of the present application is contained in an electronic device and can be used to execute the corresponding video processing method provided in the above-described embodiments, and has the corresponding functions and beneficial effects.
[0195] It should be noted that, in the embodiments of the above-described video processing apparatus, each unit and module included is only logically divided according to functions, but is not limited to the above-described division, as long as the corresponding functions can be implemented; in addition, the specific names of each functional unit are only for the convenience of mutual differentiation, and do not limit the protection scope of the present application.
[0196] FIG. 20 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. As shown in FIG. 20, the electronic device comprises a processor 310 and a memory 320, and can further comprise an input device 330, an output device 340 and a communication device 350; the number of the processor 310 in the electronic device can be one or more, and one processor 310 is taken as an example in FIG. 20; the processor 310, the memory 320, the input device 330, the output device 340 and the communication device 350 in the electronic device can be connected through a bus or other manners, and connection through a bus is taken as an example in FIG. 20.
[0197] The memory 320, as a kind of computer readable storage medium, can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the video processing method in the embodiments of the present application. The processor 310 executes various function applications and data processing of the electronic device by running the software programs, instructions and modules stored in the memory 320, that is, the above-described video processing method is implemented.
[0198] The memory 320 can include a program storage area and a data storage area, where the program storage area can store an operating system, application programs required by at least one function, and the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 320 can include a high-speed random access memory, and can further include a nonvolatile memory such as at least one of a magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some examples, the memory 320 can further include a memory remotely disposed with respect to the processor 310, which can be connected to the electronic device through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0199] The input device 330 can be used to receive input digital or character information, and to generate signal input related to user settings and function control of the electronic device. The output device 340 can include a display device such as a display screen.
[0200] The above electronic device contains a video processing device, which can be used to execute any video processing method, and has corresponding functions and beneficial effects.
[0201] The embodiments of the present application also provide a computer readable storage medium, which has a computer program stored thereon, where the computer program, when executed by a processor, is configured to perform the related operations in the video processing method provided in any of the embodiments of the present application, and has corresponding functions and beneficial effects.
[0202] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product.
[0203] Accordingly, embodiments of the present application can be embodied in the form of complete hardware embodiments, complete software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present application can take the form of a computer program product on one or more computer-readable storage media (including, but not limited to, disk memory, CD-ROMs, optical storage devices, etc.) embodying computer readable program code. Embodiments of the present application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing system, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams block or blocks. These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart illustrations and / or block diagrams block or blocks. The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart illustrations and / or block diagrams block or blocks.
[0204] In one typical arrangement, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. The memory can include non-persistent memory in the form of random access memory (RAM), and / or non-volatile memory such as read only memory (ROM) or flash memory, among others, in a computer readable medium. The memory is an example of a computer readable medium.
[0205] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can implement information storage by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.
[0206] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusions, so that processes, methods, articles or devices that include a series of elements not only include those elements, but also include other elements not explicitly listed, or inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.
[0207] Note that the above are only the preferred embodiments of the present application and the principles of technology applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments herein, and those skilled in the art can make various obvious changes, readjustments and substitutions without departing from the scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the appended claims.
Claims
A video processing method, wherein, The method comprises: acquiring a to-be-processed video stream, the to-be-processed video stream being obtained by an image acquisition module performing image acquisition on a multi-person scene; confirming relative position parameters of corresponding real persons according to a person region in a video picture of the to-be-processed video stream; confirming a to-be-displayed region from the person region, confirming a display layout of a sub-picture window corresponding to each to-be-displayed region according to the relative position parameters of the to-be-displayed region, and displaying a corresponding region picture in each sub-picture window, the region picture being a local picture including the corresponding to-be-displayed region in the video picture of the to-be-processed video stream. The video processing method of claim 1, wherein, The to-be-processed video stream is a video stream captured by a monocular camera. Correspondingly, the step of confirming the relative position parameters of the corresponding real persons according to the person region in the video picture of the to-be-processed video stream comprises: identifying the person region in the video picture of the to-be-processed video stream, and confirming a first distance value and a first angle value corresponding to each person region according to a parameter value of a preset dimension in the person region and a parameter value of a pre-stored image acquisition module, the first distance value being used to represent a relative size relationship between corresponding shooting distances of each person region, the shooting distance being a horizontal distance between the corresponding real person and the image acquisition module, and the first angle value being used to represent an included angle of a projection of a line connecting the corresponding real person and the image acquisition module in a horizontal plane relative to a preset reference line; confirming the relative position parameters of the real persons corresponding to each person region according to the first distance value and the first angle value. The video processing method of claim 2, wherein, The step of identifying the person region in the video picture of the to-be-processed video stream and confirming the first distance value and the first angle value corresponding to each person region according to the parameter value of the preset dimension in the person region and the parameter value of the pre-stored image acquisition module comprises: identifying the person region in the video picture of the to-be-processed video stream, and confirming a person size of each person region, the person size being confirmed according to a distance between at least two preset person reference points in the person region; confirming the first distance value corresponding to each person region according to each person size; confirming the first angle value corresponding to each person region according to a pre-recorded field of view angle of the image acquisition module, a picture size of the to-be-processed video stream, and a position of a center reference point, the position of the center reference point being confirmed according to the distance between the at least two person reference points. The video processing method of claim 3, wherein, The step of confirming the first distance value corresponding to each person region according to each person size comprises: dividing each person size by a minimum size in the person size to obtain a first numerical sequence, sorting the numerical values in the first numerical sequence from small to large, and sequentially associating the numerical values to the person regions sorted from large to small according to the person size as the first distance value of the corresponding person region. The video processing method according to any one of claims 2-4, wherein, The step of confirming the relative position parameters of the real persons corresponding to each person region according to the first distance value and the first angle value comprises: confirming relative position coordinates of the real persons corresponding to each person region in a same rectangular coordinate system according to the first distance value and the first angle value corresponding to each person region. The video processing method of claim 1, wherein, The to-be-processed video stream is a multi-path video stream captured by a multi-lens camera. Correspondingly, the relative position parameter of the real person corresponding to the portrait area in the video picture of the to-be-processed video stream is confirmed, and the relative position parameter of the real person corresponding to the portrait area in the video picture of the to-be-processed video stream is confirmed. According to the pre-stored calibration information of the multi-view camera, the portrait area in the video picture of the corresponding multi-channel to-be-processed video stream is mapped to the three-dimensional space confirmed by the multi-view camera; Confirm the relative position parameter of the real person corresponding to the portrait area in the three-dimensional space. The video processing method of claim 6, wherein, The display area is confirmed from the portrait area, the display layout of the sub-picture window corresponding to each display area is confirmed according to the relative position parameter of the display area, and the corresponding area picture is displayed in each sub-picture window, comprising: Confirm the display area from the portrait area of the video stream collected by the main camera in the multi-view camera, confirm the display layout of the sub-picture window corresponding to each display area according to the relative position parameter of the display area, and display the corresponding area picture in each sub-picture window. The video processing method according to any one of claims 1-4, 6 and 7, wherein, The display area is confirmed from the portrait area, the display layout of the sub-picture window corresponding to each display area is confirmed according to the relative position parameter of the display area, and the corresponding area picture is displayed in each sub-picture window, comprising: Confirm at least one portrait area with a preset display condition as a display area. The video processing method according to any one of claims 1-4, 6 and 7, wherein, The display area is confirmed from the portrait area, the display layout of the sub-picture window corresponding to each display area is confirmed according to the relative position parameter of the display area, and the corresponding area picture is displayed in each sub-picture window, comprising: Confirm the display area from the portrait area, confirm the grouping result of the display area according to the number of the display area, and the sub-picture window corresponding to each group of display areas is displayed in the same row in the display layout; According to the relative position parameter of the display area in the same group, the display position of the sub-picture window corresponding to the display area in the same group in the same row is confirmed; The area picture is displayed in the corresponding sub-picture window according to the display position. The video processing method of claim 9, wherein, The display area is confirmed from the portrait area, the grouping result of the display area is confirmed according to the number of the display area, comprising: Confirm the display area from the portrait area; In the case that the number of the display area matches the target layout number in the preset array layout number, the grouping mode of the display area is confirmed as the array mode corresponding to the target layout number; In the case that the number of the display area does not match the target layout number in the preset array layout number, the grouping mode of the display area includes grouping according to the distance from the same reference point and grouping according to the distance between the display areas. The video processing method of claim 10, wherein, According to the relative position relationship of the display area in the same group, the display position of the sub-picture window corresponding to the display area in the same group in the same row is confirmed, comprising: According to the size order of the horizontal position parameter in the relative position parameter of the display area in the same group, the display position of the sub-picture window corresponding to the display area in the same group in the same row is arranged in sequence. A video processing device, wherein, Comprising: A video acquisition unit is configured to acquire a to-be-processed video stream, wherein the to-be-processed video stream is obtained by an image acquisition module performing image acquisition on a multi-person scene; A parameter confirming unit is configured to confirm a relative position parameter of a real person according to a portrait area in a video picture of the video stream to be processed. An electronic device, wherein, comprising: one or more processors; a memory storing one or more computer programs; when the one or more computer programs are executed by the one or more processors, the electronic device implements the video processing method as claimed in any one of claims 1-11. A computer-readable storage medium having stored thereon a computer program, wherein, The computer program is executed by the processor to implement the video processing method as claimed in any one of claims 1-11.
Citation Information
Patent Citations
Face space positioning method, device and system and computer storage medium
CN108875529A
Target positioning method and device, electronic equipment and storage medium
CN110889873A
Target tracking method and device and storage medium
CN112955712A
Trolley indoor positioning method based on multiple cameras
CN117635683A
Character pane layout method and device based on video picture and terminal equipment
CN117640876A