Video processing method and device, electronic equipment and storage medium

By detecting changes in the position and distance of people in the video frame, the cropping area is dynamically adjusted to highlight the main subject's image, solving the problem of unclear people in video calls caused by the wide angle of the camera and achieving better video call clarity.

CN119232998BActive Publication Date: 2025-11-04HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310786977.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2025-11-04
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

During video calls, the wide-angle lens of the camera makes people appear smaller, making it difficult to clearly see the other person's video feed.

Method used

By detecting changes in the position and distance of people in the video frame, the cropping area is dynamically adjusted to highlight the main character's image, thus achieving focus on the main character.

Benefits of technology

It makes it easier to see the main person in a video call, improving the clarity of the video call and the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119232998B_ABST
    Figure CN119232998B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a video processing method and device, electronic equipment and storage medium, relating to the technical field of video processing, which can make it easier to see the main character's portrait. The video processing method comprises: displaying a first video picture in a first period, the distance between a first user and a camera is less than the distance between a second user and the camera in the first period, and the distance between the first user and the camera is less than the distance between a third user and the camera, and the first video picture comprises a portrait of the first user; displaying a second video picture after the first user moves away from the camera, the second video picture comprising portraits of the second user and the third user; and displaying a third video picture after a fourth user moves into the camera capture range, the third video picture comprising portraits of the first user, the second user, the third user and the fourth user.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video processing, in particular to a video processing method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the development of technology, video shooting applications are used in more and more scenarios. In some scenarios, a single video picture cannot meet the requirements. For example, when making a video call through a television, the wide-angle of the camera makes the picture of the camera large, and the person in the picture small, so that the other party cannot clearly see the person in the video call. SUMMARY

[0003] The technical solution of the present application provides a video processing method, device, electronic equipment and storage medium, which can more easily see the main character.

[0004] In a first aspect, a video processing method is provided, including: displaying a first video picture in a first period, the distance between a first user and a camera being less than the distance between a second user and the camera in the first period, and the distance between the first user and the camera being less than the distance between a third user and the camera in the first period, the first video picture including a portrait of the first user; displaying a second video picture after the first user moves away from the camera, the second video picture including portraits of the second user and the third user; displaying a third video picture after a fourth user moves into a capture range of the camera, the distance between the fourth user and the first user being less than the distance between the second user and the first user, and the distance between the fourth user and the second user being greater than the distance between the second user and the third user, the third video picture including portraits of the first user, the second user, the third user and the fourth user.

[0005] In a possible implementation, in the first period, the first user is located in a third area, and the second user and the third user are located in a fourth area; displaying the first video picture in the first period includes: based on the distance between the first user and the camera being less than the distance between the second user and the camera, and the distance between the first user and the camera being less than the distance between the third user and the camera in the first period, displaying the first video picture.

[0006] In a possible implementation, displaying the second video picture after the first user moves away from the camera includes: after the first user moves away from the camera, in a second period, based on the distance between the first user and the camera being equal to the distance between the second user and the camera, and the distance between the first user and the camera being equal to the distance between the third user and the camera in the second period, the number of users in the fourth area being greater than the number of users in the third area, displaying the second video picture.

[0007] In a possible implementation, during a process from the first time period to the second time period, the video display interface gradually changes from displaying the first video picture to displaying the second video picture.

[0008] In a possible implementation, during the first time period, displaying the first video picture comprises: obtaining a second target region including at least one person image in an original video picture captured by the camera in the first time period; obtaining a third target region in the first time period based on the second target region in the first time period and a third target region in a time period adjacent to the first time period, wherein coordinates of the third target region in the first time period are closer to coordinates of the second target region in the first time period relative to coordinates of the third target region in the time period adjacent to the first time period, or the coordinates of the third target region in the first time period are equal to the coordinates of the second target region in the first time period; performing cropping on the original video picture captured by the camera in the first time period based on the third target region in the first time period to obtain a cropped first video picture, and displaying the first video picture.

[0009] In a possible implementation, the video processing method further comprises: determining a target person image in the person image in the original video picture captured by the camera in the first time period, and obtaining a first target region in the first time period, wherein the first target region in the first time period is a minimum rectangular frame including all target person images; obtaining the second target region including at least one person image in the original video picture captured by the camera in the first time period comprises: if a non-calculation condition is met, taking the second target region in a time period adjacent to the first time period as the second target region in the first time period; or if the non-calculation condition is not met, obtaining the second target region including the first target region in the original video picture in the first time period, wherein a size of the second target region in the first time period is greater than a size of the first target region in the first time period; the non-calculation condition comprises at least one or any combination of the following conditions, when the non-calculation condition comprises multiple conditions, any one of the multiple conditions not being met means that the non-calculation condition is not met, and all of the multiple conditions being met means that the non-calculation condition is met: a preset resolution level to which the first target region in the time period adjacent to the first time period belongs is the same as a preset resolution level to which the first target region in the first time period belongs; the first target region in the first time period is covered by the second target region in the time period adjacent to the first time period; and a number of target person images in the first time period is equal to a number of target person images in the time period adjacent to the first time period.

[0010] In a possible implementation, in the first time period, the first user is located in the third area, and the second user and the third user are located in the fourth area; the closest user to the camera in the original video picture corresponds to a first portrait, and other portraits are second portraits; the first portrait and all the second portraits with a distance difference from the first portrait less than a preset distance are to-be-determined portraits; the original video picture includes a first area corresponding to the third area and a second area corresponding to the fourth area; and determining the target portrait from the portraits in the original video picture collected by the camera in the first time period includes: if the number of to-be-determined portraits in the first area and the second area of the original video picture collected by the camera in the first time period is not equal, determining all the to-be-determined portraits in the area with more to-be-determined portraits as the target portrait; or if the number of to-be-determined portraits in the first area and the second area of the original video picture collected by the camera in the first time period is equal, determining all the to-be-determined portraits in the original video picture collected by the camera in the first time period as the target portrait.

[0011] In a possible implementation, based on the second target area of the first time period and the third target area of the last time period adjacent to the first time period, the process of obtaining the third target area of the first time period includes: if the second target area of the first time period is reduced relative to the third target area of the last time period adjacent to the first time period, reducing the third target area of the last time period adjacent to the first time period to obtain the third target area of the first time period; if the second target area of the first time period is the same in size relative to the third target area of the last time period adjacent to the first time period, and the center coordinate position of the second target area of the first time period is different relative to the third target area of the last time period adjacent to the first time period, moving the third target area of the last time period adjacent to the first time period to the direction close to the second target area of the first time period to obtain the third target area of the first time period; or if the second target area of the first time period is enlarged relative to the third target area of the last time period adjacent to the first time period, and the center coordinate position of the second target area of the first time period is the same relative to the third target area of the last time period adjacent to the first time period, enlarging the third target area of the last time period adjacent to the first time period to obtain the third target area of the first time period.

[0012] In a possible implementation, the process of reducing the third target region of the last time period adjacent to the first time period to obtain the third target region of the first time period includes: reducing the 10 times of the coordinate values of the third target region of the last time period adjacent to the first time period by the 10 times of the step value by means of the integer int to obtain the 10 times of the coordinate values of the third target region of the first time period; the process of moving the third target region of the last time period adjacent to the first time period to the direction close to the second target region of the first time period to obtain the third target region of the first time period includes: moving the 10 times of the coordinate values of the third target region of the last time period adjacent to the first time period to the direction close to the 10 times of the coordinate values of the second target region of the first time period by the 10 times of the step value by means of the integer int to obtain the 10 times of the coordinate values of the third target region of the first time period; the process of enlarging the third target region of the last time period adjacent to the first time period to obtain the third target region of the first time period includes: enlarging the 10 times of the coordinate values of the third target region of the last time period adjacent to the first time period by the 10 times of the step value by means of the integer int to obtain the 10 times of the coordinate values of the third target region of the first time period; the process of cropping the original video picture of the first time period based on the third target region of the first time period to obtain the cropped picture includes: dividing the 10 times of the coordinate values of the third target region of the first time period by 10 to obtain the coordinate values of the third target region of the first time period, and cropping the original video picture of the first time period based on the coordinate values of the third target region of the first time period by means of the floating point to obtain the cropped picture.

[0013] In a possible implementation, the video processing method further includes: obtaining distances between each user and the camera in the original video picture of the current time period captured by the camera; obtaining the number of the first region and the number of the second region in the original video picture of the current time period; and any one of displaying the first video picture, displaying the second video picture, and displaying the third video picture includes: displaying the video picture based on the distance between each user and the camera, the number of the first region, and the number of the second region.

[0014] In a possible implementation, the video processing method further includes: obtaining a second target region including at least one person image in the original video picture of the current period; obtaining a third target region of the current period based on the second target region of the current period and the third target region of the previous period, the coordinates of the third target region of the current period are closer to the coordinates of the second target region of the current period relative to the coordinates of the third target region of the previous period, or the coordinates of the third target region of the current period are equal to the coordinates of the second target region of the current period; performing cropping on the original video picture of the current period based on the third target region of the current period to obtain a cropped picture, any one of the first video picture, the second video picture and the third video picture is the cropped picture; in the continuous multiple periods, if the coordinates of the second target region are unchanged, the coordinates of the third target region gradually approach the second target region.

[0015] In a possible implementation, the video processing method further includes: determining a target person image in the person image in the original video picture of the current period, obtaining a first target region of the current period, the first target region of the current period being a minimum rectangular frame including all target person images; obtaining a second target region including at least one person image in the original video picture of the current period includes: if a non-calculation condition is met, taking the second target region of the previous period as the second target region of the current period; if the non-calculation condition is not met, obtaining the second target region including the first target region of the current period in the original video picture of the current period, the size of the second target region of the current period being greater than the size of the first target region of the current period; the non-calculation condition includes at least one or any combination of the following, when the non-calculation condition includes multiple items, any one of the multiple items not being met means that the non-calculation condition is not met, and all of the multiple items being met means that the non-calculation condition is met: the preset resolution level to which the first target region of the previous period belongs is the same as the preset resolution level to which the first target region of the current period belongs; the first target region of the current period is covered by the second target region of the previous period; the number of target person images of the current period is equal to the number of target person images of the previous period.

[0016] In a possible implementation, in the original video picture, the portrait of the user closest to the camera corresponds to a first portrait, and other portraits correspond to second portraits; the first portrait and all the second portraits with a distance difference from the first portrait less than a preset distance are to-be-determined portraits; and the target portrait is determined from the portraits in the original video picture of the current time period, including: if the number of to-be-determined portraits in the first region and the second region of the original video picture of the current time period is not equal, all the to-be-determined portraits in the region with more to-be-determined portraits are determined as the target portrait; and if the number of to-be-determined portraits in the first region and the second region of the original video picture of the current time period is equal, all the to-be-determined portraits in the original video picture of the current time period are determined as the target portrait.

[0017] In a possible implementation, based on the second target region of the current time period and the third target region of the last time period, the third target region of the current time period is obtained, including: if the second target region of the current time period is reduced relative to the third target region of the last time period, the third target region of the last time period is reduced to obtain the third target region of the current time period; if the second target region of the current time period has the same size relative to the third target region of the last time period, and the center coordinate position of the second target region of the current time period relative to the third target region of the last time period is different, the third target region of the last time period is moved to the direction close to the second target region of the current time period to obtain the third target region of the current time period; and if the second target region of the current time period is enlarged relative to the third target region of the last time period, and the center coordinate position of the second target region of the current time period relative to the third target region of the last time period is the same, the third target region of the last time period is enlarged to obtain the third target region of the current time period.

[0018] In a possible implementation, the video processing method further includes: the process of reducing the third target region of the last time period to obtain the third target region of the current time period includes: reducing the 10 times of the coordinate values of the third target region of the last time period by the 10 times of the step value by an integer int to obtain the 10 times of the coordinate values of the third target region of the current time period; the process of moving the third target region of the last time period to the direction close to the second target region of the current time period to obtain the third target region of the current time period includes: moving the 10 times of the coordinate values of the third target region of the last time period to the direction close to the 10 times of the coordinate values of the second target region of the current time period by the integer int based on the 10 times of the step value to obtain the 10 times of the coordinate values of the third target region of the current time period; the process of enlarging the third target region of the last time period to obtain the third target region of the current time period includes: enlarging the 10 times of the coordinate values of the third target region of the last time period by the integer int based on the 10 times of the step value to obtain the 10 times of the coordinate values of the third target region of the current time period; and the process of cropping the original video picture of the current time period based on the third target region of the current time period to obtain the cropped picture includes: dividing the 10 times of the coordinate values of the third target region of the current time period by 10 to obtain the coordinate values of the third target region of the current time period, and cropping the original video picture of the current time period based on the coordinate values of the third target region of the current time period by a floating point to obtain the cropped picture.

[0019] In a possible implementation, the width-length ratio of the third target region is the same as the width-length ratio of the original video picture.

[0020] In a possible implementation, the displaying the first video picture includes displaying the first video picture in at least one or any combination of the video call notification interface, the video call dialing interface, and the video call interface.

[0021] In a second aspect, a video processing apparatus is provided, including: a first display module configured to display a first video picture in a first time period, in the first time period, a distance between a first user and a camera is less than a distance between a second user and the camera, and the distance between the first user and the camera is less than a distance between a third user and the camera, and the first video picture includes a portrait of the first user; a second display module configured to display a second video picture after the first user moves away from the camera, the second video picture including portraits of the second user and the third user; and a third display module configured to display a third video picture after a fourth user moves into a capturing range of the camera, a distance between the fourth user and the first user is less than a distance between the second user and the first user, and a distance between the fourth user and the second user is greater than a distance between the second user and the third user, and the third video picture includes portraits of the first user, the second user, the third user, and the fourth user.

[0022] In a third aspect, an electronic device is provided, comprising a processor and a memory, the memory being configured to store at least one instruction, the instruction being loaded and executed by the processor to enable the electronic device to perform the video processing method described above.

[0023] In a fourth aspect, a computer readable storage medium is provided, comprising a program or instructions, the method described above being performed when the program or instructions are run on a computer.

[0024] The video processing method, device, electronic device and storage medium in the embodiments of the present application re-determine the main character based on the number of people corresponding to the human image in the video picture and the change of the camera distance, and track and enlarge the newly determined main character through the cropped picture, thereby highlighting the main human image in the video picture and realizing the focusing on the main character, which can make it easier to see the main human image. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application;

[0026] Figure 2 FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application;

[0027] Figure 3 FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application;

[0028] Figure 4 FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application;

[0029] Figure 5 FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application; Figure 3 FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application;

[0030] Figure 6 FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application; Figure 3 FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application;

[0031] Figure 7 FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application; Figure 6 FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application;

[0032] Figure 8 FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application; Figure 6 FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application;

[0033] Figure 9 FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application; Figure 8 FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application;

[0034] Figure 10 FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application;

[0035] FIG. 1 is a structural block diagram of an electronic device in an embodiment of the present application;Figure 11 Figure 8 is a flowchart illustrating a specific process of another part of the steps in the embodiment of the present application; Figure 6 Figure 9 is a flowchart illustrating a specific process of another part of the steps in the embodiment of the present application;

[0036] Figure 12 Figure 10 is a schematic diagram of one of the embodiments of the present application in which the third target area gradually changes in a plurality of continuous time periods;

[0037] Figure 13 Figure 11 is a schematic diagram of a user interface change in the embodiment of the present application;

[0038] Figure 14 Figure 12 is a schematic diagram of an original video screen in the embodiment of the present application;

[0039] Figure 15 Figure 13 is a schematic diagram of another original video screen in the embodiment of the present application;

[0040] Figure 16 Figure 14 is a schematic diagram of yet another original video screen in the embodiment of the present application;

[0041] Figure 17 Figure 15 is a schematic diagram of still another original video screen in the embodiment of the present application. DETAILED DESCRIPTION

[0042] The terms used in the embodiment part of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.

[0043] Figure 1 A structural schematic diagram of an electronic device 100 is shown.

[0044] The electronic device 100 can include a processor 110, an internal memory 121, an audio module 170, a speaker 170A, a microphone 170C, a camera 193, a display screen 194, etc.

[0045] It can be understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than illustrated, or combine certain components, or split certain components, or different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0046] The processor 110 can include one or more processing units such as: an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices or integrated in one or more processors.

[0047] The controller can generate operation control signals according to instruction operation codes and timing signals to complete the control of fetching and executing instructions.

[0048] The communication function of the electronic device 100 can be realized in a wired or wireless manner.

[0049] The electronic device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.

[0050] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 can include 1 or N display screens 194, and N is a positive integer greater than 1.

[0051] The electronic device 100 can implement a photographing function through an ISP, a camera 193, a video codec, a GPU, a display 194, and an application processor, etc.

[0052] The camera 193 is used to capture still images or videos. An object projects an optical image through a lens to a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to an ISP to convert into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB, YUV, etc. In some embodiments, the electronic device 100 can include one or N cameras 193, where N is a positive integer greater than 1.

[0053] The internal memory 121 can be used to store computer executable program code, which includes instructions.

[0054] The audio module 170 is used to convert digital audio information into an analog audio signal output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.

[0055] The speaker 170A, also known as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or make a call through the speaker 170A.

[0056] The microphone 170C, also known as a "microphone", "sounder", is used to convert a sound signal into an electrical signal. When making a call, a user can input a sound signal to the microphone 170C.

[0057] The following describes the process of applying the embodiments of the present application in the video call scenario of the electronic device, taking a smart television as an example.

[0058] First, the embodiments of the present application are described in combination with a software architecture, and the software structure of the electronic device 100 is described by way of example. Figure 2 is a software structure block diagram of the electronic device 100 of the embodiments of the present application.

[0059] The layered architecture divides software into several layers, each of which has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, a software system includes, from top to bottom, an Application layer, a framework layer, and a Hardware Abstraction Layer (HAL). The Application layer can include applications such as a camera, a video call application, and the like. The framework layer can include CameraService and the like. The Hardware Abstraction Layer is used to process video streams captured by a camera, and can include an image acquisition module, an algorithm detection module, and an image cropping module, each of which can be a process or a thread.

[0060] As Figure 2 And Figure 3As shown, the embodiment of the present application provides a video processing method, which can be applied to the electronic device described above. The method can include the following steps. Step 101, the video call application sends a stream configuration request to the camera service module in response to a video call request or a video call notification request. The video call request is generated by the event that the user initiates a call actively, and the video call notification request is generated by the event that a call from the opposite party is received. The stream configuration request is used to request to open the camera and configure the frame rate and resolution of the camera and other video stream parameters. In some embodiments, the stream configuration request is sent to the camera service module in response to the video call notification request, so as to trigger the subsequent acquisition of the video stream shot by the camera to realize the video preview function. In other embodiments, the video call notification request in step 101 can be replaced by a video call connection request, that is, the video stream shot by the camera does not need to be acquired before the user connects the video call, but is acquired when the call is connected. In step 101, the video call application can specifically call the configstream function to send the stream configuration request to the camera service module. Step 102, the camera service module sends the stream configuration information to the image acquisition module in the hardware abstraction layer in response to the configuration request. Step 103, the image acquisition module performs stream configuration in response to the stream configuration information. The stream configuration process includes configuring the frame rate and resolution of the camera. After step 103, step 104 is performed, in which the video call application sends a video stream acquisition request to the camera service module. The video call application can specifically call the RepeatRequest function to issue the video stream acquisition request to the camera service module. Step 105, the camera service module sends a video stream acquisition request to the image acquisition module in response to the video stream acquisition request. Step 106, the image acquisition module sends a video stream acquisition request to the camera in response to the video stream acquisition request. Step 107, the image acquisition module acquires the original video picture, specifically acquires the original video picture collected by the camera periodically in time periods, for example, every 30 ms is a time period, the image acquisition module sends a video stream acquisition request to the camera every 30 ms, and acquires the original video picture of the current time period from the camera, that is, one image acquisition is performed. The collected image is subjected to image decoding, and the decoded original video picture is provided to the algorithm detection module, that is, step 108 is performed, in which the algorithm detection module acquires the original video picture, for example, as shown in the figure Figure 4As shown, the original video picture of the current period acquired from the camera is a 1920*1024 h264 encoded image, and is a 1920*1024 YUYV format image after decoding. Step 109, the algorithm detection module performs portrait detection based on the original video picture of the current period to obtain portrait detection information, and sends the portrait detection information to the image cropping module. Step 110, the image cropping module crops the original video picture based on the portrait detection information. In each period, the image cropping module acquires a third target region C of the current period based on the portrait detection information of the current period, where the third target region C refers to an actual cropping frame of the current picture, and the acquisition of the third target region C refers to the acquisition of the coordinates of the left frame, the right frame, the upper frame and the lower frame corresponding to the third target region C, and then the original video picture of the current period is cropped based on the third target region C. Step 110, the camera service module acquires the cropped picture. Step 111, the video call application acquires the cropped picture. The picture finally displayed by the video call application is the picture obtained by cropping the original video picture based on the third target region of the current period, and the size of the cropped picture is usually smaller than that of the original video picture, so the size is adjusted to 1080P, and the cropped picture is transmitted to the video call application for display. That is, in each period, the image cropping module crops the original video picture acquired in the current period based on the third target region of the current period and transmits the cropped picture to the video call application for display. The above steps 107-111 can be periodically executed, that is, portrait detection is performed based on the original video picture of the current period in each period, and cropping is performed based on the portrait detection information of the current period.

[0061] The above steps 101-106 will be described in detail below. As shown in Figure 5 As shown, the electronic device can include an application package 401 of a first application, a camera manager (such as CameraManager) 402, a camera device instance (such as CameraDeviceImpl) 403, a camera device client (such as CameraDeviceClient) 404, a camera device (such as Camera3Device) 405, and a camera device session (such as CameraDeviceSession) 406.

[0062] In one embodiment, the first application can be a video call application corresponding to a video call scenario. In another embodiment, the first application can be a camera application corresponding to a camera preview scenario or a camera recording scenario.

[0063] In one embodiment, the application package 401 can be an APK (Android application package) of the first application.

[0064] In one embodiment, the Camera 2 API (Application Programming Interface) can include a camera manager 402 and a camera device instance 403.

[0065] In one embodiment, the CameraService (Camera Service module) in the application framework layer of the electronic device can include a camera device client 404 and a camera device 405.

[0066] In one embodiment, the CameraHal (Camera Hardware Abstraction Layer) of the electronic device can include a camera device session 406.

[0067] In one embodiment of the present application, the camera device session 406 can include an image acquisition module, an algorithm detection module, and an image cropping module. The module functions of the three modules are described in the related technical description of the embodiment shown in Figure 3

[0068] As shown in Figure 5 The flow of steps 101-106 described above can include the following contents:

[0069] Step 4.1, the application package 401 sends a request for obtaining a camera identity document (id) to the camera manager 402 in response to the first instruction, so that the camera manager 402 returns a list of camera identity documents.

[0070] The camera manager 402 can provide the application package 401 with a list of camera identity documents according to the received request for obtaining a camera identity document, so that the application package 401 obtains the camera identity document accordingly.

[0071] In one embodiment, the related implementation code for sending the request for obtaining the camera identity document to the camera manager 402 can include CameraManager.getCameraIdList().

[0072] In one embodiment of the present application, in the case where the first application is a video call application, the first instruction can be an instruction to start a video call. In one feasible implementation, the user can trigger a video initiation control in the video call application to send the instruction to start the video call.

[0073] ​Step 4.2, the application package 401 acquires the camera identifier according to the camera identifier list sent by the camera manager 402. In an embodiment, the code for acquiring the camera identifier can include: CameraManager.getCameraCharacteristics(CameraId).

[0074] In a feasible implementation, the acquired camera identifier can be the identifier of the camera used for image acquisition, and the image acquisition module in the camera device session 406 can acquire the image captured by the camera.

[0075] Step 4.3, the application package 401 determines whether the key (Key) of the corresponding camera exists according to the acquired camera identifier, and if the key exists, step 4.4 is executed.

[0076] In a feasible implementation, if it is determined that the key of the corresponding camera does not exist, step 4.4 can not be executed.

[0077] In an embodiment, the code for determining whether the key exists can include: mCharacteristics.getKeys().contains(HIHONOR_FACEAUTOTRACK_SUPPORTED)

[0078] KEY NAME: com.hihonor.device.capabilities.faceAutoTrackModeSupportedByte type.

[0079] Step 4.4, the application package 401 determines whether the corresponding application supports portrait tracking according to the acquired camera identifier, and if the corresponding application supports portrait tracking, the camera manager 402 opens the corresponding camera according to the key of the camera. The camera in the open state can sequentially acquire each frame of image in the image acquisition range of the camera.

[0080] In another embodiment of the present application, step 4.4 can be executed before step 4.3, that is, step 4.4 is executed when the determination of step 4.4 is passed, and step 4.5 is executed when the determination of step 4.3 is passed.

[0081] In an embodiment, the code for determining whether the corresponding application supports portrait tracking can include: mCharacteristics.get(HIHONOR_FACEAUTOTRACK_SUPPORTED) 0 does not support 1 supports byte type.

[0082] Since subsequent portrait detection can be performed on the image captured by the camera to realize user portrait tracking, and then image cropping is performed according to the portrait detection information to display the cropped image, if the judgment result is "not supported", the camera manager 402 is not triggered to open the corresponding camera, and if the judgment result is "supported", the camera manager 402 is triggered to open the corresponding camera.

[0083] In one embodiment, the related implementation code for triggering the camera manager 402 to open the corresponding camera can include: CameraManager.openCamera().

[0084] Step 4.5, the application package 401 sends an instruction for creating a capture request to the camera device instance 403. In one embodiment, the related implementation code of step 4.5 can include: CameraDevice.createCaptureRequest().

[0085] Step 4.6, the camera device instance 403 sends an instruction for creating a system setting request to the camera device client 404 according to the instruction for creating a capture request. In one embodiment, the related implementation code of step 4.6 can include: createDefaultRequest(templateType).

[0086] Step 4.7, the camera device client 404 passes the received instruction for creating a system setting request to the camera device 405.

[0087] Step 4.8, the camera device 405 sends an instruction for setting a system setting request to the camera device session 406 according to the instruction for creating a system setting request. In one embodiment, the related implementation code of step 4.8 can include: constructDefaultRequestSettings.

[0088] Step 4.9, the camera device session 406 obtains camera metadata and the current application package name according to the instruction for setting a system setting request, and returns the obtained camera metadata and the current application package name to the camera device 405.

[0089] In one embodiment, the related implementation code of step 4.9 can include: com.hihonor.capture.metadata.packageName.

[0090] In one embodiment, the camera metadata obtained by the camera device session 406 can be the metadata of the opened camera.

[0091] In one embodiment, the current application package name can be the package name of the application package 401 of the first application.

[0092] Step 4.10, the camera device 405 adds the current application package name into the camera metadata, and sends the camera metadata added with the current application package name to the camera device client 404. By adding the current application package name into the camera metadata, the scene information of the current period is brought in.

[0093] Step 4.11, the camera device client 404 delivers the camera metadata added with the current application package name to the application package 401.

[0094] By delivering the camera metadata added with the current application package name to the application package 401, the application package 401, which is a CaptureRequest.Builder, can contain the package name and the device capability of the camera.

[0095] After the application package 401 obtains the camera metadata with the scene information, the process of flow allocation can be started (see steps 4.12-4.16).

[0096] Step 4.12, the application package 401 generates an instruction of creating a capture session according to the camera metadata added with the current application package name, and sends the instruction of creating a capture session to the camera device instance 403.

[0097] In one embodiment, the related implementation code of step 4.12 can include CameraDevice.CreateCaptureSession(SessionConfiguration config).

[0098] Step 4.13, the camera device instance 403 issues interface flow allocation information and information of configuring a session key to the camera device client 404 according to the instruction of creating a capture session.

[0099] In one embodiment, the flow allocation information can include frame rate and resolution. The flow allocation module of the Hal layer can configure the flow according to the flow allocation information, and configure the frame rate and the resolution for the camera.

[0100] Step 4.14, the camera device client 404 processes the received interface flow allocation information and information of configuring a session key, and sends the processed session key and the flow allocation information corresponding to the camera Hal layer to the camera device 405.

[0101] Step 4.15, the camera device 405 sends the session key and the flow allocation information corresponding to the camera Hal layer to the camera device session 406.

[0102] Step 4.16, the camera device session 406 performs pipeline processing according to the session key and the pipeline information of the corresponding camera hardware abstraction layer, and sends a pipeline completion notification to the application package 401 after the pipeline processing is completed.

[0103] In an embodiment, the pipeline processing can include mode selection and pipeline configuration.

[0104] In an embodiment, information "an implementation of an embodiment" can be sent to the application package 401 after the pipeline processing is completed, so that the application package 401 knows that the pipeline processing has been completed.

[0105] In an embodiment, the technical implementation of pipeline processing can refer to the related technical content in Figure 3 The embodiment will not be repeated here.

[0106] After the camera device session 406 completes the pipeline processing, the application package 401 can issue a repeated request to trigger the camera device session 406 to perform image processing (such as algorithm detection and cropping processing) on the image captured by the camera through the camera service module.

[0107] Step 4.17, the application package 401 sends an instruction for setting a repeated request to the camera device instance 403. In an embodiment, the related implementation code of step 4.17 can include: captureSession.setRepeatingRequest.

[0108] Step 4.18, the camera device instance 403 sends an instruction for submitting a request list (such as marked as submitRequesList) to the camera device client 404 according to the instruction for setting a repeated request.

[0109] Step 4.19, the camera device client 404 sends an instruction for setting a streaming request list (such as marked as setStreamingRequestList) to the camera device 405 according to the instruction for submitting a request list.

[0110] Step 4.20, the camera device 405 issues an image capture request to the camera device session 406 according to the instruction for setting a streaming request list.

[0111] Step 4.21, the camera device session 406 captures an image frame captured by the camera according to the image capture request, so as to subsequently process the captured image to obtain a processed image, and sends the processed image to the application package 401 for display.

[0112] In an embodiment of the present application, the camera device session 406 can include an image acquisition module, an algorithm detection module and an image cropping module. The image acquisition module acquires images from the opened camera, and the acquired images are put into an image buffer queue. The algorithm detection module and the image cropping module asynchronously acquire images from the image buffer queue.

[0113] The algorithm detection module performs portrait detection on the acquired images, and the generated portrait detection information is stored in a portrait detection information buffer queue. After the image cropping module acquires an image, it attempts to acquire portrait detection information from the portrait detection information buffer queue. If the portrait detection information is acquired, the image is cropped according to the portrait detection information, and the cropped image is taken as the capture result (for example, denoted as CaptureResult). If the portrait detection information is not acquired, the acquired image is directly taken as the capture result without cropping, or the image is cropped according to the portrait detection information used in the last cropping, and the cropped image is taken as the capture result.

[0114] The image cropping module further returns the capture result to the application package 401 for display. In this way, the camera device session 406 completes the processing (for example, denoted as processcapturerequest) of an image acquisition request. The specific implementation of image processing is described in detail in the subsequent content.

[0115] The specific processes of steps 109 and 110 are described below. Before step 109 is executed, the original video picture of the current period has been acquired, as shown in FIG. 1. Figure 4 As shown, a coordinate system is defined based on the boundaries of the original video picture. The coordinates of any position in the coordinate system are (x, y), x represents the horizontal coordinate, and y represents the vertical coordinate. The upper left corner of the picture is the origin coordinate (0, 0), the upper right corner of the picture is the coordinate (1920, 0), the lower left corner of the picture is the coordinate (0, 1080), and the lower right corner of the picture is the coordinate (1920, 1080). The coordinate unit can be pixels, that is, the size of the original video picture is 1920*1080 pixels, or the resolution of the original video picture is 1920*1080.

[0116] Step 109 can specifically perform portrait detection based on the original video picture of the current period to obtain at least one portrait detection information. Each portrait detection information includes a portrait ID, portrait frame coordinates and a distance. In each period, portrait detection is first performed based on the currently acquired original video picture, or face detection, to determine the portrait in the original video picture of the current period. The face in the picture facing the camera or the display screen within a certain range can be determined as a portrait. For example Figure 5The original video picture shown can obtain the portrait detection information corresponding to each person through portrait detection. One person corresponds to one portrait detection information. The portrait detection information can be used in the form of an array. For example, the middle person is A, and the portrait detection information corresponding to A is the portrait ID of A, the left frame coordinate, the right frame coordinate, the upper frame coordinate, and the lower frame coordinate of the portrait frame of A, and the distance of A. Similarly, the person on the left is B, and the person on the right is C. The ID is used as the identification of the portrait. The portrait frame refers to a rectangular frame area reflecting the position and size of the portrait. The portrait frame coordinates can include four coordinates corresponding to the upper, lower, left, and right four edges. The distance, i.e., the camera distance, refers to the distance between the person corresponding to the portrait in the actual physical space and the camera. That is, in step 109, the distance between each user and the camera corresponding to each portrait in the original video picture collected by the camera in the current period is obtained. Figure 5 The distance of A in the middle is the closest, the distance of B on the right is the second closest, and the distance of C on the left is the farthest. It can be seen that the farther the distance is, the smaller the size of the portrait frame is.

[0117] As shown in Figure 6 The above step 110 includes steps 200-304.

[0118] Step 200, remove the portrait detection information whose portrait frame coordinates exceed the boundary of the original video picture in at least one portrait detection information. For example, if the left frame coordinate of the portrait frame of the portrait detection information obtained in step 109 is less than 0, it indicates that the portrait frame exceeds the edge of the original video picture. Therefore, the portrait detection information is removed to avoid the problem that the subsequent cropped picture exceeds the edge of the original picture.

[0119] Step 201, sort at least one portrait detection information in the original video picture in the current period according to the distance from small to large to obtain the sorted portrait detection information. The portrait detection information in the array can be sorted by a sorting algorithm such as the bubble algorithm. The specific sorting algorithm is not described in detail in the embodiments of the present application. Figure 4 The sorted portrait detection information obtained is the arrangement order of A, B, and C.

[0120] Step 202, in the sorted portrait detection information, remove the portrait detection information whose distance is greater than the distance of the first portrait detection information by 0.5 m, to obtain the removed portrait detection information. The portrait corresponding to the first portrait detection information in the original video picture is the first portrait, and the other portraits are the second portraits. The first portrait and all the second portraits within a preset distance whose distance difference with the first portrait is less than or equal to 0.5 m are the to-be-determined portraits. The portraits corresponding to the removed portrait detection information in step 202 are the to-be-determined portraits. For example, the distance of A is close to the distance of B, and the distance of C is greater than the distance of A by 0.7 m. Therefore, the portrait detection information corresponding to C is removed from the array, and the removed portrait detection information obtained only includes the portrait detection information of A and B.

[0121] Step 203, determine whether the number of people in the removed portrait detection information is less than or equal to 2. If yes, execute step 204, if not, execute step 205.

[0122] Step 204, the removed portrait detection information is used as valid portrait detection information. For example, in the foregoing example, the removed portrait detection information only includes the portrait detection information of A and B. Therefore, the portrait detection information of A and B is directly used as valid portrait detection information for subsequent use. That is, if the number of to-be-determined portraits in the first region and the second region of the original video picture of the current period is equal, all the to-be-determined portraits of the original video picture of the current period are determined as target portraits.

[0123] Step 205, obtain the number of people whose portrait frame coordinate center points are located on the left side of the original video picture and the number of people whose portrait frame coordinate center points are located on the right side of the original video picture. The original video picture has 1920 pixels in the horizontal direction. If the portrait frame coordinate center point is less than or equal to 960, it is located on the left side of the picture, and the left side number is incremented by 1. If the portrait frame coordinate center point is greater than 960, it is located on the right side of the picture, and the right side number is incremented by 1. The region on the left side of the original video picture is the first region, and the region on the right side is the second region. Step 205 obtains the number of to-be-determined portraits in the first region and the number of to-be-determined portraits in the second region of the original video picture of the current period. In the subsequent process, the video picture is displayed based on the distance between each user and the camera, the number of to-be-determined portraits in the first region, and the number of to-be-determined portraits in the second region.

[0124] Step 206, determine whether the number of people on the left side is equal to the number of people on the right side. If yes, execute step 204, if not, execute step 207.

[0125] Step 207, taking the larger one of the left and right number of people as the effective portrait detection information corresponding to the portrait, assuming that there are 3 people on the left side of the picture and 2 people on the right side of the picture, then taking the portrait detection information corresponding to the 3 people on the left side of the picture as the effective portrait detection information for subsequent use. That is, if the number of to-be-determined portraits in the first region and the second region of the original video picture of the current period is not equal, then all to-be-determined portraits in the region with more to-be-determined portraits in the first region and the second region of the original video picture of the current period are determined as target portraits.

[0126] The portrait in the effective portrait detection information determined in the above step 204 or step 207 is the target portrait, and the process of steps 200-204 or steps 200-407 is to determine the target portrait in the portrait of the original video picture of the current period. One effective portrait detection information corresponds to one portrait, so the number of effective portrait detection information is the number of target portraits. After determining the effective portrait detection information in the above step 204 or step 207, step 301 is performed to obtain the first target region A of the current period, step 302 is performed to obtain the second target region B of the current period, and step 303 is performed to obtain the third target region C of the current period. As shown in the figure, Figure 4 The first target region A refers to the smallest rectangular frame range including the portrait frame corresponding to all effective portrait detection information of the current period. The second target region B of the current period is the frame after the first target region A of the current period is expanded by a certain range, that is, the second target region B of the current period is the enlarged and centered frame of the portrait frame corresponding to the effective portrait detection information of the current period, and the second target region B refers to the final target region of the current picture cropping. The third target region C is the actual cropping frame gradually approaching the second target region B. The above process of obtaining each target region is actually the process of obtaining the coordinates of each target region, including the coordinates of the left, right, top and bottom four edge frames. In some cases, if the frame after the first target region A of the current period is expanded by a certain range exceeds the edge of the original video picture, it is necessary to ensure that the second target region does not exceed the edge of the original video picture, at this time the second target region B will cover the first target region A, but it does not necessarily have a significant centering effect.

[0127] As shown in the figure, Figure 7As shown, the step 301 specifically includes steps 3011-3014. The step 3011 determines whether the number of valid portrait detection information is 0. If yes, the step 3012 is executed. If no, the step 3013 is executed. The step 3012 takes the boundary coordinate set of the original video picture of the current period as the coordinate of the first target region of the current period. When it is determined that the number of valid portrait detection information is 0, it means that no portrait in the picture is detected in the step 109, or after the portrait detection information beyond the boundary is removed in the step 200, there is no valid portrait detection information in the array. In this case, the boundary coordinate of the original video picture is directly taken as the coordinate of the first target region. The step 3013 takes the portrait frame coordinate of one of the valid portrait detection information as the coordinate of the first target region of the current period. The step 3014 traverses other valid portrait detection information. If the coordinate of any side of the portrait frame of the other valid portrait detection information is located outside the first target region of the current period, the coordinate of the side is assigned to the first target region of the current period. For example, there are three portraits in the original video picture. The three portraits are all the portraits corresponding to the portrait frame in the valid portrait detection information. The circle or ellipse is the portrait. In the step 3013, the portrait frame coordinate of any portrait is first taken as the coordinate of the first target region A. Then, in the step 3014, the valid portrait detection information corresponding to the second portrait is first traversed. It is determined that the left frame coordinate and the upper frame coordinate of the portrait frame corresponding to the valid portrait detection information are both located outside the first target region of the current period. Therefore, the coordinate of the upper portrait frame is assigned to the first target region of the current period. Then, the valid portrait detection information corresponding to the third portrait is traversed. It is determined that the left frame coordinate of the portrait frame corresponding to the valid portrait detection information is located outside the first target region of the current period. Therefore, the left frame coordinate of the left portrait frame is assigned to the first target region of the current period. In this way, the first target region covers the portrait frame corresponding to each valid portrait detection information through the traversal manner. Finally, the first target region of the current period is the minimum rectangular frame covering all the portrait frames of the valid portrait detection information. The acquisition of the first target region is realized.

[0128] As shown in Figure 8 The step 302 includes steps 3021-3029.

[0129] Step 3021, among a plurality of preset resolution levels, determine a preset resolution level to which the first target region of the current period belongs. For example, a plurality of resolution levels are preset based on the resolution of the original video picture, the resolution of the original picture is 1920*1080, the first level is 1920*1080, the second level is 1760*990, the third level is 1600*900, the fourth level is 1440*810, and the fifth level is 1280*720. In order to make the final cropped picture have the same proportion as the original video picture, and avoid the problem of picture stretching during cropping, the width-length ratio of each preset resolution level can be set to equal to the width-length ratio of the original video picture, which is 16:9. If the resolution of the first target region of the current period is 1880*1000, which is between 1920*1080 and 1760*990, it belongs to the first level 1920*1080; if the resolution of the first target region of the current period is 1880*880, which is between 1920*1080 and 1440*810, it also belongs to the first level 1920*1080; if the resolution of the first target region of the current period is 640*480, which is below 1280*720, it belongs to the fifth level 1280*720.

[0130] Step 3022, determine whether the preset resolution level to which the first target region of the last period belongs is the same as the preset resolution level to which the first target region of the current period belongs, if yes, execute step 3023, if no, execute step 3024. If the current period is the first cycle, the boundary coordinates of the original video picture can be used as the coordinates of the first target region of the last period.

[0131] Step 3023, determine whether the first target region of the current period is covered by the second target region of the last period, if yes, execute step 3025, if no, execute step 3024. If the current period is the first cycle, the frame of the original video picture can be used as the second target region of the last period.

[0132] Step 3025, determine whether the number of valid portrait detection information of the current period is the same as the number of valid portrait detection information of the last period, if yes, execute step 3026, if no, execute step 3024.

[0133] Step 3026, the second target area of the last period is taken as the second target area of the current period. That is, if the number of the valid portrait detection information in the picture does not change and the moving range of the person does not exceed the previous level and the range of the second target area compared with the last time, the new second target area does not need to be calculated, and the second target area of the last period is still used to avoid the frequent change of the second target area due to the small change of the portrait frame, so as to improve the picture jitter problem.

[0134] Step 3024, it is determined whether the coordinates of the first target area of the current period are the same as the boundary coordinates of the original video picture of the current period, if not, step 3027 is executed, if yes, step 3028 is executed. Whether the coordinates are the same here means whether the coordinates of the four sides are all the same, if all the coordinates of the four sides are the same, step 3028 is executed, if the coordinates of any one side of the four sides are different, step 3027 is executed.

[0135] Step 3027, the coordinates of the second target area of the current period are calculated, that is, the coordinates of the four sides of the second target area of the current period are calculated, and the specific calculation process will be described in detail later. That is, if the number of the valid portrait detection information changes, the change of the resolution level to which the person belongs or the first target area exceeds the range of the second target area of the last period compared with the last time, and the first target area of the current period is different from the boundary of the original picture, the coordinates of the second target area of the current period need to be calculated.

[0136] Step 3028, it is determined whether the first target area of the current period is the same as the second target area of the last period, if not, step 3027 is executed, if yes, step 3029 is executed.

[0137] Step 3029, the original video picture of the current period is taken as the second target area of the current period. That is, if the boundary of the first target area is the boundary of the original picture and is the same as the second target area of the last period, the coordinates of the second target area of the current period do not need to be calculated, and only the boundary coordinates of the original picture are assigned to the second target area of the current period.

[0138] As shown in Figure 9 , the above step 3027 includes steps 401-414.

[0139] Step 401, it is determined whether the left frame coordinate of the first target area A of the current period is ≤16, if yes, step 402 is executed, if not, step 403 is executed.

[0140] Step 402: Set the left border coordinate of the second target region B in the current time period to 0, and the right border coordinate to W, where W is the width of the preset resolution level to which the first target region A belongs in the current time period. Assume the resolution of the first target region A in the current time period is 640*480, and its preset resolution level is 1280*720, i.e., W = 1280. If the left border coordinate of the first target region A is ≤ 16, it is considered that the first target region A is close to or has exceeded the left edge of the original image. Therefore, the width of the second target region B in the current time period is directly set to the left side of the image based on its preset resolution level.

[0141] Step 403: Determine whether the right frame coordinates of the first target area A in the current time period are ≥MW-16. If yes, proceed to step 404; otherwise, proceed to step 405. MW is the width of the original video frame, i.e., MW = 1920.

[0142] Step 404: Set the left border coordinates of the second target area B in the current time period to MW-W = 1920-1280 = 840, and the right border coordinates to MW, i.e., 1920. If the right border coordinates of the first target area A in the current time period with a resolution of 640*480 are ≥1920-16, then it is considered that the first target area A is close to or has exceeded the right edge of the original video frame. Therefore, the width of the second target area B in the current time period is directly set to the right side of the frame based on the preset resolution level.

[0143] Step 405: Determine if mleft-(W-mW) / 2≤0 is satisfied. If yes, proceed to step 402; otherwise, proceed to step 406. Figure 10 As shown, mletf is the left border coordinate of the first target area A in the current time period, mright is the right border coordinate of the first target area A in the current time period, W = 1280, and mW is the width of the first target area A in the current time period, i.e., mW = mrght - mleft = 640. (W - mW) / 2 = (1280 - 640) / 2 = 320. If mleft - 320 ≤ 0, it means that if the left border coordinate of the first target area A is directly shifted to the left based on the 1280*720 resolution level, it will exceed the left edge of the original screen. Therefore, step 402 is executed to directly set the width of the second target area B in the current time period based on its preset resolution level to the left side of the screen.

[0144] Step 406, setting the left border coordinate of the second target region B of the current time period as mleft-(W-mW) / 2=mleft-320, and setting the right border coordinate as mleft-(W-mW) / 2+W=mleft-320+1280=mleft+960. That is, if the first target region A of the current time period is far away from the left edge and the right edge of the original picture, the left border coordinate and the right border coordinate of the second target region B of the current time period are set based on the position of the first target region A of the current time period and the width of the preset resolution level to which the first target region A belongs.

[0145] Step 407 is performed after step 406, that is, determining whether the right border coordinate of the second target region B of the current time period is ≥ MW. If yes, step 404 is performed, and if no, step 408 is performed. Since only the left border coordinate of the second target region B is ensured not to exceed the boundary of the original picture in step 406, the right border coordinate of the second target region B is determined in step 407. If it is ≥ 1920, it is possible to exceed the right edge of the original picture, so step 404 is returned, and the second target region B of the current time period is directly set on the right side of the picture based on the width of the preset resolution level to which the second target region B belongs. If the right border coordinate of the second target region B is less than 1920, it is indicated that it will not exceed the right edge of the original picture, so the subsequent process of determining the upper border coordinate and the lower border coordinate can be continued.

[0146] After steps 402, 404 and 407 are determined as no, the setting of the left border coordinate and the right border coordinate of the second target region B of the current time period is completed, and step 408 is performed to continue setting the upper border coordinate and the lower border coordinate.

[0147] Step 408, determining whether the upper border coordinate of the first target region A of the current time period is ≤ 9. If yes, step 409 is performed, and if no, step 410 is performed.

[0148] Step 409, setting the upper border coordinate of the second target region B of the current time period as 0, and setting the lower border coordinate as H, where H is the height of the preset resolution level to which the first target region A of the current time period belongs, that is, H=720. If the upper border coordinate of the first target region A of the current time period is ≤ 9, it is considered that the first target region A is close to the upper edge of the original picture or has exceeded the upper edge, so the second target region B of the current time period is directly set on the upper side of the picture based on the height of the preset resolution level to which the second target region B belongs.

[0149] Step 410, determining whether the lower border coordinate of the first target region A of the current time period is ≥ MH-9. If yes, step 411 is performed, and if no, step 412 is performed, where MH is the height of the original video picture, that is, MH=1080.

[0150] Step 411, set the upper edge coordinate of the second target region B of the current time period as MH-H = 1080-720, and the lower edge coordinate as mH = 1080. If the lower edge coordinate of the first target region A of the current time period is greater than or equal to 1080-9, it is considered that the first target region A is close to the lower edge of the original picture or has exceeded the lower edge, and thus the second target region B of the current time period is directly set at the lower side of the picture based on the height of the preset resolution level to which it belongs.

[0151] Step 412, determine whether mtop-(H-mH) / 2≤0 is satisfied, if yes, execute step 409, if not, execute step 413. mtop is the upper edge coordinate of the first target region A of the current time period, mbottom is the lower edge coordinate of the first target region A of the current time period, H = 720, and mH is the height of the first target region A of the current time period, i.e. mH = mbottom-mtop = 480. (H-mH) / 2 = (720-480) / 2 = 120. If mtop-120≤0, it is considered that if the upper edge coordinate of the first target region A after being directly moved up based on the resolution level of 1280*720 is used as the upper edge coordinate of the second target region, it will exceed the upper edge of the original picture, and thus step 409 is executed to directly set the second target region B of the current time period at the upper side of the picture based on the height of the preset resolution level to which it belongs.

[0152] Step 413, set the upper edge coordinate of the second target region B of the current time period as mtop-(H-mH) / 2 = mtop-120, and the lower edge coordinate as mtop-(H-mH) / 2+H = mtop-120+720 = mtop+600. That is, if the first target region A of the current time period is far away from the upper and lower edges of the original picture, the upper edge coordinate and the lower edge coordinate of the second target region B of the current time period are set based on the position of the first target region A of the current time period and the height of the preset resolution level to which it belongs.

[0153] After step 413, step 414 is executed to determine whether the lower border coordinates of the second target area B in the current time period are ≥MH. If yes, step 409 is executed; otherwise, step 303 is executed. Since step 413 only ensures that the left border coordinates of the second target area B do not exceed the edge of the original image, step 414 judges the lower border coordinates of the second target area B. If it is ≥1080, it may exceed the lower edge of the original image. Therefore, the process returns to step 409, and the height of the second target area B in the current time period is directly set to the bottom of the image based on the preset resolution level. If the lower border coordinates of the second target area B are less than 1080, it means that it will not exceed the lower edge of the original image. Therefore, the process of obtaining the third target area in the current time period can continue, i.e., step 303 is executed.

[0154] After confirming the above steps 409, 411 and 414 as no, the coordinates of the left border, right border, top border and bottom border of the second target area B in the current time period are set, which completes the process of obtaining the second target area B in the current time period. Next, step 303 is executed to continue to obtain the third target area in the current time period.

[0155] like Figure 11 As shown, step 303 above includes steps 501 to 508.

[0156] Step 501: Obtain the left border gap lg, right border gap rg, top border gap tg, and bottom border gap bg. Here, the gaps refer to the distance between the third target region C' of the previous time period and the second target region B of the current time period. The left border gap lg is the distance between their left borders, the right border gap rg is the distance between their right borders, the top border gap tg is the distance between their top borders, and the bottom border gap is the distance between their bottom borders. If the current time period is the first cycle, the border of the original video frame can be used as the third target region C' of the previous time period.

[0157] Step 502: Determine if the sum of the left and right border gaps (lg + rg) is less than the width speed (VW), and the sum of the top and bottom border gaps (rg + bg) is less than the height speed (VH). If yes, proceed to step 503; otherwise, proceed to step 504. The width speed (VW) and height speed (VH) can be preset values, for example, width speed (VW) of 48 pixels and height speed (VH) of 27 pixels. To ensure the final cropped image has the same aspect ratio as the original video image and avoid image stretching during cropping, the ratio of height speed (VH) to height speed (VH) can be set to equal the aspect ratio of the original video image, both being 16:9.

[0158] Step 503, taking the second target region B of the current time period as the third target region of the current time period. That is, if the difference between the frame coordinates of the second target region B of the current time period and the frame coordinates of the third target region C' of the last time period is within the preset range, it means that the two are close enough, so the coordinates of the second target region B of the current time period can be directly taken as the coordinates of the third target region C of the current time period. If the difference between the coordinates of the second target region B of the current time period and the coordinates of the third target region C' of the last time period is outside the preset range, it means that the two are not close enough, and the process of gradual transition is still needed, so in the subsequent process of 505-508, the coordinates of the third target region of the current time period are made closer to the coordinates of the second target region of the current time period relative to the coordinates of the third target region of the last time period. In this way, the tracking of the target portrait can be realized through the gradual animation effect.

[0159] Step 504, determining whether the second target region of the current time period is enlarged or reduced relative to the third target region of the last time period, if reduced, step 505 is executed, if enlarged, step 506 is executed, and if the size is the same, i.e. not enlarged or reduced, step 506 is executed.

[0160] Step 505, reducing the third target region C' of the last time period to obtain the third target region C of the current time period. The reduction here can take the width speed VW and the height speed VH as the step size, i.e. in the continuous multiple periods, step 505 is executed, then the width of the third target region C' of the last time period is reduced by 48 pixels and the height is reduced by 27 pixels in each period. It can be understood that the step size of the reduction can also be different from the above-mentioned VW and VH, but if the reduction is kept in the proportion of 16:9, the cropping proportion can be guaranteed to be the same as the original picture. In addition, the step size of the reduction can be a fixed value or a dynamically changing value.

[0161] Step 506, determining whether the position of the second target region B of the current time period is the same as the position of the third target region C' of the last time period, if not, step 507 is executed, and if yes, step 508 is executed.

[0162] Step 507, moving the third target region C' of the last time period to the direction close to the second target region B of the current time period to obtain the third target region C of the current time period. The step size of the movement can be a fixed value or a dynamically changing value.

[0163] Step 508, the third target area C' of the last period is enlarged to obtain the third target area C of the current period. The logic of enlargement is similar to that of reduction. The width speed VW and the height speed VH can be used as the step for enlargement. The step for enlargement can also be different from the above-mentioned VW and VH. Understandably, the step for enlargement can also be different from the above-mentioned VW and VH, but if the reduction is kept in the proportion of 16:9, the cutting proportion can be guaranteed to be the same as the original picture. In addition, the step for enlargement can be a fixed value or a dynamically changing value.

[0164] After steps 505, 507 or 508, the third target area C of the current period has been obtained, i.e. the coordinates of the four edges of the third target area C have been obtained, so step 304 can be performed, i.e. the original video picture of the current period is cut based on the third target area C of the current period. For example, the size of the original video picture is 1080P, and the size of the third target area C of the current period is 1900*900, so the cutting process is to adjust the third target area C to the size of 1080P. Finally, the picture of the third target area C is displayed in the size of 1080P.

[0165] In the above process of obtaining the third target area C, the width-height ratio is guaranteed to be the width-height ratio of the original video picture, so as to avoid the problem of stretching and deformation of the finally cut picture.

[0166] The processes of reduction and enlargement will be described below respectively.

[0167] As Figure 12As shown, in the continuous multiple periods t11-t18, if the second target region B is unchanged, the size and position of the third target region C gradually approach the second target region B. The third target region is the actual cropping frame of the picture, and is also a frame that can gradually change. The second target region B is the final cropping frame, and is also the centered and enlarged frame of the person in the valid person detection information. The third target region C gradually tracks the person in the valid person detection information in the continuous multiple periods, that is, the third target region C gradually approaches the second target region B. If the coordinates of the last period third target region C' and the current period second target region B are close to a certain degree, the coordinates of the second target region B are directly assigned to the current period third target region C, so as to realize the acquisition of the third target region C. In the t11 period, it is determined in the step 504 that the second target region B of the current period is reduced relative to the third target region C' of the last period, so the step 505 is executed, the center point of the third target region C' of the last period is unchanged, and the frame coordinates after the size is reduced are taken as the frame coordinates of the third target region C of the current period, and then the picture is cropped and displayed by using the third target region C; then the t12 period is entered, and the third target region C of the last period is still reduced to be taken as the third target region C of the current period; in the t11-t15 periods, the third target region C of the last period is reduced to be taken as the third target region C of the current period; in the t16 period, it is determined in the step 504 that the size of the second target region B of the current period is the same as that of the third target region C' of the last period, that is, it has been reduced to the right position, so the step 506 is executed, it is determined that the position of the second target region B of the current period is different from that of the third target region C' of the last period, that is, it has not been moved to the right position, and the position can be determined based on the center point coordinates of the two, so the step 507 is executed, the third target region C' of the last period is moved to the direction close to the second target region B of the current period to obtain the third target region C of the current period; similarly, in the t17 period, the third target region C' of the last period is also moved to the direction close to the second target region B of the current period to obtain the third target region C of the current period; in the t18 period, it is determined in the step 502 that the sum lg+rg of the left frame gap and the right frame gap is less than the width speed VW, and the sum rg+bg of the upper frame gap and the lower frame gap is less than the height speed VH, that is, it is determined that the size and position of the third target region C' of the last period and the second target region B of the current period have been adjusted to the preset range, so the second target region B of the current period is directly taken as the third target region C of the current period. In the process of t11-t18, the process that the third target region C gradually changes to the second target region B in multiple periods is realized, that is, the tracking of the valid target person is realized by using the animation.In addition, in the process of t11-t18, the previous third target region needs to be reduced and moved, and in the overall process, it is realized by the way of reducing first and then moving. This gradual change can avoid the problem that the third target region C exceeds the display edge when moving before reducing.

[0168] In addition, in the continuous multiple periods t21-t26, the third target region C presents a gradual change of moving first and then enlarging. In the period t21, in step 504, it is determined that the second target region B of the current period is enlarged relative to the third target region C' of the last period, so step 506 is performed to determine that the position of the second target region B of the current period is different from that of the third target region of the last period, which can be determined based on the coordinates of the center points of the two, so step 507 is performed to move the third target region C' of the last period to the right to approach the second target region B of the current period, and the third target region C of the current period is obtained after moving; then enter the period t22, and similar logic is continued to move the third target region C' of the last period to the right to obtain the third target region C of the current period; then enter the period t23, and similar logic is continued to move the third target region C' of the last period to the right to obtain the third target region C of the current period; then enter the period t24, in step 506, it is determined that the position of the second target region B of the current period is the same as that of the third target region of the last period, i.e., the movement is in place, so step 508 is performed to enlarge the third target region C' of the last period to obtain the third target region C of the current period; then enter the period t25, in step 506, it is determined that the position of the second target region B of the current period is the same as that of the third target region of the last period, and step 508 is performed to enlarge the third target region C' of the last period to obtain the third target region C of the current period; in the period t26, in step 502, it is determined that the sum of the left and right frame gaps lg+rg is less than the width speed VW, and the sum of the top and bottom frame gaps rg+bg is less than the height speed VH, i.e., it is determined that the size and position of the third target region C' of the last period and the second target region B of the current period have been adjusted to within the preset range, so the second target region B of the current period is directly taken as the third target region C of the current period. In the process of t21-t26, it is realized by the way of moving first and then enlarging. This gradual change can avoid the problem that the third target region C exceeds the display edge when enlarging before moving.

[0169] In a possible implementation, in the process of obtaining the third target region C of the current period in steps 503, 505, 507 and 508, the coordinate value of the third target region C of the current period is obtained by multiplying the relevant values by 10 and then by means of an integer int, where the relevant values include the coordinates of the four sides and the moving step, and in step 304, the coordinate value of the third target region C of the current period whose coordinate value is enlarged by 10 times is divided by 10, and the original video picture of the current period is cropped based on the recovered third target region C of the current period by means of a floating point. For the calculation of the coordinate value of the third target region C, if the calculation is directly based on the floating point, the error will be accumulated in the process of the gradual change of the third target region C, which may cause a large final cropping error. In the embodiment of the application, for example, the upper left corner coordinate of the third target region C of the last period is (100, 100), the lower right corner coordinate is (420, 280), and the moving step of the third target region C is (4.8, 2.7), which are all relevant values. First, the relevant values are enlarged by 10 times, that is, multiplied by 10, the upper left corner coordinate becomes (1000, 1000), the lower right corner coordinate becomes (4200, 2800), and the moving step of the third target region C becomes (48, 27). It can be seen that after being enlarged by 10 times, the calculation of the moving third target region C can be performed by using the integer int, that is, the third target region C of the current period is (1048, 1027) and (4248, 2827). Then the coordinate is divided by 10, that is, the correct coordinate value is recovered to continue the subsequent floating point operation, that is, the third target region C of the current period is restored to (104.8, 102.7) and (424.8, 282.7). In this way, the error in the calculation process can be reduced and the accuracy can be improved.

[0170] It should be noted that the above is only described by taking the application of the video processing method to the video call process as an example, but the application of the video processing method is not limited to the application scenario, for example, in other possible implementations, the video processing method can also be applied to the video shooting scene of the camera application and other scenes requiring video recording or video preview. In addition, the electronic device is a smart television, which is only an example, and the electronic device can also be any electronic device such as a mobile phone, a tablet computer, a notebook computer or a vehicle-mounted device. The camera can be a camera of the electronic device itself or an external camera.

[0171] The video processing method will be described below by taking some specific scenarios as examples.

[0172] In a possible implementation, as Figure 13As shown, the display of the cropped picture includes: displaying the cropped picture in any one or any combination of the video call notification interface, the video call dialing interface and the video call interface, and the cropped picture is the picture after adjusting the picture of the third target area C to the preset size.

[0173] In a possible implementation, as shown in Figure 13 As shown, when receiving a video call, the above-mentioned video processing method is performed, including acquiring the original video picture collected by the camera, and displaying a video call notification interface, and the video call notification interface displays the picture cropped in the step 304, and the cropped picture is used as the preview picture of the video call; and / or, when dialing a video call, the above-mentioned video processing method is performed, including acquiring the original video picture collected by the camera, and displaying a video call dialing interface, and the video call dialing interface displays the picture cropped in the step 304, and the cropped picture is used as the preview picture of the video call; and / or, when making a video call, the above-mentioned video processing method is performed, including periodically acquiring the original video picture collected by the camera, and displaying a video call interface, and the video call interface displays the picture cropped in the step 304, and the cropped picture is used as the picture of the video call.

[0174] It should be noted that the above only takes the video through scene as an example for description, and other camera shooting scenes such as camera preview and camera recording can also apply the video processing method of the embodiments of the present application.

[0175] In a possible implementation, as shown in Figure 14 As shown, in the first time period, a first video picture is displayed, and in the first time period, the distance between the first user P1 and the camera is less than the distance between the second user P2 and the camera, and the distance between the first user P1 and the camera is less than the distance between the third user P3 and the camera, and the first video picture includes the portrait of the first user P1. The steps 101-111 are performed in the first time period and each time period before the first time period, and the process of performing the above steps in the first time period is described below. The original video picture is the picture corresponding to the larger rectangular frame in Figure 14 The first video picture is the picture corresponding to the smaller rectangular frame in Figure 14 The first video picture is the picture corresponding to the smaller rectangular frame in Figure 14The original video frame is shown, and the portrait detection is performed based on the frame in step 108, and 3 portrait detection information is obtained. The circle or ellipse in the figure represents the portrait, and the size of the portrait is positively correlated with the corresponding distance. In step 201, the portrait corresponding to the first user on the right is the closest portrait to the camera, which is the first portrait, and the corresponding portrait detection information is the first portrait detection information after sorting. Among them, the camera distance of the second user P2 and the third user P3 is larger, the difference between the camera distance of the second user P2 and the camera distance of the first user P1 is greater than 0.5m, and the difference between the camera distance of the third user P3 and the camera distance of the first user P1 is greater than 0.5m, so in step 202, the portrait detection information corresponding to the second user P2 and the third user P3 is removed, and the portrait detection information after removal only corresponds to the first user P1, and the portrait corresponding to the first user P1 is the portrait to be determined. Then in step 204, the portrait corresponding to the first user P1 is determined as the target portrait, and then in steps 301 and 302, the first target area and the second target area B of the current period are obtained based on the portrait corresponding to the first user P1, Figure 14 Only the second target area B is shown, and the second target area B covers the portrait corresponding to the first user P1. Then in step 303, the third target area of the current period is obtained based on the second target area B of the current period and the third target area of the last period. Assuming that the third target area of the last period is the second target area B of the current period, then in the current period, the cropped frame includes the portrait corresponding to the first user P1, and does not include the portraits corresponding to the second user P2 and the third user P3; Assuming that the third target area of the last period is not the second target area B of the current period, then the cropped frame will gradually become to include only the portrait corresponding to the first user P1. The third area refers to the area in the actual physical space corresponding to the first area in the original video frame, and the fourth area refers to the area in the actual physical space corresponding to the second area in the original video frame. In the first period, the first user P1 is located in the third area, and the second user P2 and the third user P3 are located in the fourth area. In the process of steps 201-304, based on the distance between the first user P1 and the camera in the first period being less than the distance between the second user P2 and the camera, and the difference being greater than 0.5m, and the distance between the first user P1 and the camera being less than the distance between the third user P3 and the camera, and the difference being greater than 0.5m, the first video frame is displayed, that is, the frame with the first user P1 as the main character is displayed.

[0176] After the first user P1 moves away from the camera, the original video frame changes from Figure 14 to Figure 15 In the second period, the second video frame is displayed, and the second video frame is Figure 15The second video picture includes the portraits of the second user P2 and the third user P3. After the first user P1 moves away from the camera, steps 101-111 are performed in the second period. In step 201, the portrait corresponding to the first user P1 is still the closest to the camera, which is the first portrait, and the portrait detection information corresponding thereto is the first portrait detection information in the order. After the first user P1 moves away from the camera, the difference between the camera distance of the second user P2 and the camera distance of the first user P1 is less than 0.5 m, and the difference between the camera distance of the third user P3 and the camera distance of the first user P1 is less than 0.5 m. Therefore, in step 202, the portrait detection information corresponding to the second user P2 and the third user P3 is not removed, and the portraits corresponding to the first user P1, the second user P2 and the third user P3 are the to-be-determined portraits. Then, in step 207, the portraits corresponding to the second user P2 and the third user P3 in the left first region are determined as the target portraits. Then, in steps 301 and 302, the first target region and the second target region B of the current period are obtained based on the portraits corresponding to the second user P2 and the third user P3, Figure 15 Only the second target region B is shown, which covers the portraits corresponding to the second user P2 and the third user P3. Then, in step 303, the third target region of the current period is obtained based on the second target region B of the current period and the third target region of the last period, that is, in the second period after the first user P1 moves away from the camera, the cropped picture includes the portraits corresponding to the second user P2 and the third user P3, but does not include the portrait corresponding to the first user P1. In the second period, the positions of the portraits of the three users in the original video picture do not change, but the distances between the three users and the camera are equal. Since the number of users in the fourth region is greater than the number of users in the third region, the picture with the second user P2 and the third user P3 as the main characters in the region with more users.

[0177] During the process from the first period to the second period, the video display interface gradually changes from displaying the first video picture to displaying the second video picture. In each period between the first period and the second period, steps 101-111 are performed. Assuming that the distance between the second user P2 and the camera and the distance between the third user P3 and the camera are equal and do not change, during the process in which the first user P1 moves away from the camera, the distance between the first user P1 and the camera gradually decreases. When the difference between the distance between the first user P1 and the camera and the distance between the second user P2 and the camera is less than or equal to 0.5 m, in step 202, the portrait detection information corresponding to the second user P2 and the third user P3 is no longer removed, that is, the second target region B changes from Figure 14 to Figure 15 However, the mutation of the second target region B does not cause the mutation of the picture displayed by the video display interface. The picture displayed by the video display interface gradually changes fromFigure 14 The second target region B in the text changes to Figure 15 The second target region B in the [theory / information].

[0178] like Figure 16 As shown, after the fourth user P4 moves into the camera's capture range, a third video frame is displayed in the third time period. During this third time period, the distance between the fourth user P4 and the first user P1 is less than the distance between the second user P2 and the first user P1 (i.e., the fourth user P4 is closer to the first user P1, while the second user P2 is farther from the first user P1). Furthermore, the distance between the fourth user P4 and the second user P2 is greater than the distance between the second user P2 and the third user P3 (i.e., the fourth user P4 is farther from the second user P2, while the second user P2 is closer to the third user P3). The third video frame includes the images of the first user P1, the second user P2, the third user P3, and the fourth user P4. The distances mentioned in this paragraph do not refer to the camera distance, but rather to the distances between the users. When the camera capture range is expanded to include a fourth user, P4, steps 101-111 are executed in the third time period, and in each time period between the second and third time periods. Taking the third time period as an example, in step 201, the image corresponding to the first user, P1, is still the image closest to the camera, and is the first image. The corresponding image detection information is the first image detection information after sorting. When the camera capture range is expanded to include a fourth user, P4, the difference in camera distance between other users and the first user, P1, is less than 0.5m. Therefore, in step 202, the image detection information corresponding to the second user, P2, the third user, P3, and the fourth user, P4, is not removed, and the images corresponding to the first user, P1, the second user, P2, the third user, P3, and the fourth user, P4, are images to be determined. Then, in step 204, the images corresponding to the first to fourth users are determined as target images. Then, in steps 301 and 302, the first target area and the second target area B for the current time period are obtained based on the images corresponding to these four users. Figure 16 The image only shows the second target area B, which covers the portraits of the four users. Then, in step 303, the third target area for the current time period is obtained based on the second target area B for the current time period and the third target area for the previous time period. That is, in multiple time periods after the camera's capture range is expanded to include the portraits of the fourth user P4, the cropped image gradually changes from including the portraits of the second user P2 and the third user P3 to including the portraits of all four users.

[0179] When the second user P2, the third user P3, and the fourth user P4 leave the camera's capture range, and the first user P1 moves, the original video feed changes from... Figure 16 The result is as shown. Figure 17As shown, in the fourth time period, a fourth video picture is displayed, and the fourth video picture includes the portrait of the first user P1. When the second user P2, the third user P3 and the fourth user P4 move out of the camera capturing range, steps 101-111 are performed in the fourth time period, and steps 101-111 are performed in each time period between the third time period and the fourth time period. Taking the fourth time period as an example, in step 201, the portrait of the first user P1 is determined as the target portrait in step 204, and then the first target region and the second target region B of the current time period are obtained based on the portrait of the first user P1 in steps 301 and 302, Figure 17 Only the second target region B is shown, and the second target region B covers the portrait of the first user P1. Then, in step 303, the third target region of the current time period is obtained based on the second target region B of the current time period and the third target region of the last time period, that is, in the fourth time period after the second user P2, the third user P3 and the fourth user P4 move out of the camera capturing range, the cropped picture becomes to include the portrait of the first user P1.

[0180] As can be seen, in the continuous multiple time periods, with the change of the camera distance corresponding to the portrait in different regions of the original video picture, and the change of the target portrait, the size and position of the second target region B also change, and the third target region tends to approach the second target region B, and therefore, the actually displayed cropped picture gradually tracks and changes to cover the target portrait. Based on the change of the number of people and the camera distance corresponding to the portrait in the video picture, the main character is re-determined, and the latest determined main character is tracked and enlarged through the cropped picture, so as to highlight the main character portrait in the video picture, realize the focusing on the main character, and make it easier to see the main character portrait.

[0181] The embodiment of the present application also provides a video processing device, comprising: a first display module, configured to display a first video picture in a first time period, in the first time period, a distance between a first user and a camera is less than a distance between a second user and the camera, and the distance between the first user and the camera is less than a distance between a third user and the camera, and the first video picture includes a portrait of the first user; a second display module, configured to display a second video picture after the first user moves away from the camera, the second video picture including portraits of the second user and the third user; and a third display module, configured to display a third video picture after a fourth user moves into a camera capturing range, the distance between the fourth user and the first user is less than the distance between the second user and the first user, and the distance between the fourth user and the second user is greater than the distance between the second user and the third user, and the third video picture includes portraits of the first user, the second user, the third user and the fourth user.

[0182] The video processing device can apply the video processing methods in any of the above embodiments. The specific process and principle are the same as those in the above embodiments, and will not be repeated here.

[0183] It should be understood that the above division of the video processing device is merely a logical functional division. In actual implementation, all or part of these modules can be integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely through software calls from processing elements; they can be implemented entirely in hardware; or some modules can be implemented through software calls from processing elements, while others are implemented in hardware. For example, any one of the display modules can be a separate processing element, or it can be integrated into the video processing device, such as being integrated into a chip within the video processing device. Alternatively, it can be stored as a program in the memory of the video processing device, and its functions can be called and executed by a processing element within the video processing device. The implementation of other modules is similar. Moreover, these modules can be integrated together, or they can be implemented independently. The processing element mentioned here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.

[0184] For example, each display module can be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs). As another example, when one of the above modules is implemented through a processing element scheduler, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling programs. Furthermore, these modules can be integrated together to implement a system-on-a-chip (SOC).

[0185] like Figure 1 As shown, this application embodiment also provides an electronic device 100, wherein the internal memory 121 is used to store at least one instruction, which, when loaded and executed by the processor 110, causes the electronic device 100 to perform the video processing method in any of the above embodiments.

[0186] The electronic device involved in the present application can be any product such as a smart television, a mobile phone, a tablet computer, a personal computer (PC), a personal digital assistant (PDA), a smart watch, a wearable electronic device, an augmented reality (AR) device, a virtual reality (VR) device, an in-vehicle device, a drone device, a smart car, a smart speaker, a robot, smart glasses, and the like.

[0187] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program. When the computer program is run on a computer, the computer is caused to execute the video processing method in any of the above embodiments.

[0188] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, solid state disk), etc.

[0189] In the embodiments of the present application, "at least one" refers to one or more, and "multiple" refers to two or two more. The "and / or" describes the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or the second signal line, which can mean that A exists alone, A and the second signal line exist at the same time, and the second signal line exists alone. Wherein A, the second signal line can be singular or plural. The character "or" generally indicates that the associated objects before and after are in a "or" relationship. "At least one of the following" and the like means any combination of these items, including any combination of single or multiple items. For example, a, at least one of the second signal line and the third signal line can mean: a, the second signal line, the third signal line, a-the second signal line, a-the third signal line, the second signal line-the third signal line, or a-the second signal line-the third signal line, wherein a, the second signal line, the third signal line can be single or multiple.

[0190] The above is only the preferred embodiment of the present application and is not intended to limit the present application. Those skilled in the art can make various changes and modifications to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A video processing method, characterized in that, include: In the first time period, a first video frame is displayed. In the first time period, the distance between the first user and the camera is less than the distance between the second user and the camera, and the distance between the first user and the camera is less than the distance between the third user and the camera. The first video frame includes the image of the first user. After the first user moves away from the camera, a second video frame is displayed, which includes the images of the second user and the third user. After the fourth user moves into the camera's capture range, a third video image is displayed. The distance between the fourth user and the first user is less than the distance between the second user and the first user, and the distance between the fourth user and the second user is greater than the distance between the second user and the third user. The third video image includes the images of the first user, the second user, the third user, and the fourth user.

2. The video processing method according to claim 1, characterized in that, During the first time period, the first user was located in the third area, and the second and third users were located in the fourth area; The first time period, displaying the first video frame includes: Based on the fact that, within the first time period, the distance between the first user and the camera is less than the distance between the second user and the camera, and the distance between the first user and the camera is less than the distance between the third user and the camera, the first video image is displayed.

3. The video processing method according to claim 2, characterized in that, After the first user moves away from the camera, the second video feed is displayed, including: After the first user moves away from the camera, in the second time period, based on the fact that during the second time period, the distance between the first user and the camera is equal to the distance between the second user and the camera, the distance between the first user and the camera is equal to the distance between the third user and the camera, and the number of users in the fourth area is greater than the number of users in the third area, the second video image is displayed.

4. The video processing method according to claim 3, characterized in that, During the period from the first time period to the second time period, the video display interface gradually changes from displaying the first video frame to displaying the second video frame.

5. The video processing method according to claim 1, characterized in that, The first time period, displaying the first video frame includes: Obtain a second target area including at least one human figure from the original video footage captured by the camera in the first time period; Based on the second target region of the first time period and the third target region of the previous time period adjacent to the first time period, the third target region of the first time period is obtained. The coordinates of the third target region of the first time period are closer to the coordinates of the second target region of the first time period than the coordinates of the third target region of the previous time period adjacent to the first time period, or the coordinates of the third target region of the first time period are equal to the coordinates of the second target region of the first time period. The original video footage captured by the camera during the first time period is cropped based on the third target area of ​​the first time period to obtain the cropped first video footage, which is then displayed.

6. The video processing method according to claim 5, characterized in that, Also includes: In the original video footage captured by the camera in the first time period, the target human image is determined, and the first target region in the first time period is obtained. The first target region in the first time period is the smallest rectangular box that includes all the target human images. The step of obtaining a second target region including at least one human image from the original video footage captured by the camera in the first time period includes: If the condition of not calculating is met, then the second target region of the previous time period adjacent to the first time period is taken as the second target region of the first time period; If the non-calculation condition is not met, then a second target region including the first target region of the first time period is obtained in the original video frame of the first time period, and the size of the second target region of the first time period is larger than the size of the first target region of the first time period. The non-calculation condition includes at least one or any combination of the following. When the non-calculation condition includes multiple conditions, if any one of the multiple conditions is not satisfied, the non-calculation condition is not satisfied; if all terms in the multiple conditions are satisfied, the non-calculation condition is satisfied: The preset resolution level of the first target area in the previous time period adjacent to the first time period is the same as the preset resolution level of the first target area in the first time period. The first target area in the first time period is covered by the second target area in the previous time period adjacent to the first time period; The number of target portraits in the first time period is equal to the number of target portraits in the previous time period adjacent to the first time period.

7. The video processing method according to claim 6, characterized in that, During the first time period, the first user was located in the third area, and the second and third users were located in the fourth area; The image of the user closest to the camera in the original video frame is the first image, and the other images are the second images. The first image and all second images whose camera distance difference with the first image is ≤ a preset distance are the images to be determined. The original video frame includes a first region corresponding to the third region and a second region corresponding to the fourth region. The step of identifying the target human image from the human images captured by the camera in the first time period includes: If the number of undetermined human figures in the first region and the second region of the original video footage captured by the camera in the first time period is not equal, then all the undetermined human figures in the region with the larger number of undetermined human figures in the first region and the second region of the original video footage captured by the camera in the first time period are determined as target human figures. If the number of undetermined human figures in the first region and the second region of the original video footage captured by the camera in the first time period is equal, then all the undetermined human figures in the original video footage captured by the camera in the first time period are determined as target human figures.

8. The video processing method according to claim 5, characterized in that, The process of obtaining the third target region of the first time period based on the second target region of the first time period and the third target region of the previous time period adjacent to the first time period includes: If the second target region of the first time period is shrunk relative to the third target region of the previous time period adjacent to the first time period, then the third target region of the previous time period adjacent to the first time period is shrunk to obtain the third target region of the first time period. If the second target area of ​​the first time period has the same size as the third target area of ​​the previous time period adjacent to the first time period, and the center coordinates of the second target area of ​​the first time period are different from those of the third target area of ​​the previous time period adjacent to the first time period, then the third target area of ​​the previous time period adjacent to the first time period is moved towards the direction closer to the second target area of ​​the first time period to obtain the third target area of ​​the first time period. If the second target region of the first time period is enlarged relative to the third target region of the previous time period adjacent to the first time period, and the center coordinates of the second target region of the first time period are the same as those of the third target region of the previous time period adjacent to the first time period, then the third target region of the previous time period adjacent to the first time period is enlarged to obtain the third target region of the first time period.

9. The video processing method according to claim 8, characterized in that, The process of reducing the third target region of the previous time period adjacent to the first time period to obtain the third target region of the first time period includes: reducing the coordinate value of the third target region of the previous time period adjacent to the first time period by 10 times based on the step size value using an integer int method to obtain the coordinate value of the third target region of the first time period by 10 times. The process of moving the third target region of the previous time period adjacent to the first time period toward the second target region of the first time period to obtain the third target region of the first time period includes: moving 10 times the coordinate value of the third target region of the previous time period adjacent to the first time period in a direction with a step size of 10 times the coordinate value of the second target region of the first time period in a direction with an integer int, to obtain 10 times the coordinate value of the third target region of the first time period. The process of magnifying the third target region of the previous time period adjacent to the first time period to obtain the third target region of the first time period includes: magnifying the coordinate value of the third target region of the previous time period adjacent to the first time period by 10 times based on the step size value using an integer int method to obtain the coordinate value of the third target region of the first time period by 10 times. The process of cropping the original video frame of the first time period based on the third target region of the first time period to obtain the cropped frame includes: The coordinates of the third target region in the first time period are obtained by dividing 10 times the coordinates of the third target region in the first time period by 10. The original video frame of the first time period is then cropped based on the coordinates of the third target region in the first time period using floating-point methods to obtain the cropped frame.

10. The video processing method according to claim 5, characterized in that, The aspect ratio of the third target region is the same as that of the original video frame.

11. The video processing method according to claim 1, characterized in that, The display of the first video screen includes displaying the first video screen in at least one or any combination of a video call notification interface, a video call dialing interface, and a video call interface.

12. A video processing apparatus, characterized in that, include: The first display module is used to display a first video frame during a first time period, wherein the distance between the first user and the camera is less than the distance between the second user and the camera, and the distance between the first user and the camera is less than the distance between the third user and the camera, and the first video frame includes the image of the first user. The second display module is used to display a second video frame after the first user moves away from the camera. The second video frame includes images of the second user and the third user. The third display module is used to display a third video image after the fourth user moves into the camera's capture range. The distance between the fourth user and the first user is less than the distance between the second user and the first user, and the distance between the fourth user and the second user is greater than the distance between the second user and the third user. The third video image includes the images of the first user, the second user, the third user, and the fourth user.

13. An electronic device, characterized in that, include: A processor and a memory, the memory being used to store at least one instruction, which, when loaded and executed by the processor, causes the electronic device to perform the video processing method as described in any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, Includes a program or instructions that, when run on a computer, execute the method as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Video recording method and device and storage medium

    CN116095465A

  • Facial signature methods, systems and software

    WO2016183380A1