A video image processing method and apparatus
By combining the character information of the current frame and historical frame, the subject character is determined and cropped and scaled, the problem of discontinuity in video calls and monitoring scenes is solved, and the high-accuracy "painting with people" effect is achieved.
Patent Information
- Application Number
- CN201910819774.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-08-31
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2039-08-31
AI Technical Summary
In video calls and surveillance scenarios, it is difficult for the prior art to achieve continuous "painting with people" effect of the picture, especially when the environment is complex, the positioning of the characters is inaccurate, resulting in discontinuous display screen.
By obtaining the identity information and position information of the video image frame, combining the identity information of the video image frames in the previous N frames, M subject characters are determined, and the video images are cropped and scaled according to their position information to ensure that the subject characters are fully displayed in small-resolution images.
It improves the accuracy of the character perception process, ensures the continuity of the subject character's picture, and realizes the effect of "painting with people" through software during the image acquisition and display process.
Smart Images

Figure CN112446255B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and particularly to a method and device for processing video images. Background Art
[0002] With the rapid development of image technology, users have higher requirements for the display of video images. For example, the display of video images during video calls and in monitoring scenarios. In the conventional video acquisition and display process, a video image is acquired by an acquisition device, and the acquired video image is correspondingly cropped and scaled according to the display specifications, and then encoded and sent to a display device for display.
[0003] Usually, the acquisition and display are implemented based on a fixed hardware platform, and a video image with a fixed field of view is acquired by an acquisition camera. When the position of a person at the acquisition end changes, since the acquisition camera does not sense the person, the display screen at the display end always maintains a fixed field of view display, and the effect of "the picture follows the person" cannot be achieved, resulting in poor user experience.
[0004] Based on this, the industry applies human perception technology to the image acquisition and display process. The specific solution is as follows: The camera performs high-resolution acquisition according to a fixed field of view, and uses human body perception technology to detect and track the human body in the acquired video image, and the position of the person is located in real time. When the position of the person moves, the high-resolution video image can be correspondingly cropped and scaled according to the real-time located position of the person (the position of the person after movement) to obtain a small-resolution image that adapts to the display specifications and the person is located in a specific area of the image, so as to realize the real-time adjustment of the display screen according to the position of the person and achieve the effect of "the picture follows the person".
[0005] However, when the device environment at the acquisition end is complex (for example, the background picture is complex or there are other people frequently entering and leaving the picture), the above method may have false detections and missed detections, resulting in inaccurate person positions located in some frames, and the person cannot be displayed or cannot be completely displayed in the cropped and scaled small-resolution image, making the presented picture of the main person discontinuous. Summary of the Invention
[0006] This application provides a method and device for processing video images, which realizes the picture following the person with continuous display screen during video calls.
[0007] To achieve the above object, this application adopts the following technical solutions:
[0008] In a first aspect, a video image processing method is provided. The method may include: obtaining identity information and location information of each person in the i-th frame of video image, where i > 1; determining M main characters from the i-th frame of video image according to the identity information of the characters in the N video image frames before the i-th frame of video image, where M, N ≥ 1; cropping the i-th frame of video image according to the location information of the main characters, and the cropped i-th frame of video image includes M main characters; reducing or enlarging the cropped i-th frame of video image so that the display screen displays the cropped i-th frame of video image according to a preset display specification.
[0009] Through the video image processing method provided by this application, when determining the main characters of the video image, the identity information of the characters in this frame of image and the identity information of the characters in the N video image frames before this frame are combined, which greatly improves the accuracy of the character perception process, and accordingly improves the accuracy of the determined position of the main characters. In this way, it can be ensured that the main characters can be completely displayed in the small-resolution image cropped and scaled according to the main characters, so as to ensure the continuity of the presented picture of the main characters, so as to achieve the effect of the picture following the person and being continuous through software during the image acquisition and display process.
[0010] Among them, the identity information of the characters is used to uniquely indicate the same person in different frames. The identity information can be the marker information of the character obtained through the detection and tracking algorithm, that is, each character has its own different characteristic information.
[0011] The i-th frame of video image is any frame of video image in the video stream, and i ≤ the total number of frames of the video stream. When executing the video image processing method provided by this application, the video image processing method provided by this application is executed for each frame of image in the video stream to ensure that each frame of image can completely display the main characters after cropping, and other details will not be elaborated one by one.
[0012] Optionally, the N video image frames before the i-th frame of video image can be the first N video image frames consecutive with the i-th frame of video image in the video stream, or can also be the first N video image frames non-consecutive with the i-th frame of video image in the video stream, or can also be the video image frames within a preset time period in the video stream.
[0013] Combining the first aspect and any of the above possible implementation manners, in another possible implementation manner, determining M main characters from the i-th frame of video image according to the identity information of the characters in the N video image frames before the i-th frame of video image can be specifically implemented as: determining the characters that appear in the N video image frames for a number of frames greater than or equal to a first preset threshold and also appear in the i-th frame of video image as M main characters. Determining the main characters by accumulating the number of frames avoids the interference of the entry and exit of people who are not participating in the video call on the character recognition and improves the accuracy of the character recognition.
[0014] Specifically, the process of determining whether a person in the i-th video image frame is the main person may include: counting the cumulative number of frames in which the person appears in N video image frames. If the cumulative number of frames is greater than or equal to the first preset threshold, then the person is determined to be the main person. Whether the person appears in a video image frame can be specifically implemented as: whether the video image frame contains a person with the same identity information as the person.
[0015] Among them, the cumulative number of frames in which a person appears is the number of consecutive video image frames in which the person appears in N video image frames before the i-th video image frame; the consecutive video image frames may include S video image frames in which the person does not appear; S is greater than or equal to 0 and less than or equal to the preset number of frames.
[0016] Combined with the first aspect and any of the above possible implementation manners, in another possible implementation manner, the video image processing method provided in this application may further include: dividing the i-th video image into Y regions; configuring a preset threshold corresponding to each region; the preset threshold corresponding to the k-th region is the k-th preset threshold; the k-th region is any one of the Y regions; Y is greater than or equal to 2; k is greater than or equal to 1 and less than or equal to Y. Among them, the preset thresholds corresponding to different regions may be different. Correspondingly, according to the identity information of the people in N video image frames before the i-th video image frame, M main people are determined from the i-th video image, which is specifically implemented as: determining the people who appear in the i-th video image and whose number of frames of appearance in N video image frames is greater than or equal to the preset threshold corresponding to the region where they are located as M main people. By configuring different preset thresholds for different regions, the accuracy of determining the main people is improved, and thus the accuracy of person recognition is improved.
[0017] Combined with the first aspect, in a possible implementation manner, the above method further includes: obtaining the person information of each person in the i-th video image, and the person information may include one or more of the following information: whether speaking information, priority information. Correspondingly, according to the identity information of the people in N video image frames before the i-th video image frame, M main people are determined from the i-th video image, which can be specifically implemented as: determining the people who speak in N video image frames and whose number of frames is greater than or equal to the second preset threshold and who appear in the i-th video image as M main people. Or, determining the people whose priority information in N video image frames is greater than the third preset threshold and who appear in the i-th video image as M main people. Or, determining the M most important people according to the priority information among the people who speak in N video image frames and whose number of frames is greater than or equal to the second preset threshold and who appear in the i-th video image as M main people.
[0018] Among them, the speaking information is used to indicate whether the person in the video image is speaking or not. Audio processing technology can be combined with the lip movements of the person in the video image to obtain the speaking information of the person, or the speaking information of the person can be directly obtained through the lip movements of the person in the video image.
[0019] The priority information is used to indicate the importance of the person in the video image. The priority information of different people using the device can be pre-configured to correspond to the identity information of the person. Then, when processing each frame of the video image, when the identity information of the person is obtained, the pre-configured priority information is searched to obtain the priority information of the person. Alternatively, the priority information input by the user for different people in the video image can be received.
[0020] Combined with the first aspect and any of the above possible implementation manners, in another possible implementation manner, the video image processing method provided by this application may further include: receiving the priority information input by the user. To achieve real-time configuration of the person priority level by the user and improve the person recognition accuracy.
[0021] Combined with the first aspect and any of the above possible implementation manners, in another possible implementation manner, according to the position information of the main person, the i-th frame of the video image is cropped, which can be specifically implemented as: determining a cropping frame, where the cropping frame includes the minimum circumscribed rectangle of M main persons; cropping the i-th frame of the video image with the determined cropping frame.
[0022] Among them, the cropping frame can be the minimum circumscribed rectangle of M main persons plus a cropping margin, and the cropping margin can be greater than or equal to 0.
[0023] It should be noted that the cropping frame including the minimum circumscribed rectangle of M main persons can be understood as: the determined cropping frame tries to completely contain the minimum circumscribed rectangle of M main persons.
[0024] Combined with the first aspect and any of the above possible implementation manners, in another possible implementation manner, determining the cropping frame can be specifically implemented as: obtaining the distance between the center point of the candidate cropping frame and the center point of the cropping frame of the previous frame of the video image, where the candidate cropping frame includes the minimum circumscribed rectangle of M main persons; if the distance is greater than or equal to the distance threshold, expanding the candidate cropping frame until the distance from the center point of the candidate cropping frame to the center point of the cropping frame of the previous frame of the video image is less than the preset threshold, and taking the expanded candidate cropping frame as the determined cropping frame.
[0025] Among them, the candidate cropping frame can be the minimum circumscribed rectangle of M main persons plus a cropping margin, and the cropping margin can be greater than or equal to 0.
[0026] Combined with the first aspect and any of the above possible implementation manners, in another possible implementation manner, determining the cropping frame can be specifically implemented as follows: obtaining the distance between the center point of the first candidate cropping frame and the center point of the cropping frame of the previous frame of video image, where the first candidate cropping frame includes the minimum bounding rectangle of M main characters; if the distance is greater than or equal to the distance threshold, determining a second cropping frame, the center point of the second cropping frame is the center point of the cropping frame of the previous frame of video image plus an offset, and the size of the second cropping frame is the same as the size of the cropping frame of the previous frame of video image; if the second cropping frame contains the minimum bounding rectangle of M main characters, taking the third cropping frame as the cropping frame; where the third cropping frame is the second cropping frame, or the third cropping frame is the second cropping frame reduced to a cropping frame that contains the minimum bounding rectangle; if the second cropping frame does not completely contain the minimum bounding rectangle, expanding the second cropping frame to contain the minimum bounding rectangle, and taking the expanded second cropping frame as the cropping frame.
[0027] Wherein, the offset can be a preset value, or can also be the distance between the center point of the first candidate cropping frame and the center point of the cropping frame of the previous frame of video image multiplied by a weighting value, or others.
[0028] Combined with the first aspect and any of the above possible implementation manners, in another possible implementation manner, when the character information includes priority information, the candidate cropping frame or the first candidate cropping frame can be an outer bounding rectangle centered on the character with the highest priority among the M main characters and including a cropping margin for the M main characters.
[0029] Combined with the first aspect and any of the above possible implementation manners, in another possible implementation manner, when the character information includes whether the character is speaking information, the candidate cropping frame or the first candidate cropping frame can be an outer bounding rectangle centered on the speaking character among the M main characters and including a cropping margin for the M main characters.
[0030] Combined with the first aspect and any of the above possible implementation manners, in another possible implementation manner, the video image processing method provided in this application can further include: displaying the cropped i-th frame of video image according to a preset display specification. Wherein, the preset display specification can be adapted to the specification of the display screen, or can also be a preset display screen ratio.
[0031] Combined with the first aspect and any of the above possible implementation manners, in another possible implementation manner, the video image processing method provided in this application can further include: saving at least one of the following information of each character in the i-th frame of video image: identity information, location information, character information.
[0032] Combined with the first aspect and any of the above possible implementation manners, in another possible implementation manner, the video image processing method provided in this application may further include: obtaining the j-th frame of video image, where j is less than or equal to X and X is greater than 1; obtaining and saving the identity information and position information of each person in the j-th frame of video image; directly reducing the j-th frame of video image to an image with a preset display specification.
[0033] Combined with the first aspect or any of the above possible implementation manners, in another possible implementation manner, the video image processing method provided in this application is applied to the sending-end device in a video call. The video image processing method provided in this application may further include: sending the reduced or enlarged i-th frame of video image to the receiving-end device.
[0034] In a second aspect, this application provides a video image processing device. This device may be an electronic device, or a device or chip system in an electronic device, or a device that can be used in matching with an electronic device. The video image processing device may implement the functions executed in the above aspects or various possible designs. The functions may be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. For example, the video image processing device may include: an obtaining unit, a determining unit, a cropping unit, and a scaling unit.
[0035] Among them, the obtaining unit is used to obtain the identity information and position information of each person in the i-th frame of video image, where i is greater than 1; the determining unit is used to determine M main characters from the i-th frame of video image according to the identity information of the characters in the N video image frames before the i-th frame of video image, where M and N are greater than or equal to 1; the cropping unit crops the i-th frame of video image according to the position information of the main characters, and the cropped i-th frame of video image includes M main characters; the scaling unit reduces or enlarges the cropped i-th frame of video image so that the display screen displays the cropped i-th frame of video image according to a preset display specification.
[0036] It should be noted that the video image processing device provided in the second aspect is used to execute the video image processing method provided in the first aspect above, and the specific implementation may refer to the specific implementation of the first aspect above.
[0037] In a third aspect, an embodiment of this application provides an electronic device. The electronic device may include: a processor and a memory; the processor is coupled to the memory, and the memory may be used to store computer program code. The computer program code includes computer instructions. When the computer instructions are executed by the electronic device, the electronic device executes the video image processing method as described in the first aspect or any of the possible implementation manners.
[0038] Fourthly, an embodiment of the present application provides a computer-readable storage medium, which may include: computer software instructions; when the computer software instructions run on an electronic device, the electronic device is caused to execute the video image processing method described in any one of the first aspect or possible implementation manners of the first aspect.
[0039] Fifthly, an embodiment of the present application provides a computer program product, which when running on a computer, causes the computer to execute the video image processing method described in any one of the first aspect or any possible implementation manner of the claims.
[0040] Sixthly, an embodiment of the present application provides a chip system, which is applied to an electronic device; the chip system includes an interface circuit and a processor; the interface circuit and the processor are interconnected by a line; the interface circuit is configured to receive a signal from a memory of the electronic device and send the signal to the processor, and the signal includes computer instructions stored in the memory; when the processor executes the computer instructions, the chip system executes the video image processing method described in any one of the first aspect or any possible implementation manner.
[0041] Seventhly, an embodiment of the present application provides a graphical user interface (GUI), which is stored in an electronic device, and the electronic device includes a display, a memory, and one or more processors; the one or more processors are configured to execute one or more computer programs stored in the memory, and the graphical user interface includes: the GUI displayed on the display, and the GUI includes a video screen, and the video screen includes the i-th frame video image processed by the above-mentioned first aspect or any possible implementation manner, and the video screen is transmitted to the electronic device by another electronic device (such as a second electronic device), and the second electronic device includes a display screen and a camera.
[0042] It should be understood that the description of technical features, technical solutions, beneficial effects or similar languages in the present application does not imply that all features and advantages can be achieved in any single embodiment. On the contrary, it can be understood that the description of features or beneficial effects means that specific technical features, technical solutions or beneficial effects are included in at least one embodiment. Therefore, the description of technical features, technical solutions or beneficial effects in this specification does not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions and beneficial effects described in this embodiment can be combined in any appropriate manner. Those skilled in the art will understand that an embodiment can be implemented without one or more specific technical features, technical solutions or beneficial effects of a specific embodiment. In other embodiments, additional technical features and beneficial effects can also be identified in specific embodiments that do not embody all embodiments. Brief Description of the Drawings
[0043] Figure 1 A schematic diagram of a video scene provided by an embodiment of the present application;
[0044] Figure 2 A schematic diagram of the system architecture of a video call scene provided by an embodiment of the present application;
[0045] Figure 3 A schematic diagram of a video image provided by an embodiment of the present application;
[0046] Figure 4 A schematic diagram of video image processing provided by an embodiment of the present application;
[0047] Figure 5 A schematic diagram of the result of video image processing provided by an embodiment of the present application;
[0048] Figure 6 Another schematic diagram of the result of video image processing provided by an embodiment of the present application;
[0049] Figure 7 A schematic diagram of the system architecture of a video surveillance scene provided by an embodiment of the present application;
[0050] Figure 8 A schematic diagram of the structure of an electronic device provided by an embodiment of the present application;
[0051] Figure 9 A schematic diagram of the flowchart of a video image processing method provided by an embodiment of the present application;
[0052] Figure 10 A schematic diagram of a video call interface provided by an embodiment of the present application;
[0053] Figure 11 Another schematic diagram of a video call interface provided by an embodiment of the present application;
[0054] Figure 12 Another schematic diagram of a video call interface provided by an embodiment of the present application;
[0055] Figure 13 Another schematic diagram of video image processing provided by an embodiment of the present application;
[0056] Figure 14 Another schematic diagram of video image processing provided by an embodiment of the present application;
[0057] Figure 15 Another schematic diagram of video image processing provided by an embodiment of the present application;
[0058] Figure 16Schematic flowchart of another video image processing method provided by an embodiment of the present application;
[0059] Figure 17A Schematic diagram of another video call interface provided by an embodiment of the present application;
[0060] Figure 17B Schematic diagram of another video call interface provided by an embodiment of the present application;
[0061] Figure 18A Schematic diagram of another video call interface provided by an embodiment of the present application;
[0062] Figure 18B Schematic diagram of another video call interface provided by an embodiment of the present application;
[0063] Figure 18C Schematic diagram of another video call interface provided by an embodiment of the present application;
[0064] Figure 19 Schematic diagram of another video image processing provided by an embodiment of the present application;
[0065] Figure 19A Schematic diagram of another video image processing provided by an embodiment of the present application;
[0066] Figure 19B Schematic diagram of another video call interface display provided by an embodiment of the present application;
[0067] Figure 20 Schematic diagram of another video image processing provided by an embodiment of the present application;
[0068] Figure 20A Schematic diagram of another video call interface display provided by an embodiment of the present application;
[0069] Figure 21 Schematic diagram of another video image processing provided by an embodiment of the present application;
[0070] Figure 21A Schematic diagram of another video image processing provided by an embodiment of the present application;
[0071] Figure 21B Schematic diagram of another video call interface display provided by an embodiment of the present application;
[0072] Figure 22 Schematic diagram of another video image processing provided by an embodiment of the present application;
[0073] Figure 22A Schematic diagram of another video call interface display provided by an embodiment of the present application;
[0074] Figure 23A schematic diagram of video image processing in a monitoring scenario provided by an embodiment of the present application;
[0075] Figure 24 Another schematic diagram of video image processing in a monitoring scenario provided by an embodiment of the present application;
[0076] Figure 25 Another schematic diagram of video image processing in a monitoring scenario provided by an embodiment of the present application;
[0077] Figure 26 Another schematic diagram of video image processing in a monitoring scenario provided by an embodiment of the present application;
[0078] Figure 27 A schematic diagram of the structure of a video image processing device provided by an embodiment of the present application;
[0079] Figure 28 Another schematic diagram of the structure of a video image processing device provided by an embodiment of the present application. Detailed implementation manners
[0080] Terms such as "first", "second", and "third" in the specification, claims, and the above-mentioned drawings of the present application are used to distinguish different objects, rather than to limit a specific order.
[0081] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0082] For ease of understanding, the nouns involved in the present application are first explained.
[0083] Video stream may refer to the data transmitted in a video service, that is, a dynamic continuous image sequence in a video call, a video conference, or a monitoring scenario.
[0084] Video image may refer to a static picture, and each frame image in a video stream is called a video image.
[0085] Person may refer to a person who is active or stationary in a video image. Of course, in the application scenarios of the present application, it is not only applicable to people who are active or stationary in a video image, but also applicable to other main objects in a video image, such as active or static animals or other things. Hereinafter, people in a video image will be used as an example for illustration, and it should not limit the application scenarios.
[0086] Identity information may refer to the feature identifiers of each person recognized by the human body detection and tracking algorithm in the video image, which is used to uniquely identify the same person in different frames to distinguish different individuals. The identity information may include, but is not limited to, appearance information, annotation information, or other recognized feature information. The expression form of the identity information may include text, serial numbers, person numbers, or other information related to individual characteristics.
[0087] Location information can be used to indicate the relative position or area of a person in the video image. The form of the location information can be the pixel positions of one or more points of the person in the video image, or the pixel positions of the person's contour, or the pixel positions of the area where the person is located, etc. The pixel positions can be indicated by pixel coordinates or others. The location information is used to indicate the relative position of the person in the video image and is not limited to a specific location.
[0088] Person information may refer to the additional information of each person in the video image obtained through the recognition algorithm or the marking algorithm to better perform person recognition and determine the main person. The person information may include, but is not limited to, one or more of the following information: whether the person is speaking information, person priority information, etc.
[0089] Currently, in order to achieve the function of the picture following the person during the video acquisition and display process, there are two solutions in the industry.
[0090] One is the hardware implementation solution, which uses a camera with a pan-tilt unit and auxiliary additional person positioning devices (such as locating the position of the speaker through voice) to locate the person's position, and then controls the pan-tilt unit to direct the camera towards the speaker's direction for acquisition. The hardware solution of the pan-tilt camera is large in size and high in cost, which is not conducive to large-scale popularization.
[0091] The other is the software algorithm implementation solution. The camera performs large-resolution acquisition according to a fixed field of view, and the person detection and tracking algorithm locates the person's position in real time. Then, according to the located person's position, the large-resolution image is correspondingly cropped, reduced, or enlarged (scaled) to obtain a small-resolution image of a predetermined specification. However, the software solution may have defects such as false detection and missed detection. If the image is directly cropped after positioning, the accuracy of person perception is not high, and the continuity of the final display screen will be difficult to guarantee.
[0092] Based on this, an embodiment of the present application provides a video image processing method to achieve continuous picture following the person of the presented main person in a software manner. This method can be applied to an electronic device. In the method provided in this embodiment, after processing the video image to locate the person, the main person is determined by combining the person identity information of the current frame and the historical frames, and the currently captured video image is cropped and scaled according to the main person. This greatly improves the accuracy of the person perception process, and accordingly improves the accuracy of the determined position of the main person. In this way, it can be ensured that the main person can be completely displayed in the small-resolution image cropped and scaled according to the main person, so as to ensure the continuity of the picture of the presented main person, and to achieve continuous picture following the person through software during the image acquisition and display process.
[0093] The implementation manner of the embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0094] The video image processing method provided by the embodiment of the present application can be applied to the video image acquisition and display process of an electronic device. This image acquisition and display process can be a video call (video conference) scenario, a video surveillance scenario, or others. Exemplarily, when the video image acquisition and display process is a video call scenario, as Figure 1 shown, user A uses electronic device 1, user B uses electronic device 2, and user A and user B have a video call.
[0095] Figure 2 FIG. is a schematic system architecture diagram of the above video image processing method provided by the embodiment of the present application applied to a video call scenario. As Figure 2 shown, the system architecture may include a sending-end device 201 and a receiving-end device 202.
[0096] Specifically, the sending-end device 201 can be one end of a video call and communicate with the receiving-end device 202. For example, one or more users 1 can communicate with one or more users 2 of the receiving-end device 202 through the sending-end device 201.
[0097] Among them, the call in this embodiment can refer to a video call or a video conference. Therefore, the sending-end device 201 at least includes a camera and a display screen, and the receiving-end device 202 also at least includes a camera and a display screen. In addition, the sending-end device 201 and the receiving-end device 202 may further include a receiver (or speaker), a microphone, etc. The camera can be used to capture video images during the call. The display screen can be used to display images during the call. The receiver (or speaker) is used to play voices during the call. The microphone is used to capture voices during the call.
[0098] Specifically, as Figure 2As shown in the figure, the sending device 201 includes a video collector 2011, a video pre-processor 2012, a video encoder 2013, and a transmitter 2014. The receiving device 202 includes a video display 2021, a video post-processor 2022, a video decoder 2023, and a receiver 2024.
[0099] Among them, Figure 2 The working process of the schematic system architecture is as follows: The video collector 2011 in the sending device 201 captures video images frame by frame in a video call, and transmits the captured video images to the video pre-processor 2012 for corresponding pre-processing (including but not limited to: person recognition, cropping, scaling, etc.), and then transmits them to the video encoder 2013 for encoding and then to the transmitter 2014. The transmitter 2014 transmits the encoded video images to the receiver 2024 of the receiving device 202 through a wired or wireless medium. The receiver 2024 transmits the received video images to the video decoder 2023 for decoding. The decoded video images are transmitted to the video display 2021 for display after being processed by the video post-processor 2022.
[0100] Exemplarily, the electronic device described in the embodiments of the present application may be a television, a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer (such as a Huawei notebook computer), a desktop computer, an ultra-mobile personal computer (UMPC), a netbook, and a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) / virtual reality (VR) device, etc., which include or are connected to a display screen and a camera. The specific form of the device is not particularly limited in the embodiments of the present application.
[0101] In addition, in some embodiments, the above-mentioned sending device 201 and receiving device 202 may be the same type of electronic device. For example, both the sending device 201 and the receiving device 202 are televisions. In some other embodiments, the above-mentioned sending device 201 and receiving device 202 may be different types of electronic devices. For example, the sending device 201 is a television and the receiving device 202 is a notebook computer. Here, a specific example is combined to illustrate the video image transmission process in a video call or a video conference.
[0102] For example, in Figure 1 the shown scenario, assume that electronic device 1 is the sending device and electronic device 2 is the receiving device. The video image of the fixed field of view captured by its camera at a certain moment may be as Figure 3 shown. Electronic device 1 pairsFigure 3 The video image shown uses a person detection and tracking algorithm to identify the identity information and location information of a person. For example, the location information can be like Figure 4 the coordinates shown. Among them, the coordinates here are exemplified by the specific coordinates of each key point in the person, and the key point can include but is not limited to: head, shoulders, arms, hands, legs, feet, eyes, nose, mouth, clothes, etc. Figure 4 The coordinates are shown as different points in [reference], and each coordinate point has a determined coordinate value in the video image. The electronic device 1 determines the minimum bounding rectangle of the identified person as shown in Figure 4 [reference]. Assuming a resolution image with a width of w and a height of h for the display specification of the electronic device 2, the electronic device 1 takes the minimum bounding rectangle as the center and crops the video image shown in Figure 3 according to the width-to-height ratio of the display specification of the electronic device 2, and obtains the cropping result shown in Figure 5 [reference]. The electronic device 1 scales the cropping result shown in Figure 5 to a resolution image with a width of w and a height of h as shown in Figure 6 [reference]. The specific scaling process is as follows: if the resolution of the cropping result is less than w×h, it is enlarged; if the resolution of the cropping result is greater than w×h, it is reduced.
[0103] Figure 7 FIG. [reference] is a schematic diagram of a system architecture to which the above video image processing method provided by an embodiment of the present application is applied to a video surveillance scenario. As shown in Figure 7 [reference], the system architecture may include a collection device 701, a processing device 702, a storage device 703, and a display device 704.
[0104] It should be noted that Figure 7 the devices included in the system architecture shown in [reference] can be centrally deployed or distributedly deployed. Figure 7 the devices included in the system architecture shown in [reference] can be deployed in at least one electronic device.
[0105] Among them, Figure 7 the working process of the system architecture shown in [reference] is: the collection device 701 collects video images frame by frame, and transmits the collected video images to the processing device 702 for corresponding preprocessing (including but not limited to: person recognition, cropping, scaling, etc.) and then stores them in the storage device 703. The display device 704 obtains the video images from the storage device 703 and displays them.
[0106] Figure 8 FIG. [reference] is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. The structures of the above-mentioned transmitting-end device 201, receiving-end device 202, Figure 7 the electronic devices where the devices included in the system architecture shown in [reference] are located can be as shown in Figure 8 [reference].
[0107] As shown Figure 8 , the electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0108] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more or fewer components than shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0109] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0110] The controller may be the nerve center and command center of the electronic device. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0111] A memory can also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the said memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0112] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include one or more of an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM interface, a USB interface, etc.
[0113] The charging management module 140 is configured to receive a charging input from a charger. Herein, the charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 can receive the charging input from the wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 can receive the wireless charging input through the wireless charging coil of the electronic device. While charging the battery 142, the charging management module 140 can also supply power to the electronic device through the power management module 141.
[0114] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives the inputs from the battery 142 and the charging management module 140 and supplies power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, the wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as the battery capacity, the number of battery charge cycles, and the battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be provided in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 can also be provided in the same device.
[0115] The wireless communication function of the electronic device can be implemented by Antenna 1, Antenna 2, Mobile Communication Module 150, Wireless Communication Module 160, Modulation and Demodulation Processor, and Baseband Processor, etc.
[0116] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, Antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0117] Mobile Communication Module 150 can provide wireless communication solutions applied to the electronic device, including the second generation mobile communication technology (2G) / the third generation mobile communication technology (3G) / the fourth generation mobile communication technology (4G) / the fifth generation mobile communication technology (5G), etc. Mobile Communication Module 150 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. Mobile Communication Module 150 can receive electromagnetic waves by Antenna 1, perform filtering, amplification and other processing on the received electromagnetic waves, and transmit them to the Modulation and Demodulation Processor for demodulation. Mobile Communication Module 150 can also amplify the signal modulated by the Modulation and Demodulation Processor and convert it into electromagnetic waves through Antenna 1 for radiation. In some embodiments, at least some functional modules of Mobile Communication Module 150 can be disposed in Processor 110. In some embodiments, at least some functional modules of Mobile Communication Module 150 and at least some modules of Processor 110 can be disposed in the same device.
[0118] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.), or displays an image or video through the display screen 194. In some embodiments, the modulation and demodulation processor may be an independent device. In other embodiments, the modulation and demodulation processor may be independent of the processor 110 and be provided in the same device as the mobile communication module 150 or other functional modules.
[0119] The wireless communication module 160 may provide solutions for wireless communications applied to the electronic device, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and transmits the processed signals to the processor 110. The wireless communication module 160 may also receive the signal to be transmitted from the processor 110, perform frequency modulation and amplification on it, and convert it into electromagnetic waves through the antenna 2 and radiate it out.
[0120] In some embodiments, antenna 1 of the electronic device is coupled to the mobile communication module 150, and antenna 2 is coupled to the wireless communication module 160, enabling the electronic device to communicate with the network and other devices through wireless communication technologies. For example, the electronic device can conduct a video call or video conference with other electronic devices through antenna 1 and the mobile communication module 150. The wireless communication technologies may include one or more of global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, IR technology, etc. The GNSS may include one or more of global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), satellite based augmentation systems (SBAS), etc.
[0121] The electronic device implements the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change the display information.
[0122] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device may include one or N display screens 194, where N is a positive integer greater than 1. For example, in the embodiments of the present application, during the process of the user using the electronic device to make a video call or a video conference with the user of another electronic device, the display screen 194 can display a video answering interface, or a video reminder interface, or a video call interface, or a video monitoring interface (such as including the video image sent by the peer device and the video image collected by this device).
[0123] The electronic device can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.
[0124] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera photosensitive element. The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the noise, brightness, and skin color of the image through algorithms. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0125] The camera 193 is used to capture still images or videos. For example, in the embodiments of the present application, the camera 193 can be used to collect video images during a video call or a video conference. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB or YUV. In some embodiments, the electronic device may include one or N cameras 193, where N is a positive integer greater than 1. In this embodiment, the camera 193 can be arranged in the electronic device in a hidden manner or not in a hidden manner, and this embodiment does not make specific limitations here.
[0126] The digital signal processor is used to process digital signals. For example, by using a human body monitoring and tracking algorithm for digital video images, after determining the main person in the video image, the video image is cropped and scaled accordingly to obtain an image that adapts to the display specifications of the receiving device.
[0127] The video codec is used to compress or decompress digital videos. The electronic device can support one or more video codecs. In this way, the electronic device can play or record videos in multiple coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0128] The NPU is a neural-network (NN) computing processor. By referring to the structure of a biological neural network, for example, referring to the transmission mode between human brain neurons, it can quickly process the input information and can also continuously learn by itself. Through the NPU, applications such as intelligent cognition of the electronic device can be realized, such as image recognition, face recognition, speech recognition, text understanding, etc.
[0129] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0130] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 121. For example, in the embodiment of the present application, the processor 110 can process the video image by executing the instructions stored in the internal memory 121, locate the person, determine the main person by combining the current frame person information and the historical frame person information, and crop and scale the currently captured video image according to the main person to ensure the continuity of the display screen of the receiving end device, so as to achieve the effect that the displayed image follows the person during a video call. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device (such as audio data, phone book, etc.). In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. In this embodiment, the internal memory 121 can also be used to store the original large-resolution video image captured by the camera 193, the small-resolution video image that has been recognized, screened, cropped, and scaled by the processor 110, and the person information of each frame of the video image, etc.
[0131] The electronic device can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc. Such as calls, music playback, recording, etc.
[0132] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.
[0133] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device can listen to music or hands-free calls through the speaker 170A.
[0134] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device answers a call or a voice message, the voice can be listened to by placing the receiver 170B close to the ear.
[0135] The microphone 170C, also known as the "microphone" or "transmitter", is used to convert sound signals into electrical signals. When making a call, sending a voice message, or triggering an electronic device to perform certain functions through a voice assistant, the user can speak close to the microphone 170C with their mouth to input the sound signal into the microphone 170C. The electronic device can be provided with at least one microphone 170C. In some other embodiments, the electronic device can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device can also be provided with three, four or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.
[0136] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0137] The pressure sensor 180A is used to sense pressure signals and can convert pressure signals into electrical signals. In some embodiments, the pressure sensor 180A can be provided on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can include at least two parallel plates with conductive materials. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device determines the intensity of the pressure based on the change in capacitance. When a touch operation acts on the display screen 194, the electronic device detects the intensity of the touch operation according to the pressure sensor 180A. The electronic device can also calculate the position of the touch based on the detection signal of the pressure sensor 180A. In some embodiments, touch operations with the same touch position but different touch operation intensities can correspond to different operation instructions. For example: when a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, the instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, the instruction to create a new short message is executed.
[0138] The gyro sensor 180B can be used to determine the motion posture of the electronic device. In some embodiments, the angular velocity of the electronic device around three axes (i.e., x, y, and z axes) can be determined by the gyro sensor 180B. The gyro sensor 180B can be used for anti-shake shooting. For example, when the shutter is pressed, the gyro sensor 180B detects the angle of the electronic device shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the electronic device through reverse movement to achieve anti-shake. The gyro sensor 180B can also be used for navigation and somatosensory game scenes.
[0139] The air pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device calculates the altitude through the air pressure value measured by the air pressure sensor 180C to assist positioning and navigation.
[0140] The magnetic sensor 180D includes a Hall sensor. The electronic device can use the magnetic sensor 180D to detect the opening and closing of the flip leather case. In some embodiments, when the electronic device is a flip phone, the electronic device can detect the opening and closing of the flip cover according to the magnetic sensor 180D. Then, according to the detected opening and closing state of the leather case or the opening and closing state of the flip cover, the flip cover can be automatically unlocked.
[0141] The acceleration sensor 180E can detect the magnitude of the acceleration of the electronic device in all directions (generally three axes). When the electronic device is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.
[0142] The distance sensor 180F is used to measure the distance. The electronic device can measure the distance by infrared or laser. In some embodiments, when shooting a scene, the electronic device can use the distance sensor 180F to measure the distance to achieve fast focusing.
[0143] The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode may be an infrared light emitting diode. The electronic device emits infrared light outward through the light emitting diode. The electronic device uses a photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device. When insufficient reflected light is detected, the electronic device can determine that there is no object near the electronic device. The electronic device can use the proximity light sensor 180G to detect when the user holds the electronic device close to the ear to talk, so as to automatically turn off the screen to save power. The proximity light sensor 180G can also be used in leather case mode and pocket mode to automatically unlock and lock the screen.
[0144] The ambient light sensor 180L is used to sense the ambient light brightness. The electronic device can adaptively adjust the brightness of the display screen 194 according to the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance during photography. The ambient light sensor 180L can also cooperate with the proximity light sensor 180G to detect whether the electronic device is in a pocket to prevent accidental touch.
[0145] The fingerprint sensor 180H is used to collect fingerprints. The electronic device can use the collected fingerprint characteristics to achieve fingerprint unlocking, access the application lock, fingerprint photography, fingerprint answering calls, etc.
[0146] The temperature sensor 180J is used to detect the temperature. In some embodiments, the electronic device uses the temperature detected by the temperature sensor 180J to execute a temperature processing strategy. For example, when the temperature reported by the temperature sensor 180J exceeds the threshold, the electronic device reduces the performance of the processor near the temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the electronic device heats the battery 142 to prevent abnormal shutdown of the electronic device caused by low temperature. In other embodiments, when the temperature is lower than yet another threshold, the electronic device boosts the output voltage of the battery 142 to prevent abnormal shutdown caused by low temperature.
[0147] The touch sensor 180K, also known as the "touch panel". The touch sensor 180K can be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also known as the "touch screen". The touch sensor 180K is used to detect touch operations acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device, at a different position from the display screen 194.
[0148] The bone conduction sensor 180M can obtain vibration signals. In some embodiments, the bone conduction sensor 180M can obtain the vibration signals of the vibrating bone mass of the human vocal part. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure pulsation signals. In some embodiments, the bone conduction sensor 180M can also be disposed in the earphone to form a bone conduction earphone. The audio module 170 can parse the voice signal based on the vibration signal of the vibrating bone mass of the vocal part obtained by the bone conduction sensor 180M to implement the voice function. The application processor can parse the heart rate information based on the blood pressure pulsation signal obtained by the bone conduction sensor 180M to implement the heart rate detection function.
[0149] The button 190 includes a power-on button, volume buttons, etc. The button 190 can be a mechanical button or a touch button. The electronic device can receive button inputs and generate key signal inputs related to the user settings and function controls of the electronic device.
[0150] The motor 191 can generate vibration prompts. The motor 191 can be used for incoming call vibration prompts and also for touch vibration feedback. For example, touch operations on different applications (such as taking pictures, audio playing, etc.) can correspond to different vibration feedback effects. For touch operations on different areas of the display screen 194, the motor 191 can also correspond to different vibration feedback effects. Different application scenarios (such as time reminder, receiving messages, alarm clock, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0151] The indicator 192 can be an indicator light and can be used to indicate the charging state, power change, and can also be used to indicate messages, missed calls, notifications, etc.
[0152] The SIM card interface 195 is used to connect the SIM card. The SIM card can be in contact with and separated from the electronic device by inserting into or removing from the SIM card interface 195. The electronic device can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The electronic device interacts with the network through the SIM card to implement functions such as calls and data communication. In some embodiments, the electronic device uses an eSIM, that is, an embedded SIM card. The eSIM card can be embedded in the electronic device and cannot be separated from the electronic device.
[0153] The methods in the following embodiments can all be implemented in an electronic device with the above hardware structure.
[0154] Figure 9 It is a schematic flowchart of a video image processing method provided by an embodiment of the present application. In the present application, the electronic device processes the video stream in a video call or video surveillance frame by frame. For each acquired video image, it is processed according to the image processing method provided by the present application. The processing method for each frame of image by the electronic device is the same. The following embodiments only describe the detailed process of the electronic device processing the i-th frame of video image, and others will not be elaborated one by one. The i-th frame of video image is any frame of video image in the video stream. As Figure 9 shown, the method may include:
[0155] S901. The electronic device obtains the identity information and location information of each person in the i-th frame of video image.
[0156] Wherein, i is greater than 1 and less than or equal to the total number of frames of the video stream.
[0157] For example, i can be greater than or equal to X, where X is a frame number threshold value pre-configured for starting to execute the video image processing method provided by the embodiments of the present application in the video stream.
[0158] Specifically, in S901, the electronic device can use a human detection and tracking algorithm to identify the people in the i-th frame of video image. The identified people are one or more. While identifying the people, the identity information and location information of each person can be obtained.
[0159] It should be noted that the human detection and tracking algorithm is an image processing technology used to identify people in images. The embodiments of the present application do not limit the specific implementation of the human detection and tracking algorithm. For example, the human detection and tracking algorithm can be the YOLO algorithm or the SSD algorithm or others.
[0160] Specifically, the identity information of a person can be used to uniquely indicate the same person in different frames. The identity information can be the marker information of the person obtained through the detection and tracking algorithm, that is, each person has their own different characteristic information. Or, the identity information can also be the person number corresponding to the characteristic information.
[0161] The location information of a person can be the unique coordinate values of one or more key points of the person in the video image.
[0162] Further, as Figure 16 shown, the video processing method provided by the embodiments of the present application can further include S901a.
[0163] S901a. The electronic device obtains the person information of each person in the i-th frame of video image.
[0164] Wherein, the person information can include one or more of the following information: whether speaking information, priority information. In practical applications, the content included in the person information can not be limited by the content of this article and can be configured according to actual needs.
[0165] Wherein, the whether speaking information is used to indicate whether the person in the video image is speaking or not speaking. The audio processing technology can be combined with the lip shape of the person in the video image to obtain the whether speaking information of the person, or the whether speaking information of the person can be directly obtained through the lip shape of the person in the video image.
[0166] Priority information is used to indicate the importance of people in video images. The priority information of different people using the device can be pre-configured to correspond to the identity information of the people. Then, when processing each frame of video image, when the identity information of the person is obtained, the pre-configured priority information is searched to obtain the priority information of the person. Alternatively, the priority information input by the user for different people in the video image can be received. Alternatively, the priority information can be obtained by converting the speaking information. For example, the priority of a person who is speaking is higher than that of a person who is not speaking, and the priority of a person who speaks for a longer time is higher than that of a person who speaks for a shorter time.
[0167] Exemplarily, photo information of different people and corresponding priority information are stored in the electronic device. When performing video image processing, if the similarity between the person recognized in the video image and a certain stored photo is greater than the similarity threshold, the priority information corresponding to the stored photo is used as the priority information of the recognized person.
[0168] Among them, the photo information of different people and the corresponding priority information stored in the electronic device can be manually input by the user into the function configuration interface of the electronic device, and the electronic device stores the photos and priority information of different people; or, the electronic device can record the photo information of different people and the corresponding priority information obtained during the historical video acquisition and display process; or, the user can manually input the photos and priority information of different people, and at the same time, the electronic device dynamically updates the photos of different people and the corresponding priority information each time video acquisition and display are performed.
[0169] Optionally, when the priority information is input by the user of the electronic device, the video image processing method provided in this application may further include: receiving the priority information input by the user.
[0170] Here, the process of the user inputting priority information is illustrated by examples.
[0171] For example, when the user configures the priority information for a certain person recognized in the video image, the user can long-press on the screen of the electronic device to display the configuration menu for configuration. As Figure 10 shown, assuming that the video image collected by the electronic device is the Figure 10 picture, the user long-presses on the position of a certain person in this picture (the finger position in Figure 10 is used to indicate the position where the user long-presses, which is only an example and does not constitute a limitation), and the electronic device displays the Figure 11 configuration menu shown in the figure. The user can select "Configure person priority information" in the Figure 11 configuration menu shown in the figure for priority configuration. When the user selects Figure 11 "Configure person priority information" in the figure, the electronic device displays Figure 12The interactive interface shown, where the user inputs the priority information of the person, and the electronic device simultaneously captures a photo of the person and stores the photo together with the importance level recorded by the user in the Figure 12 interface.
[0172] S902. The electronic device determines M main characters from the i-th video image based on the identity information of the characters in the N video image frames before the i-th video image frame.
[0173] Among them, the identity information of the characters in the N video image frames before the i-th video image frame is saved after the electronic device performs the S901 process on the corresponding video images. The specific process is the same as S901 and will not be elaborated here.
[0174] Specifically, N is greater than or equal to 1. Optionally, N can be less than or equal to i - 1. In practical applications, the specific value of N can be configured according to actual needs.
[0175] Optionally, the N video image frames before the i-th video image frame can be the first N video image frames adjacent to the i-th video image frame in the video stream, or can also be the first N video image frames not adjacent to the i-th video image frame in the video stream, or can also be the video image frames within a preset time period in the video stream. The embodiments of the present application do not limit the specific position of the N video image frames before the i-th video image frame in the video stream.
[0176] In a possible implementation, during the process of processing a video stream, the value of N can also be a dynamic value. When i is less than the configured threshold, N takes the value equal to i - 1. When i is greater than the configured threshold, N takes a fixed value less than i - 1. When i is equal to the configured threshold, N can take the value equal to i - 1 or can also take a fixed value less than i - 1. The present application does not make specific limitations.
[0177] Among them, when N takes a fixed value less than i - 1, the specific value of the fixed value can be configured according to experience, and the present application does not make specific limitations.
[0178] Among them, M can be one or more. The embodiments of the present application do not make specific limitations on the value of M.
[0179] In a possible implementation, M can be the total number of main characters determined in each video image frame.
[0180] In another possible implementation, M can be a pre-configured fixed value.
[0181] Specifically, S902 can be implemented in but not limited to the following several possible ways.
[0182] Implementation 1: The electronic device determines M principal characters as those who appear in at least the first preset threshold number of frames among N video image frames and appear in the i-th video image frame.
[0183] Specifically, the process of determining whether a character in the i-th video image frame is a principal character may include: counting the cumulative number of frames in which the character appears among N video image frames. If the cumulative number of frames is greater than or equal to the first preset threshold, then the character is determined as a principal character. Whether the character appears in a video image frame can be specifically implemented as: whether the video image frame contains a character with the same identity information as this character.
[0184] Among them, the cumulative number of frames in which a character appears is the number of consecutive video image frames in which the character appears among N video image frames before the i-th video image frame; the consecutive video image frames may include S video image frames in which the character does not appear; S is greater than or equal to 0 and less than or equal to the preset number of frames.
[0185] Implementation 2: The electronic device divides the i-th video image frame into Y regions; configures a preset threshold corresponding to each region; the preset threshold corresponding to the k-th region is the k-th preset threshold; the k-th region is any one of the Y regions; Y is greater than or equal to 2; k is greater than or equal to 1 and less than or equal to Y. The electronic device determines M principal characters as those who appear in at least the preset threshold corresponding to their region among N video image frames and appear in the i-th video image frame.
[0186] Among them, in Implementation 2, the preset thresholds corresponding to different regions may be different.
[0187] For example, when Y is equal to 3, the video image is divided into Figure 13 the three preset regions of left, middle, and right as shown, respectively recorded as Region 1, Region 2, and Region 3. The preset thresholds configured for each region are respectively recorded as Threshold 1, Threshold 2, and Threshold 3, and Threshold 1, Threshold 2, and Threshold 3 are different. Then, if it is recognized that character A is located in Region 2 in the i-th video image frame and the cumulative number of frames in which character A appears is greater than Threshold 2, then character A is determined as a principal character. If it is recognized that character B is located in Region 3 in the i-th video image frame and the cumulative number of frames in which character B appears is less than Threshold 3, then character B is not a principal character.
[0188] It should be noted that Y can also be 1. In this case, the specific implementation of Implementation 2 is the same as that of the above Implementation 1 and will not be elaborated here.
[0189] Implementation 3: Corresponding to S901a obtaining the character information of each character in the i-th video image frame, S902 is specifically implemented as:
[0190] Determine the M main characters as those whose number of speaking frames in N video image frames is greater than or equal to a second preset threshold and who appear in the i-th video image frame. Or, determine the M main characters as those whose priority information in N video image frames is greater than a third preset threshold and who appear in the i-th video image frame; or, among the characters whose number of speaking frames in N video image frames is greater than or equal to the second preset threshold and who appear in the i-th video image frame, select the most important M according to the priority information to determine them as the M main characters.
[0191] It should be noted that the values of the above-mentioned preset thresholds can be configured according to actual needs, and the embodiments of the present application do not specifically limit this. The cumulative appearance frames can also be converted into cumulative appearance durations, and the content of the corresponding preset thresholds can be time thresholds.
[0192] S903. The electronic device crops the i-th video image according to the main character position information.
[0193] Among them, the cropped i-th video image includes M main characters. It should be understood that the cropped i-th video image can completely display the M main characters.
[0194] Specifically, the electronic device crops the i-th video image according to the main character position information, which can be specifically implemented as: determining a cropping frame, which is the minimum bounding rectangle of the M main characters; cropping the i-th video image with the cropping frame.
[0195] Among them, the aspect ratio of the cropping frame should adapt to the preset display specification.
[0196] It should be noted that the minimum bounding rectangle of the cropping frame containing the M main characters can be understood as: the determined cropping frame is the minimum bounding rectangle that tries to completely contain the M main characters.
[0197] Optionally, the specific implementation of determining the cropping frame may include but is not limited to the following implementation schemes.
[0198] Implementation Scheme 1. The electronic device determines the candidate cropping frame as the cropping frame.
[0199] In a possible implementation, the candidate cropping frame can be the minimum bounding rectangle of the M main characters plus a cropping margin, and the cropping margin can be greater than or equal to 0.
[0200] For example, the specific process of the electronic device using the minimum bounding rectangle as the determined cropping frame to crop the video image can refer to Figure 4 and Figure 5 for illustration.
[0201] In another possible implementation, when the person information includes priority information, the candidate cropping frame can be an outer circumscribed rectangle centered on the person with the highest priority among the M main persons and including the M main persons, plus a cropping margin.
[0202] For example, Figure 14 It is shown that the determined cropping frame is an outer circumscribed rectangle centered on the person with the highest priority among the M main persons and including the M main persons, and the i-th frame of the video image is cropped to completely display the scene of the main person.
[0203] In another possible implementation, when the person information includes information on whether a person is speaking, the candidate cropping frame can be an outer circumscribed rectangle centered on the speaking person among the M main persons and including the M main persons, plus a cropping margin.
[0204] For example, Figure 15 It is shown that the determined cropping frame is an outer circumscribed rectangle centered on the speaking person among the M main persons and including the M main persons, and the i-th frame of the video image is cropped to completely display the scene of the main person.
[0205] Of course, the range of the candidate cropping frame can be configured according to actual needs, and the embodiments of the present application do not specifically limit this.
[0206] Implementation solution 2: The electronic device determines the cropping frame of the i-th frame of the video image according to the first candidate cropping frame and the cropping frame of the previous frame of the video image.
[0207] Among them, the first candidate cropping frame in implementation solution 2 is the same as the candidate cropping frame in implementation solution 1.
[0208] Specifically, in implementation solution 2, the electronic device first obtains the distance between the center point of the first candidate cropping frame and the center point of the cropping frame of the previous frame of the video image. The first candidate cropping frame includes the minimum outer circumscribed rectangle of the M main persons. If the distance is greater than or equal to the distance threshold, a second cropping frame is determined. The center point of the second cropping frame is the center point of the cropping frame of the previous frame of the video image plus an offset. The size of the second cropping frame is the same as the size of the cropping frame of the previous frame of the video image. If the second cropping frame contains the minimum outer circumscribed rectangle of the M main persons, the third cropping frame is used as the cropping frame. Among them, the third cropping frame is the second cropping frame, or the third cropping frame is the second cropping frame reduced to a cropping frame that contains the minimum outer circumscribed rectangle. If the second cropping frame does not completely contain the minimum outer circumscribed rectangle, the second cropping frame is expanded to contain the minimum outer circumscribed rectangle, and the expanded second cropping frame is used as the cropping frame.
[0209] Among them, the offset can be a preset value, or it can also be the distance between the center point of the first candidate cropping frame and the center point of the cropping frame of the previous frame of the video image multiplied by a weighting value, or obtained according to a preset algorithm. The embodiments of the present application do not specifically limit this.
[0210] Exemplarily, expanding or shrinking the candidate cropping frame can be implemented as expanding one or more sides of the candidate cropping frame outward or shrinking them inward.
[0211] Further, if the distance is less than the distance threshold, the electronic device can directly use the candidate cropping frame as the determined cropping frame.
[0212] Wherein, the distance between the center point of the candidate cropping frame and the center point of the cropping frame of the previous frame of video image can be the straight-line distance or others, and the embodiments of the present application do not specifically limit this.
[0213] S904. The electronic device shrinks or enlarges the cropped i-th frame of video image.
[0214] Specifically, the electronic device executes S904 so that the display screen displays the cropped i-th frame of video image according to the preset display specification. In S904, the electronic device shrinks or enlarges the cropped i-th frame of video image in S903 according to the preset display specification.
[0215] Wherein, the preset display specification can be the specification adapted to the display screen or a fixed screen-to-body ratio.
[0216] For example, if the resolution of the cropped i-th frame of video image in S903 is less than the preset display specification, in S904, the electronic device enlarges the cropped i-th frame of video image to an image with the preset display specification; if the resolution of the cropped i-th frame of video image in S903 is greater than the preset display specification, in S904, the electronic device shrinks the cropped i-th frame of video image to an image with the preset display specification; if the resolution of the cropped i-th frame of video image in S903 is equal to the preset display specification, in S904, the electronic device uses the cropped i-th frame of video image as an image with the preset display specification.
[0217] Further, after S904, for subsequent frames of video images, the electronic device can continue to execute the processes of S901 to S904, that is, traverse each frame of video image in the video stream with i+1, process frame by frame, and obtain and process one frame until the end of the video stream.
[0218] Through the video image processing method provided by the present application, when determining the main character of the video image, the character identity information of this frame of image and the character identity information of N video image frames before this frame are combined, so that the accuracy of the character perception process is greatly improved, and the accuracy of the determined position of the main character is correspondingly improved. In this way, it can be ensured that the main character can be completely displayed in the small-resolution image cropped and scaled according to the main character, so as to ensure the continuity of the presented picture of the main character, so as to achieve the effect of the picture following the person continuously through software during the image acquisition and display process.
[0219] Further, the video image processing method provided by this application may further include: The electronic device acquires the j-th frame of video image, where j is less than or equal to X; X is greater than 1. Acquire and save the identity information and / or location information of each person in the j-th frame of video image; directly reduce the j-th frame of video image to an image with a preset display specification. Among them, the identity information and / or location information of the j-th frame of video image can be used as reference information for subsequent frames of video image.
[0220] Of course, the electronic device may also acquire and save the personal information of each person in the j-th frame of video image.
[0221] Further, as Figure 16 shown, the image processing method provided by the embodiment of this application may further include S905.
[0222] S905. The electronic device displays the cropped i-th frame of video image according to the preset display specification.
[0223] In a possible implementation, the electronic device that executes the Figure 9 or Figure 16 shown video image processing method may be the sending device in a video call. The video image processing method provided by this application may further include: The electronic device encodes the image with the preset display specification obtained by reduction or enlargement and sends it to the receiving device, and the receiving device displays the cropped i-th frame of video image according to the preset display specification. For the specific process, refer to the Figure 2 workflow of the system architecture shown.
[0224] In a possible implementation, the electronic device that executes the Figure 9 or Figure 16 shown video image processing method may be the sending device in a video call. The video image processing method provided by this application may further include: The electronic device displays the cropped i-th frame of video image according to the preset display specification, and at the same time displays the cropped video image of the other party according to the preset display specification.
[0225] In a possible implementation, the electronic device that executes the Figure 9 or Figure 16 shown video image processing method may be the receiving device in a video call. The video image processing method provided by this application may further include: The electronic device displays the image with the preset specification obtained by reduction or enlargement through a display device. For the specific process, refer to the Figure 2 workflow of the system architecture shown.
[0226] Next, taking a specific video call scenario as an example, the video image processing method provided by the embodiment of this application will be described in detail.
[0227] Video call applications are installed in electronic device 1701 and electronic device 1702. The video call application is a client that can provide video call services for users. The video call applications installed in electronic device 1701 and electronic device 1702 can access a video call server through the Internet for data interaction to complete a video call and provide video call services for users using electronic device 1701 and electronic device 1702.
[0228] For example, as Figure 17A shown, on the main interface (i.e., the desktop) of electronic device 1701, there is an application icon 17011 of the video call application. As Figure 17B shown, on the desktop of electronic device 1702, there is an application icon 17021 of the video call application. Electronic device 1701 invokes the video call application to make a video call with electronic device 1702, and performs video image processing described in the embodiments of the present application on the video images during the video call.
[0229] For example, electronic device 1701 can receive a click operation (such as a touch click operation or an operation through a remote control device) by the user on Figure 17A the application icon 17011 shown, and display Figure 18A the video call application interface 1801 shown. The video call application interface 1801 includes a "New Friend" option 1802 and at least one contact option. For example, the at least one contact option includes a contact option 1803 for Bob and a contact option 1804 for user 311. Among them, the "New Friend" option 1802 is used to add new contacts. In response to a click operation (such as a click operation or an operation through a remote control device) by the user on the contact option 1804 for user 311, electronic device 1701 sends a video call request to electronic device 1702 logged in with the account of user 311 and makes a video call with electronic device 1702.
[0230] Exemplarily, in response to a click operation on the contact option 1804 by the user, electronic device 1701 can activate its own camera to collect an image with a fixed field of view as a scene image, and the display screen of electronic device 1701 displays a video call interface 1805 including the scene image collected by the camera as Figure 18B shown. The video call interface 1805 includes a prompt message "Waiting for the other party to respond!" 1806 and a "Cancel" button 1807. The "Cancel" button 1807 is used to trigger electronic device 1701 to cancel the video call with electronic device 1702.
[0231] Correspondingly, electronic device 1702 receives the video call request sent by electronic device 1701 from the video call server, and the display screen of electronic device 1702 displays a video call interface 1808 as Figure 18CAs shown in the figure. The video call interface 1808 includes a "Receive" button 1809 and a "Reject" button 1810. Among them, the "Receive" button 1809 is used for the electronic device 1702 to establish a video call connection with the electronic device 1701. The "Reject" button 1810 is used to trigger the electronic device 1702 to reject the video call request of the electronic device 1701.
[0232] The electronic device 1702 can receive the user's click operation on the "Receive" button 1809 (such as a touch click operation or an operation through a remote control device), and establish a video call connection with the electronic device 1701. After the connection is established, the electronic device 1701 and the electronic device 1702 are the two parties of the video call. The electronic device 1701 and the electronic device 1702 can respectively use their own cameras to collect images with a fixed field of view as scene images, and after cropping, scaling, and encoding frame by frame, send the scene images to the opposite end for display. The electronic device 1701 and the electronic device 1702 can display the cropped video images of the local end while displaying the cropped video images of the opposite end. Among them, during the video call process, in the process of the electronic device 1701 sending video images to the electronic device 1702, the electronic device 1701 is the sending end device and the electronic device 1702 is the receiving end device. In the process of the electronic device 1702 sending video images to the electronic device 1701, the electronic device 1702 is the sending end device and the electronic device 1701 is the receiving end device. The specific process of video image transmission between electronic devices can refer to Figure 2 the working process of the system architecture shown in the figure.
[0233] Among them, for the first X (for example, X is equal to 120) frames of video images, the electronic device 1701 and the electronic device 1702 can directly reduce the original image to an image with the display specification of the opposite end and encode it and send it to the opposite end. For the i-th frame (i is greater than 120) of video images, the electronic device 1701 and the electronic device 1702 can process it according to the video image processing method provided in the embodiments of the present application.
[0234] Exemplarily, at a certain moment during the video call between the electronic device 1701 and the electronic device 1702, the video image with a fixed field of view collected by the camera of the electronic device 1701 is as shown in Figure 19 (a) in the figure. The electronic device 1701 processes according to the video image processing method provided in the embodiments of the present application to determine the main character and crops and scales it into an image with the display specification of the electronic device 1702 as shown in Figure 19 (b) in the figure. The electronic device 1701 encodes the image shown in Figure 19 (b) in the figure and sends it to the electronic device 1702. At the same time, at this moment, the video image with a fixed field of view collected by the camera of the electronic device 1702 is as shown in Figure 19AAs shown in (a) of [Figure], the electronic device 1702 processes the video image according to the video image processing method provided in the embodiments of the present application to determine the main character and crops and scales it into an image with the display specifications of the electronic device 1701, such as Figure 19A As shown in (b) of [Figure], the electronic device 1702 will Figure 19A The image shown in (b) of [Figure] is encoded and sent to the electronic device 1701. At this time, the display interfaces of the electronic devices 1701 and 1702 are as shown in Figure 19B . As shown in Figure 19B , the large pictures on the main interfaces of the electronic devices 1701 and 1702 are the images of the other end after being cropped and scaled, and the small pictures are the images of the main characters determined by processing according to the video image processing method provided in the embodiments of the present application and cropped and scaled to their own display specifications. It should be noted that when the electronic device displays the image collected by itself, it can display the original image collected by itself or the image of the main character determined by processing according to the video image processing method provided in the embodiments of the present application and cropped and scaled to its own display specifications.
[0235] At another moment during the video call between the electronic device 1701 and the electronic device 1702, in the collection scene of the electronic device 1701, the position of the person changes. At this time, the video image of the fixed field of view collected by the camera of the electronic device 1701 is as shown in Figure 20 As shown in (a) of [Figure], the electronic device 1701 processes the video image according to the video image processing method provided in the embodiments of the present application to determine the main character and crops and scales it into an image with the display specifications of the electronic device 1702, such as Figure 20 As shown in (b) of [Figure]. The electronic device 1701 will Figure 20 The image shown in (b) of [Figure] is encoded and sent to the electronic device 1702. At the same time, at this moment, it is assumed that the position of the person in the collection scene of the electronic device 1702 is the same as that shown in Figure 19A and has not changed. At this time, the display interfaces of the electronic devices 1701 and 1702 are as shown in Figure 20A . As shown in Figure 20A , the large pictures on the main interfaces of the electronic devices 1701 and 1702 are the images of the other end after being cropped and scaled, and the small pictures are the images of the main characters determined by processing according to the video image processing method provided in the embodiments of the present application and cropped and scaled to their own display specifications.
[0236] At another moment during the video call between the electronic device 1701 and the electronic device 1702, in the collection scene of the electronic device 1701, the number of people increases. At this time, the video image of the fixed field of view collected by the camera of the electronic device 1701 is as shown in Figure 21As shown in (a) of [figure reference], the electronic device 1701 processes the determined main subject according to the video image processing method provided in the embodiments of the present application, crops and scales it into an image that conforms to the display specification of the electronic device 1702, such as Figure 21 as shown in (b) of [figure reference]. The electronic device 1701 encodes the Figure 21 image shown in (b) of [figure reference] and sends it to the electronic device 1702. At the same time, at this moment, the acquisition scene of the electronic device 1702 relative to Figure 19A has changed in the position of the person. At this time, the video image of the fixed field of view captured by the camera of the electronic device 1702 is as shown in Figure 21A (a) of [figure reference]. The electronic device 1702 processes the determined main subject according to the video image processing method provided in the embodiments of the present application, crops and scales it into an image that conforms to the display specification of the electronic device 1701, such as Figure 21A (b) of [figure reference]. The electronic device 1702 encodes the Figure 21A image shown in (b) of [figure reference] and sends it to the electronic device 1701. At this time, the display interfaces of the electronic device 1701 and the electronic device 1702 are as shown in Figure 21B . As shown in Figure 21B , the large pictures on the main interfaces of the electronic device 1701 and the electronic device 1702 are respectively the images that have been captured, cropped, and scaled by the other end. The small pictures are processed according to the video image processing method provided in the embodiments of the present application to determine the main subject, and are cropped and scaled into images that conform to their own display specifications.
[0237] At another moment during the video call between the electronic device 1701 and the electronic device 1702, in the acquisition scene of the electronic device 1701, the number of people increases and the position changes. At this time, the video image of the fixed field of view captured by the camera of the electronic device 1701 is as shown in Figure 22 (a) of [figure reference]. The electronic device 1701 processes the determined main subject according to the video image processing method provided in the embodiments of the present application, crops and scales it into an image that conforms to the display specification of the electronic device 1702, such as Figure 22 (b) of [figure reference]. The electronic device 1701 encodes the Figure 22 image shown in (b) of [figure reference] and sends it to the electronic device 1702. At the same time, at this moment, it is assumed that the position of the person in the acquisition scene of the electronic device 1702 is the same as that Figure 21A shown in [figure reference] and has not changed. At this time, the display interfaces of the electronic device 1701 and the electronic device 1702 are as shown in Figure 22A . As shown in Figure 22A , the large pictures on the main interfaces of the electronic device 1701 and the electronic device 1702 are respectively the images that have been captured, cropped, and scaled by the other end. The small pictures are the images processed according to the video image processing method provided in the embodiments of the present application to determine the main subject, and are cropped and scaled into images that conform to their own display specifications.
[0238] Taking a specific monitoring scenario as an example, the video image processing method provided by the embodiments of the present application will be described in detail below.
[0239] Suppose the monitoring system includes a camera 1, a server 2, and a display device 3. The camera 1 is used to collect video images of a fixed field of view. The server 2 is used to process the video images collected by the camera 1 through the video image processing method provided by the embodiments of the present application. The processed video images can be displayed in real time through the display device 3. The processed video images can also be stored in the storage device in the server 2. When the server 2 receives a reading instruction, it reads the processed video images from the storage device and displays them through the display device 3.
[0240] Exemplarily, at a certain moment during the operation of the monitoring system, the video image of the fixed field of view collected by the camera 1 is as shown in Figure 23 (a). The camera 1 sends the collected image to the server 2. The server 2 processes and determines the main character according to the video image processing method provided by the embodiments of the present application, and crops and scales it into an image with the display specification of the display device 3 as shown in Figure 23 (b). The server 2 displays the image shown in Figure 23 (b) in real time through the display device 3. At the same time, the server 2 stores the image shown in Figure 23 (b) in the storage device in the server 2. When the server 2 receives an instruction to read the video image, it reads the video image from the storage device and displays it through the display device 3.
[0241] At another moment during the operation of the monitoring system, the position of the person in the collection scene changes. At this time, the video image of the fixed field of view collected by the camera 1 is as shown in Figure 24 (a). The camera 1 sends the collected image to the server 2. The server 2 processes and determines the main character according to the video image processing method provided by the embodiments of the present application, and crops and scales it into an image with the display specification of the display device 3 as shown in Figure 24 (b). The server 2 displays the image shown in Figure 24 (b) in real time through the display device 3. At the same time, the server 2 stores the image shown in Figure 24 (b) in the storage device in the server 2. When the server 2 receives an instruction to read the video image, it reads the video image from the storage device and displays it through the display device 3.
[0242] At another moment during the operation of the monitoring system, the number of people in the collection scene increases. At this time, the video image of the fixed field of view collected by the camera 1 is as shown in Figure 25As shown in (a) therein, the camera 1 sends the captured image to the server 2. The server 2 processes the determined main character according to the video image processing method provided in the embodiment of the present application, crops and scales it into an image with the display specification of the display device 3, such as Figure 25 as shown in (b) therein. The server 2 will Figure 25 The image shown in (b) therein is displayed in real time through the display device 3. At the same time, the server 2 will Figure 25 The image shown in (b) therein is stored in the storage device in the server 2. When the server 2 receives an instruction to read the video image, it reads the video image from the storage device and displays it through the display device 3.
[0243] At another moment during the operation of the monitoring system, the number of people in the captured scene increases and their positions change. At this time, the video image of the fixed field of view captured by the camera 1 is as shown in Figure 26 as shown in (a) therein. The camera 1 sends the captured image to the server 2. The server 2 processes the determined main character according to the video image processing method provided in the embodiment of the present application, crops and scales it into an image with the display specification of the display device 3, such as Figure 26 as shown in (b) therein. The server 2 will Figure 26 The image shown in (b) therein is displayed in real time through the display device 3. At the same time, the server 2 will Figure 26 The image shown in (b) therein is stored in the storage device in the server 2. When the server 2 receives an instruction to read the video image, it reads the video image from the storage device and displays it through the display device 3.
[0244] The above mainly introduces the solution provided in the embodiment of the present application from the perspective of the electronic device. It can be understood that in order for the electronic device to implement the above functions, it includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combined with the examples described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0245] The embodiment of the present application can divide the functional modules of the electronic device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiment of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0246] In the case of dividing each functional module corresponding to each function, as Figure 27 shown in FIG. 270 is a video image processing apparatus provided by an embodiment of the present application, which is used to implement the functions of the electronic device in the above method. The video image processing apparatus 270 may be an electronic device, or a device in the electronic device, or a device that can be used in combination with the electronic device. Among them, the video image processing apparatus 270 may be a chip system. In the embodiment of the present application, the chip system may be composed of chips, or may include chips and other discrete devices. As Figure 27 shown, the video image processing apparatus 270 may include: an acquisition unit 2701, a determination unit 2702, a cropping unit 2703, and a scaling unit 2704. The acquisition unit 2701 is used to execute Figure 9 or Figure 16 S901, S901a in, the determination unit 2702 is used to execute Figure 9 or Figure 16 S902 in, the cropping unit 2703 is used to execute Figure 9 or Figure 16 S903 in, and the scaling unit 2704 is used to execute Figure 9 or Figure 16 S904 in. Among them, all relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be elaborated here.
[0247] Further, as Figure 27 shown, the video image processing apparatus 270 may further include a display unit 2705, which is used to execute Figure 16 S905 in.
[0248] As Figure 28 shown, FIG. 280 is a video image processing apparatus provided by an embodiment of the present application, which is used to implement the functions of the electronic device in the above method. The video image processing apparatus 280 may be an electronic device, or a device in the electronic device, or a device that can be used in combination with the electronic device. Among them, the video image processing apparatus 280 may be a chip system. The video image processing apparatus 280 includes at least one processing module 2801, which is used to implement the functions of the electronic device in the method provided by the embodiment of the present application. Exemplarily, the processing module 2801 may be used to execute Figure 9 or Figure 16 the processes S901, S901a, S902, S903, S904 in. For specific details, please refer to the detailed description in the method example, and will not be elaborated here.
[0249] The video image processing device 280 may further include at least one storage module 2802 for storing program instructions and / or data. The storage module 2802 is coupled to the processing module 2801. The coupling in the embodiments of the present application is an indirect coupling or communication connection between devices, units or modules, which may be electrical, mechanical or other forms for information interaction between devices, units or modules. The processing module 2801 may cooperate with the storage module 2802. The processing module 2801 may execute the program instructions stored in the storage module 2802. At least one of the at least one storage module may be included in the processing module.
[0250] The video image processing device 280 may further include a communication module 2803 for communicating with other devices through a transmission medium, so as to determine that the devices in the video image processing device 280 can communicate with other devices.
[0251] The video image processing device 280 may further include a display module 2804, which may be used to execute Figure 16 the process S905 in
[0252] When the processing module 2801 is a processor, the storage module 2802 is a memory, and the display module 2804 is a display screen, the video image processing device 280 involved in the embodiments of the present application may be Figure 28 the electronic device shown in Figure 8 the figure.
[0253] As described above, the video image processing device 270 or the video image processing device 280 provided in the embodiments of the present application may be used to implement the functions of the electronic device in the methods implemented in the above embodiments of the present application. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the embodiments of the present application.
[0254] Some other embodiments of the present application further provide a computer-readable storage medium, which may include computer software instructions. When the computer software instructions run on an electronic device, the electronic device is caused to execute each step executed by the electronic device in the above Figure 16 or
[0255] shown embodiments. Figure 9 or Figure 16 shown embodiments.
[0256] Some other embodiments of the present application further provide a chip system, which can be applied to an electronic device. The electronic device includes a display screen and a camera. The chip system includes an interface circuit and a processor; the interface circuit and the processor are interconnected through a line; the interface circuit is configured to receive a signal from a memory of the electronic device and send the signal to the processor, and the signal includes computer instructions stored in the memory; when the processor executes the computer instructions, the chip system performs each step executed by the electronic device in the embodiments as described above Figure 9 or Figure 16 shown in the embodiments.
[0257] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0258] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0259] The units described as separate components may or may not be physically separated. The components shown as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0260] In addition, each functional unit in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0261] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0262] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed rights.
Claims
1. A video image processing method, characterized in that, the method includes: Obtaining the identity information, person information, and position information of each person in the i-th frame of video image; where i>1; the person information includes one or more of the following information: speaking information, priority information; Determining M main characters from the i-th frame of video image according to the identity information of the characters in the N video image frames before the i-th frame of video image and the person information; M and N are greater than or equal to 1; Cropping the i-th frame of video image according to the position information of the main characters, and the cropped i-th frame of video image includes the M main characters; Reducing or enlarging the cropped i-th frame of video image so that the display screen displays the cropped i-th frame of video image according to a preset display specification; Wherein, the determining M main characters from the i-th frame of video image according to the identity information of the characters in the N video image frames before the i-th frame of video image includes: Determining the persons who have spoken in the N video image frames for a number of frames greater than or equal to a second preset threshold and appear in the i-th frame of video image as M main characters; or, determining the persons who have a priority information greater than a third preset threshold and appear in the i-th frame of video image as M main characters; or, selecting the most important M from the persons who have spoken in the N video image frames for a number of frames greater than or equal to a second preset threshold and appear in the i-th frame of video image according to the priority information and determining them as M main characters.
2. The method according to claim 1, characterized in that, the determining M main characters from the i-th frame of video image according to the identity information of the characters in the N video image frames before the i-th frame of video image includes: Determining the persons who have appeared in the N video image frames for a number of frames greater than or equal to a first preset threshold and appear in the i-th frame of video image as M main characters.
3. The method according to claim 1, characterized in that, the method further includes: dividing the i-th frame of video image into Y regions; configuring a preset threshold corresponding to each region; the preset threshold corresponding to the k-th region is the k-th preset threshold; the k-th region is any one of the Y regions; Y is greater than or equal to 2; k is greater than or equal to 1 and less than or equal to Y; the determining M main characters from the i-th frame of video image according to the identity information of the characters in the N video image frames before the i-th frame of video image includes: Determining the persons who have appeared in the N video image frames for a number of frames greater than or equal to the preset threshold corresponding to the region where they are located and appear in the i-th frame of video image as M main characters.
4. The method according to any one of claims 1-3, characterized in that, the cropping the i-th frame of video image according to the position information of the main characters includes: Determining a cropping frame, the cropping frame includes the minimum circumscribed rectangle frame of the M main characters; Cropping the i-th frame of video image with the cropping frame.
5. The method according to claim 4, wherein, the determining the cropping frame includes: obtaining the distance between the center point of the first candidate cropping frame and the center point of the cropping frame of the previous video image, where the first candidate cropping frame includes the minimum bounding rectangle of the M main characters; if the distance is greater than or equal to a distance threshold, determining a second cropping frame, where the center point of the second cropping frame is the center point of the cropping frame of the previous video image plus an offset, and the size of the second cropping frame is the same as the size of the cropping frame of the previous video image; if the second cropping frame contains the minimum bounding rectangle, taking a third cropping frame as the cropping frame; wherein, the third cropping frame is the second cropping frame, or the third cropping frame is the second cropping frame reduced to a cropping frame that contains the minimum bounding rectangle; if the second cropping frame does not completely contain the minimum bounding rectangle, expanding the second cropping frame to contain the minimum bounding rectangle, and taking the expanded second cropping frame as the cropping frame.
6. The method according to any one of claims 1-3, wherein, the method further includes: displaying the cropped i-th frame of video image according to a preset display specification.
7. A video image processing apparatus, wherein, the apparatus includes: an obtaining unit, configured to obtain the identity information, person information, and position information of each person in the i-th frame of video image; i>1; the person information includes one or more of the following information: speaking information, priority information; a determining unit, configured to determine M main characters from the i-th frame of video image according to the person identity information and the person information in the N video image frames before the i-th frame of video image; M, N≥1; a cropping unit, configured to crop the i-th frame of video image according to the position information of the main characters determined by the determining unit, and the cropped i-th frame of video image includes the M main characters; a scaling unit, configured to reduce or enlarge the cropped i-th frame of video image so that the display screen displays the cropped i-th frame of video image according to a preset display specification; the determining unit is specifically configured to: determine as M main characters the persons who have spoken in the N video image frames for a number of frames greater than or equal to a second preset threshold and appear in the i-th frame of video image; or determine as M main characters the persons whose priority information in the N video image frames is greater than a third preset threshold and appear in the i-th frame of video image; or determine as M main characters the M most important persons selected according to the priority information among the persons who have spoken in the N video image frames for a number of frames greater than or equal to a second preset threshold and appear in the i-th frame of video image.
8. The apparatus according to claim 7, wherein, the determining unit is specifically configured to: determine as M main characters the persons who have appeared in the N video image frames for a number of frames greater than or equal to a first preset threshold and appear in the i-th frame of video image.
9. The apparatus according to claim 7, wherein, The determining unit is specifically configured to: Divide the i-th frame of video image into Y regions; configure a preset threshold corresponding to each region; the preset threshold corresponding to the k-th region is the k-th preset threshold; the k-th region is any one of the Y regions; Y is greater than or equal to 2; k is greater than or equal to 1 and less than or equal to Y; Determine M main characters as those who appear in the i-th frame of video image and whose number of appearances in the N video image frames is greater than or equal to the preset threshold corresponding to their region.
10. The apparatus according to any one of claims 7-9, wherein, The cropping unit is specifically configured to: Determine a cropping frame, which includes the minimum bounding rectangle of the M main characters; Crop the i-th frame of video image with the cropping frame.
11. The apparatus according to claim 10, wherein, The cropping unit is specifically configured to: Obtain the distance between the center point of the first candidate cropping frame, which includes the minimum bounding rectangle of the M main characters, and the center point of the cropping frame of the previous frame of video image; If the distance is greater than or equal to a distance threshold, determine a second cropping frame, the center point of the second cropping frame is the center point of the cropping frame of the previous frame of video image plus an offset, and the size of the second cropping frame is the same as the size of the cropping frame of the previous frame of video image; If the second cropping frame contains the minimum bounding rectangle, use the third cropping frame as the cropping frame; wherein, the third cropping frame is the second cropping frame, or the third cropping frame is the second cropping frame shrunk to contain the minimum bounding rectangle; If the second cropping frame does not completely contain the minimum bounding rectangle, expand the second cropping frame to contain the minimum bounding rectangle, and use the expanded second cropping frame as the cropping frame.
12. The apparatus according to any one of claims 7-9, wherein, The apparatus further includes: A display unit, configured to display the cropped i-th frame of video image according to a preset display specification.
13. An electronic device, wherein, The electronic device includes: a processor and a memory; the processor is coupled with the memory, the memory is used to store computer program code, the computer program code includes computer instructions, when the computer instructions are executed by the electronic device, the electronic device is caused to execute the video image processing method according to any one of claims 1-6.
14. A computer-readable storage medium, wherein, It includes: Computer software instructions; When the computer software instructions run in an electronic device, the electronic device is caused to execute the video image processing method according to any one of claims 1-6.
15. A computer program product, wherein, When the computer program product runs on a computer, the computer is caused to execute the video image processing method according to any one of claims 1-6.
Citation Information
Patent Citations
Event-based media grouping, playback, and sharing
CN103023965A
Automated image cropping and sharing
CN105659286A