Close-up image determination method, apparatus, device, and storage medium
The method addresses the issue of participants appearing in multiple close-up screens by managing close-up frames to ensure each participant appears in only one frame, improving user experience and communication efficiency.
Patent Information
- Application Number
- JP2024568617
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2026-02-03
AI Technical Summary
In video communication, participants close to each other repeatedly appear in multiple close-up images, affecting the user's viewing experience and communication efficiency.
A method to identify and manage close-up frames for each participant, ensuring they appear in only one close-up screen by deleting redundant frames and reconstructing frames based on stable areas around the subject objects.
Reduces the likelihood of participants appearing in multiple close-up screens, enhancing user experience and communication efficiency by rationalizing the layout of close-up images.
Smart Images

Figure 2026503916000001_ABST
Abstract
Description
[Technical Field]
[0001] TECHNICAL FIELD The embodiments of the present application relate to the technical field of video calling, and in particular to a method, apparatus, device and storage medium for determining a close-up view. [Background technology]
[0002] In a scene where an electronic device is used to realize video communication (which can also be understood as a video call), in order to ensure that each participant in the video screen can be clearly seen, the electronic device creates a separate composition for each participant and then integrates and displays these compositions together. For example, FIG. 1 is a first schematic diagram of a video screen in a video communication scene in the related art. Referring to FIG. 1, there are four participants in the current video screen. The electronic device creates a separate composition for each of the four participants. For the screen in each composition, the screen surrounded by the close-up frame 11 including each participant in FIG. 1 can be referenced. The electronic device then integrates and displays the screens in the four close-up frames 11. FIG. 2 is a first schematic diagram of an integrated screen in a video communication scene in the related art. Referring to FIG. 2, the screen in the four close-up frames 11 in FIG. 1 is displayed so that the user of the electronic device can clearly see each participant in the video communication.
[0003] However, in a video communication process, if two participants are close to each other, creating a composition for one participant alone will include the other participant in the composition, and the other participant will appear not only in his or her own composition screen, but also in the composition screen of the participant who is close to him or her. For example, FIG. 3 is a second schematic diagram of a video screen in a video communication scene in the related art. Referring to FIG. 3, there are four participants in the current video screen. The electronic device creates a composition for each of the four participants. For the screen within each composition, the user can refer to the screen surrounded by the close-up frame 12 including each participant in FIG. 3. At this time, based on FIG. 3, it can be seen that participant 13 appears in the close-up frame 12 of participant 14 because the distance between participant 13 and participant 14 is close. Therefore, when the electronic device integrates and displays the screens within the four close-up frames 12, participant 13 appears repeatedly. FIG. 4 is a second schematic diagram of an integrated screen in a video communication scene in the related art. Referring to FIG. 4, the screens within the four close-up frames 12 in FIG. 3 are displayed. Based on FIG. 4, it can be seen that participant 13 appears in two close-up screens. This affects the user's viewing experience in video communication and also affects the user's communication efficiency in video communication. Summary of the Invention [Problem to be solved by the invention]
[0004] The embodiments of the present application provide a method, apparatus, device, and storage medium for determining a close-up image to solve a technical problem in the related art in that, when creating a close-up composition of participants in a video communication, participants who are close to the participant repeatedly appear in multiple close-up images. [Means for solving the problem]
[0005] In a first aspect, a method for determining a close-up image according to one embodiment of the present application includes: identifying a close-up frame to which each head object in a current frame screen of the video stream data belongs, each close-up frame corresponding to at least one subject object, the subject object being the head object determined to be surrounded by the corresponding close-up frame; selecting one of the close-up frames as a current close-up frame; if it is determined that there is another head object that can be added to the current close-up frame, determining the other head object as a subject object corresponding to the current close-up frame, and deleting the close-up frame constructed for the other head object; selecting another close-up frame from the close-up frame and updating it as the current close-up frame until all the current close-up frames have been traversed, and if it is determined that there is another head object that can be added to the current close-up frame, determining the other head object as a subject object corresponding to the current close-up frame and deleting the close-up frame constructed for the other head object; and displaying a close-up screen surrounded by each of the close-up frames in the video stream data.
[0006] As described above, the close-up frame to which each head object in the current frame screen of the video stream data belongs is identified, the close-up frame corresponds to at least one subject object, and the subject object is a head object that is determined to be surrounded by the corresponding close-up frame, then one close-up frame is selected from the close-up frames as the current close-up frame, and if it is determined that there are other head objects that can be added to the current close-up frame, the other head objects are determined as the subject objects corresponding to the current close-up frame and the close-up frames to which the other head objects belong are deleted, then another close-up frame is selected from the close-up frames and updated as the current close-up frame until all close-up frames have been traversed, and it is determined whether there are other head objects that can be added to the current close-up frame, and the close-up screens surrounded by each close-up frame in the video stream data are displayed. This technical means solves the technical problem in the related art that when video communication participants are composed in close-up, participants who are close to each other repeatedly appear in multiple close-up screens. When it is determined that a certain participant (i.e., a head object) can be added to another close-up frame, the close-up frame of the participant is deleted so that the participant appears in only one close-up screen, thereby avoiding the participant from appearing in multiple close-up screens as much as possible and improving the communication efficiency of users in video communication.
[0007] Based on the above embodiment, when it is determined that there is another head object that can be added to the current close-up frame, the step of determining the other head object as a subject object corresponding to the current close-up frame and deleting the close-up frame constructed for the other head object may include: determining whether the subject object in the current close-up frame has changed; if the object in the current close-up frame has changed, selecting the object in the current close-up frame with the largest area as the current object, and deleting the current close-up frame; reconstructing a close-up frame for the current subject object and using it as a current close-up frame; The method includes the steps of searching for other head objects that can be added to the current close-up frame, determining the other head objects as subject objects corresponding to the current close-up frame, and deleting the close-up frames constructed for the other head objects.
[0008] As described above, it is determined whether the subject object in the current close-up frame has changed, and if it has changed, the current subject object is searched for again to construct the current close-up frame, and then other head objects that can be added to the current close-up frame are searched for and the head objects are added to the current close-up frame. Thus, if the subject object in the close-up frame has changed, the close-up frame to be applied to each subject object can be reconstructed, and the rationality of the layout of the subject objects in the close-up screen can be ensured.
[0009] According to the above embodiment, the step of searching for another head object that can be added to the current close-up frame, determining the other head object as a subject object corresponding to the current close-up frame, and deleting the close-up frame constructed for the other head object, constructing a stable frame around the current subject object; moving the stable frame and searching for other head objects that can be added to the current close-up frame in the moving process, where the other head objects that can be added to the current close-up frame overlap with the stable frame at a corresponding point in the moving process; determining the retrieved other head object as a subject object corresponding to the current close-up frame, and deleting the close-up frame constructed for the other head object; and again executing the operation of selecting the object with the largest area in the current close-up frame, updating it as the current object, and constructing a stable frame with the current object as the center, until no other head object that can be added to the current close-up frame can be detected.
[0010] As described above, by constructing a stable frame, moving the stable frame, and searching for other head objects that can be added to the current close-up frame, the searched head objects can all be located in the surrounding area of the current subject object, i.e., close to the current subject object, thereby ensuring the rationality of the searched head objects; and by updating the subject object with the largest area to the current subject object, it can be ensured that the surrounding area of the subject object with the largest area is used as the reference when searching for a head object, and further ensure that the layout of the subject objects in the close-up screen is more rational.
[0011] Based on the above embodiment, after the step of determining whether the subject object in the current close-up frame has changed, if the subject object in the current close-up frame does not change, determining whether another head object is present in the current close-up frame; If there is another head object, perform the following operation: determining the other head object as a subject object corresponding to the current close-up frame, deleting the close-up frame constructed for the other head object, and selecting another close-up frame from the close-up frame to update it as the current close-up frame; If no other head object exists, the step of selecting another close-up frame from the close-up frame and updating it as the current close-up frame is included.
[0012] As described above, if the subject object in the current close-up frame does not change, by further searching whether there is another head object in the current close-up frame, it is possible to avoid missing a head object that can be added to the current close-up frame.
[0013] Based on the above embodiment, when another head object exists, the step of determining the other head object as a subject object corresponding to the current close-up frame and deleting the close-up frame constructed for the other head object includes: If there is another head object, selecting the object with the largest area in the current close-up frame and updating it as the current object; constructing a stable frame around the current subject object; moving the stable frame and searching for other head objects that can be added to the current close-up frame in the moving process, where the other head objects that can be added to the current close-up frame overlap with the stable frame at a corresponding point in the moving process; determining the retrieved other head object as a subject object corresponding to the current close-up frame, and deleting the close-up frame constructed for the other head object; and again executing the operation of selecting the object with the largest area in the current close-up frame, updating it as the current object, and constructing a stable frame with the current object as the center, until no other head object that can be added to the current close-up frame can be detected.
[0014] As described above, it is possible to avoid missing a head object that can be added to the current close-up frame, and to ensure the rationality of the searched head object.
[0015] Based on the above embodiment, in the process of moving the stable frame, the position of the corresponding current object does not change and is always kept within the stable frame.
[0016] As mentioned above, no matter which direction the stable frame is moved, it ensures that the current subject object does not exceed the stable frame, and at this time, the area covered in the stable frame moving process may all be considered as the surrounding area of the current subject object, and further ensures the accuracy of other head objects searched based on the stable frame.
[0017] Based on the above embodiment, the size of the stable frame is determined by the size of the corresponding current subject object in the current frame screen.
[0018] As described above, as long as the display size of the current subject object does not change, even if the aspect ratio of the close-up frame changes, the size of the corresponding stable frame will not change. Therefore, if the head object is relatively stable in the video stream data (i.e., does not move or moves slightly), the head object searched based on the stable frame will also be relatively fixed, which can further reduce the problem of shaking in the close-up screen caused by changes in the aspect ratio of the close-up frame.
[0019] According to the above embodiment, the step of moving the stable frame and searching for other head objects that can be added to the current close-up frame in the moving process includes: moving the stable frame, and in the moving process, searching for other head objects located within the current close-up frame that overlap with the stable frame; determining a minimum rectangular area including the retrieved other head objects and the subject object in the current close-up frame; and determining that if the width and height of the minimum rectangular area are smaller than the width and height of the stable frame, the other head object found can be added to the current close-up frame.
[0020] As described above, when searching for a head object based on a stable frame, the width and height of the smallest rectangular area including other head objects and each subject object in the current close-up frame are compared with the width and height of the stable frame to further determine whether the searched head object can be added to the current close-up frame, thereby ensuring that the distance between the head object that can be added to the current close-up frame and each subject object is close, and further ensuring that the layout of the subject objects in the close-up screen is more reasonable.
[0021] Based on the above embodiment, the step of determining whether the subject object in the current close-up frame has changed includes: determining whether the number of subject objects in the current close-up frame has changed; or determining whether the area change range of the object with the largest area in the current close-up frame exceeds an area threshold; or The method includes determining whether there is a subject object in the current close-up frame that exceeds the stable frame created for the subject object with the largest area.
[0022] As described above, it is possible to ensure that changes in the subject object are effectively identified.
[0023] Based on the above embodiment, the step of displaying a close-up screen surrounded by each close-up frame in the video stream data includes: If a close-up frame corresponds to a plurality of subject objects, creating a rectangular area corresponding to the close-up frame, the rectangular area being the smallest rectangular area that includes all of the subject objects in the corresponding close-up frame, creating a virtual object centered on a center point of the rectangular area, and updating an enclosing position of the close-up frame in the video stream data based on the virtual object so that the area of the virtual object is equal to the area of the subject object with the largest area in the corresponding close-up frame, and the updated close-up frame is centered on the virtual object; If the close-up frame corresponds to one object, updating the enclosing position of the close-up frame in the video stream data with the object object as the center; and displaying a close-up screen surrounded by each updated close-up frame in the video stream data.
[0024] As described above, when there are multiple subject objects corresponding to the close-up frame, by constructing a virtual object and adjusting the enclosing position of the close-up frame, it is possible to ensure that the distribution of each subject object in the close-up screen is more reasonable, without highlighting a certain subject object as the center.
[0025] According to the above embodiment, the step of identifying a close-up frame to which each head object in a current frame screen of the video stream data belongs includes: obtaining a most recently obtained close-up frame identification result for the video stream data; determining a close-up frame to which each head object in a current frame screen of the video stream data belongs based on the close-up frame identification result; After traversing all current close-up frames, The method further includes updating the close-up frame identification result.
[0026] In a second aspect, a close-up screen determination device according to an embodiment of the present application includes: a close-up frame identification unit used to identify a close-up frame to which each head object in a current frame screen of the video stream data belongs, each close-up frame corresponding to at least one subject object, and the subject object is a head object determined to be surrounded by the corresponding close-up frame; a close-up frame selection unit for selecting one of the close-up frames as a current close-up frame; a first object determination unit for determining, when determining that there is another head object that can be added to the current close-up frame, the other head object as a subject object corresponding to the current close-up frame, and deleting the close-up frame constructed for the other head object; a first close-up frame updating unit for selecting another close-up frame from the close-up frame and updating it as the current close-up frame until all the current close-up frames have been traversed, and when it is determined that there is another head object that can be added to the current close-up frame, determining the other head object as a subject object corresponding to the current close-up frame and deleting the close-up frame constructed for the other head object; and a close-up display unit for displaying a close-up screen surrounded by each of the close-up frames in the video stream data.
[0027] In a third aspect, a close-up screen determining device according to one embodiment of the present application includes: one or more processors; a display screen for displaying a close-up image; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the close-up screen determination method according to the first aspect.
[0028] In a fourth aspect, a computer-readable storage medium according to an embodiment of the present application stores a computer program, which, when executed by a processor, implements the close-up image determination method according to the first aspect.
[0029] For the beneficial effects of the above-mentioned close-up image determination apparatus, device, and storage medium, reference can be made to the beneficial effects of the close-up image determination method. [Brief explanation of the drawings]
[0030] [Figure 1] FIG. 1 is a first schematic diagram of a video screen in a video communication scene in the related art; [Figure 2] FIG. 1 is a first schematic diagram of an integrated screen in a video communication scene in the related art; [Figure 3] FIG. 2 is a second schematic diagram of a video screen in a video communication scene in the related art; [Figure 4] FIG. 2 is a second schematic diagram of an integrated screen in a video communication scene in the related art; [Figure 5] 1 is a schematic structural diagram of a close-up scene determining device according to one embodiment of the present application; [Figure 6] 1 is a flowchart of a close-up scene determination method according to one embodiment of the present application; [Figure 7] 1 is a schematic diagram of a close-up screen arrangement according to one embodiment of the present application; [Figure 8] FIG. 10 is a schematic diagram of another close-up screen arrangement according to one embodiment of the present application. [Figure 9] 10 is a flowchart of another close-up scene determination method according to an embodiment of the present application. [Figure 10] 1 is a schematic diagram of a close-up view and a stable frame according to one embodiment of the present application; [Figure 11] FIG. 10 is a schematic diagram of another close-up view and stable frame according to one embodiment of the present application. [Figure 12] 1 is a schematic diagram of a close-up frame according to one embodiment of the present application. [Figure 13] FIG. 1 is a schematic diagram of a virtual object according to one embodiment of the present application. [Figure 14]FIG. 10 is a schematic diagram of another close-up frame according to one embodiment of the present application. [Figure 15] 1 is an exemplary flowchart of a close-up scene determination method according to one embodiment of the present application. [Figure 16] 1 is a schematic structural diagram of a close-up scene determining device according to one embodiment of the present application; DETAILED DESCRIPTION OF THE INVENTION
[0031] The present application will be described in more detail below in combination with drawings and examples. It can be understood that the specific examples described here are for the purpose of interpreting the present application, but are not intended to limit the present application. In addition, for the purpose of explanation, the drawings only show a portion of the structure related to the present application, not all of it, in order to facilitate the explanation.
[0032] It should be noted that due to space limitations, the present specification does not cover all possible embodiments, and after reading the present specification, a person skilled in the art should be able to understand that any combination of technical features can constitute a possible embodiment as long as the technical features are not mutually contradictory.
[0033] Each example will be described in detail below.
[0034] Video communication can be understood as a method of realizing a video call over a network so that users in different locations can have the effect of face-to-face communication. Video communication is widely used in everyday communication, meetings, classes, and so on.
[0035] Currently, in order for a user to clearly view each participant in a video communication, an electronic device for video communication independently creates a composition for each participant on the video screen to obtain a close-up image of each participant. When independently creating a composition, a close-up frame is created for each participant with reference to the close-up frames shown in FIGS. 1 and 3. Here, the close-up frame may be understood as a rectangular area created around the corresponding participant on the video screen, and the image surrounded by the close-up frame on the video screen may be considered as the close-up image of the corresponding participant. When creating the close-up frames, each close-up frame may have the same aspect ratio but may have different sizes. Here, the larger the area of the corresponding participant on the video screen, the larger the size of the close-up frame required to surround the participant, and the smaller the area of the participant on the video screen, the smaller the size of the close-up frame required to surround the participant. Then, the electronic device displays the close-up image in each close-up frame according to a certain arrangement so that the user can clearly view each participant.
[0036] However, in a video communication scene, some participants may be close to each other. In this case, a close-up frame created for one participant may surround another participant who is close to the other participant, but the other participant may have a corresponding close-up frame. In this case, the other participant may appear in two close-up frames, such as participant 13 in FIG. 4. This may affect the user's viewing experience in the video communication and also affect the user's video communication efficiency. Specifically, when a user is paying attention to participant 13 in a video communication and participant 13 appears in two close-up frames, the user may not know which close-up frame to focus on when watching the video. The user may hesitate and miss the content of the speech during the video conference. Alternatively, when the user first discovers that participant 13 is in the close-up frame in the upper right corner of FIG. 4, the user may focus on participant 13 in the close-up frame in the upper right corner for a while. However, because participant 13 does not occupy a large proportion of the close-up frame, the user must expend effort to watch participant 13, which may affect the efficiency of the video conference. As can be seen, if a participant repeatedly appears in multiple close-up views, the efficiency of the video conference decreases, so there is a need to reduce the probability of this situation occurring.
[0037] Based on this, an embodiment of the present application provides a close-up screen determination method, which can fit multiple video communication participants who are close to each other into the same close-up screen, so as to avoid situations where a participant repeatedly appears in multiple close-up screens as much as possible and reduce the probability of such situations occurring.
[0038] The close-up screen determination method according to an embodiment of the present application may be performed by a close-up screen determination device. The close-up screen determination device may be implemented by software and / or hardware, and may be composed of two or more physical entities, or may be composed of one physical entity. Currently, the close-up screen determination device may be an electronic device capable of video communication, such as an interactive smart tablet, a tablet computer, or a notebook computer.
[0039] Fig. 5 is a schematic structural diagram of a close-up screen determining device according to an embodiment of the present application. Referring to Fig. 5, the close-up screen determining device includes a processor 21, a memory 22, and a display screen 23. Here, the processor 21, the memory 22, and the display screen 23 may be connected via a bus or other means, and Fig. 5 takes the connection via a bus as an example.
[0040] The number of processors 21 may be one or more, and Fig. 5 illustrates one processor 21 as an example. The processor 21 may include processing units such as an application processor (AP), a graphics processing unit (GPU), and a central processing unit (CPU).
[0041] The memory 22 may be used as a computer-readable storage medium to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the close-up screen determination method in the embodiments of the present application. The memory 22 may mainly include a program storage area and a data storage area, where the program storage area may store an operating system and / or an application program required for at least one function, and the data storage area may store data generated in response to use of the close-up screen determination device. The memory 22 may also include high-speed random access memory and non-volatile memory, such as at least one magnetic disk memory device, flash memory device, or other non-volatile solid-state memory device. In some embodiments, the memory 22 may further include memory located remotely from the processor 21, and these remote memories may be connected to the close-up screen determination device via a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. When the processor 21 has a storage function, it is understood that the memory 22 and the processor 21 may be integrated into one physical entity.
[0042] The number of display screens 23 (also referred to as screens or displays) may be one or more. FIG. 5 illustrates one display screen 23. The display screen 23 may be a liquid crystal display (LCD), an LED display, an organic light-emitting diode (OLED) display, a flexible light-emitting diode (FLED) display, or the like, but the embodiments are not limited thereto. The display screen may also be a linear screen or a curved screen, but the embodiments are not limited thereto. The display screen 23 may realize the display function of a close-up screen determination device. In one embodiment, the display screen 23 may be integrated with a touch function, and in this case, the display screen 23 includes a display panel and a touch panel. The display panel is used to complete visual output. The touch panel may be a touch component that supports infrared touch, electromagnetic touch, capacitive touch, resistive touch, or the like. The touch panel may pass detected touches to the processor 21 for further processing, and the display panel may provide visual output related to the touches.
[0043] In addition, the close-up screen determination device may further include one or more communication interfaces (not shown), through which it can communicate with other electronic devices, for example, to transmit data required for video communication with other electronic devices. The type of the communication interface is not currently limited.
[0044] The close-up screen determining device may further include devices such as a power supply, a speaker, and a physical button, and the embodiment is not limited thereto.
[0045] Based on the above hardware structure, the close-up screen determining device supports at least one kind of operating system, which may be an operating system such as an Android system, a Windows system, a Linux system, and so on.
[0046] The close-up screen determination device may have at least one application program installed in the operating system. The installed application program may be an application program provided by the operating system, or an application program downloaded from a back-end server or a third-party device. The close-up screen determination device can realize corresponding functions by executing each application program. Currently, the close-up screen determination device is installed with at least an application program for realizing video communication and an application program for realizing the current close-up screen determination method. In actual application, these two application programs may be integrated into one application program.
[0047] FIG. 6 is a flowchart of a close-up scene determination method according to an embodiment of the present application. Referring to FIG. 6, when a close-up scene determination device executes the close-up scene determination method, it specifically includes steps 310 to 350.
[0048] Step 310: Identify the close-up frame to which each head object in the current frame screen in the video stream data belongs, and each close-up frame corresponds to at least one subject object, and the subject object is the head object determined to be surrounded by the corresponding close-up frame.
[0049] Video stream data may be understood as video data used in current video communication. The video stream data may be data obtained by a close-up screen determination device after collecting the screens of local participants, or may be data including the screens of remote participants remotely transmitted and received by the close-up screen determination device. Here, participants refer to people participating in video communication. In the video communication process, the close-up screen determination device receives video stream data in real time and obtains video screens for each frame to realize video communication based on the video stream data. The video screen includes the screens of participants engaged in video communication. Generally, when the heads and faces of participants appear on the video screen, the heads of each participant in the video screen are used as head objects, and one head object can represent one participant. Currently, each head object can be obtained by detecting and identifying the head or face region of each person in the video screen. Existing image identification means can be used for the detection and identification means, and will not be described separately at present.
[0050] In a video communication process, a current frame screen refers to a video screen of one frame in currently received video stream data. A close-up frame refers to a rectangular frame generated based on a video screen, and at least one head object exists in the area enclosed by the close-up frame. Currently, the close-up frame is generated based on coordinates, for example, a rectangular close-up frame is generated based on four coordinates (head or face coordinates obtained by head detection and identification or face detection and identification). The close-up frame may be visible or invisible. In this case, the screen within the area enclosed by the close-up frame on the video screen may be referred to as a close-up screen enclosed by the close-up frame in the video stream data, and one or more head objects included in the close-up frame may be considered as subject objects corresponding to the close-up frame. That is, each close-up frame has at least one corresponding subject object, and the subject object may be considered as a head object belonging to the corresponding close-up frame, and the close-up screen corresponding to the close-up frame may be considered as a close-up screen when a subject object is close-up. At this time, it is determined whether a participant is surrounded by a close-up, and primarily whether the participant's head region (ie, head object) is surrounded by the close-up frame.
[0051] It can be understood that when the number of close-up frames changes, the aspect ratio corresponding to the close-up frames also changes, thereby ensuring that when close-up scenes in the close-up frames are displayed, each close-up scene can be adaptively arranged to ensure a good viewing effect. For example, Fig. 7 is a schematic diagram of the arrangement of close-up scenes according to one embodiment of the present application, and Fig. 8 is a schematic diagram of another arrangement of close-up scenes according to one embodiment of the present application. Fig. 7 shows a schematic diagram of the arrangement of close-up scenes when displaying close-up scenes in two close-up frames, where the two close-up scenes are arranged horizontally in a 1:1 ratio, and Fig. 8 shows a schematic diagram of the arrangement of close-up scenes when displaying close-up scenes in three close-up frames, where the three close-up scenes are arranged horizontally in a 1:1:1 ratio. When a close-up scene is displayed, the aspect ratio of the close-up scene is the same as the aspect ratio of the close-up frame and is related to the number of close-up frames. For example, the required aspect ratio when there are two close-up frames is different from the required aspect ratio when there are three close-up frames. Referring to Fig. 7, when there are two close-up frames, the aspect ratio of the close-up frames is 8:9, and referring to Fig. 8, when there are three close-up frames, the aspect ratio of the close-up frames is 16:27. Currently, when creating or updating close-up frames, a close-up screen determination device determines the aspect ratio of each close-up frame based on the number of close-up frames, and can further determine the size of the close-up frame on the video screen by combining the aspect ratio and the display size of the corresponding object on the video screen. When there are multiple object on the close-up frame, the size of the close-up frame is determined based on the object with the largest display size.Optionally, the size relationship between the close-up frame and the subject object is: height of close-up frame = p * height (or width) of the subject object on the video screen, width of close-up frame = q * height (or width) of the subject object on the video screen, where p and q are coefficients whose values are related to the current aspect ratio of the close-up frame, and whether the width or height of the subject object on the video screen is used to calculate the close-up frame may be preset by the close-up screen determination device.
[0052] When the close-up screen determination device receives a current frame screen, it identifies each head object therein and determines the close-up frame to which the head object belongs, where each head object has a close-up frame to which it belongs. In one embodiment, the close-up screen determination device determines the close-up frame to which each head object in the current frame screen belongs based on the close-up frame identification result of the previously received video stream data. Here, the close-up frame identification result describes each close-up frame created (i.e., existing) in the previous video screen and the subject object corresponding to each close-up frame. When the close-up screen determination device receives video stream data, it tracks the subject object and each close-up frame corresponding to each close-up frame in the video screen of each frame in the video stream data based on the previously obtained close-up frame identification result, so that when the video screen changes (from the previous frame screen to the next frame screen), it can determine the close-up frame to which each head object in the current frame screen belongs based on the tracking result. It can be understood that for the video screen of the first frame in the video stream data, the previous close-up frame identification result must be empty, and at this time, the close-up screen determination device can identify each head object and create a close-up frame for each head object, so as to obtain the close-up frame to which each head object belongs.When one or several new head objects appear in the video screen of the current frame in the video stream data, there is no close-up frame for the newly appearing head object in the close-up frame identification result, and at this time, one close-up frame can be created for each newly appearing head object to ensure that each head object in the video screen has a close-up frame to which it belongs.
[0053] Optionally, the close-up scene determining device identifies the close-up frame to which each head object belongs each time a frame of video screen is received, and executes subsequent steps (i.e., steps 320 to 350), that is, updates the close-up frame identification result for each frame; alternatively, the close-up scene determining device identifies the close-up frame to which each head object belongs each time a frame of video screen is received, and performs close-up display according to the subject object corresponding to the close-up scene of the close-up frame, and sets the number of frames or time length at regular intervals, and identifies the close-up frame to which each head object belongs in the current frame screen, and then executes subsequent steps (i.e., steps 320 to 350), and updates the close-up frame identification result at regular intervals.
[0054] Step 320: Select one close-up frame from the close-up frames as the current close-up frame.
[0055] There are usually a plurality of close-up frames, and one close-up frame is selected from the plurality of close-up frames as the current close-up frame.
[0056] Illustratively, the current close-up frame refers to the close-up frame that is currently updating the subject object. The close-up screen determination device may randomly select one close-up frame from each close-up frame in the current frame screen as the current close-up frame, or may select one close-up frame according to a set order (such as from top to bottom, left to right, etc.) as the current close-up frame, or may select one close-up frame according to the order of creating the close-up frames, where each close-up frame is created according to the order of identifying each face or each head in head identification or face identification.
[0057] Step 330: If it is determined that there is another head object that can be added to the current close-up frame, the other head object is determined as the subject object corresponding to the current close-up frame, and the close-up frame constructed for the other head object is deleted.
[0058] The existence of another head object that can be added to the current close-up frame can be understood as the other head object being included in the close-up screen of the current close-up frame, but the close-up frame to which the other head object belongs is not the current close-up frame, i.e., the other head object is currently being used as a subject object in another close-up frame.
[0059] The manner in which other head objects that can be selectively added to the current close-up frame are determined is currently unlimited.
[0060] For example, if it is detected that another head object appears in the close-up screen of the current close-up frame, it is determined that the other head object can be added to the current close-up frame. It can be understood that when identifying the head object and creating the close-up frame, the coordinates of the head object and the coordinates of the close-up frame can be known, and then, based on both coordinates, it can be determined whether the head object appears in the close-up frame and in which close-up frame.
[0061] Also, for example, a stable frame is created based on an existing subject object in a current close-up frame, and the stable frame refers to a rectangular region created around the corresponding subject object. The stable frame differs from the close-up frame in that the close-up frame is used to close up the subject object, and its aspect ratio changes depending on the number of close-up frames. However, the stable frame is used to determine whether there are other head objects around the corresponding subject object, i.e., whether there are any head objects in the vicinity. Similar to the close-up frame, the stable frame is generated based on coordinates, for example, a rectangular stable frame is generated based on four coordinates (coordinates of the head or face based on head identification or face identification). The stable frame may be visible or invisible. The size of the stable frame is determined by the display size of the corresponding object in the video stream data. In one embodiment, the height of the stable frame = n * the height (or width) of the object in the video screen, and the width of the stable frame = m * the height (or width) of the object in the video screen, where n and m are coefficients whose values are fixed. Whether the width or height of the object in the video screen is used to calculate the stable frame may be preset by the close-up frame determination device. By setting n and m to reasonable coefficients, the stable frame can cover as much area as possible in close-up frames of various aspect ratios to ensure the rationality and accuracy of the searched head object that is close. After the stable frame is created, other head objects that can be added to the current close-up frame can be determined based on the stable frame. Here, if only one object is present in the current close-up frame, a corresponding stable frame can be created based on the object. If multiple object objects are present in the current close-up frame, a corresponding stable frame can be created based on the object with the largest area in the video screen.When determining other head objects that can be added to the current close-up frame based on the stable frame, a head object other than the subject object with the largest area covered by the stable frame can be used as the head object that can be added to the current close-up frame, and the stable frame is moved around the subject object with the largest area corresponding to the stable frame as the center. In the process of moving the stable frame, other head objects covered by the stable frame can be searched for and used as the head objects that can be added to the current close-up frame. The other head objects currently covered by the stable frame are generally located within the current close-up frame.
[0062] In one embodiment, if it is detected that the subject objects in the current close-up frame change—for example, if the number of subject objects changes, if the subject objects move too far and the display area changes significantly, or if the subject objects move outside a designated stable frame (e.g., a stable frame created based on the subject object with the largest area)—this indicates that the current close-up frame may no longer be applicable to the subject objects within the current close-up frame. In this case, a close-up frame can be reconstructed for each subject object. To reconstruct the current close-up frame, the current close-up frame is deleted, a close-up frame is reconstructed for the subject object with the largest area, and this is used as the current close-up frame. Then, other head objects that can be added to the current close-up frame are determined until no other head objects remain. If there are still head objects without close-up frames, a close-up frame is created for the head object, and when the close-up frame is updated to the current close-up frame, other head objects that can be added to the close-up frame are searched for. Correspondingly, if it is detected that the subject object in the current close-up frame does not change, it indicates that the current close-up frame also applies to the internal subject object, and at this time, it only needs to determine whether there is another head object (the close-up frame to which the head object belongs is not the current close-up frame) in the current close-up frame, and if there is another head object, it determines that there is another head object that can be added to the current close-up frame; if not, there is no need to update the subject object corresponding to the current close-up frame.
[0063] For example, if it is determined that there is another head object that can be added to the current close-up frame, the other head object is used as the subject object of the current close-up frame. At this time, the number of subject objects corresponding to the current close-up frame increases by one. After using the other head object as the subject object of the current close-up frame, in an embodiment, to prevent the other head object from appearing in other close-up frames, the close-up frame constructed for the other head object is deleted. Optionally, if other subject objects still exist in the close-up frame after the close-up frame is deleted, the close-up frame can be recreated for the other subject object. At this time, if the number of close-up frames changes, the aspect ratio of the close-up frame also changes.
[0064] Also, if a head object is selectively added to the current close-up frame as a corresponding subject object first, when subsequent processing is performed on other close-up frames, if it is determined that the head object can also be added to other close-up frames to be processed thereafter, the head object is not processed, i.e., the close-up frame in which it is first determined that the head object can be added is used as the reference.
[0065] Optionally, each time one other head object is added to the current close-up frame, another other head object that can be added to the current close-up frame is searched for until no other head object can be found, at which point the current close-up frame and the corresponding subject object are updated, after which step 340 is performed.
[0066] Step 340: Select another close-up frame from the close-up frame and update it as the current close-up frame until all the current close-up frames have been traversed; if it is determined that there is another head object that can be added to the current close-up frame, determine the other head object as the subject object corresponding to the current close-up frame, and delete the close-up frame constructed for the other head object; and perform this operation again.
[0067] Illustratively, another unupdated close-up frame is selected as the current close-up frame, and returning to perform step 330 to update the current close-up frame and its corresponding subject object again.
[0068] It can be understood that every time one current close-up frame is updated, a search is made to see if there is currently a close-up frame that is not being used as the current close-up frame, and if there is, one close-up frame is selected from the close-up frames that are not being used as the current close-up frame and updated as the current close-up frame, and then the process returns to execute step 330. If there is not, all close-up frames and their corresponding subject objects may be considered to have been updated, and then step 350 can be executed.
[0069] Optionally, after traversing all the close-up frames, the currently determined close-up frame and the subject object corresponding to the close-up frame can be saved, i.e., the close-up frame identification result can be updated for subsequent use.
[0070] Step 350 displays the close-up screen surrounded by each close-up frame in the video stream data.
[0071] For example, the close-up images surrounded by the close-up frames in the current frame image of the video stream data are acquired, and each close-up frame has a corresponding close-up image, and then the close-up images are displayed on the display screen according to the corresponding arrangement (related to the number of the close-up frames).
[0072] In one embodiment, when the subject objects in the close-up frames are updated, the composition difference between different close-up frames is reduced to ensure that the layout of each subject object in the close-up screen is more reasonable, and when displaying the close-up screen of each close-up frame, the enclosed area of the close-up frame is adjusted based on the subject objects in the close-up frame. When making the adjustment, first determine the smallest rectangular area that includes all the subject objects in the close-up frame, then create one virtual object centered on the center point of the smallest rectangular area, where the virtual object refers to a virtual head area, and the area of the virtual object is equal to the area of the subject object with the largest area in the close-up frame. At this time, it may be considered that a head object with the same area is created as a virtual object based on the subject object with the largest area, and the center of the virtual object is the center of the smallest rectangular area. Then, determine the position of the close-up frame based on the position of the virtual object, and adjust the size / ratio of the close-up frame according to the aspect ratio at which the close-up frame needs to be displayed (the aspect ratio at which the display needs to be displayed is set in advance, for example, in the case of the above-mentioned two close-up frames, the aspect ratio of the close-up frame is 8:9, and in the case of three close-up frames, the aspect ratio of the close-up frame is 16:27) and size (determined based on the size of the virtual object, for example, the height of the close-up frame is determined by multiplying the height of the virtual object by a preset coefficient). In this case, since each close-up frame surrounds the close-up screen based on the position and size of the virtual object, the difference in composition between different close-up screens is reduced, and the area difference between the largest object in each close-up screen is small when the close-up screen is displayed. If there is only one object in the close-up frame, there is no need to create a virtual object, and it is only necessary to update the surrounding position of the close-up frame in the video stream data around the object.Then, the close-up image for each close-up frame can be displayed.
[0073] As described above, the close-up frame to which each head object in the current frame screen of the video stream data belongs is identified, the close-up frame corresponds to at least one subject object, and the subject object is the head object determined to be surrounded by the corresponding close-up frame; then, one close-up frame is selected from the close-up frames as the current close-up frame; if it is determined that there are other head objects that can be added to the current close-up frame, the other head objects are determined to be the subject objects corresponding to the current close-up frame and the close-up frames to which the other head objects belong are deleted; then, another close-up frame is selected and updated as the current close-up frame until all close-up frames have been traversed; and it is determined whether there are other head objects that can be added to the current close-up frame; and close-up images surrounded by each close-up frame in the video stream data are displayed. This technical means solves the technical problem in the related art that when video communication participants are close up to compose a composition, participants who are close to each other repeatedly appear in multiple close-up images, and improves user communication efficiency in video communication. If it is determined that a certain participant (i.e., a head object) can be added to another close-up frame, the close-up frame of the participant is deleted so that the participant appears in only one close-up screen, thereby avoiding the participant appearing in multiple close-up screens as much as possible.
[0074] 9 is a flowchart of another close-up scene determination method according to one embodiment of the present application, which exemplarily describes a process of determining other head objects that can be added to the current close-up frame based on the above embodiment, a process of determining a close-up frame to which each head object in the current frame screen belongs, and a process of displaying a close-up scene of each close-up frame in the video stream data. Referring to FIG. 9, the close-up scene determination method includes steps 410 to 4310.
[0075] Step 410: obtain the most recently obtained close-up frame identification result for the video stream data.
[0076] The close-up frame identification result describes each close-up frame created in the video screen and the subject object corresponding to each close-up frame. Every time the close-up frame identification result is updated, the currently updated close-up frame identification result is saved as the most recently obtained close-up frame identification result.
[0077] In one embodiment, for example, a time length or a frame number is set to update the close-up frame identification result at a fixed interval. In this case, after updating the close-up frame identification result once, start counting or start recording the frame number. When a current frame screen in the video stream data is obtained, determine whether the corresponding counting time reaches the set time length or the corresponding frame number reaches the set frame number. If so, determine that the close-up frame identification result needs to be updated. At this time, obtain the latest close-up frame identification result, i.e., obtain the most recently updated close-up frame identification result, and execute step 420. If not, track each close-up frame and each object within the video screen of each frame in the video stream data according to the latest close-up frame identification result, and display each close-up screen surrounded by the close-up frame.
[0078] Step 420: Based on the close-up frame identification result, determine the close-up frame to which each head object in the current frame screen of the video stream data belongs.
[0079] For example, after identifying each head object in the current frame screen, it can be determined which close-up frame each head object should be a subject object of based on the close-up frame identification result. Optionally, since the positions of the same head object in the video screen in consecutive frames are the same or very close, it can be determined which subject object each head object should be by combining the position of the head object and the position of the subject object. Then, based on the close-up frame to which each subject object belongs in the close-up frame identification result, it can be determined which close-up frame each head object belongs to in the current frame screen.
[0080] For the first frame screen in the video stream data, the latest close-up frame identification result obtained must be empty, and at this time, one close-up frame can be created for each head object in the screen.
[0081] Step 430: Select one close-up frame from the close-up frames as the current close-up frame.
[0082] Step 440: determining whether the object of the subject in the current close-up frame has changed; if the object of the subject in the current close-up frame has changed, execute step 450. If the object of the subject in the current close-up frame has not changed, execute step 480.
[0083] For example, in a process of tracking video stream data and displaying a close-up screen based on the close-up frame and the corresponding object recorded in the close-up frame identification result, when the close-up frame identification result needs to be updated, it can be determined whether the object in the current close-up frame in the current frame screen has changed. If it has changed, it indicates that the current close-up frame may not be applicable to the corresponding object. In this case, step 450 is performed. If it has not changed, it indicates that the current close-up frame is also applicable to the corresponding object. In this case, step 480 is performed.
[0084] In one embodiment, determining whether the subject objects in the current close-up frame have changed may be determining whether the number of subject objects in the current close-up frame has changed, or determining whether the area change range of the subject object with the largest area in the current close-up frame has exceeded an area threshold, or determining whether there is a subject object in the current close-up frame that has exceeded the stability frame created for the subject object with the largest area.
[0085] Here, when determining whether the number of subject objects in the current close-up frame has changed, it mainly determines whether the number of subject objects has decreased. That is, it determines whether the subject objects have moved away from the current close-up frame. If the number has changed, it indicates that the subject objects have moved away from the current close-up frame, and therefore it is necessary to reconstruct the close-up frame applicable to each subject object. In one embodiment, the current close-up frame and its corresponding subject objects are tracked based on the close-up frame identification result, and it is determined whether the subject objects have moved away from the current close-up frame based on the position of the subject objects in the video stream data and the position of the current close-up frame in the video stream data, and it is further determined whether the number of subject objects in the current close-up frame has changed. If the number has changed, it is determined that the subject objects have changed, and step 450 is executed; if not, step 480 is executed.
[0086] When determining whether the area change range of the object with the largest area in the current close-up frame exceeds the area threshold, the area change range may be understood as the change range of the display area of the object in the close-up frame, and the area threshold is a preset value. If the area change range of the object with the largest area in the current close-up frame exceeds the area threshold, it may indicate that the object with the largest area may be changed to another object (i.e., the display area of the original object with the largest area may become smaller), or it may indicate that the area of the object with the largest area may become larger. Since the size of the close-up frame is related to the area of the object with the largest area, if the size of the object with the largest area changes, the size of the close-up arm also needs to be adjusted in a timely manner. Therefore, when the size is adjusted, the object to be applied to the close-up arm may also change. Therefore, the close-up arm needs to be re-updated based on the object with the largest area, and the object to be applied needs to be re-determined. In one embodiment, when tracking each object in the current close-up frame, the display area of each object (which may be obtained from the width and height of the object) can be detected, and then the area change range can be determined based on the area of the object with the largest area in the current close-up frame on the current frame screen and the display area of the object in the previous close-up frame identification result determination. If the area change range exceeds the area threshold, it is determined that the object has changed, and step 450 is executed; otherwise, step 480 is executed.
[0087] It is determined whether the current close-up frame contains an object whose area exceeds the stability frame created for the object with the largest area. Specifically, a stability frame is created based on the object with the largest area in the current close-up frame. In the process of tracking each object and its corresponding close-up frame in the video stream data, it is identified whether other objects close to the object with the largest area are moving away through the stability frame. If other objects exceed the stability frame, it indicates that the other objects may be moving away from the object with the largest area. A single close-up frame may not be able to simultaneously encompass these two objects; that is, the other objects and the object with the largest area may need to belong to different close-up frames. Therefore, when tracking each object in the video stream data, a stability frame may be created for the object with the largest area in the current close-up frame, and it is determined whether other objects in the current close-up frame have moved away from the stability frame. If the current close-up frame contains an object whose area exceeds the stability frame created for the object with the largest area, it is determined that the object has changed, and step 450 is executed; otherwise, step 480 is executed.
[0088] In practical applications, if any one of the above three schemes is met, the subject object in the current close-up frame may be considered to have changed.
[0089] Step 450: Select the object with the largest area in the current close-up frame as the current object, and delete the current close-up frame.
[0090] For example, the display area of each subject object corresponding to the current close-up frame in the current frame screen can be determined based on the close-up screen surrounded by the current close-up frame in the current frame screen. Then, the object with the largest area is selected as the current object, and a close-up frame to be applied is reconstructed based on the current object. Alternatively, an existing current close-up frame can be deleted, and in this case, each object corresponding to the current close-up frame is not used as a subject object of the current close-up frame. Optionally, if the deleted current close-up frame contains other object objects other than the object with the largest area, close-up frames that are not used as the current close-up box can be created for the other object objects to ensure that each head object in the current frame screen has a close-up frame to which it belongs.
[0091] It can be understood that the larger the area of the subject object in the current frame screen, the larger the area of the subject object in the close-up screen when the close-up screen is displayed.
[0092] Step 460 reconstructs a close-up frame for the current subject object and uses it as the current close-up frame.
[0093] For example, a close-up frame is reconstructed for the current object and used as the current close-up frame. When constructing a close-up frame, the aspect ratio of the close-up frame is determined based on the number of close-up frames, and the size of the close-up frame is determined based on the aspect ratio of the close-up frame and the height (or width) of the current object in the current frame screen. The close-up frame is then created centered on the current object and according to the size. It can be understood that when the number of close-up frames changes, the aspect ratio of the close-up frame also adaptively changes. For example, when a subject object in a close-up frame is added to another close-up frame, the close-up frame is deleted, and the number of close-up frames changes. In this case, the aspect ratio of the close-up frame is determined based on the number of the latest close-up frames, and the size of each current close-up frame is then changed in combination with the aspect ratio of the close-up frame.
[0094] After the creation is completed, the close-up frame is used as the current close-up frame, and at this time, the current close-up frame includes one object, that is, the current object is used as the object of the current close-up frame.
[0095] 470, if another head object that can be added to the current close-up frame is detected, the other head object is determined as the subject object corresponding to the current close-up frame, and the close-up frame constructed for the other head object is deleted. Step 4100 is performed.
[0096] For example, other head objects that can be added to the current close-up frame are searched for. Here, the search method may be to use other head objects located in the current close-up frame as head objects that can be added to the current close-up frame. The search method may be to determine head objects that are close to the current subject object by constructing a stable frame, and then obtain head objects that can be added to the current close-up frame.
[0097] In one embodiment, a stable frame is constructed to search for a head object that can be added to the current close-up frame. In this case, step 470 includes steps 471 to 474.
[0098] Step 471: Build a stable frame centered on the current subject object.
[0099] In one embodiment, the size of the stable frame is determined by the size of the corresponding current object in the current frame screen, where height of the stable frame = n * height (or width) of the object in the video screen, width of the stable frame = m * height (or width) of the object in the video screen, n and m are coefficients with fixed values, and whether the width or height of the object in the video screen is used to calculate the stable frame may be preset by the close-up screen determination device.
[0100] For example, a preset coefficient for determining the size of the stable frame can be obtained, and the size of the stable frame can be determined based on the width (or height) of the current object in the current frame screen and the preset coefficient. Then, a stable frame of a corresponding size is constructed with the current object as the center, and the stable frame is the stable frame of the current object. The area where the stable frame is located can be considered as the area close to the current object. At this time, as long as the display size of the current object does not change, the size of the stable frame will not change regardless of whether the aspect ratio of the close-up frame changes.
[0101] Generally, the stable frame is located inside the close-up frame, and rarely exceeds the close-up frame.
[0102] Step 472: Move the stable frame, and in the moving process, search for other head objects that can be added to the current close-up frame, and the other head objects that can be added to the current close-up frame overlap with the stable frame at the corresponding time in the moving process.
[0103] After the stable frame is constructed, the stable frame is moved. In one embodiment, during the process of moving the stable frame, the position of the corresponding current object remains unchanged and always remains within the stable frame. That is, regardless of the direction of moving the stable frame, the current object does not exceed the stable frame, so that the stable frame always moves around the current object. In this case, the area covered during the process of moving the stable frame can be considered as the peripheral area of the current object, and the accuracy of other head objects searched based on the stable frame is further ensured. Specifically, during the process of moving the stable frame, if a certain other head object is covered, the other head object can be considered close to the current object, and the other head object can be added to the current close-up frame corresponding to the current object. Here, the other head object that can be covered by the stable frame refers to the display area of the other head object in the current frame screen being completely located within the stable frame. Other head objects currently covered in the stable frame may be considered as head objects that can be added to the current close-up frame, and generally, the head objects are located within the current close-up frame and at some point (i.e., a corresponding point) in the stable frame movement process, the head objects overlap with the stable frame.
[0104] When moving the stability frame, the moving direction of the stability frame may be preset, and it is only necessary to ensure that the surrounding area of the current subject object is covered in the stability frame moving process.
[0105] In one embodiment, in the process of moving the stable frame, multiple other head objects may be covered, and in this case, to avoid that many head objects are searched and the area consisting of many head objects exceeds the area of the stable frame, only one head object among them is selected as the currently searched other head object. Selectably, from the multiple covered other head objects, the other head object closest to the current subject object is selected as the currently searched head object.
[0106] In one embodiment, in the process of moving the stable frame, there may be overlapping with multiple head objects, but not every head object can be added to the current close-up frame, for example, in the process of moving the stable frame, there may be overlapping with a head object, but the head object is far from another subject object in the current close-up frame, so it is not suitable to be added to the current close-up frame, whereby this step can further include steps 4721 to 4723.
[0107] Step 4721 moves the stable frame, and in the moving process, searches for other head objects that are located in the current close-up frame and overlap with the stable frame.
[0108] For example, based on the coordinates of each head object and the coordinates of the current close-up frame, other head objects located in the current close-up frame can be identified, and then the stable frame can be used to search for these other head objects located in the current close-up frame. Here, in the process of moving the stable frame, if it overlaps with some other head objects in the current close-up frame, it may be considered that the currently searched other head objects may be added to the current close-up frame.
[0109] Optionally, when multiple head objects overlap in the moving process, the head object closest to the subject object corresponding to the stable frame can be selected, or the head object that is searched earliest can be selected.
[0110] In step 4722, the smallest rectangular area that includes the other head objects found and the subject object in the current close-up frame is determined.
[0111] After the other head objects are found in the stable frame, it can be further determined whether the other head objects can be added to the current close-up frame. If so, first determine the smallest rectangular area that includes the other head objects found and the subject object in the current close-up frame.
[0112] Specifically, after other head objects are found in the stable frame, a rectangular area is determined, which includes the other head objects found and each subject object corresponding to the current close-up frame, and is the smallest rectangular area that includes the objects. The smallest rectangular area may be understood as the smallest area required to include the objects.
[0113] Step 4723 determines that if the width and height of the minimum rectangular area are smaller than the width and height of the stable frame, the other head objects found can be added to the current close-up frame.
[0114] After the minimum rectangular region is obtained, it is determined whether the width of the minimum rectangular region is smaller than the width of the stable frame and whether the height of the minimum rectangular region is smaller than the height of the stable frame. If both are smaller, it indicates that the stable frame can completely cover the minimum rectangular region, that is, other head objects searched for in the stable frame may be covered by the stable frame, that is, the detected other head objects are close to the subject objects corresponding to the current close-up frame, and one close-up frame can be used to surround each subject object and the searched head object, and therefore the searched other head objects can be added to the current close-up frame, and then step 473 is performed. If the width of the minimum rectangular region is larger than the width of the currently used stable frame or the height of the minimum rectangular region is larger than the height of the currently used stable frame, it indicates that the searched other head objects are far from one or more subject objects corresponding to the current close-up frame, and using one close-up frame may not be able to surround all subject objects and the searched other head objects, and therefore the currently searched head object is discarded. At this time, it is determined that a head object that can be added to the current close-up frame cannot be detected, and step 4100 can be performed.
[0115] It should be noted that, as long as the display size of the current object does not change, even if the aspect ratio of the close-up frame changes, the size of the corresponding stable frame will not change. Therefore, if the head object is relatively stable in the video stream data (i.e., does not move or moves slightly), the head object searched based on the stable frame will also be relatively fixed, thereby mitigating the problem of close-up image shaking caused by the change in the aspect ratio of the close-up frame. For example, FIG. 10 is a schematic diagram of close-up images and stable frames according to an embodiment of the present application, and FIG. 11 is a schematic diagram of another close-up image and stable frame according to an embodiment of the present application. Currently, when changing from two close-up images shown in FIG. 10 to three close-up images shown in FIG. 11, the aspect ratio of the close-up image obviously changes, but the size of the stable frame 41 does not change, and therefore the object corresponding to the close-up frame finally searched based on the stable frame 41 also does not obviously change, i.e., the object in the close-up image does not obviously change, and further, the shaking of the close-up image caused by the change in the aspect ratio of the close-up frame can be effectively avoided.
[0116] In step 473, the searched other head object is determined as the subject object corresponding to the current close-up frame, and the close-up frame constructed for the other head object is deleted.
[0117] If another head object that can be added to the current close-up frame is found, the other head object is used as a subject object of the current close-up frame. At this time, the number of subject objects corresponding to the current close-up frame is increased by one. After the other head object is used as the subject object of the current close-up frame, in an embodiment, the close-up frame constructed for the other head object is deleted to prevent the other head object from appearing in other close-up frames. Optionally, after the close-up frame is deleted, if other subject objects other than the other head object still exist in the close-up frame, the close-up frame can be recreated for the other subject object.
[0118] Step 474: Select the object with the largest area in the current close-up frame, update it as the current object, and build a stable frame around the current object, and repeat this process until no other head object that can be added to the current close-up frame can be detected.
[0119] For example, when the searched head object is added to the current close-up frame, the object in the current close-up frame is updated, and at this time, the object with the largest area may change. Based on the principle of searching for a head object around the object with the largest area, when the object in the current close-up frame is updated, the object with the largest area is reselected and updated as the current object, and the process returns to step 471. That is, after the stable frame is created, a stable frame is re-created for the current object until no other head objects that can be added to the current close-up frame are found, and other head objects close to the current object are searched for and added to the current close-up frame. At this time, if it is determined that no other head objects can be detected through the stable frame, another close-up frame may be selected as the current close-up frame, and it may be determined again whether another head object can be added to the current close-up frame, that is, step 4100 may be executed.
[0120] Step 480 determines whether there is another head object in the current close-up frame. If there is another head object, execute step 490; if there is no other head object, execute step 4100.
[0121] For example, if the subject object in the current close-up frame does not change, it is possible to further determine other head objects in the current close-up frame, i.e., to further determine whether a head object as a non-subject object is added to the framing area corresponding to the current close-up frame.
[0122] If another head object is present in the current close-up frame, the display area of the other head object in the current frame screen may be considered to be located in the current close-up frame.
[0123] In one embodiment, if there is another head object in the current close-up frame, then execute step 490 to determine the other head object as the subject object of the current close-up frame; otherwise, execute step 4100.
[0124] Step 490: Determine the other head object as the subject object corresponding to the current close-up frame, and delete the close-up frame constructed for the other head object. Step 4100 is executed.
[0125] For example, if another head object exists in the current close-up frame, the other head object is used as the subject object of the current close-up frame. At this time, the number of subject objects corresponding to the current close-up frame is increased by one. After the other head object is used as the subject object of the current close-up frame, in order to prevent the other head object from appearing in other close-up frames, in an embodiment, the close-up frame constructed for the other head object is deleted. Optionally, if another subject object other than the other head object still exists in the close-up frame after the close-up frame is deleted, a close-up frame can be recreated for the other subject object.
[0126] In one embodiment, if another head object in the current close-up frame is located at the edge of the current close-up frame and is far from each subject object in the current close-up frame, adding the other head object to the current close-up frame may result in a poor layout of the subject objects in the close-up screen of the current close-up frame, which may prevent the user from clearly seeing the head object located at the edge. Therefore, it is better to include the other head object in another close-up frame (the other head object is located at a more central position in the other close-up frame). In this case, the other head object may appear in two close-up frames, but it may be located at the center of one close-up frame and at the edge of the other close-up frame (only a part of its area may be located within the close-up frame), which may have a small impact on the user's visual impression when viewing the close-up screen, and the head object with a large area may not be visible within the two close-up frames. Based on this, if another head object is present in the current close-up frame, it can be further determined whether the other head object can be added to the current close-up frame.At this time, if there is another head object, the other head object is determined as the subject object corresponding to the current close-up frame, and the close-up frame constructed for the other head object is deleted. Specifically, if there is another head object, the subject object with the largest area in the current close-up frame is selected and updated as the current subject object, a stable frame is constructed with the current subject object as the center, the stable frame is moved, and in the moving process, other head objects that can be added to the current close-up frame are searched for, and the other head objects that can be added to the current close-up frame overlap with the stable frame at the corresponding time point in the moving process, the searched other head object is determined as the subject object corresponding to the current close-up frame, the close-up frame constructed for the other head object is deleted, and the subject object with the largest area in the current close-up frame is selected and updated as the current subject object, and a stable frame is constructed with the current subject object as the center. This operation is repeated until no other head objects that can be added to the current close-up frame are detected.
[0127] For example, if there is another head object in the current close-up frame, the object with the largest area in the current close-up frame is selected and updated as the current object, and then a stable frame is created, and an object that can be added to the current close-up frame is searched for based on the stable frame. For this part, please refer to the related description of steps 471 to 474. This continues until there is no other head object in the current close-up frame, or until it is determined whether each other head object in the current close-up frame can be used as the object corresponding to the current close-up frame, i.e., until it is determined that there is no object that can be added to the current close-up frame. Then, step 4100 can be executed.
[0128] Step 4100: Determine whether there are any close-up frames that have not been traversed. If all close-up frames in the current video stream data have been traversed, execute step 4110. If all close-up frames in the current video stream data have not been traversed, execute step 4130.
[0129] Illustratively, when each close-up frame in the current frame screen is used as the current close-up frame and the subject objects therein are confirmed, it can be determined that each current required close-up frame and their corresponding subject objects have been obtained, and therefore step 4100 can be performed. If there is a close-up frame in the current frame screen that has not been traversed as the current close-up frame, the untraversed close-up frame is updated to the current close-up frame, i.e., step 4130 is performed to continue determining the subject object corresponding to the current close-up frame.
[0130] Step 4110: Update the close-up frame identification result. Step 4120 is executed.
[0131] For example, each currently obtained close-up frame and the object object corresponding to each close-up frame are used as the currently obtained close-up frame identification result, and replace the previously obtained close-up frame identification result to realize updating of the close-up frame identification result.The close-up frame identification result is used to track each head object in the video stream data and the close-up frame to which the head object belongs, and display the close-up screen of each close-up frame.When the close-up frame identification result needs to be updated next time, the currently obtained close-up frame identification result can be used to determine the close-up frame to which each head object in the video screen belongs.
[0132] Step 4120 displays the close-up screen of each close-up frame in the video stream data.
[0133] In one embodiment, step 4120 includes the steps of: creating a rectangular area corresponding to the close-up frame, if there are multiple subject objects corresponding to the close-up frame, the rectangular area being the smallest rectangular area that includes all subject objects in the corresponding close-up frame; creating a virtual object centered on the center point of the rectangular area; updating an enclosing position of the close-up frame in the video stream data based on the virtual object, such that the area of the virtual object is equal to the area of the subject object with the largest area in the corresponding close-up frame, and the updated close-up frame is centered on the virtual object; if there is only one subject object corresponding to the close-up frame, updating the enclosing position of the close-up frame in the video stream data to center on the subject object; and displaying a close-up screen enclosed by each updated close-up frame in the video stream data.
[0134] For example, when multiple object images are present in a close-up frame, a rectangular area is created to position each object in the central area as much as possible and to more rationally display each object in the close-up screen. The rectangular area is the smallest rectangular area that includes each object (mainly the head area of the object). The rectangular area can clarify the area in which each object is concentrated and distributed. Then, a virtual object is created centered on the center point of the rectangular area and designated as a virtual object. The area of the virtual object is equal to the area of the object with the largest area in the close-up frame. In this case, the virtual object may be considered as the simulated object with the largest area displayed at the center of the close-up frame. Then, the position of the close-up frame in the video stream data is updated with the virtual object as the center. The current position of the close-up frame in the video stream data is designated as an enclosing position, and the center of the enclosing position is the virtual object. Here, updating the enclosing position of the close-up frame in the video stream data based on the virtual object includes centering the virtual object in the close-up frame and updating the enclosing position of the close-up frame in the video stream data based on the size of the close-up frame. The size of the close-up frame is related to the aspect ratio of the close-up frame and the maximum area of each subject object, so that each subject object can be distributed as close to the center of the close-up frame as possible. For example, Fig. 12 is a schematic diagram of a close-up frame according to one embodiment of the present application. Referring to Fig. 12, there are two subject objects in the close-up frame 51, which are currently marked as subject object 52 and subject object 53, respectively. At this time, a minimum rectangular area 54 including the two head areas is created based on the head area of the subject object 52 and the head area of the subject object 53.Then, a virtual object is created with the center point of the minimum rectangular region 54 as its center, and the display area of the virtual object is equal to the display area of the subject object 52. FIG. 13 is a schematic diagram of a virtual object according to one embodiment of the present application. Referring to FIG. 13, a virtual head region is created as virtual object 55 within the minimum rectangular region 54 shown in FIG. 12. Then, the position of the close-up frame is changed based on the position of virtual object 55. FIG. 14 is a schematic diagram of another close-up frame according to one embodiment of the present application. Referring to FIG. 14, a close-up frame 56 centered on virtual object 55 and a stable frame 57 centered on virtual object 55 are shown. At this time, when subsequently determining whether the subject object in the close-up frame 56 has changed, the close-up frame 56 and the stable frame 57 can be used as references.
[0135] If one object is present in the close-up frame, the object is positioned at the center of the video stream data, and therefore the enclosing position of the close-up frame in the video stream data is updated with the object at the center.
[0136] Then, the close-up screen in each close-up frame can be displayed according to the set size. Here, the set size is related to the arrangement of the close-up screen. For example, referring to FIG. 7, when the close-up screens are arranged in a 1:1 ratio, the size of the current display area of each close-up screen can be determined, and then the close-up screen in each close-up frame can be displayed in each display area.
[0137] Step 4130: Select another close-up frame from the close-up frames and update it as the current close-up frame. Return to execute step 440.
[0138] As described above, by determining whether the object in the current close-up frame has changed, and if so, re-searching for the current object to construct a current close-up frame, and then searching for other head objects that can be added to the current close-up frame, the system can ensure that the close-up frames applicable to each object are reconstructed when the object in the close-up frame has changed, and further ensure the rationality of the layout of the object in the close-up scene. Also, by constructing a stable frame and moving the stable frame to search for other head objects that can be added to the current close-up frame, all of the searched head objects can be located in the peripheral area of the current object, i.e., close to the current object, thereby ensuring the rationality of the searched head objects. Furthermore, by updating the object with the largest area as the current object, the system can ensure that the peripheral area of the object with the largest area is used as a reference when searching for a head object, and further ensure the rationality of the layout of the object in the close-up scene. In addition, when searching for a head object based on a stable frame, the width and height of the smallest rectangular area including other head objects and each subject object in the current close-up frame are compared with the width and height of the stable frame to further determine whether the searched head object can be added to the current close-up frame, thereby ensuring that the distance between the head object added to the current close-up frame and each subject object is close, and further ensuring that the layout of the subject objects in the close-up screen is more reasonable.Furthermore, if the subject object in the current close-up frame does not change, it is determined whether another head object exists in the current close-up frame. If another head object exists, it is further determined whether another head object can be added to the current close-up frame, thereby avoiding overlooking a head object that can be added to the current close-up frame. Furthermore, by detecting changes in the number, area, etc. of the subject object, it is determined whether the subject object in the current close-up frame has changed, thereby effectively identifying changes in the subject object. Furthermore, if there are multiple subject objects corresponding to the close-up frame, it is possible to ensure a more reasonable distribution of the subject objects in the close-up screen by constructing a virtual object and adjusting the enclosing position of the close-up frame, without highlighting one subject object as the center.
[0139] The following is an exemplary description of a close-up scene determination method according to an embodiment of the present application, in which the close-up frame identification result is updated every second. Figure 15 is an exemplary flowchart of a close-up scene determination method according to one embodiment of the present application.
[0140] Referring to FIG. 15, the close-up image determination method includes steps 510 to 5170.
[0141] Step 510: based on the close-up frame identification result, track each head object in the video stream data, and display the close-up screen of the close-up frame to which each head object belongs.
[0142] In step 520, after an interval of 1 s, the close-up frame to which each head object in the current frame screen of the video stream data belongs is determined based on the close-up frame identification result.
[0143] Step 530: Select one close-up frame from the close-up frames as the current close-up frame.
[0144] Step 540, determining whether the subject object in the current close-up frame has changed, and if so, performing step 550; If there is no change, step 5120 is executed.
[0145] Step 550: Use the object with the largest display area in the current close-up frame as the current object, and delete the current close-up frame.
[0146] Step 560: construct a close-up frame and a stable frame for the current subject object, and use the constructed close-up frame as the current close-up frame.
[0147] Step 570: Move the stable frame, and in the moving process, search for other head objects located in the current close-up frame that overlap with the stable frame. If no other head objects are found, execute step 5130. If other head objects are found, execute step 580.
[0148] Step 580 determines whether the height and width of the smallest rectangular area including the other head objects found and each subject object corresponding to the current close-up frame are smaller than the height and width of the stable frame. If so, execute step 590; if not, execute step 5110.
[0149] Step 590: Determine the other head object found as the subject object corresponding to the current close-up frame, and delete the close-up frame constructed for the other head object. Step 5100 is executed.
[0150] Step 5100 selects the subject object with the largest area in the current close-up frame to update it as the current subject object, builds a stable frame for the current subject object, and returns to execute step 570 .
[0151] Step 5110: Abandon determining the currently searched other head object as the subject object corresponding to the current close-up frame. Step 5130 is executed.
[0152] Step 5120 determines whether there is another head object in the current close-up frame. If there is another head object, execute step 5100. If there is no other head object, execute step 5130.
[0153] Step 5130: Determine whether there are any untraversed close-up frames. If there are any untraversed close-up frames, execute step 5140. If there are no untraversed close-up frames, execute step 5150.
[0154] Step 5140: Select another close-up frame from the close-up frames and update it as the current close-up frame. Return to execute step 540.
[0155] Step 5150: if there are multiple subject objects corresponding to the close-up frame, create a rectangular area corresponding to the close-up frame, create a virtual object centered on the center point of the rectangular area, and update the enclosing position of the close-up frame in the video stream data based on the virtual object; if there is one subject object corresponding to the close-up frame, update the enclosing position of the close-up frame in the video stream data centered on the subject object.
[0156] In step 5160, the close-up screen surrounded by each updated close-up frame in the video stream data is displayed.
[0157] That is, the close-up frame identification result is updated, and then the process can return to step 510. This process continues until the current video communication ends or the display of the close-up screen ends.
[0158] The above example makes it possible to avoid, as much as possible, a single head object appearing repeatedly in multiple close-up screens, making the layout of the head objects in each close-up screen more rational and improving the user's viewing experience of the close-up screens.
[0159] An embodiment of the present application further provides a close-up scene determining device. Figure 16 is a schematic structural diagram of a close-up scene determining device according to an embodiment of the present application. Referring to Figure 16, the close-up scene determining device includes: a close-up frame identifying unit 601, a close-up frame selecting unit 602, a first object determining unit 603, a first close-up frame updating unit 604 and a close-up display unit 605.
[0160] Here, the close-up frame identification unit 601 is used to identify a close-up frame to which each head object in a current frame screen in the video stream data belongs, and each of the close-up frames corresponds to at least one subject object, and the subject object is a head object determined to be surrounded by the corresponding close-up frame; the close-up frame selection unit 602 is used to select one of the close-up frames as a current close-up frame; if the first object determination unit 603 determines that there is another head object that can be added to the current close-up frame, it determines the other head object as a subject object corresponding to the current close-up frame, and selects the other head object from the close-up frames. the first close-up frame updating unit 604 is used to select another close-up frame from the close-up frame and update it as the current close-up frame until all the current close-up frames have been traversed; if it is determined that there is another head object that can be added to the current close-up frame, it is used to determine the other head object as the subject object corresponding to the current close-up frame and delete the close-up frame constructed for the other head object; and the close-up display unit 605 is used to display a close-up screen surrounded by each of the close-up frames in the video stream data.
[0161] Based on the above embodiment, the first object determination unit 603 includes: a change determination subunit for determining whether the subject object in the current close-up frame has changed; a first object selection subunit for selecting the subject object with the largest area in the current close-up frame as the current subject object and deleting the current close-up frame if the subject object in the current close-up frame has changed; a close-up frame reconstruction subunit for reconstructing a close-up frame for the current subject object and using it as the current close-up frame; and a second object determination subunit for searching for other head objects that can be added to the current close-up frame, determining the other head objects as subject objects corresponding to the current close-up frame and deleting the close-up frames constructed for the other head objects.
[0162] Based on the above embodiment, the second object determination subunit includes: a first stable frame construction subunit for constructing a stable frame around the current object; a first stable frame movement subunit for moving the stable frame and searching for other head objects that can be added to the current close-up frame in the moving process, where the other head objects that can be added to the current close-up frame overlap with the stable frame at the corresponding time in the moving process; a third object determination subunit for determining the searched other head objects as the object object corresponding to the current close-up frame and deleting the close-up frame constructed for the other head objects; and a first object updating subunit for selecting the object with the largest area in the current close-up frame and updating it as the current object until no other head objects that can be added to the current close-up frame can be found, and then re-performing the operation of constructing a stable frame around the current object.
[0163] Based on the above embodiment, the close-up screen determination device further includes: an object existence determination unit for determining whether another head object exists in the current close-up frame after determining whether the subject object in the current close-up frame has changed, if the subject object in the current close-up frame has not changed; a fourth object determination unit for performing the following operation if the other head object exists: determine the other head object as the subject object corresponding to the current close-up frame, delete the close-up frame constructed for the other head object, and select another close-up frame from the close-up frame to update it as the current close-up frame; and a second close-up frame updating unit for performing the operation if the other head object does not exist: select another close-up frame from the close-up frame to update it as the current close-up frame.
[0164] Based on the above embodiment, the fourth object determining unit includes: a second object selecting subunit, for selecting the object with the largest area in the current close-up frame and updating it as the current object if there are other head objects; a second stable frame constructing subunit, for constructing a stable frame around the current object; a second stable frame moving subunit, for moving the stable frame and searching for other head objects that can be added to the current close-up frame in the moving process, so that the other head objects that can be added to the current close-up frame overlap with the stable frame at a corresponding time in the moving process; a fifth object determining subunit, for determining the searched other head objects as the object corresponding to the current close-up frame and deleting the close-up frame constructed for the other head objects; and a second object updating subunit, for repeating the operation of selecting the object with the largest area in the current close-up frame and updating it as the current object and constructing a stable frame around the current object until no other head objects that can be added to the current close-up frame can be found.
[0165] Based on the above embodiment, in the process of moving the stable frame, the position of the corresponding current object does not change and is always kept within the stable frame.
[0166] Based on the above embodiment, the size of the stable frame is determined by the size of the corresponding current subject object in the current frame screen.
[0167] Based on the above embodiment, the first stable frame moving subunit and the second stable frame moving subunit are specifically used to move the stable frame, and in the moving process, search for other head objects located in the current close-up frame and overlapping with the stable frame, determine a minimum rectangular area including the searched other head objects and the subject object in the current close-up frame, and if the width and height of the minimum rectangular area are smaller than the width and height of the stable frame, determine that the searched other head objects can be added to the current close-up frame.
[0168] Based on the above embodiment, the change determination subunit is specifically used to determine whether the number of object objects in the current close-up frame has changed, or to determine whether the area change range of the object with the largest area in the current close-up frame has exceeded an area threshold, or to determine whether there is an object in the current close-up frame that exceeds the stable frame created for the object with the largest area.
[0169] Based on the above embodiment, the close-up display unit 605 includes: a first enclosing position updating subunit for creating a rectangular area corresponding to the close-up frame when there are multiple subject objects corresponding to the close-up frame, the rectangular area being the smallest rectangular area that includes all subject objects in the corresponding close-up frame; creating a virtual object centered on the center point of the rectangular area, and updating the enclosing position of the close-up frame in the video stream data based on the virtual object, such that an area of the virtual object is equal to an area of the subject object with the largest area in the corresponding close-up frame, and the updated close-up frame is centered on the virtual object; a second enclosing position updating subunit for updating the enclosing position of the close-up frame in the video stream data around the subject object when there is only one subject object corresponding to the close-up frame; and a close-up screen display subunit for displaying the close-up screen enclosed by each updated close-up frame in the video stream data.
[0170] According to the above embodiment, the close-up frame identification unit 601 includes a result obtaining subunit for obtaining a close-up frame identification result most recently obtained for the video stream data, and a close-up identification subunit for determining a close-up frame to which each head object in a current frame screen of the video stream data belongs according to the close-up frame identification result. Correspondingly, the close-up screen determination device further includes a result updating unit for updating the close-up frame identification result after traversing all current close-up frames.
[0171] The above close-up scene determining device may be used to implement the close-up scene determining method according to any of the above embodiments, and has corresponding functions and beneficial effects.
[0172] It should be noted that in the above embodiment of the close-up screen determination device, each unit and module provided is only divided according to functional logic, and is not limited to the above divisions, as long as it can realize the corresponding function, and the specific names of each functional unit are only for facilitating mutual division, and are not used to limit the protection scope of the present invention.
[0173] An embodiment of the present application further provides a close-up screen determination device. Referring to Fig. 5, the close-up screen determination device includes a processor 21, a memory 22, and a display screen 23. Here, the display screen 23 is used to display a close-up screen, and the memory 22 is used to store one or more programs. When the one or more programs are executed by the one or more processors 21, the one or more processors 21 realize the close-up screen determination method described in any of the above embodiments. For the related content of each component, please refer to the above description.
[0174] The above-mentioned close-up screen determination device can be used to perform any close-up screen determination method, and comprises a close-up screen determination apparatus having corresponding functions and beneficial effects, and for details not currently described, reference can be made to the relevant description of the above-mentioned close-up screen determination method.
[0175] One embodiment of the present application provides a storage medium containing computer-executable instructions, which, when executed by a processor of a close-up scene determination device, are used to perform relevant operations in the close-up scene determination method provided by any embodiment of the present application and have corresponding functions and beneficial effects.
[0176] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program.
[0177] Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining hardware and software. Furthermore, the present application may take the form of a computer program product embodied in one or more computer-usable storage media (including, but not limited to, magnetic disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code. The present application is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or a processing module of another programmable data processing device to generate a machine, whereby the instructions, executed by the processing module of the computer or other programmable data processing device, generate an apparatus for implementing the function specified in one or more flows in the flowcharts and / or one or more blocks in the block diagrams. These computer program instructions may be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, whereby the instructions stored in the computer-readable memory produce an article of manufacture that includes an instruction apparatus that implements the functions specified in one or more flows in the flowcharts and / or one or more blocks in the block diagrams.These computer program instructions may be loaded into a computer or other programmable data processing device and cause the computer or other programmable device to perform a series of operational steps to generate a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows in the flowcharts and / or one or more blocks in the block diagrams.
[0178] Computer-readable media includes persistent and non-persistent, removable and non-removable media capable of storing information in any manner or technology, such as computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only disks (CD-ROM), digital versatile disks (DVD) or other optical memory, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic memory, or any other non-transmission medium available for storing information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0179] It should also be understood that the terms "comprise," "include," or any variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, product, or device comprising a set of elements includes not only those elements but also other elements not expressly stated, or includes the inherent elements of such process, method, product, or device. Unless further limited, elements defined by the phrase "comprise one or more" do not exclude the presence of other identical elements in the process, method, product, or device that includes the elements. It should be noted that the above is merely a description of the preferred embodiments and applied technical principles of the present application. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious modifications, adjustments, and substitutions can be made without departing from the scope of protection of the present application. Therefore, although the present application will be described in more detail through the above embodiments, the present application is not limited only to the above embodiments and can further include many other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the appended claims.
Claims
1. A close-up image determination method, comprising: identifying a close-up frame to which each head object in a current frame screen of the video stream data belongs, each close-up frame corresponding to at least one subject object, the subject object being the head object determined to be surrounded by the corresponding close-up frame; selecting one of the close-up frames as a current close-up frame; if it is determined that there is another head object that can be added to the current close-up frame, determining the other head object as a subject object corresponding to the current close-up frame, and deleting the close-up frame constructed for the other head object; selecting another close-up frame from the close-up frame and updating it as the current close-up frame until all the current close-up frames have been traversed, and if it is determined that there is another head object that can be added to the current close-up frame, determining the other head object as a subject object corresponding to the current close-up frame and deleting the close-up frame constructed for the other head object; and displaying a close-up screen surrounded by each of the close-up frames in the video stream data.
2. When it is determined that there is another head object that can be added to the current close-up frame, the step of determining the other head object as a subject object corresponding to the current close-up frame and deleting the close-up frame constructed for the other head object includes: determining whether the subject object in the current close-up frame has changed; if the object in the current close-up frame has changed, selecting the object in the current close-up frame with the largest area as the current object, and deleting the current close-up frame; reconstructing a close-up frame for the current subject object and using it as a current close-up frame; 2. The close-up screen determination method according to claim 1, further comprising the steps of: searching for another head object that can be added to the current close-up frame; determining the other head object as a subject object corresponding to the current close-up frame; and deleting the close-up frame constructed for the other head object.
3. the step of searching for another head object that can be added to the current close-up frame, determining the other head object as a subject object corresponding to the current close-up frame, and deleting the close-up frame constructed for the other head object, constructing a stable frame around the current subject object; moving the stable frame and searching for other head objects that can be added to the current close-up frame in the moving process, where the other head objects that can be added to the current close-up frame overlap with the stable frame at a corresponding point in the moving process; determining the retrieved other head object as a subject object corresponding to the current close-up frame, and deleting the close-up frame constructed for the other head object; and re-executing the operation of selecting the object with the largest area in the current close-up frame, updating it as the current object, and constructing a stable frame around the current object, until no other head object that can be added to the current close-up frame can be detected.
4. After the step of determining whether the subject object in the current close-up frame has changed, if the subject object in the current close-up frame does not change, determining whether another head object is present in the current close-up frame; If there is another head object, perform the following operation: determining the other head object as a subject object corresponding to the current close-up frame, deleting the close-up frame constructed for the other head object, and selecting another close-up frame from the close-up frame to update it as the current close-up frame; and if no other head object exists, performing an operation of selecting another close-up frame from the close-up frames and updating it as the current close-up frame.
5. If another head object exists, the step of determining the other head object as a subject object corresponding to the current close-up frame and deleting the close-up frame constructed for the other head object includes: If there is another head object, selecting the object with the largest area in the current close-up frame and updating it as the current object; constructing a stable frame around the current subject object; moving the stable frame and searching for other head objects that can be added to the current close-up frame in the moving process, where the other head objects that can be added to the current close-up frame overlap with the stable frame at a corresponding point in the moving process; determining the retrieved other head object as a subject object corresponding to the current close-up frame, and deleting the close-up frame constructed for the other head object; and repeating the operation of selecting the object with the largest area in the current close-up frame, updating it as the current object, and constructing a stable frame around the current object, until no other head object that can be added to the current close-up frame can be detected.
6. The method for determining a close-up scene according to claim 3 or 5, characterized in that in the process of moving the stable frame, the position of the corresponding current subject object does not change and is always maintained within the stable frame.
7. The close-up scene determining method according to claim 3 or 5, wherein the size of the stable frame is determined by the size of the corresponding current object in the current frame scene.
8. The step of moving the stable frame and searching for other head objects that can be added to the current close-up frame in the moving process includes: moving the stable frame, and in the moving process, searching for other head objects located within the current close-up frame that overlap with the stable frame; determining a minimum rectangular area including the retrieved other head objects and the subject object in the current close-up frame; and determining that the searched other head object can be added to the current close-up frame if the width and height of the smallest rectangular area are smaller than the width and height of the stable frame.
9. The step of determining whether the subject object in the current close-up frame has changed includes: determining whether the number of subject objects in the current close-up frame has changed; or determining whether the area change range of the object with the largest area in the current close-up frame exceeds an area threshold; or 3. The method of claim 2, further comprising the step of determining whether the current close-up frame includes an object whose area exceeds the stable frame created for the object with the largest area.
10. The step of displaying a close-up screen surrounded by each close-up frame in the video stream data includes: If a close-up frame corresponds to a plurality of subject objects, creating a rectangular area corresponding to the close-up frame, the rectangular area being the smallest rectangular area that includes all of the subject objects in the corresponding close-up frame, creating a virtual object centered on a center point of the rectangular area, and updating an enclosing position of the close-up frame in the video stream data based on the virtual object so that the area of the virtual object is equal to the area of the subject object with the largest area in the corresponding close-up frame, and the updated close-up frame is centered on the virtual object; If the close-up frame corresponds to one object, updating the enclosing position of the close-up frame in the video stream data with the object object as the center; 2. The close-up image determining method according to claim 1, further comprising the step of: displaying a close-up image surrounded by each updated close-up frame in the video stream data.
11. The step of identifying a close-up frame to which each head object in a current frame screen of the video stream data belongs includes: obtaining a most recently obtained close-up frame identification result for the video stream data; determining a close-up frame to which each head object in a current frame screen of the video stream data belongs based on the close-up frame identification result; After traversing all current close-up frames, The close-up frame determining method according to claim 1 , further comprising the step of updating the close-up frame identification result.
12. A close-up image determination device, a close-up frame identification unit used to identify a close-up frame to which each head object in a current frame screen of the video stream data belongs, each close-up frame corresponding to at least one subject object, the subject object being a head object determined to be surrounded by the corresponding close-up frame; a close-up frame selection unit for selecting one of the close-up frames as a current close-up frame; a first object determination unit for, when determining that there is another head object that can be added to the current close-up frame, determining the other head object as a subject object corresponding to the current close-up frame and deleting the close-up frame constructed for the other head object; a first close-up frame updating unit for selecting another close-up frame from the close-up frame and updating it as the current close-up frame until all the current close-up frames have been traversed, and when it is determined that there is another head object that can be added to the current close-up frame, determining the other head object as a subject object corresponding to the current close-up frame and deleting the close-up frame constructed for the other head object; a close-up display unit for displaying a close-up screen surrounded by each of the close-up frames in the video stream data.
13. A close-up screen determination device, comprising: one or more processors; a display screen for displaying a close-up image; a memory for storing one or more programs; A close-up screen determination device, characterized in that when the one or more programs are executed by the one or more processors, the one or more processors realize the close-up screen determination method described in any one of claims 1 to 11.
14. 12. A computer-readable storage medium storing a computer program, the computer-readable storage medium being characterized in that, when the program is executed by a processor, the method for determining a close-up image according to any one of claims 1 to 11 is realized.
Citation Information
Patent Citations
Dynamic adjustment method and device of video conference viewing frame and computer equipment
CN116506720A
Video Stream Operations
JP2023544627A
Apparatus and method of detecting and displaying video conferencing groups
US11350029B1
System and method for not displaying duplicate images in a video conference
US20150138302A1
Automatic Video Framing of Conference Participants
US20180063482A1