Background position allocation in multi-attendee video communication
The proposed position allocating mechanism for video communication systems optimally positions attendee representations based on attributes like size, light, and orientation, enhancing realism and immersion in multi-person interactions.
Patent Information
- Application Number
- PCT/US2025/016181
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-25
- Filing Date
- 2025-02-16
- Publication Date
- 2025-10-30
Smart Images

Figure US2025016181_30102025_PF_FP_ABST
Abstract
Description
BACKGROUND POSITION ALLOCATION IN MULTI-ATTENDEE VIDEO COMMUNICATIONBACKGROUND
[0001] Video communication is a hotspot technology in the field of communications, which is for enabling remote communication services through transmitting voice and image data in realtime. Through providing real-time voice and image data transmission capabilities, the video communication technology enables attendees of video communication to have the same real experience as face-to-face communication in realistic scenes. At present, various video communication technologies have been widely concerned and applied in many fields such as online meeting, remote teaching, telemedicine and the like.SUMMARY
[0002] This Summary is provided to introduce a selection of concepts that are further described below in the Detailed Description. It is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0003] Embodiments of the present disclosure propose methods, apparatuses, and non- transitory computer readable media for video communication. One or more video streams for one or more video communication attendees may be received, wherein each video stream is associated with at least one video communication attendee. At least one visual representation of the at least one video communication attendee and an attribute of each visual representation of the at least one visual representation may be obtained for each video stream. A background comprising a plurality of candidate positions may be obtained. Position allocating information may be generated based on the background and attributes of one or more visual representations of the one or more video communication attendees, wherein the position allocating information indicates a target position corresponding to each visual representation of the one or more visual representations among the plurality of candidate positions. A target image may be generated through arranging each visual representation of the one or more visual representations at a corresponding target position according to the position allocating information.
[0004] It should be noted that the above one or more aspects comprise the features hereinafter fully described and particularly pointed out in the claims. The following description and the drawings set forth in detail certain illustrative features of the one or more aspects. These features are only indicative of the various ways in which the principles of various aspects may be implemented, and the present disclosure is intended to include all such aspects and theirequivalents.BRIEF DESCRIPTION OF THE DRAWINGS
[0005] The disclosed aspects will hereinafter be described in conjunction with the appended drawings that are provided to illustrate and not to limit the disclosed aspects.
[0006] FIG. 1 illustrates an exemplary application scenario of background position allocation in multi-attendee video communication according to an embodiment.
[0007] FIG. 2 illustrates an exemplary process for background position allocation in multiattendee video communications according to an embodiment.
[0008] FIG. 3 illustrates a schematic diagram of determining one or more target frames in a video stream according to an embodiment.
[0009] FIG. 4 illustrates a schematic diagram of generating a visual representation and an attribute of the visual representation at a server according to an embodiment.
[0010] FIG. 5 illustrates a schematic diagram of receiving a visual representation by a server and generating an attribute of the visual representation at the server according to an embodiment.
[0011] FIG. 6 illustrates a schematic diagram of receiving a visual representation and an attribute of the visual representation by a server according to an embodiment.
[0012] FIG. 7 illustrates a schematic diagram of an example background according to an embodiment.
[0013] FIG. 8 illustrates a schematic diagram of allocating a target position to a visual representation according to a rule that is set for a size of the visual representation according to an embodiment.
[0014] FIG. 9 illustrates a schematic diagram of allocating a target position to a visual representation according to a rule that is set for ambient light information of the visual representation according to an embodiment.
[0015] FIG. 10 illustrates an exemplary process of determining a light source position and a light intensity of a background virtual light source according to an attribute of a visual representation according to an embodiment.
[0016] FIG. 11 illustrates a schematic diagram of allocating a target position to a visual representation according to a rule that is set for a facial orientation of the visual representation according to an embodiment.
[0017] FIG. 12 illustrates a schematic diagram of allocating a target position to a visual representation according to a rule that is set for an occlusion state of the visual representation according to an embodiment.
[0018] FIG. 13 illustrates an exemplary process of performing position allocation according to priorities of attributes according to an embodiment.
[0019] FIG. 14 illustrates a schematic diagram of generating position allocating information with a position allocating model according to an embodiment.
[0020] FIG. 15 illustrates a schematic diagram of a visual representation filling region associated with a candidate position according to an embodiment.
[0021] FIG. 16 illustrates a schematic diagram of a target image according to an embodiment.
[0022] FIG. 17 illustrates a flowchart of an exemplary method for video communication according to an embodiment.
[0023] FIG. 18 illustrates an exemplary apparatus for video communication according to an embodiment.
[0024] FIG. 19 illustrates an exemplary apparatus for video communication according to an embodiment.DETAILED DESCRIPTION
[0025] The present disclosure will now be discussed with reference to several exemplar)’ implementations. It is to be understood that these implementations are discussed only for enabling those skilled in the art to better understand and thus implement the embodiments of the present disclosure, rather than suggesting any limitations on the scope of the present disclosure.
[0026] Some video communication services may provide a functionality to aggregate portraits of multiple video communication attendees into the same virtual space, which functionality is simply referred to in the present disclosure as an aggregation function. The application scenario of the aggregation function may be, for example, a multi-person video conference, and the like. When a conference attendee enables the aggregation function at his client, it can be seen that the portraits of the multiple conference attendees are aggregated into a background representing the virtual conference room. The background may include a plurality of seats, wherein a portrait of each conference attendee may be disposed at a seat. Thus, realism may be created during a video conference as if multiple conference attendees w ere interacting face-to-face in the realistic world. In addition to the multi-person video conference, an example of application scenarios of the aggregation function may further include, for example, online teaching, remote interview, and the like. In the present disclosure, for the sake of clarity, a portrait to be arranged into the background is referred to as a visual representation of a video communication attendee, and a position for arranging the portrait in the background, for example, the position of the seat in the virtual conference room, is referred to as a candidate position.
[0027] When using the aggregation function, a candidate position in the background needs to be allocated for the visual representation of the video communication attendee. Each visual representation may be arranged to a corresponding allocated position in the background. Some existing position allocating mechanisms for aggregation function allow for manually allocating ofcorresponding candidate positions for visual representations of video communication attendees. For example, in a multi-person video conference application scenario, a conference attendee may be allowed to manually select a desired candidate position after joining the video conference, or a moderator of the conference may be allowed to manually allocate a corresponding candidate position for the conference attendee, etc. Additionally, some existing position allocating mechanisms for aggregation function also allow for automatically allocating corresponding candidate positions for visual representations of video communication attendees. When automatically allocating candidate positions, these existing mechanisms typically allocate corresponding candidate positions to the visual representations of the video communication attendees in a random selection manner, or according to an order in which the video communication attendees join the video communication or an order in which the video communication attendees turn on the cameras.
[0028] The present disclosure proposes an improved position allocating mechanism for the aggregation function of video communications. The mechanism proposed by the present disclosure performs candidate position allocation in the case of taking attributes of visual representations of video communication attendees into account. Thus, the position allocated according to the mechanism of the present disclosure to each visual representation and the attribute of the visual representation are matched with each other. The technical effect of this matching is that the video communication attendee enabling the aggregation function can see, at his client, that each visual representation is adaptively arranged to a corresponding position in the background. For video communication attendees, images generated through such an adaptative arrangement appear to be harmonious, natural, real, and reasonable in entirety. Thus, stronger immersion and realism can be provided for video communication attendees during video communication.
[0029] In one aspect, embodiments of the present disclosure may obtain a visual representation of each video communication attendee and an attribute of the visual representation. The visual representation and the attribute of the visual representation may be obtained with a mask segmenting model and an attribute detecting model, respectively. The attribute may be used to characterize the characteristics, natures, or properties possessed by the visual representation. For example, the attribute may be used to characterize a size, ambient light information, a facial orientation, an occlusion state, etc., of the visual representation. Embodiments of the present disclosure may also obtain a background including a plurality of candidate positions. The candidate position is a position in the background for arranging the visual representation. The background may be, for example, selected from a background library based on attributes of one or more visual representations. After obtaining the background and the attributes of the visualrepresentations, position allocating information may be generated based on the background and the attributes. The generated position allocating information may indicate each visual representation is adapted to which target position in the background. In one aspect, position allocating information may be generated according to the background and the attributes based on predetermined rules. In another aspect, the position allocating information may be generate according to the background and the attributes with a position allocating model. Then, a target image may be generated according to the position allocating information. In the generated target image, each visual representation is arranged at an adapted target position that is allocated to the visual representation. The target image may be displayed at the client of the video communication attendee with the aggregation function enabled, to provide the video communication attendee with the same immersion and realism as multi-person interaction in the realistic world.
[0030] Exemplary embodiments of the present disclosure will be described below with reference to the accompanying drawings.
[0031] FIG. 1 illustrates an exemplary application scenario 100 of background position allocation in multi-attendee video communication according to an embodiment. FIG. 1 illustrates a server 110. The server 110 may be a server for providing aggregation function in multi -attendee video communication, and may be configured to implement the position allocating mechanism for the aggregation function proposed by the present disclosure. The server 110 may be any type or configuration of remote server, cloud server, or the like. The server 110 may represent a server that works independently, a server cluster composed of a plurality of servers that cooperate with each other, and the like. When the server 110 includes multiple servers, the servers may be centrally arranged in the same geographic position, may be dispersedly arranged at multiple different geographic positions, and the like.
[0032] The server 110 may be in remote communication with a plurality of client devices 120- 1 to 120-N, wherein each client device is used by at least one of a plurality7of video communication attendees 130-1 to 130-M for participating in video communication. In one example, one client device may be used by one video communication attendee, as shown in FIG. 1. client device 120- 1 may be used by video communication attendee 130-1. In another example, one client device may be used by two or more video communication attendees (not shown for clarity in FIG. 1), such an example scenario may occur for example when two or more video communication attendees are located in the same realistic space and using one client device for participating in video communication. Each of the plurality of client devices 120-1 to 120-N may be any type or configuration of client devices, such as a desktop computer, a notebook, a tablet, a smart phone, or the like. The client device may be configured with an image capture module. The image capture module may be built into or attached to the client device for capturing images of at least one videocommunication atendee that is participating in video communication using the client device and providing the captured images to the client device.
[0033] In a video communication process, the server 110 may receive a plurality' of video streams 140-1 to 140-N from the plurality’ of client devices 120-1 to 120-N. Each video stream may include image data for at least one video communication atendee participating in video communication with a respective client device. As shown in FIG. 1, for example, the server 110 may receive a video stream 140-1 from the client device 120-1, wherein the video stream 140-1 includes image data for the video communication atendee 130-1. A video stream, also referred to as a bitstream, refers to image data that is streamed between each client and the server 110. The video stream may be generated by the client device through using any video encoding techniques to perform video encoding on one or more captured images of the video communication atendee. The image data generated through performing video encoding for each image may be referred to as encoded video frame data. In the present disclosure, for simplicity, the encoded video frame data is simply referred to as a frame. Thus, each of the plurality of video streams 140-1 to 140-N may include one or more frames. After receiving the video stream, the server 110 may decode one or more frames in the video stream into one or more corresponding images respectively through using any video decoding technique.
[0034] Through using the received plurality of video streams 140-1 to 140-N, the server 110 may aggregate visual representations of the plurality of video communication atendees 130-1 to 130-M into the same virtual space. During performing the aggregation operation, the server 110 may implement the position allocating mechanism discussed in the present disclosure, such that visual representations of the plurality of video communication atendees 130-1 to 130-M may be respectively arranged at adapted target positions in the background. In the present disclosure, for simplicity, an image that is generated by the server 110 in which the visual representations are respectively arranged at the adapted target positions in the background is referred to as a target image.
[0035] Among the plurality of client devices 120-1 to 120-N as shown in FIG. 1, there may be one or more requesting clients. In the present disclosure, a requesting client refers to a client device that enables the aggregation function. Through enabling the aggregation function, the requesting client may request, to the server 110, a target image generated with the position allocating mechanism discussed in the present disclosure for presenting to the video communication atendee. As shown in FIG. 1, for example, the client device 120-1 may be a requesting client. In response to the request of the client device 120-1, the server 110 may transmit a target image 150 to the client device 120-1. The client device 120-1 may present the target image 150 to the video communication atendee 130-1 to cause the video communication atendee 130-1 to obtain, during the video communication, the same immersion and realism as multi-person interaction in the realistic world.
[0036] Although a plurality of video streams 140-1 to 140-N are shown in the exemplary application scenario 100 of FIG. 1, in one example, a situation in which only one of the plurality of video streams 140-1 to 140-N is received by the server 110 may occur. Such a situation may occur, for example, when only one video communication attendee using one client device successfully joins the video communication and is waiting for other video communication attendees to join.
[0037] FIG. 2 illustrates an exemplary process 200 for background position allocation in multi-attendee video communications according to an embodiment. Process 200 may be performed by a server that provides aggregation function in multi-attendee video communication, such as the server 110 discussed above in connection with FIG. 1.
[0038] Process 200 may begin with receiving one or more video streams 210. Each of the video streams 210 may correspond to one of the video streams 140-1 to 140-N as discussed above in connection with FIG. 1. Each video stream may be associated with at least one video communication attendee and include one or more frames.
[0039] A visual representation 220 of the video communication attendee and an attribute 230 of the visual representation may be obtained. Each visual representation and the attribute of the visual representation may be obtained for one particular frame in each video stream. In the present disclosure, a particular frame in a video stream for obtaining a visual representation and its attribute is referred to as a target frame. The content about determining the target frame in the video stream will be discussed in further detail below in conjunction with FIG. 3. In an example where a video stream is associated with one video communication attendee, a visual representation of the video communication attendee and an attribute of the visual representation may be obtained for the target frame of the video stream. In an example where a video stream is associated with two or more video communication attendees, a visual representation of each of the two or more video communication attendees and an attribute of each visual representation may be obtained for the target frame of the video stream.
[0040] The visual representation 220 of the video communication attendee refers to a digital representation of a portrait of the video communication attendee. A mask segmentation, or referred to as an instance segmentation, may be performed for the target frame, to detect a portrait in the target frame and to segment the portrait from the target frame to form a visual representation. In one example, the mask segmentation may be performed through using a mask segmenting model.
[0041] The attribute 230 of the visual representation may be used to characterize thecharacteristics, natures, or properties possessed by the visual representation 220. The attribute of the visual representation may be generated through performing an attribute detection for the target frame from which the visual representation is segmented. In one example, the attribute detection may be performed through using an attribute detecting model. The target frame or at least one visual representation generated for the target frame may be provided to the attribute detecting model to generate the attribute for each visual representation. The attribute 230 may include a variety of attributes, such as size, ambient light information, facial orientation, occlusion state, and the like.
[0042] The size is used to characterize the size of an area occupied by the visual representation. In one example where a predetermined resolution is set for the plurality of visual representations, the size of each visual representation may be represented as the number of pixels included in the visual representation under the predetermined resolution, etc. The ambient light information is used to characterize the condition of a virtual light source determined for the visual representation. The virtual light source is a simulation of a real light source in a realistic scene where the video communication attendee is located. In one example, the virtual light source may emit scattered virtual light to produce a change in brightness at different regions of the visual representation. In one example, the ambient light information of the visual representation may be represented, for example, as a light intensity of the virtual light source, relative position information between the visual representation and the virtual light source, etc. In an example in which ambient light information is considered for allocating a target position in the background to the visual representation, a background virtual light source needs to be set for the background. The background virtual light source is a simulation of a real light source in the virtual space represented by the background and is used for generating virtual light in the background. In one example, the background virtual light source may emit scattered virtual light to produce a change in brightness at different regions of the background. The process of setting the background virtual light source may include an operation of determining a light source position and a light intensity of the background virtual light source. Various example manners may be adopted to determine the light source position and light intensity of the background virtual light source. An exemplary manner may include determining the light source position and light intensity of the background virtual light source with a predetermined light source configuration list. Another exemplary' manner may include determining the light source position and light intensity of the background virtual light source according to attributes of one or more visual representations to be arranged into the background. These two exemplary manners will be discussed in detail below. The facial orientation is used to characterize a direction in which a face of the visual representation is oriented. In one example, the facial orientation of the visual representation may be represented asa direction category (e.g., top left, top right, etc.) of the facial orientation of the visual representation. In another example, the facial orientation of the visual representation may be represented as an angle of the facial orientation of the visual representation. The occlusion state is used to characterize a conditon of the visual representation being occluded. In one example, the occlusion state may be represented as whether the visual representation is occluded, which portion of the visual representation is occluded, or the like.
[0043] The visual representation 220 and / or the attribute 230 of the visual representation may be generated at the server performing the process 200, or received by the server, as discussed in further detail below in connection with FIGS. 4-6.
[0044] Process 200 may include performing, at 250, a background obtaining operation based on the attribute of the visual representation. Through performing the background obtaining operation, a background 260 may be obtained. The background 260 may represent the virtual space into which the visual representations of the plurality’ of video communication attendees are to be aggregated. In one example, the background 260 may be embodied as a data format of a two-dimensional background image. In another example, the background 260 may be embodied as a data format of a three-dimensional background model. A plurality’ of candidate positions may be included in the background 260. The candidate position is a position in the background 260 for arranging a visual representation.
[0045] In one example, the background obtaining operation at 250 may be performed through background selection. In this example, a background may be selected, according to the attribute 230 of the visual representation, from a background library' 240 storing a plurality of backgrounds, as the background 260. In an aspect, the background 260 may be manually selected from the background library' 240. For example, a requesting client may provide a background selection function, such that the video communication attendee may utilize the function to manually select a desired background from the background library' 240 as the background 260. In another aspect, the background 260 may be automatically selected from the background library 240. In one example, the background 260 in which the number of candidate positions matches the number of video communication attendees may be automatically selected based on the number of video communication attendees. In another example, the background 260 may be automatically selected according to the attribute 230 of the visual representation.
[0046] In another example, the background obtaining operation at 250 may be performed through background modification. In this example, an initial background may be selected from the background library' 240 storing a plurality of backgrounds. The initial background may be selected manually or automatically as discussed above. The initial background may then be modified according to the attribute 230 of the visual representation to generate the background260. In one example, the background modification may be performed in cases where it is not possible to select a suitable background from the background 11 bran 240 according to the attribute 230 of the visual representation.
[0047] In yet another example, the background obtaining operation at 250 may be performed through background generation. In this example, the background 260 may be automatically generated with a background generating model according to the attribute 230 of the visual representation. The content relating to obtaining the background 260 will be discussed in further detail below in conjunction with FIG. 7.
[0048] Based on the background 260 and the attribute 230 of the visual representation, a position allocation may be performed at 270. The position allocation refers to a process to allocate, to each visual representation, a position in the background 260 that matches the attribute of the visual representation. In the present disclosure, for the sake of clarity, a matched position allocated to a visual representation is referred to as a target position corresponding to the visual representation. By performing the position allocation, position allocating information 280 may be generated. The position allocating information 280 may indicate, for each visual representation, which candidate position in the background 260 is the corresponding target position of the visual representation. For example, as shown in FIG. 2, the position allocating information 280 may include a plurality of data entries, where a first data entry may indicate that the corresponding target position of the visual representation 1 is the candidate position 2, the second data entry may indicate that the corresponding target position of the visual representation 2 is the candidate position 3, the third data entry may indicate that the corresponding target position of the visual representation 3 is the candidate position 1, and the like.
[0049] In an aspect, a rule-based position allocating mechanism 271 may be employed to generate the position allocating information 280. The rule-based position allocating mechanism 271 may allocate a corresponding target position to each visual representation according to a predetermined rule. The predetermined rule may specify a visual representation with what attribute is matched with which candidate position. In one example, the position allocation may be performed for each attribute, respectively, through employing the rule-based position allocating mechanism 271, as will be discussed in further detail below in connection with FIGS. 8-12. In one example, the position allocation may be performed in the case of jointly considering multiple attributes through employing the rule-based position allocating mechanism 271. as will be discussed in further detail below in connection with FIG. 13.
[0050] In another aspect, the position allocating information 280 may be generated through employing a position allocating mechanism 272 that utilizes a model. The model may be an Al model for performing a background position allocating task, as will be discussed in further detailbelow in connection with FIG. 14.
[0051] The target image may then be generated at 290 based on the visual representation 220, the background 260, and the position allocating information 280. In the generated target image, each visual representation is respectively arranged at a matched target position in the background 260 indicated by the position allocating information 280 for the visual representation. In one example, an image enhancing operation may be performed on the generated target image, for example, by an image enhancing Al model, to further improve the realism of the target image. Contents about generating the target image will be discussed in further detail below in conjunction with FIGS. 15-16.
[0052] FIG. 3 illustrates a schematic diagram 300 of determining one or more target frames in a video stream according to an embodiment. FIG. 3 illustrates a video stream 310, which may be a video stream associated with at least one video communication attendee, e.g., one of the video streams 140-1 to 140-N as discussed above in connection with FIG. 1. The video stream 310 may include a plurality of frames arranged in a transmission time order. One frame or some frames in the video stream 310 may be selected to be used as the target frame, wherein the selected target frame may be used to obtain a visual representation of each video communication attendee and the attribute of each visual representation as discussed above in connection with FIG. 2.
[0053] In one example, one target frame may be selected from frames of the video stream 310. The target frame may be selected in response to a first trigger condition being met. For example, the first trigger condition may be the detection of a first frame of the video stream 310 containing a portrait of a video communication attendee. The first trigger condition may also be the detection of an instruction to execute position allocation from the video communication attendee, or other trigger conditions associated with the selection of the target frame. FIG. 3 illustrates a first target frame 320-1 selected in response to the first trigger condition being met. For the first target frame 320-1, a visual representation 330-1 and an attribute 340-1 may be obtained. The position allocation may then be performed once according to the operations in process 200 discussed above in connection with FIG. 2, in order to allocate a target position in a background to the visual representation 330-1 based on the attribute 340-1. A server may perform similar operations for other video streams not shown to allocate target positions to visual representations for other video streams, respectively . Thus, in the position allocation, the visual representation for each video stream is allocated a matched target position in the background.
[0054] The above example corresponds to a case in which the position allocation is performed only once in the video communication process and no subsequent position allocation is reperformed. For example, the position allocation may be performed once at the beginning of the video communication. Through performing the position allocation only once, technical effects ofavoiding frequently performing the position allocation causing occupying computing resources and data transmission resources can be achieved. In this example, although the position allocation is no longer re-performed in the subsequent process, the server may still update the respective visual representations in the target image with subsequent frames after the first target frame in each video stream. Through continuously updating the respective visual representations in the target image, it may facilitate presenting a video in which the visual representation is updated in real-time to a video communication attendee at a requesting client.
[0055] In another example, a plurality of target frames may be selected from frames of the video stream 310. The target frames may be respectively selected in response to sequentially satisfying each of a set of second trigger conditions. For example, the second trigger condition may be reaching a preset time interval for performing the position allocation, for example, 30 minutes. The second trigger condition may also be the detection of an instruction to execute position allocation from the video communication attendee, or other trigger conditions associated with the selection of the target frame, etc. FIG. 3 illustrates a first target frame 320-1, a second target frame 320-2,- • •, and an Nth target frame 320-N selected respectively in response to sequentially satisy ing the set of second trigger conditions. For the first target frame 320-1, a visual representation 330-1 and an attribute 340-1 may be obtained. A first position allocation may then be performed according to the operations in process 200 discussed above in connection with FIG. 2, in order to allocate a target position in the background to the visual representation 330-1 based on the attribute 340-1. The server may perform similar operations for other video streams not shown to allocate target positions to visual representations for other video streams, respectively. Thus, in the first position allocation, the visual representation for each video stream is allocated a matched target position in the background. A second position allocation may then be performed for the second target frame 320-2 in a similar manner until a Nth position allocation is performed for the Nth target frame 320-N.
[0056] The above example corresponds to a case in which the position allocation is performed for multiple times during the video communication. For example, the position allocation may be performed once each time when a preset time interval is reached. Through performing the position allocation for multiple times, the allocated target position of the visual representation can be adjusted in time according to the change in the attribute of the visual representation, thereby producing a technical effect of dynamically adapting the allocated position to the change of the attribute of the corresponding visual representation. Similar to the example discussed above in which the position allocation is performed only once, in this example, between two position allocations, the server may utilize frames between two target frames in each video stream to update respective visual representations in the target image, respectively. Through continuously updatingrespective visual representations in the target image between two position allocations, it may facilitate presenting a video in which the visual representation is updated in real-time to the video communication attendee at the requesting client.
[0057] Although only the obtained visual representation of one video communication attendee and the attribute of the visual representation are shown and discussed for each target frame in FIG. 3 and above description with respect to FIG. 3, such description is merely served as an example. In another example, the video stream 310 may be associated with two or more video communication attendees. In this example, for each target frame, a visual representation of each of the two or more video communication attendees may be obtained, and the attribute of each visual representation may be obtained accordingly.
[0058] FIG. 4 illustrates a schematic diagram 400 of generating a visual representation and an attribute of the visual representation at a server according to an embodiment. FIG. 4 illustrates a requesting client 410, one or more other clients 420, and a server 430. Each of the one or more other clients 420 may have a similar structure as the requesting client 410 and may perform similar operations as the requesting client 410. A difference between the other clients and the requesting client 410 may be that the requesting client 410 may issue a request for the target image, and the serv er 430 may provide a target image 440 to the requesting client 410 in response to the request.
[0059] The server 430 may receive video streams from the requesting client 410 and one or more other clients 420, respectively. Each video stream may include one or more target frames. A mask segmenting model 431 may be deployed at the server 430 for performing mask segmentation for target frames in each video stream, respectively, to generate the visual representations 432 of all video communication attendees. An attribute detecting model 433 may also be deployed at the server 430 for performing attribute detection for target frames in each video stream, respectively, to generate the attributes 434 for the visual representations of all video communication attendees. The server 430 may perform a position allocation in accordance with the operations in the process 200 discussed above in connection with FIG. 2 based on the visual representations 432 and the attributes 434 of the visual representations to generate the target image 440. The target image 440 may be communicated to the requesting client 410 for being presented to the corresponding video communication attendee.
[0060] Through deploying both the mask segmenting model 431 and the attribute detecting model 433 at the server, the application scenario shown in FIG. 4 avoids the requirement for the ability of performing mask segmentation and attribute detection by the client devices, so that a technical effect of making the mechanism presented in the disclosure flexibly applicable to the client devices with various configurations can be achieved.
[0061] FIG. 5 illustrates a schematic diagram 500 of receiving a visual representation by aserver and generating an attribute of the visual representation at the server according to an embodiment. FIG. 5 illustrates a requesting client 510, one or more other clients 520, and a server 530. Each of the one or more other clients 520 may have a similar structure as the requesting client 510 and may perform similar operations as the requesting client 510. A difference between the other clients and the requesting client 510 may be that the requesting client 510 may issue a request for the target image, and the server 530 may provide a target image 540 to the requesting client 510 in response to the request.
[0062] The server 530 may receive video streams from the requesting client 510 and one or more other clients 520, respectively. A mask segmenting model 511 may be deployed at each of the requesting client 510 and the one or more other clients 520 for performing a mask segmentation for the target frame to generate a visual representation 512 of at least one video communication attendee associated with the client. In one example, the visual representation 512 and the target frame may be transmitted in the video stream. In another example, only the visual representation 512 may be transmitted in the video stream without transmitting the target frame. The server 530 may receive each video stream from the requesting client 510 and one or more other clients 520, respectively, to obtain the visual representations 512 of all video communication attendees. An attribute detecting model 533 may be deployed at the server 530 for performing attribute detection for the target frame in each video stream, respectively, to generate the attributes 534 for the visual representations of all video communication attendees. In an example where the visual representation 512 and the target frame are transmitted together in the video stream, the attribute detecting model 533 may use the target frame in the video stream as an input to generate the attributes 534. In an example where only the visual representation 512 is transmitted in the video stream without transmitting the target frame, the attribute detecting model 533 may use the visual representation 512 as an input to generate the attributes 534. In the example where the target frame is not transmitted, since the amount of data transmitted in the video stream is reduced, the achieved technical effect may be saving data transmission resources, for example, communication bandwidth, etc. The server 530 may perform the position allocation in accordance with the operations in the process 200 discussed above in connection with FIG. 2 based on the obtained visual representations 512 and the generated attributes 534 of the visual representations to generate the target image 540. The target image 540 may be communicated to the requesting client 510 for being presented to the corresponding video communication attendee.
[0063] Through deploying the mask segmenting model 51 1 at the client, the application scenario shown in FIG. 5 can fully utilize the computing resources of the client, thereby bringing a technical effect of reducing the computational burden at the server 530.
[0064] FIG. 6 illustrates a schematic diagram 600 of receiving a visual representation and anatribute of the visual representation by a server according to an embodiment. FIG. 6 illustrates a requesting client 610, one or more other clients 620, and a server 630. Each of the one or more other clients 620 may have a similar structure as the requesting client 610 and may perform similar operations as the requesting client 610. A difference between the other clients and the requesting client 610 may be that the requesting client 610 may issue a request for the target image, and the server 630 may provide a target image 640 to the requesting client 610 in response to the request.
[0065] The server 630 may receive video streams from the requesting client 610 and one or more other clients 620, respectively. A mask segmenting model 611 may be deployed at each of the requesting client 610 and the one or more other clients 620 for performing a mask segmentation for the target frame to generate a visual representation 612 of at least one video communication atendee associated with the client. An atribute detecting model 613 may also be deployed at each client for performing an atribute detection for the target frame to generate an attribute 614 of the visual representation of at least one video communication atendee associated with the client. The visual representation 612 and the atribute 614 of the visual representation may be communicated in each video stream. In one example, the target frame may not be transmited in the video stream because the server 630 may directly receive the visual representation 612 and the atribute 614 of the visual representation in each video stream. In the example in which the target frame is not transmited, since the amount of data transmited in the video stream is reduced, the achieved technical effect may be saving data transmission resources, for example, communication bandwidth. The server 630 may obtain the visual representations 612 of all video communication atendees and the attributes 614 of the visual representations through receiving the plurality of video streams. The server 630 may perform the position allocation in accordance with the operations in the process 200 discussed above in connection with FIG. 2 based on the obtained visual representations 612 and attributes 614 of the visual representations to generate the target image 640. The target image 640 may be communicated to the requesting client 610 for being presented to the corresponding video communication atendee.
[0066] Thorough deploying both the mask segmenting model 611 and the atribute detecting model 613 at the client, the application scenario shown in FIG. 6 can further fully utilize the computing resources of the client, thereby bringing a technical effect of reducing the computational burden at the server 630.
[0067] In an aspect, the mask segmenting model, e.g., 431, 511, 611 discussed above in connection with FIGS. 4-6, is an Al model for performing a mask segmenting task to segment a visual representation from a target frame. The mask segmenting model may adopt any known neural network architecture for performing a mask segmenting task, for example, a U Net architecture, a Mask-RCNN architecture, or the like.
[0068] In another aspect the attribute detecting model, e.g., 433, 533, 613 discussed above in connection with FIGS. 4-6, may be an Al model for performing an attribute detecting task to detect one or more attributes possessed by the visual representation.
[0069] In one example, the attribute detecting model may be implemented as a size detecting model, a facial orientation detecting model, an occlusion state detecting model, or the like. The size detecting model may detect the size of the visual representation. The detected size of the visual representation may be represented as the number of pixels included in the visual representation, etc. The facial orientation detecting model may be used to detect a facial orientation of the visual representation. The facial orientation detecting model may be a multiclassification model to classify the facial orientation of the visual representation as one of a plurality of direction categories. For example, the plurality of direction categories may be “top”, “top left”, “top right”, “left”, “right”, “bottom left”, “bottom”, “bottom right”, and the like. The facial orientation detecting model may also be a normalization model for outputting an angle of the facial orientation of the visual representation, for example, a specific angle in a range of [0 °, 360 °], The occlusion state detecting model may be used for detecting an occlusion state of the visual representation, for example, whether the visual representation is occluded, which portion of the visual representation is occluded, or the like. In this example, the size detecting model, the facial orientation detecting model, and the occlusion state detecting model may employ any known neural network architecture for performing an object segmenting task, a mask or instance segmenting task, for example, a Resnet architecture, a MobileNet architecture, a RetinaNet architecture, a Faster RCNN architecture, a Mask-RCNN architecture, or the like.
[0070] In one example, one or more of the size detecting model, the facial orientation detecting model, and the occlusion state detecting model may share the same core network architecture as the mask segmenting model. For example, one or more of the size detecting model, the facial orientation detecting model, and the occlusion state detecting model may be implemented through adding at least one additional layer on the basis of the network structure of the mask segmenting model. At least one additional layer may be added after one intermediate layer of network or last layer of network of the mask segmenting model. In one example, the added additional layer may be an output layer, which may be, for example, a fully connected layer, an activation layer implemented with a softmax function, or the like. In one example, the added additional layer may be one or more additional feature extraction layers and auxiliary layers between the mask segmenting model and the aforementioned output layer, e.g., alternately arranged convolutional layers, pooling layers, etc., for performing further feature extraction for the output of the mask segmenting model.
[0071] In one example, the attribute detecting model may be implemented as an ambient lightdetecting model to detect ambient light information of the visual representation. As discussed above in connection with FIG. 2, the ambient light information is used to characterize the condition of a virtual light source that simulates a real light source in the realistic scene where the video communication attendee is located. The ambient light detecting model may determine some basic information of the virtual light source according to the light distribution at individual pixel points of the visual representation, for example, an angle offset between the virtual light source and the visual representation, a distance between the virtual light source and the visual representation, a light intensity of the virtual light source, and the like. In this example, the ambient light detecting model may employ any known method for performing a virtual light source detecting task, such as an image content-based analysis method; or employ any known neural network model for performing a virtual light source detecting task, such as a model with CNN architecture, or the like.
[0072] In one example, the attribute detecting model may include a plurality of sub-modules, for example, a size detecting sub-module, a facial orientation detecting sub-module, an occlusion state detecting sub-module, an ambient light detecting sub-module, and the like. The structures and functions of these sub-modules may be similar to the size detecting model, facial orientation detecting model, occlusion state detecting model, ambient light detecting model discussed above, respectively.
[0073] FIG. 7 illustrates a schematic diagram 700 of an example background according to an embodiment. A background 710 is schematically illustrated in FIG. 7. The background 710 is a two-dimensional background image and may represent a virtual space such as a virtual conference room in which multiple rows of seats are disposed.
[0074] The background 710 includes a plurality of candidate positions for arranging visual representations of video communication attendees. These candidate positions correspond to the multiple rows of seats shown in FIG. 7, respectively. Three exemplary7candidate positions of the plurality of candidate positions are shown in FIG. 7 with solid dots: a candidate position 720 at the third seat from the left of the first row, a candidate position 730 at the second seat from the left of the second row, and a candidate position 740 at the first seat from the left of the third row. It should be understood that the background 710 shown in FIG. 7 that includes a plurality of candidate positions merely serves as an exemplary7representation of the background. In other examples, the background 710 may represent other virtual spaces other than the virtual conference room, and the candidate positions in the background 710 may be arranged differently, e.g., arranged in a circle, in a horizontal row, or the like.
[0075] In addition, although FIG. 7 shows the background 710 in the form of a two- dimensional background image, in one example, the background 710 may also take a data formof a three-dimensional background model. Compared with a two-dimensional background image, the three-dimensional background model has the advantages of strong controllability, convenience for modification, and the like.
[0076] In an aspect, as discussed above in connection with FIG. 2. the background 710 may be automatically selected from a background library storing a plurality of backgrounds according to attributes of one or more visual representations to be arranged into the background 710. In one example, when the background is selected according to the size, a background providing large seats may be selected from the background library in response to determining that most of the visual representations have a large size. In one example, when the background is selected according to the ambient light information, a background providing a background virtual light source located on the left may be selected from the background library in response to determining that most of the visual representations have virtual light sources located on the left. In one example, when the background is selected according to the facial orientation, a background providing seats deviated to a sideway direction may be selected from the background library in response to determining that the faces of most of the visual representations are oriented to a sideway direction rather than a straight-ahead direction. In one example, when the background is selected according to the occlusion state, a background providing a plurality of occluded seats may be selected from the background library in response to determining that some of the visual representations are occluded. Through selecting the background 710 according to the attribute of the visual representation, such as size, ambient light information, facial orientation, occlusion status, etc., it can further facilitate to adaptively fill the visual representations into the background 710.
[0077] In another aspect, as discussed above in connection with FIG. 2, the background 710 may be obtained through modifying an initial background according to the attributes of one or more visual representations. The initial background may be selected from a background library including a plurality of backgrounds. In one example, when the initial background is modified according to the size, the initial background may be modified to provide large seats in response to determining that most of the visual representation has a large size. In one example, when the background is modified according to the ambient light information, the initial background may be modified to provide a background virtual light source located on the left in response to determining that most of the visual representations have virtual light sources located on the left. In one example, when the background is modified according to the facial orientation, the initial background may be modified to provide seats deviated to a sideway direction in response to determining that the faces of most of the visual representations are oriented to a sideway direction rather than a straight-ahead direction. In one example, when the background is modified accordingto the occlusion state, the initial background may be modified to provide a plurality of occluded seats in response to determining that some of the visual representations are occluded. Through modifying the initial background according to the attribute of the visual representation, e.g., size, ambient light information, facial orientation, occlusion status, etc., to obtain the background 710, it can further facilitate to adaptively fill the visual representations into the background 710.
[0078] In one example, when the initial background is a two-dimensional background image, the background modification operations discussed above may be performed through utilizing any known image modification method, e.g., utilizing an image modifying model. The image modifying model may be a generative model for generating a modified image with an image as input, e.g., Midjoumey, etc. In one example, when the initial image is a three-dimensional background model or a two-dimensional rendered picture of a three-dimensional background model, the background modification operations discussed above may be performed through utilizing any known three-dimensional model modification method.
[0079] In yet another aspect, as discussed above in connection with FIG. 2, the background 710 may be automatically generated utilizing a background generating model according to attributes of the one or more visual representations. In this example, the background generating model may be a generative model for generating a background image with the attributes of the one or more visual representations as input, e.g., Midjoumey, DALL-E, etc. In one example, when the background is generated according to the size, a background providing large seats may be generated in response to determining that most of the visual representations have a large size. In one example, when the background is generated according to the ambient light information, a background providing a background virtual light source located on the left may be generated in response to determining that most of the visual representations have virtual light sources located on the left. In one example, when the background is generated according to the facial orientation, a background providing seats deviated to a sideway direction may be generated in response to determining that the faces of most of the visual representations are oriented to a sideway direction rather than a straight-ahead direction. In one example, when the background is generated according to the occlusion state, a background providing a plurality of occluded seats may be generated in response to determining that some of the visual representations are occluded. Through generating the background 710 according to the attribute of the visual representation, such as size, ambient light information, facial orientation, occlusion status, etc., it can further facilitate to adaptively fill the visual representations into the background 710.
[0080] FIG. 8 illustrates a schematic diagram 800 of allocating a target position to a visual representation according to a rule that is set for a size of the visual representation according to an embodiment. FIG. 8 illustrates a background 810. FIG. 8 also illustrates a first visualrepresentation 820 and a second visual representation 840 for being adaptively arranged into the background 810. The first visual representation 820 has an attribute 830 that indicates that the size of the first visual representation 820 is, for example, 300 pixels. The second visual representation 840 has an attribute 850 that indicates that the size of the second visual representation is, for example. 400 pixels.
[0081] The rule set for the size of the visual representation may specify the visual representation with what size is matched with which candidate position in the background 810. In one example, the rule that is set for the size of the visual representation may specify that a large visual representation matches a near candidate position, and a small visual representation matches a far candidate position. This rule makes the arrangement of the visual representation in the background 810 conform to the “near large, far small” principle.
[0082] In one example, according to the exemplary7rule described above, a candidate position at the front rows in the background 810, e.g., a candidate position 860 at the first row. may be allocated to the second visual representation 840 which is large as the target position of the second visual representation 840. Meanwhile, according to the rule described above, a candidate position at the back rows in the background 810, for example, a candidate position 870 at the last row, may be allocated to the first visual representation 820 which is small as the target position of the first visual representation 820.
[0083] FIG. 9 illustrates a schematic diagram 900 of allocating a target position to a visual representation according to a rule that is set for ambient light information of the visual representation according to an embodiment. FIG. 9 illustrates a background 910. FIG. 9 also illustrates a visual representation 920 for being adaptively arranged into the background 910. The visual representation 920 has an attribute 930 that indicates the ambient light information of the visual representation 920. The ambient light information of the visual representation 920 may be represented, for example, as a light intensity of a virtual light source determined for the visual representation 920 and relative position information between the visual representation 920 and the virtual light source. The ambient light information may be generated through performing a light analysis on the illuminated region of the visual representation 920 utilizing the ambient light detecting model discussed above or any other ambient light detection algorithm.
[0084] In one example, the relative position information between the visual representation 920 and the virtual light source may be represented as an angular offset between the visual representation 920 and the virtual light source and a distance between the visual representation 920 and the virtual light source. When the angular offset between the visual representation 920 and the virtual light source is fixed, the distance between the visual representation 920 and the virtual light source is proportional to the light intensity of the virtual light source. For example,when the distance between the visual representation 920 and the virtual light source is larger, the light intensity of the virtual light source needed to illuminate a particular region of the visual representation 920 is greater. When the distance between the visual representation 920 and the virtual light source is smaller, the light intensity of the virtual light source needed to illuminate a particular region of the visual representation 920 is lesser. Therefore, when both the angle offset between the visual representation 920 and the virtual light source and the distance between the visual representation 920 and the virtual light source are determined, the light intensity of the virtual light source required by the visual representation 920 may be calculated based on the foregoing proportional relationship. In one example, other manners may be used to represent relative position information between the visual representation 920 and the virtual light source. For example, a coordinate system may be established with a virtual camera capturing the target frame containing the visual representation 920 as the origin, and the relative position information may be determined from the coordinates of the visual representation 920 and the virtual light source in the coordinate system.
[0085] A background virtual light source 940 may be set for the background 910, the background virtual light source 940 is for generating virtual light in the virtual space represented by the background 910. A light source position and light intensity of the background virtual light source 940 may affect the condition of the generated virtual light. Thus, the process of setting the background virtual light source 940 includes the operation of determining the light source position and light intensity of the background virtual light source 940, as discussed in detail below.
[0086] After setting the background virtual light source 940, a target position in the background may be allocated to the visual representation according to a rule that is set for the ambient light information of the visual representation. The rule that is set for the ambient light information of the visual representation may specify the ambient light information of the visual representation is matched to the light generated at which candidate position in the background 910 by the background virtual light source 940. First, coordinates of the visual representation 920 in the background 910 may be determined according to the light intensity and the light source position of the background virtual light source 940 and the ambient light information of the visual representation 920. For example, the direction in which the visual representation 920 should be located relative to the background virtual light source 940 may be calculated according to the coordinates of the background virtual light source 940 in the background 910 and the angular offset between the visual representation 920 and the virtual light source. The point at which the visual representation 920 should be located in the determined direction may be calculated according to the light intensity' of the background virtual light source 940 and the proportional relationship between the distance from the visual representation 920 to the virtual light source andthe light intensity of the virtual light source as discussed above. FIG. 9 illustrates a point 950 at which the visual representation 920 should be located that is determined in the manner described above. The light generated by the background virtual light source 940 at point 950 is close to the light generated by the virtual light source of the visual representation 920 at the visual representation. The target position of the visual representation 920 may then be determined from the coordinates of the point 950 in the background 910. For example, a candidate position 960 in the background 910 closest to the coordinate of the point 950 may be determined as the target position of the visual representation 920.
[0087] It should be understood that although the background virtual light source 940 is explicitly shown in FIG. 9, in one example, the background virtual light source 940 may not be presented in the background 910. In this example, the presence of a background virtual light source may be implicitly represented by changes in light at different candidate positions.
[0088] In one example, one or more optional auxiliary’ light sources may be set in the background 940 in addition to the background virtual light source 940. This example applies to a scenario in which a particular visual representation includes multiple illuminated regions. In this example, the ambient light information discussed above may be generated for a region of the multiple illuminated regions having the highest average brightness. To increase the realism of the generated target image, respective auxiliary light sources may be set for other illuminated regions of the particular visual representation after the particular visual representation is arranged in the background 910. Light analysis may be similarly performed on other illuminated regions to generate respective ambient light information, and then a light source position and light intensity' of an auxiliary light source may be calculated based on the generated ambient light information and the position of the particular visual representation in the background 910. In one example, the auxiliary light source may be defined to be set within a predetermined distance from the particular visual representation such that the virtual light of the auxiliary light source does not affect other visual representations.
[0089] In one example, after the visual representation 920 is arranged in the background 910, a virtual light effect may be added in the background 910 according to the determined light source position and light intensity7of the background virtual light source 940, to generate the target image with illumination information. As exemplarily shown in FIG. 9, the virtual light generated by the background virtual light source 940 forms such a virtual light effect: the upper right comer of the first seat from the left in the second row is illuminated brighter relative to other portions of the seat, the upper half of the second seat from the left in the second row- is illuminated brighter relative to other portions of the seat, the upper left comer of the third seat from the left in the second row is illuminated brighter relative to other portions, etc. In one example, the virtual lighteffect may also be added according to the light source position and the light intensity of the one or more optional auxiliary light sources. Through adding the virtual light effect, the realism of the generated target image may be further increased.
[0090] Contents related to the operations of determining the light source position and light intensity of the background virtual light source will be discussed in detail below. Various example manners may be adopted to determine the light source position and the light intensity of the background virtual light source.
[0091] An exemplary' manner of determining the light source position and the light intensity' of the background virtual light source may include determining the light source position and the light intensity of the background virtual light source with a predetermined light source configuration list. The predetermined light source configuration list may include a plurality of optional light source positions, for example, a first light source position, a second light source position, a third light source position, and the like; and a plurality of optional light intensities, for example, a first light intensity', a second light intensity, a third light intensity, and the like. A variety of ways may be employed to determine the optional light source positions in the predetermined light source configuration list. In one example, a plurality' of optional light source positions having a predetermined spatial interval may be set. For example, the background may be grid divided, and one optional light source position may be determined for every several grids. In one example, a plurality of optional light source positions may be pre-designated in other manners, for example, the first light source position may be located at the center of the background, the second light source position may be located at the upper left comer of the background, and the third light source position may be located at the upper middle of the background, etc. A variety of manners may be employed to determine the optional light intensities in the predetermined light source configuration list. In one example, a plurality of optional light intensities may be pre-designated in various manners, for example, through setting a plurality of predetermined light intensity values.
[0092] According to this example, an initial light source position and an initial light intensity may be first selected from the predetermined light source configuration list. The coordinates in the background of each of the one or more visual representations to be arranged into the background may then be calculated for the initial light source position and the initial light intensity, e.g., in a similar manner as discussed in connection with FIG. 9. Then, whether the distribution of the one or more visual representations in the background satisfies a predetermined criterion may be determined according to the calculated coordinates of one or more visual representations. In one example, the predetermined criterion described above may be associated with whether the distribution of one or more visual representations in the background is uniform.For example, an average of the coordinates of the one or more visual representations in the background may be calculated. If the calculated average of the coordinates approaches the coordinate of the center of the background, e.g., the difference between the two is less than a predetermined threshold, then the distribution of the visual representations may be considered to be uniform, and thus the predetermined criterion is met. If the calculated average of the coordinates is away from the coordinate of the center of the background, for example, the difference between the two is greater than the predetermined threshold, the distribution of the visual representations may be considered to be not uniform, and thus the predetermined criterion is not met. If the predetermined criterion is not met, an updated light source position and / or an updated light intensity may be selected in the predetermined light source configuration list. The above operations may then be re-performed for the updated light source position and / or the updated light intensity. In one example, the above operations may be performed multiple times until the predetermined criterion is met. In response to the predetermined criterion being met, the currently selected light source position and light intensity may be determined as the light source position and the light intensity of the background virtual light source. In one example, the background discussed above is a two-dimensional background image.
[0093] Another exemplary' manner of determining the light source position and the light intensity of the background virtual light source may include determining the light source position and the light intensity of the background virtual light source according to attributes of one or more visual representations to be arranged into the background. FIG. 10 illustrates an exemplary process 1000 of determining a light source position and a light intensity of a background virtual light source according to an attribute of a visual representation according to an embodiment.
[0094] An initial light source position of the background virtual light source may first be determined. In one example, the initial light source position may be the center of the background.
[0095] At 1010, the light intensity of the background virtual light source may be determined for the light source position of the background virtual light source. In one example, it may first be determined, based on the size of the one or more visual representations, e.g., according to the rule that is set for the size as discussed above, each visual representation should be arranged at which row of positions in the background. Then, for each visual representation, it may be calculated, according to an angular offset between the visual representation and its respective virtual light source and according to the coordinate of the background virtual light source, which direction the visual representation should be located with respect to the background virtual light source. The coordinate of the visual representation in the background may be calculated according to the determined coordinate of the row in which the visual representation is located and the direction as described above. Next, the distance between the visual representation and the background virtuallight source may be calculated according to the coordinate of the visual representation and the coordinate of the background virtual light source. Then, the light intensity required for the visual representation may be calculated according to the distance between the visual representation and the background virtual light source, and the proportional relationship between the distance from the visual representation to the virtual light source and the light intensity of the virtual light source as discussed above. Based on the required light intensity calculated for each visual representation, the light intensity that the background virtual light source should have may be calculated. In one example, the light intensity that the background virtual light source should have may be the median of the required light intensities for the one or more visual representations.
[0096] At 1020, a distribution of the one or more visual representations may be calculated for the light source position of the background virtual light source and the light intensity of the background virtual light source calculated at 1010. In one example, the coordinates of the one or more visual representations in the background may be calculated in a similar manner as discussed above in connection with FIG. 9 for the light source position and light intensity of the background virtual light source. The distribution of the visual representations may then be calculated through calculating an average of the coordinates of the one or more visual representations.
[0097] At 1030, whether an iteration stop condition is met may be determined. In one example, the iteration stop condition may be associated with whether the distribution of the one or more visual representations in the background is uniform. A distance threshold d may be preset, and the difference between the average of the calculated coordinates and the coordinate of the center of the background may be compared with the distance threshold d. If the difference between the calculated average of the coordinates and the coordinate of the center of the background is less than the distance threshold d, then the distribution of the visual representations may be considered to be uniform, and thus the iteration stop condition is met. If the difference between the calculated average of the coordinates and the coordinate of the center of the background is greater than the distance threshold d, then the visual representations may be considered to be not uniformly distributed, and thus the iteration stop condition is not met. In one example, the iteration stop condition may be associated with a count that the iteration is performed. In this example, each time the light source position of the background virtual light source is updated, it may be referred to that one iteration is performed. An iteration count threshold N may be preset, such that after N iterations are performed, the iteration stop condition is met.
[0098] If it is determined at 1030 that an iteration stop condition is not met (as indicated by the branch “No”), the process 1000 may proceed to 1040. At 1040, the light source position of the background virtual light source may be updated. The updated light source position makes the average of the coordinates of the visual representations to be closer to the center of the backgroundas compared with the previous iteration. This manner for updating makes the average of the coordinates of the visual representations to be increasingly closer to the center of the background. Thus, it is facilitated to generate a position allocating result in which the visual representations are uniformly distributed in the background. In one example, each time the light source position of the background virtual light source is updated, the step size that the coordinate of the background virtual light source moves may be determined based on a difference between the average of the coordinates of the visual representations in the previous iteration and the coordinate of the center of the background. For example, assume that the coordinate of the background virtual light source in the background in the previous iteration is (x. y). Meanwhile, the average of the coordinates of the visual representations in the background in the previous iteration is (xa, ), and the coordinate of the center of the background is (xc, w). Thus, the difference between the average of the coordinates of the visual representations and the coordinate of the center of the background is dx= Xa-Xc. dy= ya-yc. In response to the iteration stop condition being not met, the coordinate of the background virtual light source may be updated to x-dx, y-dy). The operations at 1010 through 1030 may then be repeated for the updated light source position of the background virtual light source.
[0099] After the operation at 1040 is performed, or if the iteration stop condition is determined to be met at 1030 (as shown by the branchL‘Yes”), process 1000 may proceed to 1050. At 1050, the current light source position and the current light intensity of the background virtual light source may be determined as the light source position and the light intensity for which the subsequent position allocation operation may be performed.
[0100] In one example, the background in the exemplary process 1000 of determining the light source position and the light intensity of the background virtual light source according to the attribute of the visual representation discussed above in connection with FIG. 10 is a two- dimensional background image.
[0101] FIG. 11 illustrates a schematic diagram 1100 of allocating a target position to a visual representation according to a rule that is set for a facial orientation of the visual representation according to an embodiment. FIG. 11 illustrates a background 1110. FIG. 11 also illustrates a visual representation 1120 for being adaptively arranged into the background 1110. The visual representation 1120 has an attribute 1130 that indicates that a facial orientation angle of the visual representation 1120 is. for example, 45 °.
[0102] A reference position in the background 1 110 may be determined. The reference position is a position in the background 1110 where it is desired that the face of the visual representation can be oriented as much as possible. The rule that is set for the facial orientation of the visual representation may specify the angular offset produced by the face orientation relativeto the reference position is matched to which candidate position in the background 1110.
[0103] The reference position in the background may be preset. A reference position 1140 located in the middle of the lower edge of the background 1110 is show n in the example of FIG. 11. A coordinate system defining 0 ° to 360 ° angle may be established with the reference position 1140 as a center. The matched candidate position may then be selected according to the facial orientation angle of the visual representation 1120. For example, a facial orientation 1150 at the 450direction may be determined in the coordinate system, and then a candidate position near the reverse extension line of the facial orientation 1150, e.g., a candidate position 1160, may be selected as the target position of the visual representation 1120.
[0104] FIG. 12 illustrates a schematic diagram 1200 of allocating a target position to a visual representation according to a rule that is set for an occlusion state of the visual representation according to an embodiment. FIG. 12 illustrates a background 1210. FIG. 12 also illustrates a visual representation 1220 for being adaptively arranged into the background 1210. The visual representation 1220 has an attribute 1230 that indicates that the right side of the visual representation 1220 is occluded.
[0105] The rule that is set for the occlusion state of the visual representation may specify the visual representation with what occlusion state is matched with which candidate position in the background 1210. In one example, the rule that is set for the occlusion state may specify that the visual representation that is occluded on one side is matched with the candidate position at the edge of the same side in the background. According to the above exemplary rule, the visual representation 1220 that is occluded on the right side may be allocated a candidate position at the right edge in the background 1210, e.g., a candidate position 1240, as the target position of the visual representation 1220.
[0106] Exemplary operations for performing position allocation according to the rule that is set for each attribute are discussed in detail above in connection with FIGS. 8-12. In one example, the position allocation may be performed according to a rule that is set in a case of jointly considering multiple attributes. In this example, the rule that is set may specify priorities of the attributes in order to perform the position allocation for each attribute, respectively, as per the priorities. FIG. 13 illustrates an exemplary process 1300 of performing position allocation according to priorities of attributes according to an embodiment.
[0107] At 1310, a priority of each attribute may be determined. In one example, the priority of the attribute may be specified by a preset rule. The attribute priority may be determined according to the degree of influence of each attribute on the position allocating result. In one example, if the visual representation is not arranged at a matched candidate position, e.g., does not conform to the ‘"near large, far small'’ principle, it may cause the generated target image toappear very unrealistic. Thus, a higher atribute priority may be allocated to size. For example, the preset rule may specify that the atribute priorities are in order of size, occlusion state, ambient light information, facial orientation, and the like, from high to low.
[0108] At 1320, a corresponding range of candidate positions may be determined for the current atribute. The atribute with a first priority, e.g., size, may be taken as the current atribute. Based on a size of each visual representation, a range of candidate positions of a first level for the visual representation may be determined. For example, in response to determining that the first visual representation and the second visual representation have a large size and the third visual representation has a small size, it may be determined that the range of candidate positions of the first level of the first visual representation and the second visual representation are the first two rows of seats, while the range of candidate positions of the first level of the third visual representation is the last two rows of seats.
[0109] At 1330, a target position may be allocated, within the range of candidate positions of the first level, to the visual representation based on another attribute having a lower priority than the priority of the current atribute. For example, the other atribute may be occlusion state having a second priority. For example, in response to determining that the occlusion state of the first visual representation is that the right side is occluded and the occlusion state of the second visual representation is not occluded, it may be determined that the two right-side seats of the first two rows of seats are suitable as the target position of the first visual representation, while the other seats of the first two rows of seats other than the two right-side seats are suitable as the target position of the second visual representation.
[0110] In one example, the operations at 1320 through 1330 may be performed iteratively until each attribute is fully considered or it is able to produce an allocation result. In this example, when performing the operation at 1320 through 1330 for the second time, the atribute with the second priority, for example, the occlusion state, may be used as the current atribute, and the corresponding range of candidate positions may be determined. For the first visual representation, the two right-side seats of the first two rows of seats may be determined as the range of candidate positions of the second level within the range of candidate positions of the first level. Similarly, for the second visual representation, other seats in the first two rows of seats other than the two right-side seats may be determined as the range of candidate positions of the second level within the range of candidate positions of the first level. Then, the target position may be allocated, within the range of candidate positions of the second level, to the visual representation based on the atribute with the third priority, for example, the ambient light information. Next, when performing the operation at 1320 through 1330 for the third time, the ambient light information may be used as the current atribute to determine the corresponding range of candidate positions.The target position may then be allocated to the visual representation based on the facial orientation. Through iteratively performing the operations at 1320 through 1330, a matched target position for each visual representation may be determined in a way of stepwise refinement based on a variety' of attributes.
[0111] In one example, the background may be modified in conjunction with the background modification operations discussed above, in cases for example where a suitable range of candidate positions cannot be determined for any one attribute or a suitable position allocating result cannot be generated, e.g., when the same position is allocated for two visual representations, or when there is no suitable position for a particular visual representation.
[0112] At 1340, position allocating information may be generated based on a result of the position allocation performed at 1330.
[0113] As discussed above in connection with FIG. 2, in addition to generating the position allocating information from the background and attributes based on predetermined rules, the position allocating information may also be generated from background and attributes utilizing a position allocating model. FIG. 14 illustrates a schematic diagram 1400 of generating position allocating information with a position allocating model according to an embodiment.
[0114] The position allocating model 1410 may be implemented as a multimodal generative model that takes a background 1420 and attributes 1430-1 through 1430-N of one or more visual representations as input, and outputs position allocating information 1440. The position allocating information 1440 may include a plurality of data entries, each of which may indicate, for each visual representation, which candidate position in the background is a corresponding target position of the visual representation.
[0115] In one example, the position allocating model 1410 may be implemented by any general multimodal generative model, such as a large language model (LLM) that supports multimodal. LLM is a type of large-scale language model with the ability7to efficiently perform understanding and generating of general language. LLM can obtain this capability through learning a large number of parameters using a large amount of data during training. In one example, the LLM that supports multimodal may include, for example, various GPT variants.
[0116] In one example, the position allocating model 1410 may be implemented by any multimodal generative model specifically trained to perform a position allocating task. In this example, the position allocating model 1410 may employ any suitable model architecture, such as SORA, Transformer architecture, or the like. The training data used to train the position allocating model 1410 may include a plurality of training data pairs. The input of each training data pair may be a background and attributes of one or more visual representations, and the output may be a candidate position in the background allocated for each visual representation.
[0117] After generating the position allocating information, a target image may be generated through arranging each visual representation at a corresponding target position according to the position allocating information. FIG. 15 illustrates a schematic diagram 1500 of a visual representation filling region associated with a candidate position according to an embodiment. FIG. 15 illustrates a background 1510. Each candidate position in the background 1510 may be associated with a visual representation filling region. For example, FIG. 15 illustrates a visual representation filling region 1525 associated with a candidate position 1520. If the candidate position 1520 is selected as the corresponding target position of a particular visual representation, the arrangement of the particular visual representation at the candidate position 1520 may be achieved through filling the particular visual representation into the visual representation filling region 1525. In one example, the size of each visual representation filling region may be predetermined. In one example, the size of each visual representation filling region may be determined through being adjusted according to the size of the visual representation to be filled. In one example, the relative position between the candidate position and the associated visual representation filling region may be predefined, e.g., the candidate position is at the center of the associated visual representation filling region.
[0118] FIG. 16 illustrates a schematic diagram 1600 of a target image according to an embodiment. As shown in FIG. 16. the target image 1610 may be generated through arranging each visual representation at a corresponding target position. In the target image 1610, a visual representation 1620 and a visual representation 1630 are filled in a visual representation filling region 1625 and a visual representation filling region 1635, respectively, e.g., according to a position allocating result determined for size in connection with FIG. 8. A visual representation 1640 is filled in a visual representation filling region 1645, e.g., according to a position allocating result determined for ambient light information in connection with FIG. 9. A visual representation 1650 is filled in a visual representation filling region 1655, e.g., according to a position allocating result determined for facial orientation in connection with FIG. 11. A visual representation 1660 is filled in a visual representation filling region 1665, e.g.. according to a position allocating result determined for occlusion state in connection with FIG. 12.
[0119] In one example, when the background is a two-dimensional background image, generating the target image 1610 may involve the operation of covering the original pixels in the background with pixels of the visual representation in the respective visual representation filling region. In one example, when the background is a three-dimensional background model, generating the target image 1610 may involve the operation of performing data fusion on the two- dimensional data of the visual representation and the three-dimensional data in the respective visual representation filling region in the three-dimensional background model, and the operationof converting the fused three-dimensional model into a two-dimensional target image 1610.
[0120] In one example, an image enhancing operation may be performed on the target image 1610 to further improve the realism of the target image 1610. In one example, the image enhancing operation may include processing pixels at edges of the visual representation in the target image 1610 to make the edges of the visual representation being smoother. In one example, the image enhancing operation may include performing color processing on the target image 1610, for example, color saturation processing, etc., to make the color of the target image 1610 being more realistic. In one example, the image enhancing operation may include any image processing operation for making the two-dimensional target image 1610 obtained through performing data conversion on the three-dimensional model being more realistic.
[0121] In one example, an image enhancement model may be utilized to perform the image enhancing operation for the target image 1610. The image enhancing model may be implemented as an Al model of any architecture for performing an image enhancing task, for example, a Midjoumey architecture, a SORA architecture, or the like may be used.
[0122] FIG. 17 illustrates a flowchart of an exemplary method 1700 for video communication according to an embodiment.
[0123] At 1710, one or more video streams for one or more video communication attendees may be received, wherein each video stream is associated with at least one video communication attendee.
[0124] At 1720, at least one visual representation of the at least one video communication attendee and an attribute of each visual representation of the at least one visual representation may be obtained for each video stream.
[0125] At 1730, a background comprising a plurality of candidate positions may be obtained.
[0126] At 1740, position allocating information may be generated based on the background and attributes of one or more visual representations of the one or more video communication attendees, wherein the position allocating information may indicate a target position corresponding to each visual representation of the one or more visual representations among the plurality of candidate positions.
[0127] At 1750, a target image may be generated through arranging each visual representation of the one or more visual representations at a corresponding target position according to the position allocating information.
[0128] The obtaining a visual representation of the at least one video communication attendee and an attribute of each visual representation of the at least one visual representation may comprise: determining one or more target frames in the video stream; and obtaining, for each target frame of the one or more target frames, the at least one visual representation and the attributeof each visual representation of the at least one visual representation.
[0129] In one implementation, the obtaining, for each target frame, the at least one visual representation and the attribute of each visual representation of the at least one visual representation may comprise at least one of: generating the at least one visual representation through performing, for the target frame, mask segmenting with a mask segmenting model: and generating the attribute of each visual representation of the at least one visual representation through performing, for the target frame, attribute detecting with an attribute detecting model.
[0130] In one implementation, the obtaining, for each target frame, the at least one visual representation and the attribute of each visual representation of the at least one visual representation may comprise at least one of: receiving the at least one visual representation associated with the target frame in the video stream; and receiving the attribute of each visual representation of the at least one visual representation associated with the target frame in the video stream.
[0131] In one implementation, the attribute detecting model may be implemented through adding at least one additional layer on the basis of the mask segmenting model.
[0132] In one implementation, the attribute of each visual representation of the one or more visual representations may comprise at least one of a size, ambient light information, facial orientation, and occlusion state of the visual representation. The ambient light information may indicate a light intensity of a virtual light source of the visual representation and relative position information between the visual representation and the virtual light source.
[0133] In one implementation, when the attribute includes the ambient light information, the method 1700 may further include: determining a light source position and a light intensity of a background virtual light source for producing virtual light in the background. The determining a light source position and a light intensity of a background virtual light source may comprise: determining, with a predetermined light source configuration list, the light source position and the light intensity of the background virtual light source; or determining the light source position and the light intensity of the background virtual light source according to the attributes of the one or more visual representations.
[0134] In one implementation, the determining, with a predetermined light source configuration list, the light source position and the light intensity of the background virtual light source may comprise: selecting, from the predetermined light source configuration list, the light source position and the light intensity of the background virtual light, such that a distribution of the one or more visual representations calculated for the light source position and the light intensity of the background virtual light meets a predetermined criterion.
[0135] In one implementation, the determining the light source position and the light intensityof the background virtual light source according to the attributes of the one or more visual representations may comprise: a group of iterative operations: calculating, for a current light source position of the background virtual light source, a current light intensity7of the background virtual light source according to the attributes of the one or more visual representations; calculating, for the current light source position and the current light intensity of the background virtual light source, a distribution of the one or more visual representations; determining whether the distribution meets an iteration stop condition; and in response to the iteration stop condition is not met, repeatedly performing the group of iterative operations; or in response to the iteration stop condition is met, determining the current light source position and the current light intensity of the background virtual light source as the light source position and the light intensity of the background virtual light source.
[0136] In one implementation, the generating position allocating information may comprise: allocating a corresponding target position to each visual representation of the one or more visual representations according to predetermined rules, wherein the predetermined rules include at least one of: a size of the visual representation being matched with the target position; ambient light information of the visual representation being matched with light produced at the target position by the background virtual light source; an angular offset of a facial orientation of the visual representation relative to a reference position in the background being matched with the target position; and an occlusion state of the visual representation being matched with the target position.
[0137] In one implementation, the generating position allocating information may comprise: determining a priority of each attribute of multiple attributes; and allocating, within a range of candidate positions corresponding to a priority7of a current attribute of the multiple attributes, the target position to each visual representation of the one or more visual representations based on at least another attribute having a lower priority than the priority7.
[0138] In one implementation, the generating position allocating information may comprise: generating, with a position allocating model, the position allocating information based on the background and the attributes of the one or more visual representations.
[0139] In one implementation, the obtaining a background may comprise: selecting, according to the attributes of the one or more visual representations, the background from a plurality of backgrounds or generating the background with a background generating model.
[0140] In one implementation, the obtaining a background may comprise: selecting an initial background from a plurality of backgrounds; and modifying the initial background according to the attributes of the one or more visual representations, to generate the background.
[0141] In one implementation, the generating a target image may comprise: adding a virtual light effect in the background at least according to the light source position and the light intensity7of the background virtual light source.
[0142] In one implementation, the method 1700 may further include: performing image enhancing to the target image with an image enhancing model.
[0143] It should be understood that the method 1700 may further include any steps / processes for video communication according to the above embodiments of the present disclosure.
[0144] FIG. 18 illustrates an exemplary apparatus 1800 for video communication according to an embodiment.
[0145] The apparatus 1800 may include: a video stream receiving module 1810 for receiving one or more video streams for one or more video communication attendees, each video stream being associated with at least one video communication attendee; a visual representation and attribute obtaining module 1820 for obtaining, for said each video stream, at least one visual representation of the at least one video communication attendee and an attribute of each visual representation of the at least one visual representation; a background obtaining module 1830 for obtaining a background comprising a plurality of candidate positions; a position allocating information generating module 1840 for generating position allocating information based on the background and attributes of one or more visual representations of the one or more video communication attendees, the position allocating information indicating a target position corresponding to each visual representation of the one or more visual representations among the plurality of candidate positions; and a target image generating module 1850 for generating a target image through arranging each visual representation of the one or more visual representations at a corresponding target position according to the position allocating information.
[0146] In addition, the apparatus 1800 may further include any other module configured to perform any operation of the method for video communication according to the above embodiments of the present disclosure.
[0147] FIG. 19 illustrates an exemplary apparatus 1900 for video communication according to an embodiment.
[0148] The apparatus 1900 may include at least one processor 1910. The apparatus 1900 may also include a memory 1920 connected to the at least one processor 1910. The memory 1920 may store computer-executable instructions that, when executed, cause the at least one processor 1910 to: receive one or more video streams for one or more video communication attendees, each video stream being associated with at least one video communication attendee; obtain, for said each video stream, at least one visual representation of the at least one video communication attendee and an attribute of each visual representation of the at least one visual representation; obtain a background comprising a plurality of candidate positions; generate position allocating information based on the background and attributes of one or more visual representations of the one or morevideo communication atendees, the position allocating information indicating a target position corresponding to each visual representation of the one or more visual representations among the plurality of candidate positions; and generate a target image through arranging each visual representation of the one or more visual representations at a corresponding target position according to the position allocating information. In addition, the at least one processor 1910 may be further configured to perform any other operation of the method for video communication according to the above embodiments of the present disclosure.
[0149] Embodiments of the present disclosure may be implemented in a non- transitory7computer-readable medium. The non-transitory computer-readable medium may include instructions that, when executed, cause at least one processor to: receive one or more video streams for one or more video communication atendees, each video stream being associated with at least one video communication atendee; obtain, for said each video stream, at least one visual representation of the at least one video communication atendee and an atribute of each visual representation of the at least one visual representation; obtain a background compnsing a plurality of candidate positions; generate position allocating information based on the background and attributes of one or more visual representations of the one or more video communication atendees, the position allocating information indicating a target position corresponding to each visual representation of the one or more visual representations among the plurality of candidate positions; and generate a target image through arranging each visual representation of the one or more visual representations at a corresponding target position according to the position allocating information. In addition, the instructions, when executed, may also cause the at least one processor to perform any step / process of the method for video communication according to the above embodiments of the present disclosure.
[0150] Embodiments of the present disclosure provide a computer program product comprising computer programs being operable by a processor to: receive one or more video streams for one or more video communication atendees, each video stream being associated with at least one video communication atendee; obtain, for said each video stream, at least one visual representation of the at least one video communication attendee and an atribute of each visual representation of the at least one visual representation; obtain a background comprising a plurality' of candidate positions; generate position allocating information based on the background and attributes of one or more visual representations of the one or more video communication atendees, the position allocating information indicating a target position corresponding to each visual representation of the one or more visual representations among the plurality7of candidate positions; and generate a target image through arranging each visual representation of the one or more visual representations at a corresponding target position according to the position allocatinginformation. In addition, the computer programs are also operable by the processor to perform any step / process of the method for video communication according to the above embodiments of the present disclosure.
[0151] It should be appreciated that all the operations in the methods described above are merely exemplary, and the present disclosure is not limited to any operations in the methods or orders of these operations, and should cover all other equivalents under the same or similar concepts.
[0152] Moreover, the articles “a’' and "an" as used in this specification and the appended claims should generally be construed to mean “one’?or “one or more7' unless specified otherwise or clear from the context to be directed to a singular form.
[0153] It should also be appreciated that all the modules in the apparatuses described above may be implemented in various approaches. These modules may be implemented as hardware, software, or a combination thereof. Moreover, any of these modules may be further functionally divided into sub-modules or combined together.
[0154] Processors have been described in connection with various apparatuses and methods. These processors may be implemented using electronic hardware, computer software, or any combination thereof. Whether such processors are implemented as hardware or software will depend upon the particular application and overall design constraints imposed on the system. Byway of example, a processor, any portion of a processor, or any combination of processors presented in the present disclosure may be implemented with a micro-processor, micro-controller, digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic device (PLD). a state machine, gated logic, discrete hardware circuits, and other suitable processing components configured to perform the various functions described in the present disclosure. The functionality' of a processor, any portion of a processor, or any' combination of processors presented in the present disclosure may be implemented with software being executed by a microprocessor, micro-controller, DSP, or other suitable platform.
[0155] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, threads of execution, procedures, functions, etc. The software may reside on a computer-readable medium. A computer-readable medium may include, by way of example, memory such as a magnetic storage device (e.g., hard disk, floppy disk, magnetic strip), an optical disk, a smart card, a flash memory device, random access memory (RAM), read only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), a register, or a removable disk. Although a memory is shown as being separate from the processor in various aspects presented inthis disclosure, the memory may also be internal to the processor (e.g., a cache or a register).
[0156] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents to the elements of the various aspects described throughout the present disclosure that are know n or later come to be known to those of ordinary skilled in the art are intended to be encompassed by the claims.
Claims
CLAIMS1. A method for video communication, comprising: receiving one or more video streams for one or more video communication attendees, each video stream being associated with at least one video communication attendee; obtaining, for said each video stream, at least one visual representation of the at least one video communication attendee and attributes of each visual representation of the at least one visual representation; obtaining a background comprising a plurality of candidate positions; generating position allocating information based on the background and attributes of one or more visual representations of the one or more video communication attendees, the position allocating information indicating a target position corresponding to each visual representation of the one or more visual representations among the plurality of candidate positions; and generating a target image through arranging each visual representation of one or more visual representations at a corresponding target position according to the position allocating information.
2. The method of claim 1, wherein the obtaining at least one visual representation of the at least one video communication attendee and attributes of each visual representation of the at least one visual representation comprises: determining one or more target frames in the video stream; and obtaining, for each target frame of the one or more target frames, the at least one visual representation and the attributes of each visual representation of the at least one visual representation.
3. The method of claim 2. wherein the obtaining, for each target frame, the at least one visual representation and the attributes of each visual representation of the at least one visual representation comprises at least one of: generating the at least one visual representation through performing, for the target frame, mask segmenting with a mask segmenting model; and generating the attributes of each visual representation of the at least one visual representation through performing, for the target frame, attribute detecting with an attribute detecting model.
4. The method of claim 2, the obtaining, for each target frame, the at least one visual representation and the attributes of each visual representation of the at least one visual representation comprises at least one of: receiving the at least one visual representation associated with the target frame in the video stream; and receiving the attributes of each visual representation of the at least one visual representation associated with the target frame in the video stream.
5. The method of claim 1, wherein, the attributes of each visual representation of the one or more visual representations comprise at least one of a size, ambient light information, facial orientation, and occlusion state of the visual representation; and the ambient light information indicates a light intensity of a virtual light source of the visual representation and relative position information between the visual representation and the virtual light source.
6. The method of claim 5, wherein in a case that attributes include the ambient light information, the method further comprises: determining a light source position and a light intensity of a background virtual light source for producing virtual light in the background.
7. The method of claim 6, wherein the determining a light source position and a light intensity of a background virtual light source comprises: determining, with a predetermined light source configurations list, the light source position and the light intensity of the background virtual light source; or determining the light source position and the light intensity7of the background virtual light source according to the attributes of the one or more visual representations.
8. The method of claim 7, wherein the determining the light source position and the light intensity of the background virtual light source according to the attributes of the one or more visual representations comprises: a group of iterative operations: calculating, for a current light source position of the background virtual light source, a current light intensity of the background virtual light source according to the attributes of the one or more visual representations; calculating, for the current light source position and the current light intensity of the background virtual light source, a distribution of the one or more visual representations; determining whether the distribution meets an iteration stop condition; and in response to the iteration stop condition is not met, repeatedly performing the group of iterative operations; or in response to the iteration stop condition is met, determining the current light source position and the current light intensity of the background virtual light source as the light source position and the light intensity of the background virtual light source.
9. The method of claim 6, wherein the generating position allocating information comprises: allocating a corresponding target position to each visual representation of the one or more visual representations according to predetermined rules, wherein the predetermined rules include at least one of:a size of the visual representation being matched with the target position; ambient light information of the visual representation being matched with light produced at the target position by the background virtual light source; an angular offset of a facial orientation of the visual representation relative to a reference position in the background being matched with the target position; and an occlusion state of the visual representation being matched with the target position.
10. The method of claim 6, wherein the generating position allocating information comprises: determining a priority of each kind of attribute of multiple kinds of attributes; and allocating, within a range of candidate positions corresponding to a priority of a current attribute of the multiple kinds of attributes, the target position to each visual representation of the one or more visual representations based on at least another kind of attributes having a lower priority than the priority.
11. The method of claim 1, wherein the generating position allocating information comprises: generating, with a position allocating model, the position allocating information based on the background and the attributes of the one or more visual representations.
12. The method of claim 1, wherein the obtaining a background comprises: selecting, according to the attributes of the one or more visual representations, the background from a plurality of backgrounds or generating the background with a background generating model.
13. The method of claim 1, wherein the obtaining a background comprises: selecting an initial background from a plurality of backgrounds; and modifying the initial background according to the attributes of the one or more visual representations, to generate the background.
14. An apparatus for video communication, comprising: at least one processor; and a memory storing computer-executable instructions that, when executed, cause the at least one processor to: receive one or more video streams for one or more video communication attendees, each video stream being associated with at least one video communication attendee: obtain, for said each video stream, at least one visual representation of the at least one video communication attendee and attributes of each visual representation of the at least one visual representation; obtain a background comprising a plurality of candidate positions; generate position allocating information based on the background and attributes ofone or more visual representations of the one or more video communication attendees, the position allocating information indicating a target position corresponding to each visual representation of the one or more visual representations among the plurality' of candidate positions; and generate a target image through arranging each visual representation of one or more visual representations at a corresponding target position according to the position allocating information.
15. A non -transitory computer-readable medium, comprising instructions that, when executed, cause at least one processor to: receive one or more video streams for one or more video communication attendees, each video stream being associated with at least one video communication attendee; obtain, for said each video stream, at least one visual representation of the at least one video communication attendee and attributes of each visual representation of the at least one visual representation; obtain a background comprising a plurality of candidate positions; generate position allocating information based on the background and attributes of one or more visual representations of the one or more video communication attendees, the position allocating information indicating a target position corresponding to each visual representation of the one or more visual representations among the plurality of candidate positions; and generate a target image through arranging each visual representation of one or more visual representations at a corresponding target position according to the position allocating information.
Citation Information
Patent Citations
System and method for augmented reality multi-view telepresence
US11363240B2
Compositing Video Streams
US20110025819A1
Method and system for adapting a CP layout according to interaction between conferees
US20140002585A1
Systems and methods for integrating user personas with content during video conferencing
US20150029294A1