Video conference interaction method and apparatus

HK40088219BActive Publication Date: 2026-07-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
HK · HK
Patent Type
Patents
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2023-07-21
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In video conferencing, the varying distances and positions of participants from the conferencing terminal result in cluttered images in the video, impacting interaction efficiency.

Method used

By displaying adaptive human images in a virtual frame, the size of the adaptive human images is matched with the area where the virtual location is located. The virtual frame is generated and sent to the terminal by adaptively adjusting the human image parameters and face parameters.

Benefits of technology

It enhances the realism and interactive efficiency of video conferencing, avoids cluttered images of people, and improves the interactive effect of the meeting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The application relates to a method and device for video conference interaction, computer equipment, a storage medium and a computer program product. The method relates to online virtual video conference technology, and comprises the following steps: in response to a trigger operation of joining a video conference among a plurality of terminals, a virtual picture-in-picture picture of the video conference is displayed, and a virtual scene of the virtual picture-in-picture picture comprises a plurality of virtual positions for accommodating conference members; in at least two virtual positions of the virtual scene, adaptive portraits corresponding to human images in real-time pictures of the conference members collected by at least two terminals in the plurality of terminals are displayed; wherein the size of each adaptive portrait displayed in the at least two virtual positions matches the size of the position area where the corresponding virtual position is located. The method can improve the interaction efficiency of the video conference.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. 2021114425187, filed on November 30, 2021, entitled “Method and Apparatus for Video Conferencing Interaction,” the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of computer technology, and in particular to a method, apparatus, computer equipment, storage medium, and computer program product for video conferencing interaction, as well as a method, apparatus, computer equipment, storage medium, and computer program product for processing video conferencing images. Background Technology

[0003] With the development of computer technology, video conferencing technology, which allows people in different locations to communicate face-to-face via communication devices and networks, has been widely used in entertainment, education, training, marketing, and advertising. In video conferencing, participants collect video data through their respective conferencing terminals, which is then aggregated and displayed on the video conferencing interface.

[0004] However, the varying distances and positions of participants from the conference terminal result in cluttered images of people in the video displayed on the video conferencing interface, affecting the realism of the video conference and reducing the efficiency of interaction during the video conferencing process. Summary of the Invention

[0005] Therefore, it is necessary to provide a method, apparatus, computer equipment, storage medium, and computer program product for video conferencing interaction that can improve the interactive efficiency of video conferencing, as well as a method, apparatus, computer equipment, storage medium, and computer program product for processing video conferencing images, in response to the above-mentioned technical problems.

[0006] A method for video conferencing interaction, the method comprising:

[0007] In response to the triggering operation of joining a video conference between multiple terminals, a virtual frame of the video conference is displayed. The virtual scene of the virtual frame includes multiple virtual locations to accommodate the conference members.

[0008] At at least two virtual locations in a virtual scene, display adaptive human images corresponding to the real-time images of meeting members captured by at least two of multiple terminals.

[0009] At least two virtual locations each display an adaptive human portrait size that matches the size of the corresponding virtual location's location area.

[0010] In one embodiment, the method further includes: in response to a target terminal among the plurality of terminals being in an abnormal network state, displaying network abnormality prompt information about the meeting member corresponding to the adaptive portrait at the target virtual location of the adaptive portrait corresponding to the portrait in the real-time image of the meeting members collected by the target terminal.

[0011] A video conferencing interaction apparatus, the apparatus comprising:

[0012] The co-op display module is used to respond to the trigger operation of joining a video conference between multiple terminals and display the virtual co-op screen of the video conference. The virtual scene of the virtual co-op screen includes multiple virtual positions to accommodate the meeting members.

[0013] The adaptive human image display module is used to display adaptive human images corresponding to the human images in the real-time images of meeting members captured by at least two of multiple terminals at at least two virtual locations in a virtual scene.

[0014] At least two virtual locations each display an adaptive human portrait size that matches the size of the corresponding virtual location's location area.

[0015] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program performing the following steps:

[0016] In response to the triggering operation of joining a video conference between multiple terminals, a virtual frame of the video conference is displayed. The virtual scene of the virtual frame includes multiple virtual locations to accommodate the conference members.

[0017] At at least two virtual locations in a virtual scene, display adaptive human images corresponding to the real-time images of meeting members captured by at least two of multiple terminals.

[0018] At least two virtual locations each display an adaptive human portrait size that matches the size of the corresponding virtual location's location area.

[0019] A computer-readable storage medium having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0020] In response to the triggering operation of joining a video conference between multiple terminals, a virtual frame of the video conference is displayed. The virtual scene of the virtual frame includes multiple virtual locations to accommodate the conference members.

[0021] At at least two virtual locations in a virtual scene, display adaptive human images corresponding to the real-time images of meeting members captured by at least two of multiple terminals.

[0022] At least two virtual locations each display an adaptive human portrait size that matches the size of the corresponding virtual location's location area.

[0023] A computer program product, comprising a computer program that, when executed by a processor, performs the following steps:

[0024] In response to the triggering operation of joining a video conference between multiple terminals, a virtual frame of the video conference is displayed. The virtual scene of the virtual frame includes multiple virtual locations to accommodate the conference members.

[0025] At at least two virtual locations in a virtual scene, display adaptive human images corresponding to the real-time images of meeting members captured by at least two of multiple terminals.

[0026] At least two virtual locations each display an adaptive human portrait size that matches the size of the corresponding virtual location's location area.

[0027] The aforementioned video conferencing interaction methods, devices, computer equipment, storage media, and computer program products, when triggering a video conference between multiple terminals, display a virtual frame of the video conference. The virtual scene of this frame includes multiple virtual locations to accommodate conference members. At at least two virtual locations within the virtual scene, adaptive human images corresponding to the real-time images of conference members captured by at least two of the multiple terminals are displayed. The size of the displayed adaptive human images matches the size of the corresponding virtual location's area. During the video conferencing interaction, displaying adaptive human images at virtual locations within the virtual scene of the corresponding virtual frame, with a size matching the size of the corresponding virtual location's area, controls the size of the displayed adaptive human images based on the location area of ​​the virtual location in the virtual scene. This avoids cluttered images in the real-time images of conference members, enhances the realism of the video conference, and thus improves the interactive efficiency of the video conference.

[0028] A method for processing video conference footage, the method comprising:

[0029] Acquire real-time video feeds of meeting participants captured by multiple terminals that have joined the video conference;

[0030] For each meeting member's real-time video feed, based on the facial parameters and face parameters corresponding to the person's image in the real-time video feed, facial analysis is performed on the person's image in the real-time video feed to obtain the adaptive adjustment parameters corresponding to the person's real-time video feed.

[0031] Based on the adaptive adjustment parameters, the portraits of the meeting members in the real-time video of the meeting members are adaptively adjusted to obtain the adaptive portraits corresponding to the portraits in the real-time video of the meeting members.

[0032] The virtual co-op images generated based on each adaptive human portrait are sent to each terminal for display on each terminal.

[0033] A video conferencing image processing apparatus, the apparatus comprising:

[0034] The real-time video acquisition module is used to acquire the real-time video of each participant captured by multiple terminals that have joined the video conference.

[0035] The parameter acquisition module is used to perform portrait analysis on the portrait of each meeting member in the real-time image based on the portrait parameters and face parameters corresponding to the portrait in the real-time image of the meeting member, and obtain the adaptive adjustment parameters corresponding to the real-time image of the meeting member.

[0036] The adaptive adjustment module is used to adaptively adjust the portraits of meeting members in the real-time video of the meeting members based on the adaptive adjustment parameters, so as to obtain the adaptive portraits corresponding to the portraits in the real-time video of the meeting members.

[0037] The on-screen display module is used to send virtual on-screen images generated based on each adaptive human image to each terminal for display on each terminal.

[0038] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program performing the following steps:

[0039] Acquire real-time video feeds of meeting participants captured by multiple terminals that have joined the video conference;

[0040] For each meeting member's real-time video feed, based on the facial parameters and face parameters corresponding to the person's image in the real-time video feed, facial analysis is performed on the person's image in the real-time video feed to obtain the adaptive adjustment parameters corresponding to the person's real-time video feed.

[0041] Based on the adaptive adjustment parameters, the portraits of the meeting members in the real-time video of the meeting members are adaptively adjusted to obtain the adaptive portraits corresponding to the portraits in the real-time video of the meeting members.

[0042] The virtual co-op images generated based on each adaptive human portrait are sent to each terminal for display on each terminal.

[0043] A computer-readable storage medium having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0044] Acquire real-time video feeds of meeting participants captured by multiple terminals that have joined the video conference;

[0045] For each meeting member's real-time video feed, based on the facial parameters and face parameters corresponding to the person's image in the real-time video feed, facial analysis is performed on the person's image in the real-time video feed to obtain the adaptive adjustment parameters corresponding to the person's real-time video feed.

[0046] Based on the adaptive adjustment parameters, the portraits of the meeting members in the real-time video of the meeting members are adaptively adjusted to obtain the adaptive portraits corresponding to the portraits in the real-time video of the meeting members.

[0047] The virtual co-op images generated based on each adaptive human portrait are sent to each terminal for display on each terminal.

[0048] A computer program product, comprising a computer program that, when executed by a processor, performs the following steps:

[0049] Acquire real-time video feeds of meeting participants captured by multiple terminals that have joined the video conference;

[0050] For each meeting member's real-time video feed, based on the facial parameters and face parameters corresponding to the person's image in the real-time video feed, facial analysis is performed on the person's image in the real-time video feed to obtain the adaptive adjustment parameters corresponding to the person's real-time video feed.

[0051] Based on the adaptive adjustment parameters, the portraits of the meeting members in the real-time video of the meeting members are adaptively adjusted to obtain the adaptive portraits corresponding to the portraits in the real-time video of the meeting members.

[0052] The virtual co-op images generated based on each adaptive human portrait are sent to each terminal for display on each terminal.

[0053] The aforementioned video conferencing processing method, apparatus, computer equipment, storage medium, and computer program product, for each real-time image captured by multiple terminals joining the video conference, performs facial analysis on the images of the conference members in their real-time images based on facial and human face parameters. Then, based on the obtained adaptive adjustment parameters, the images of the conference members in their real-time images are adaptively adjusted to obtain adaptive human images. Virtual side-by-side images generated based on these adaptive human images are sent to each terminal for display. In processing the video conferencing images, adaptively adjusting the images of each conference member in their real-time images, based on facial and human face parameters determined through facial analysis, ensures that the size of the adaptive human image displayed in the virtual side-by-side image on the terminal matches the size of the corresponding virtual location area. This avoids cluttered images in the real-time images of conference members, enhances the realism of the video conference, and thus improves the interactive efficiency of the video conference. Attached Figure Description

[0054] Figure 1 This is an application environment diagram of a video conferencing interaction method in one embodiment;

[0055] Figure 2 This is a flowchart illustrating a video conferencing interaction method in one embodiment;

[0056] Figure 3 This is a schematic diagram of the interface of a virtual frame in one embodiment;

[0057] Figure 4 This is a schematic diagram of the interface at each virtual location in a virtual scene in one embodiment;

[0058] Figure 5 This is a schematic diagram of the interface distribution of virtual locations in a virtual scene in one embodiment;

[0059] Figure 6 This is a schematic diagram of an interface displaying meeting member information in one embodiment;

[0060] Figure 7 This is a schematic diagram of the interface comparing the location area and the human image area in one embodiment;

[0061] Figure 8 This is a schematic diagram of an interface comparing the uniform proportions of human figures in one embodiment.

[0062] Figure 9 This is a schematic diagram illustrating the interface changes that trigger screen sharing in one embodiment;

[0063] Figure 10This is a schematic diagram of the interface changes when ending screen sharing in one embodiment;

[0064] Figure 11 This is a schematic diagram illustrating the interface changes of a moving adaptive human figure in one embodiment;

[0065] Figure 12 This is a schematic diagram of an interface displaying a network error message in one embodiment;

[0066] Figure 13 This is a schematic diagram of an interface displaying an abnormal status prompt in one embodiment;

[0067] Figure 14 This is a flowchart illustrating the process of determining member parameters in one embodiment;

[0068] Figure 15 This is a flowchart illustrating a method for processing video conference images in one embodiment;

[0069] Figure 16 This is a schematic diagram of the architecture of a video conferencing system with multiple cameras in one embodiment.

[0070] Figure 17 This is an interactive schematic diagram of a video conferencing system with multiple cameras in one embodiment;

[0071] Figure 18 This is a flowchart illustrating a video conferencing interaction method in another embodiment;

[0072] Figure 19 This is a flowchart illustrating the process of determining adaptive adjustment parameters in one embodiment;

[0073] Figure 20 This is a flowchart illustrating the adaptive adjustment process in one embodiment;

[0074] Figure 21 This is a flowchart illustrating the process of determining the direction of portrait parameter processing in one embodiment;

[0075] Figure 22 This is a schematic diagram of the process of updating the parameter queue in one embodiment;

[0076] Figure 23 This is a schematic diagram illustrating the meaning of offsets in one embodiment;

[0077] Figure 24 This is a schematic diagram of the process for sharpening a human face mask image in one embodiment;

[0078] Figure 25 This is a flowchart illustrating the threshold determination of face parameters and portrait parameters in one embodiment;

[0079] Figure 26This is a flowchart illustrating the process of determining the width offset in one embodiment;

[0080] Figure 27 This is a flowchart illustrating the process of determining the height offset in one embodiment;

[0081] Figure 28 This is a flowchart illustrating the process of determining the face scaling factor in one embodiment;

[0082] Figure 29 This is a flowchart illustrating the process of determining the scaling factor in one embodiment;

[0083] Figure 30 This is a flowchart illustrating the unified handling of characters when they are too small in one embodiment.

[0084] Figure 31 This is a flowchart illustrating the unified handling of excessively large characters in one embodiment.

[0085] Figure 32 This is a schematic diagram of the image synthesis process in one embodiment;

[0086] Figure 33 This is a structural block diagram of a video conferencing interaction device in one embodiment;

[0087] Figure 34 This is a structural block diagram of a video conferencing screen processing device in one embodiment;

[0088] Figure 35 This is an internal structural diagram of a computer device in one embodiment;

[0089] Figure 36 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation

[0090] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0091] The video conferencing interaction method provided in this application can be applied to, for example... Figure 1In the application environment shown, multiple terminals 102 communicate with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed in the cloud or on another network server. When a video conference is triggered between multiple terminals 102, each terminal 102 displays a virtual frame of the video conference. The virtual frame includes multiple virtual locations to accommodate conference members. At at least two virtual locations within the virtual scene, adaptive portraits corresponding to the real-time images of the conference members captured by at least two of the terminals are displayed. The size of the displayed adaptive portraits matches the size of the corresponding virtual location's area.

[0092] The video conferencing image processing method provided in this application can be applied to, for example... Figure 1 In the application environment shown, server 104 acquires real-time images of conference members from multiple terminals 102 that have joined the video conference. For each real-time image of a conference member, server 104 performs facial analysis on the image of the conference member in the real-time image based on the facial parameters and face parameters corresponding to the image in the real-time image. Based on the obtained adaptive adjustment parameters, server 104 adaptively adjusts the image of the conference member in the real-time image to obtain an adaptive image corresponding to the image in the real-time image. Server 104 then sends the virtual frame-to-frame image generated based on each adaptive image to each terminal 102 for display on each terminal 102.

[0093] The terminal 102 can collect video data, and can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0094] The video conferencing technology involved in this application is a specific implementation of cloud conferencing. Cloud conferencing is an efficient, convenient, and low-cost meeting format based on cloud computing technology. Users only need to perform simple and easy-to-use operations through an internet interface to quickly and efficiently share voice, data files, and video with teams and clients around the world. The complex technologies such as data transmission and processing during the meeting are handled by the cloud conferencing service provider. Currently, domestic cloud conferencing mainly focuses on services based on the SaaS (Software as a Service) model, including telephone, internet, and video services. Video conferencing based on cloud computing is called cloud conferencing. In the era of cloud conferencing, data transmission, processing, and storage are all handled by the computer resources of the video conferencing provider. Users no longer need to purchase expensive hardware or install cumbersome software; they only need to open a browser and log in to the corresponding interface to conduct efficient remote meetings.

[0095] In one embodiment, such as Figure 2 As shown, a method for video conferencing interaction is provided, which can be applied to... Figure 1 Taking the terminal in the example, the explanation includes the following steps:

[0096] Step 202: In response to the triggering operation of joining a video conference between multiple terminals, a virtual frame of the video conference is displayed. The virtual scene of the virtual frame includes multiple virtual locations to accommodate the conference members.

[0097] In this context, the terminal refers to the user-side device participating in the video conference. Users join the video conference through the terminal; for example, users can log in to the terminal's video conference client using their corresponding account to join the conference. In practical applications, the terminal has video data acquisition capabilities. For instance, the terminal can be configured with an image sensing device, such as a camera, allowing it to display the real-time video data of the user on the terminal within the video conference interface, facilitating face-to-face communication between participants. Alternatively, even if the terminal does not support video data acquisition, users can still participate in the video conference, but they cannot display their own real-time video data within the terminal's interface. In this case, they can only view the real-time video feed from other terminals within the video conference interface. Furthermore, for terminals capable of video data acquisition, users can choose whether to enable the video acquisition function to determine whether to share their local real-time video feed during the video conference. When a video conference is triggered, the video acquisition function can be enabled by default for participants entering the video conference room, displaying the real-time video feed of each participant on each terminal within the video conference interface.

[0098] Joining a video conference can be triggered by either initiating or accepting the invitation. For example, the host can initiate a video conference from a first terminal to the corresponding second terminal of a participating user, such as by sending a video conference invitation. Alternatively, a user can receive a video conference invitation from the host via the first terminal on their second terminal; after accepting the invitation, the user can join the video conference. The triggering action for a video conference can be a single click, double click, or swipe within the trigger area.

[0099] Virtual side-by-side display refers to a virtual screen where the real-time video feeds of all participants in a video conference are displayed together within the same frame. This means that during the video conference, the surrounding environment of multiple participants is replaced with a specified image or video, essentially allowing participants to share a virtual background for real-time video display. For example... Figure 3 As shown, in the virtual side-by-side frame of a video conference, the real-time video feeds of the four conference participants are displayed side-by-side against a background of virtual landscape, creating an atmosphere of interactive meetings conducted amidst natural scenery. The virtual scene refers to the virtual background within the virtual side-by-side frame. This frame can include various virtual scenes to correspond to different video conference themes. Specifically, virtual scenes can include seminars, classrooms, and other settings. Setting different virtual scenes helps create an atmosphere that corresponds to the video conference theme, enhancing the realism of the video conference and thus improving its interactive efficiency. For example, when conducting course training via video conference, the virtual side-by-side frame could be a classroom setting, creating the atmosphere of a video conference conducted in a classroom. The virtual side-by-side frame can be displayed when the video conference is in side-by-side mode. For instance, if the conference host needs to display the real-time video feeds of each participant side-by-side, they can trigger the side-by-side mode, which then displays the virtual side-by-side frame.

[0100] The virtual scene includes multiple virtual locations to accommodate meeting participants. The distribution of these virtual locations can be flexibly configured according to the actual needs of the scenario. Furthermore, the virtual scene can also include other scene layout elements. For example, when the virtual scene is a virtual classroom, in addition to the virtual locations for meeting participants, it can also include decorations such as virtual lights and virtual banners. Meeting participants refer to users attending the video conference. Virtual locations are positions used to accommodate corresponding meeting participants, and can take various forms such as different small houses, sign boxes, or seats. The form of the virtual locations can be pre-set according to the needs of the virtual scene. For example, for a virtual classroom scene, the corresponding virtual location could be a virtual seat in the classroom; for a virtual seminar scene, the corresponding virtual location could be a virtual house dividing meeting participants into different groups. Virtual locations can accommodate corresponding meeting participants. For example, when a virtual location is a virtual seat, the image of the corresponding meeting participant in the real-time screen can occupy that virtual seat, thus indicating that the meeting participant is sitting in that seat. In a specific application, such as... Figure 4 As shown, the virtual scene includes multiple virtual locations, specifically virtual seats. These virtual seats can accommodate participants in a video conference, thus creating an atmosphere where participants are sitting in a meeting room and communicating.

[0101] Specifically, when a video conference needs to be initiated, participants can trigger it through their respective terminals. For example, the conference host can start the conference, and participants can join via their respective terminals. On each terminal's interface, a virtual frame of the video conference is displayed, with multiple virtual locations to accommodate participants. This virtual frame can be configured by the host initiating the frame-of-view mode or use a default setting. In practical applications, after a video conference is triggered between terminals, each terminal can capture video footage in real-time using its own video capture device, such as a camera connected to the terminal, and send the footage to the server. The server then aggregates the video feeds from all terminals and sends the aggregated footage back to each terminal for display. In implementation, when displaying the aggregated video feed, terminals can either display all participants' video feeds by default, or filter out the local feeds and display only the feeds of other participants.

[0102] Step 204: At at least two virtual locations in the virtual scene, display adaptive portraits corresponding to the portraits of meeting members captured in real-time images from at least two of the multiple terminals; wherein the size of the adaptive portrait displayed at each of the at least two virtual locations matches the size of the corresponding virtual location's location area.

[0103] The "real-time video feed of conference participants" refers to the live video captured by the terminal during a video conference. This feed includes the images of the participants captured by the terminal, such as the individual images of the participants themselves. The "adaptive human image" is a human image obtained by adaptively adjusting the human images in the real-time video feed. Specifically, adjustment parameters are determined based on the distribution of the human images within the video image captured by the terminal. These parameters are then used to adaptively adjust the human images in the real-time video feed. The size of the adaptive human image, obtained after adaptive adjustment, matches the size of the area containing its virtual location. This area can be a region covering the virtual location, specifically the bounding rectangle of the virtual location. This location area allows for the management and control of various virtual locations within the virtual scene. For each of the at least two virtual locations in the virtual scene, the size of the adaptive human image is matched with the size of the location area of ​​the corresponding virtual location. For example, the area ratio of the adaptive human image in the location area can be a preset ratio. That is, the adaptive human image displayed at the virtual location is obtained by adaptively adjusting the human image in the real-time screen of the meeting members according to the size of the location area of ​​the virtual location and the preset ratio.

[0104] Specifically, in the virtual frame-to-frame display on each terminal, at least two virtual locations within the virtual scene display adaptive portraits corresponding to the portraits of meeting members captured in real-time footage from at least two of the multiple terminals. The size of the adaptive portrait matches the size of the corresponding virtual location's area. This adjustment of the displayed adaptive portrait size based on the size of the area where each virtual location is located ensures harmony and unity between the size of the adaptive portrait displayed at each virtual location, enhancing the realism of the adaptive portrait's position and avoiding clutter in the real-time meeting member footage. In one embodiment, in addition to displaying the corresponding adaptive portraits at each virtual location, meeting member information corresponding to the adaptive portraits can also be displayed within the virtual locations. This includes information such as the meeting member's name, account, position, network status, etc., thus displaying rich meeting member information within the virtual frame-to-frame display, allowing meeting members to quickly and intuitively obtain this information.

[0105] In a specific application, such as Figure 5 As shown, the virtual scene of the virtual meeting room includes multiple virtual locations to accommodate meeting members. Each virtual location is divided into six rows. At each virtual location, an adaptive portrait of the meeting member is displayed, corresponding to the real-time image captured by the terminal. The size of the adaptive portrait corresponds to the size of the area where the virtual location is located. The size of the area where the virtual location is located is related to the row in front or behind it; the area where the virtual location is located in the front row is larger than the area where the virtual location is located in the back row. Correspondingly, the adaptive portrait displayed at the front row virtual location is larger than the adaptive portrait displayed at the back row virtual location, thus creating a realistic meeting room atmosphere. In a specific application, such as... Figure 6 As shown, in the virtual scene of the virtual group screen, the corresponding adaptive portrait is displayed in the virtual position, and the name of the meeting member corresponding to the adaptive portrait is further displayed in the virtual position, so that the meeting members can intuitively understand the information of each meeting member.

[0106] In the aforementioned video conferencing interaction method, when a video conference is triggered between multiple terminals, a virtual frame of the video conference is displayed. This virtual frame includes multiple virtual locations to accommodate conference members. At at least two virtual locations within the virtual scene, adaptive portraits corresponding to the real-time images of the conference members captured by at least two of the multiple terminals are displayed. The size of the displayed adaptive portraits matches the size of the corresponding virtual location's area. During the video conferencing interaction, adaptive portraits matching the size of the corresponding virtual location's area are displayed at the virtual locations within the virtual scene of the corresponding virtual frame of the video conference. This control over the size of the displayed adaptive portraits based on the location area of ​​the virtual locations within the virtual scene avoids cluttered portraits in the real-time images of conference members, enhances the realism of the video conference, and thus improves the interaction efficiency of the video conference.

[0107] In one embodiment, at at least two virtual locations in a virtual scene, an adaptive human image corresponding to the human image in the real-time video of a meeting member captured by at least two of the multiple terminals is displayed. This includes: determining the size of the human image area matching the corresponding virtual location based on the size of the location area in the virtual scene and a preset uniform proportion condition for human images; and at the virtual location in the virtual scene, displaying the target adaptive human image corresponding to the human image in the real-time video of a meeting member captured by any one of the multiple terminals according to the corresponding human image area size.

[0108] The location region can be the area covering a virtual location, specifically the bounding rectangle of that virtual location. This location region allows for the division and categorization of virtual locations within the virtual scene, enabling personalized settings for the content displayed at each location. The portrait region covers the adaptive portrait, such as the bounding rectangle of the adaptive portrait, specifically the smallest bounding rectangle of the adaptive portrait. The portrait region covers the entire area of ​​the adaptive portrait. The portrait region size refers to the size of the portrait region, specifically its area.

[0109] The portrait uniformity ratio refers to the ratio between the size of the adaptive portrait's corresponding portrait area displayed in a virtual location and the size of the location area where the virtual location is located. Specifically, it can be expressed as the area ratio. A higher portrait uniformity ratio indicates that the adaptive portrait occupies a larger area within the location area of ​​the virtual location. The portrait uniformity ratio condition refers to the conditions that must be met for the portrait uniformity ratio to be valid. This condition can be preset according to actual needs or set based on the virtual scene within the virtual frame. Different virtual scenes can correspond to different portrait uniformity ratio conditions, thus creating a corresponding realistic scene atmosphere according to the needs of different virtual scenes. For example, the portrait uniformity ratio condition can include a portrait uniformity ratio threshold, meaning that the portrait uniformity ratio between the portrait area and the location area corresponding to the virtual location meets the threshold, such as 80%, indicating that the area of ​​the adaptive portrait's portrait area reaches 80% within the location area corresponding to the virtual location. Furthermore, the portrait uniformity ratio condition can also include a preset portrait uniformity ratio range, such as [70%, 90%], meaning that the area of ​​the adaptive portrait's portrait area is between 70% and 90% within the location area corresponding to the virtual location.

[0110] The size of the portrait area can be determined based on the size of the virtual location's location area in the virtual scene and preset uniform portrait proportion conditions. For example, if the uniform portrait proportion condition includes 85% and the size of the virtual location's location area in the virtual scene is A, then the size of the portrait area matching the corresponding virtual location can be determined to be 0.85A, meaning that the area ratio of the adaptive portrait displayed in that virtual location to the area of ​​the virtual location's location is 85%. The target adaptive portrait refers to the adaptive portrait displayed at the corresponding virtual location in the virtual scene. Different meeting members' adaptive portraits are displayed at different virtual locations, meaning that different meeting members can display adaptive portraits with different portrait area sizes.

[0111] Specifically, after displaying each virtual location in the virtual scene, the terminal determines the size of the area where each virtual location is located. This can be done by querying the attribute information of each virtual location and determining the size of the area based on that information. The terminal obtains preset uniform portrait ratio conditions and, based on the size of the area where the virtual location is located and the uniform portrait ratio conditions, determines the size of the portrait area matched to each virtual location; that is, it determines the size of the area in each virtual location used to display the adaptive portrait. At the virtual location in the virtual scene, the terminal displays the target adaptive portrait corresponding to the portrait in the real-time image of the meeting members captured by any one of the multiple terminals, according to the determined portrait area size. The size of the portrait area of ​​the target adaptive portrait and the size of the area where the virtual location is located must meet the preset uniform portrait ratio conditions, such as the area ratio of the portrait area in the location area meeting the preset uniform portrait ratio threshold.

[0112] In this embodiment, the portrait area of ​​the adaptive portrait displayed at at least two virtual locations in the virtual scene meets the preset portrait uniformity ratio condition between the portrait area of ​​the corresponding virtual location and the location area of ​​the virtual location. The size of the portrait area corresponding to the adaptive portrait to be displayed can be determined based on the size of the location area and the portrait uniformity ratio condition. Specifically, in the virtual scene of the virtual co-frame screen on the terminal, the portrait area of ​​the adaptive portrait displayed at at least two virtual locations meets the preset portrait uniformity ratio threshold or falls within the preset portrait uniformity ratio range, thereby ensuring the adaptability of the adaptive portrait to the virtual location, forming a harmonious unity between the adaptive portrait and the virtual location, and creating a realistic scene atmosphere.

[0113] For example, in a virtual seminar setting with a small number of participants, the preset uniform portrait ratio can be set to a high 90% to create a realistic seminar atmosphere that fosters closer relationships among participants. Conversely, in a virtual classroom setting with a large number of participants, where the meeting is primarily led by the host, the preset uniform portrait ratio can be set to a flexible range [70%, 90%] to create a classroom atmosphere with a certain degree of distance. In practice, the uniform portrait ratio can be preset according to the virtual scene or flexibly set by authorized participants. For instance, the meeting host can customize the uniform portrait ratio to adjust the size relationship between the portrait area and the corresponding virtual location area as needed.

[0114] In a specific application, such as Figure 7As shown, the virtual scene within the virtual frame includes four virtual locations, specifically four virtual seats. The area occupied by each virtual seat is the same size, represented by a diagonal shaded area. On the first and third virtual seats, corresponding adaptive portraits are displayed. The portrait area of ​​the adaptive portrait in the third virtual seat is shown as the horizontal shaded area. The portrait area of ​​the adaptive portrait displayed in each virtual seat has a uniform proportion of 80% to the total portrait area of ​​its location. In other implementations, a uniform proportion condition for the portrait can be set for each virtual location in the virtual scene. This means the uniform proportion condition for the portrait at each virtual location can be the same or different, allowing for flexible control over the size of the adaptive portrait displayed at each virtual location.

[0115] In this embodiment, the terminal determines the size of the image area matching the corresponding virtual location based on the size of the location area in the virtual scene and the preset uniform proportion of human images. At the virtual location in the virtual scene, the terminal displays the target adaptive human image corresponding to the human image in the real-time screen of the meeting members captured by any one of the multiple terminals, according to the corresponding human image area size. This ensures the adaptability of the adaptive human image to the virtual location, forming a harmonious unity between the adaptive human image and the virtual location, creating a realistic scene atmosphere, avoiding clutter in the real-time screen of the meeting members, improving the realism of the video conference, and thus helping to improve the interactive efficiency of the video conference.

[0116] In one embodiment, at at least two virtual locations in a virtual scene, an adaptive human image corresponding to the human image in the real-time video of a meeting member captured by at least two of the multiple terminals is displayed. This includes: determining the distribution location of each virtual location in the virtual scene; determining the size of the human image area matching the corresponding virtual location according to the uniform proportion condition of the human image corresponding to each distribution location; and displaying the target adaptive human image corresponding to the human image in the real-time video of a meeting member captured by any one of the multiple terminals at the virtual location in the virtual scene, according to the corresponding human image area size.

[0117] The distribution location refers to the positional distribution information of various virtual locations in the virtual scene. Different virtual locations within the virtual scene correspond to different portrait uniformity ratio conditions. The portrait uniformity ratio condition is the condition that a portrait uniformity ratio must meet. It refers to the ratio between the size of the adaptive portrait area displayed in the virtual location and the size of the location area where the virtual location is located; specifically, it can be the area ratio. A higher portrait uniformity ratio value indicates that the adaptive portrait occupies a larger area within the location area of ​​the virtual location. The portrait uniformity ratio condition corresponds to the distribution location of each virtual location in the virtual scene. For example, the portrait uniformity ratio condition includes a portrait uniformity ratio threshold, meaning that the portrait uniformity ratio between the portrait area and the location area corresponding to the virtual location meets the threshold. When there are multiple rows of virtual locations in the virtual scene, virtual locations in the front row can be set with a higher portrait uniformity ratio threshold, while virtual locations in the back row correspond to a smaller threshold, thus creating a perspective atmosphere where objects appear larger when closer and smaller when farther away, further enhancing the realism of the virtual scene. In practical implementation, the correspondence between the distribution of each virtual location in the virtual scene and the uniform proportion condition of the human image can be preset according to actual needs. For example, the uniform proportion condition of the human image can be set based on the perspective principle, so that the adaptive human image displayed at the virtual location that is closer is larger than the adaptive human image displayed at the virtual location that is farther away, thereby realizing a perspective view of near objects appearing larger and far objects appearing smaller, so as to enhance the realism of the virtual scene.

[0118] The image region size refers to the size of the image region, specifically its area. The image region covers the adaptive image, such as the area of ​​the bounding rectangle corresponding to the adaptive image, or more specifically, the area of ​​the smallest bounding rectangle corresponding to the adaptive image. The image region can cover the entire range of the adaptive image. The image region size can be determined based on the uniform proportion condition of the image corresponding to the virtual location. Specifically, it can be determined based on the size of the area where the virtual location is located in the virtual scene, and the determined uniform proportion condition of the image. The target adaptive image refers to the adaptive image displayed at the virtual location in the virtual scene. Different meeting members' adaptive images are displayed at different virtual locations, meaning different meeting members can display adaptive images with different image region sizes.

[0119] Specifically, when the terminal displays adaptive portraits corresponding to the real-time images of meeting members at virtual locations, it determines the distribution of each virtual location within the virtual scene. This can be achieved by acquiring the attribute information of the virtual locations and determining their distribution based on this information. The terminal then determines a pre-defined uniform portrait ratio condition based on the distribution position of each virtual location. For example, it can use the mapping relationship between distribution position and uniform portrait ratio condition to determine the corresponding uniform portrait ratio condition. Finally, the terminal determines the size of the portrait area matching the corresponding virtual location based on this uniform portrait ratio condition. For instance, it can determine the size of the portrait area matching the virtual location based on the uniform portrait ratio condition and the location area where the virtual location is situated. This determined portrait area size is the size of the adaptive portrait area displayed at that virtual location. At a virtual location in a virtual scene, the terminal displays a target adaptive human image corresponding to the human image in the real-time image of a meeting member captured by any one of the multiple terminals, according to a determined human image area size. The size of the human image area of ​​the target adaptive human image and the size of the location area where the virtual location is located meet the preset uniform proportion conditions for human images, such as the area proportion of the human image area in the location area being within the preset uniform proportion range for human images.

[0120] In a specific application, such as Figure 8 As shown, the virtual positions in the virtual scene of the virtual co-op image are divided into two rows, front and back. The area occupied by each virtual position in both rows is the same size, which is the area marked by the dashed box. For the virtual positions in the front row, the threshold of the unified proportion of human portraits in the unified proportion condition is higher, that is, the human portrait area displayed at the front virtual positions is larger; while for the virtual positions in the back row, the threshold of the unified proportion of human portraits in the unified proportion condition is lower, that is, the human portrait area displayed at the back virtual positions is smaller, thus creating a perspective effect where the adaptive human portraits in the foreground are larger and the adaptive human portraits in the background are smaller.

[0121] In this embodiment, the terminal determines the size of the human image area matching the corresponding virtual location based on the uniform proportion of human images corresponding to the distribution of virtual locations in the virtual scene. At the virtual location in the virtual scene, the terminal displays the target adaptive human image corresponding to the human image in the real-time video of the meeting members captured by any one of the multiple terminals, according to the corresponding human image area size. This ensures the adaptability between the adaptive human image and the virtual locations distributed in different positions in the virtual scene, forming a harmonious unity between the adaptive human image and the virtual locations distributed in different positions, creating a realistic scene atmosphere, avoiding clutter in the real-time video of the meeting members, improving the realism of the video conference, and thus helping to improve the interactive efficiency of the video conference.

[0122] In one embodiment, the video conferencing interaction method further includes: displaying a virtual screen sharing area belonging to the virtual scene in the virtual co-op screen; and displaying the screen sharing content of the target terminal in the virtual screen sharing area in response to the target terminal triggering screen sharing among the multiple terminals joining the video conference.

[0123] The virtual screen sharing area is the area where screen sharing takes place. Screen sharing allows for content sharing during video conferences, enabling participants to interact and communicate based on the shared content, thus improving the efficiency of video conference interactions. The target terminal is the terminal that triggers the screen sharing event; specifically, it can be the terminal of the participant in the video conference whose screen needs to be shared. The screen sharing content is the content that the target terminal needs to share and display, which can include, but is not limited to, various forms of content such as text, tables, presentation files, audio and video data. The screen sharing content is displayed in the virtual screen sharing area within the virtual shared frame, so that all participants in the video conference can view it.

[0124] Specifically, in the virtual shared-screen view of a video conference, the terminal displays a virtual screen-sharing area belonging to the virtual scene. This virtual screen-sharing area corresponds to the virtual scene; different virtual scenes can have different forms of virtual screen-sharing areas. For example, in a virtual seminar scene, the virtual screen-sharing area could be a virtual electronic screen, while in a virtual classroom scene, it could be a virtual blackboard, allowing content sharing within a virtual blackboard. Furthermore, different virtual scenes have different scene layouts, and the virtual screen-sharing area can be laid out accordingly. For instance, different virtual scenes can display the corresponding virtual screen-sharing area in different locations, and the size of the displayed virtual screen-sharing area can also be set according to the virtual scene to ensure a realistic atmosphere. When a target terminal among the multiple terminals joining the video conference triggers screen sharing, the terminal can detect the corresponding screen-sharing event, indicating that the target terminal needs to share its screen. The terminal then displays the target terminal's screen-sharing content, such as a presentation file specified by the target terminal, within the virtual screen-sharing area of ​​the virtual scene.

[0125] In a specific application, such as Figure 9 As shown, a virtual screen sharing area belonging to the virtual scene is displayed in the virtual shared frame. When a meeting member of the target terminal triggers a screen sharing event through the screen sharing control, such as when a meeting member of the target terminal clicks the "Share Screen" screen sharing control, screen sharing is triggered, and the screen sharing content specified by the meeting member through the target terminal is displayed in the virtual screen sharing area. Specifically, it can be a presentation file.

[0126] In this embodiment, the terminal displays a virtual screen sharing area belonging to the virtual scene in the virtual frame-sharing screen, and displays the screen sharing content specified by the target terminal in the virtual screen sharing area, thereby realizing content sharing in the video conference, so that the conference members can interact and communicate based on the shared content, which is conducive to improving the interaction efficiency of the video conference.

[0127] In one embodiment, displaying a virtual screen sharing area belonging to the virtual scene in a virtual frame-to-frame view includes: displaying the virtual screen sharing area in the background area other than the virtual location in the virtual frame-to-frame view.

[0128] The background area refers to the area outside the virtual location within the virtual frame. By displaying the virtual screen sharing area in this background area, the problem of the adaptive human image displayed in the virtual location obscuring the content shared in the video conference can be avoided, thus preventing any impact on the content sharing effect. Specifically, when displaying the virtual screen sharing area in the virtual frame, the terminal can do so within the background area outside the virtual location.

[0129] Furthermore, the video conferencing interaction method also includes: when multiple terminals joining the video conference do not trigger screen sharing, displaying a virtual blank screen in the virtual screen sharing area; and in response to the target terminal triggering cancellation of screen sharing, displaying a virtual blank screen in the virtual screen sharing area.

[0130] The virtual blank screen is a blank virtual screen with no content displayed, serving as the background for the virtual scene. If none of the terminals joining the video conference trigger screen sharing (i.e., no screen sharing event occurs during the video conference), or the target terminal cancels screen sharing, a virtual blank screen is displayed within the virtual screen sharing area to serve as the background for the virtual scene. Specifically, when a terminal detects that multiple terminals joining the video conference have not triggered screen sharing, such as newly joined terminals that have not triggered screen sharing, the terminal displays a virtual blank screen in the virtual screen sharing area. Conversely, if the target terminal joining the video conference has already triggered screen sharing, the screen sharing content specified by the target terminal is displayed in the virtual screen sharing area. When the target terminal cancels screen sharing, indicating the end of screen sharing, the terminal displays a virtual blank screen in the virtual screen sharing area to terminate screen sharing.

[0131] In a specific application, such as Figure 10As shown, during the process of triggering screen sharing by the target terminal, the screen sharing content of the target terminal is displayed in the virtual screen sharing area. The virtual screen sharing area is located in the background area of ​​the virtual frame except for the virtual position. If the target terminal triggers the event to cancel screen sharing through the screen sharing control, such as when a meeting member of the target terminal clicks the "Cancel Sharing" screen sharing control, the screen sharing is canceled. If no other meeting member triggers screen sharing through the corresponding terminal, a virtual blank screen is displayed in the virtual screen sharing area, and the virtual blank screen is used as the background of the video conference.

[0132] In this embodiment, if no screen sharing event occurs during the video conference, or if the target terminal cancels screen sharing, the terminal displays a virtual blank screen in the virtual screen sharing area displayed in the background area other than the virtual position in the virtual frame view, as the background of the virtual scene. The background of the virtual scene can be fully utilized as needed to facilitate interaction and communication among meeting members, which is beneficial to improving the interaction efficiency of the video conference.

[0133] In one embodiment, the video conferencing interaction method further includes: responding to a first adaptive human image displayed at a first virtual location in a mobile virtual scene moving to a second virtual location, and at the second virtual location, displaying an adaptive human image corresponding to the human image in the real-time screen of the meeting member captured by the terminal corresponding to the first adaptive human image.

[0134] In this context, the first virtual location and the second virtual location are any different virtual locations within the virtual scene, and the first adaptive avatar is the adaptive avatar displayed at the first virtual location. Specifically, during a video conference, participants can update the virtual location of the adaptive avatar displayed in the virtual frame using their terminals. This can be done by participants with position adjustment permissions, such as the conference host, using their corresponding terminals. The terminal responds by moving the first adaptive avatar displayed at the first virtual location in the virtual scene to the second virtual location. Specifically, the conference host can move the first adaptive avatar displayed at the first virtual location to the second virtual location using their corresponding terminal, triggering an update to the avatar's display position. At the second virtual location, the terminal displays the adaptive avatar corresponding to the avatar in the real-time view of the participants captured by the terminal corresponding to the first adaptive avatar. In practical applications, the size of the adaptive avatar displayed at the second virtual location matches the size of the area where the second virtual location is located.

[0135] In a specific application, such as Figure 11As shown, in a video conference, participants can move their adaptive avatars, initially displayed in the sixth row, to the third row, where they will be displayed. The size of the avatar matches the size of the virtual area in the third row. In practice, position update permissions for each participant can be configured. For example, the conference host can update the display position of all participants' avatars via their terminal. They can also grant permissions to other participants, allowing those with permission to update the avatars within their assigned area. Conversely, participants without permission can only update their own avatars or cannot update the avatars within the virtual frame.

[0136] In this embodiment, when the virtual position corresponding to the adaptive portrait displayed in the virtual scene is updated, the terminal can display the corresponding adaptive portrait at the updated virtual position. This allows for flexible management and control of the display position of each adaptive portrait in the virtual scene, which helps ensure the compatibility between the adaptive portrait and the virtual position, forming a harmonious unity between the adaptive portrait and the virtual position, creating a realistic scene atmosphere, enhancing the realism of the video conference, and thus improving the interactive efficiency of the video conference.

[0137] In one embodiment, the video conferencing interaction method further includes: in response to a target terminal among multiple terminals being in an abnormal network state, displaying network abnormality prompt information about the meeting member corresponding to the adaptive portrait at the target virtual location of the adaptive portrait corresponding to the portrait in the real-time video of the meeting members collected by the target terminal.

[0138] The target terminal refers to the terminal among the multiple terminals joining the video conference that is in an abnormal network state, specifically a terminal with poor network signal or a disconnected network connection. The target virtual location is the virtual location of the adaptive human image corresponding to the real-time image of the conference member captured by the target terminal. The network anomaly prompt message is used to notify the conference member corresponding to the target terminal that a network anomaly has occurred, and they may not be able to receive conference information in the video conference or interact normally.

[0139] Specifically, the terminal or server can detect the network status of each terminal joining the video conference. When an abnormal network status is detected in a target terminal, indicating that the corresponding conference member has lost network connection to the video conference or the network connection is abnormal and unable to conduct normal video conference interaction, the terminal will display a network abnormality prompt message about the conference member corresponding to the adaptive avatar in the real-time video of the conference member captured by the target terminal. This provides timely notification of the abnormal network status of the target terminal. In practical applications, the display position of the network abnormality prompt message can be related to the adaptive avatar, both located at the corresponding virtual position. The network abnormality prompt message can also cover the corresponding adaptive avatar for prominent display. In addition, the visibility permission of the network abnormality prompt message can be set according to actual needs. For example, it can be set to be visible to all conference members in the video conference, or only to the conference host.

[0140] In a specific application, such as Figure 12 As shown, in the virtual scene of the virtual co-op screen, the terminal of the meeting member corresponding to the adaptive portrait located in the second row of the virtual position is in an abnormal network state. The terminal at this virtual position, covering the adaptive portrait, displays the prompt message "Network abnormal" to indicate the network status of the meeting member.

[0141] In this embodiment, by displaying network anomaly alerts at virtual locations to indicate abnormal network status of the terminals corresponding to those virtual locations, network anomalies in participating video conference terminals can be promptly detected, which helps ensure the normal operation of video conferences and maintains their interactive efficiency.

[0142] In one embodiment, the video conferencing interaction method further includes: in response to a target meeting member being in an abnormal participation state in the real-time video of the meeting members captured by the target terminal among multiple terminals, displaying an abnormal state prompt message about the target meeting member at the target virtual location of the adaptive human image corresponding to the human image in the real-time video of the meeting members captured by the target terminal.

[0143] In this context, the target terminal refers to the terminal among multiple terminals joining the video conference where the target participant is in an abnormal participation state in the real-time video feed. An abnormal participation state means that the participant is not in a normal participation state and may be unable to effectively communicate and interact during the video conference. Examples include participants being inattentive, fooling around, dozing off, playing games, or engaging in other behaviors unrelated to normal communication and interaction in a video conference. The participation state of a participant can be determined through behavioral analysis of the corresponding real-time video feed. For example, behavior recognition can be performed based on the real-time video feed to determine the participation state of the corresponding participant. If a participant is in an abnormal participation state, an abnormal status message will be displayed at the target virtual location corresponding to the adaptive human image in the real-time video feed of the participant captured by the target terminal. The target virtual location is the virtual location of the adaptive human image corresponding to the participant in the real-time video feed captured by the target terminal. The abnormal status message is used to inform the participant corresponding to the target terminal that they are not in a normal participation state and may not be able to know the meeting information in the video conference or interact normally.

[0144] Specifically, the terminal or server can detect the participation status of meeting members corresponding to each terminal joining the video conference. When an abnormal participation status is detected for a target meeting member corresponding to a target terminal, indicating that the target meeting member cannot interact normally in the video conference in a timely manner, the terminal displays an abnormal status prompt message for the target meeting member at the target virtual position of the adaptive avatar corresponding to the avatar in the real-time image captured by the target terminal. This provides timely notification of the abnormal participation status of the target meeting member. In practical applications, the display position of the abnormal status prompt message can be related to the adaptive avatar, both being located at the corresponding virtual position. The abnormal status prompt message can also be displayed around the corresponding adaptive avatar for prominent display. Furthermore, the abnormal status prompt message can take different forms, such as text, images, special effects, videos, etc. For example, for a target meeting member in an abnormal participation status, the terminal can add abnormal participation effects to the adaptive avatar corresponding to that member, such as displaying multiple question marks above the avatar's head, or adding text effects to the avatar, to promptly notify the target meeting member of their abnormal participation status. Furthermore, the visibility permissions of the abnormal status notification information can be set according to actual needs; for example, it can be set to be visible to all meeting members in the video conference, or only to the corresponding meeting host.

[0145] In a specific application, such as Figure 13As shown, in the virtual scene of the virtual co-op screen, the meeting member corresponding to the adaptive portrait in the third row of the virtual position is in an abnormal meeting state. The terminal adds a prompt message "Cloud Touring..." to the right of the adaptive portrait at this virtual position, thereby indicating the meeting member's meeting status.

[0146] In this embodiment, by displaying abnormal status prompts at virtual locations, the abnormal participation status of the meeting members corresponding to those virtual locations can be promptly notified. This helps ensure the normal operation of video conferences and maintains their interactive efficiency.

[0147] In one embodiment, the video conferencing interaction method further includes: displaying a video interface of the video conference in response to a triggering operation of joining the video conference; and displaying real-time images of the conference members captured by at least one of the multiple terminals in the video interface.

[0148] The "Join Video Conference" trigger operation refers to the action of joining a video conference. Specifically, this can be the action of the conference host starting the video conference, or the action of a conference participant joining the video conference. The video interface is used to display the real-time video feeds of the participating participants captured by each terminal in the video conference. Specifically, users can trigger the action of joining a video conference on their terminals, such as receiving a video conference invitation or actively requesting to join. After entering the video conference, the terminal displays the video conference interface, which shows the real-time video feeds of the participating participants captured by at least one of the multiple terminals.

[0149] Furthermore, in response to the triggering operation of a video conference between multiple terminals, a virtual side-by-side screen of the video conference is displayed, including: in response to the triggering of the side-by-side mode for the video conference, canceling the display of the real-time screen of the conference members in the video interface, and displaying the virtual side-by-side screen of the video conference in the video interface.

[0150] The "shared-frame mode" refers to a mode that displays the video feeds of all participants in a video conference within the same frame. Shared-frame display means that during a video conference, the surrounding environment of multiple participants is replaced with a specified image or video, meaning that participants share a virtual background for real-time video display. Specifically, when shared-frame mode is triggered in a video conference, such as when at least two terminals have their cameras turned on, or when the host of the video conference actively enables shared-frame mode, the real-time video feeds of the participants are not displayed on the video interface. Instead, a virtual shared-frame view of the video conference is displayed on the video interface, allowing the terminals to see the virtual scene of the video conference, which includes multiple virtual locations to accommodate the participants.

[0151] Furthermore, the video conferencing interaction method also includes: in response to the video conference meeting the conditions for ending the same-frame mode, canceling the display of the virtual same-frame screen in the video interface.

[0152] The "End Frame-on Mode" condition refers to the conditions under which the video conference's frame-on mode ends. Specifically, this could be the conference host disabling frame-on mode, or the number of conference participants not meeting the conditions for maintaining frame-on mode, such as fewer than two participants or fewer than two participants having their cameras turned on. The end-frame-on mode conditions can be set according to actual needs, such as based on time or location. For example, the condition can be considered met when a preset frame-on mode end time is reached; or it can be triggered when a preset location point is reached. Specifically, the terminal can detect whether the video conference meets the preset end-frame-on mode conditions. If the conditions are met, indicating that the frame-on mode needs to be ended, the terminal will cancel the display of the virtual frame-on screen in the video interface and instead display the real-time video feed of at least one of the multiple terminals in the video conference interface, thus safely exiting frame-on mode.

[0153] In this embodiment, when the shared-frame mode of a video conference is triggered, a virtual shared-frame view of the video conference is displayed on the corresponding video interface, allowing all conference members to be shown together in the virtual shared-frame view. When the conditions for ending the shared-frame mode are met, the terminal cancels the display of the virtual shared-frame view on the video interface. By enabling and disabling the shared-frame mode of the video conference, the conference format can be changed, enriching the interactive forms of video conferences and improving their interactive efficiency.

[0154] In one embodiment, the adaptive portrait displayed at the virtual location is obtained by adaptively adjusting the portraits of conference members in the real-time video feed captured by the terminal. This adaptive adjustment process can be performed by the server or the terminal. The adaptive adjustment process may include: acquiring real-time video feeds of conference members captured by each terminal joining the video conference; for each conference member's real-time video feed, determining the portrait parameters and face parameters of the conference member in the real-time video feed; performing portrait analysis on the conference member's portrait in the real-time video feed based on the portrait parameters and face parameters to obtain the adaptive adjustment parameters corresponding to the conference member's real-time video feed; and adaptively adjusting the conference member's portrait in the real-time video feed based on the adaptive adjustment parameters to obtain the corresponding adaptive portrait.

[0155] The "real-time video feed of conference participants" refers to the live video captured by the terminal during the video conference. This feed includes the images of the participants captured by the terminal, such as the individual portraits of the participants. Each terminal joining the video conference can activate its camera to capture video and obtain the real-time video feed of the participants. "Portrait parameters" refer to the parameters of the foreground area in the real-time video feed of the participants. These parameters can be those corresponding to the portraits of the participants, including but not limited to the width and height of the portrait rectangle and the number of pixels in the portrait area. "Face parameters" refer to the parameters corresponding to the faces of the participants in the real-time video feed, including but not limited to the width and height of the face rectangle and the number of pixels in the face area. "Adaptive adjustment parameters" are obtained by performing portrait analysis on the portraits of the participants in the real-time video feed based on the portrait and face parameters, and are used to adaptively adjust the portraits of the participants in the real-time video feed. The adaptive adjustment parameters may specifically include adjustment parameters for adjusting the width of the conference member's image in the live view, adjustment parameters for adjusting the height of the conference member's image in the live view, and scaling parameters for scaling the conference member's image in the live view.

[0156] Specifically, when the adaptive adjustment process is executed by the server, the server can acquire real-time images of conference members captured by each terminal joining the video conference. Each terminal captures its corresponding real-time image of a conference member, and the server can obtain the corresponding captured real-time image of a conference member from each terminal. For each conference member's real-time image, the server performs portrait segmentation and face segmentation to obtain the portrait parameters and face parameters of the conference member in the real-time image. In specific applications, image segmentation technology can be used to segment the real-time image of the conference members to determine the portrait parameters and face parameters of the conference members in the real-time image. Based on the obtained portrait parameters and face parameters, the server performs portrait analysis on the portraits of the conference members in the real-time image. Specifically, it can analyze the distribution of the portraits of the conference members in the real-time image and determine the adjustment parameters that need to be adjusted based on this distribution, that is, obtain the adaptive adjustment parameters corresponding to the real-time image of the conference members. The adaptive adjustment parameters include parameters for adjusting the portraits of meeting members in the live view across multiple dimensions, specifically including adjustment parameters in the width direction, height direction, and scaling parameters. After determining the adaptive adjustment parameters, the server adaptively adjusts the portraits of meeting members in the live view based on these parameters, obtaining an adaptive portrait of each meeting member. This adaptive portrait is then displayed at a virtual location within the virtual background.

[0157] In this embodiment, based on the portrait and face parameters of the meeting members in the real-time video feed collected by the terminal, the portraits of the meeting members in the real-time video feed are analyzed to determine the corresponding adaptive adjustment parameters. Based on these parameters, the portraits of the meeting members in the real-time video feed are adaptively adjusted to obtain the corresponding adaptive portraits. By using the adaptive adjustment parameters determined from the portrait and face parameters of the meeting members in the real-time video feed, the size of the resulting adaptive portrait matches the size of the corresponding virtual location area. This avoids cluttered portraits in the real-time video feed, enhances the realism of the video conference, and thus improves the interactive efficiency of the video conference.

[0158] In one embodiment, based on portrait parameters and face parameters, portrait analysis is performed on the portraits of meeting members in the real-time image of the meeting members to obtain adaptive adjustment parameters corresponding to the real-time image of the meeting members. This includes: when the real-time image of the meeting members is determined to be a valid portrait image based on the portrait parameters and face parameters, determining the height offset and width offset based on the portrait parameters; determining the scaling factor based on the portrait parameters, face parameters, and height offset; and using the height offset, width offset, and scaling factor as adaptive adjustment parameters corresponding to the real-time image of the meeting members.

[0159] The real-time video feed of meeting members indicates that their images are valid and can be displayed in the virtual frame. Specifically, the proportions of the faces in the real-time video feed are determined based on the portrait and face parameters. When both proportions exceed their respective thresholds, the real-time video feed is considered valid. Adaptive adjustments can be made based on the real-time video feed, and the resulting adaptive images are displayed at virtual positions within the video conferencing virtual scene. The height offset is the adjustment amount needed to adjust the height of the images in the real-time video feed, the width offset is the adjustment amount needed to adjust the width of the images in the real-time video feed, and the scaling factor is the scaling parameter required after adjusting the height and width of the images in the real-time video feed; specifically, it can be a scaling factor.

[0160] Specifically, when the server performs facial image analysis on the real-time images of meeting members, it can first determine whether the real-time images of meeting members are valid facial images. Specifically, the server can determine the area corresponding to the face and the area corresponding to the person in the real-time images of meeting members based on the facial image parameters and the face parameters, respectively. Based on the proportion of the area corresponding to the person in the real-time images of meeting members and the proportion of the area corresponding to the face in the area corresponding to the person, it can be determined whether the real-time images of meeting members capture enough people and enough faces. If so, then the real-time images of meeting members can be determined to be valid facial images. In practical implementation, the server can determine the number of portrait pixels and face pixels in the real-time image of the meeting members based on portrait parameters and face parameters, respectively. The server compares the number of portrait pixels with the total number of pixels in the real-time image of the meeting members to determine whether enough portrait pixels have been captured in the real-time image of the meeting members. The server compares the number of face pixels with the number of portrait pixels to determine whether the portrait in the real-time image of the meeting members includes enough face pixels. If the real-time image of the meeting members includes enough portrait pixels and the portrait in the real-time image of the meeting members also includes enough face pixels, then the real-time image of the meeting members is determined to be a valid portrait image and can be displayed at a virtual location in the virtual scene.

[0161] When the live feed of meeting members is a valid image, the server determines the height and width offsets based on the face parameters. Specifically, it determines the initial width and height offsets based on the width and height of the face mask rectangle in the face parameters, as well as the face mask image corresponding to the live feed of the meeting members. The initial width and height offsets are then corrected using previous historical images corresponding to the live feed of the meeting members to obtain the height and width offsets. Face parameters include the face width and face height of the face mask rectangle; image parameters include the image width and face width of the face mask rectangle. The server determines the face scaling factor based on the face width, face height, and height offset, and obtains the scaling factor for the image in the live feed of the meeting members based on the face scaling factor, image width, face width, and height offset. The server uses the height offset, width offset, and scaling factor as adaptive adjustment parameters for the live feed of the meeting members, and adaptively adjusts the images of the meeting members in the live feed using these parameters to obtain the corresponding adaptive images.

[0162] In this embodiment, when determining that the real-time images of meeting members are valid images based on portrait and face parameters, height and width offsets are determined according to the portrait parameters. A scaling factor is then determined based on the portrait, face, and height offsets. The adaptive adjustment parameters corresponding to the real-time images of meeting members include the obtained height offset, width offset, and scaling factor. By judging image validity using portrait and face parameters and further determining adaptive adjustment parameters such as height offset, width offset, and scaling factor, it is ensured that the size of the adaptive image processed based on the adaptive adjustment parameters matches the size of the corresponding virtual location's location area. This avoids cluttered images in the real-time images of meeting members, enhances the realism of the video conference, and thus improves the interactive efficiency of the video conference.

[0163] In one embodiment, the portrait parameters include the portrait mask rectangle parameters in the portrait mask image corresponding to the real-time image of the meeting member; determining the height offset and width offset based on the portrait parameters includes: determining the initial width offset and initial height offset based on the portrait mask rectangle parameters and the size of the portrait mask image; calculating the reference width offset corresponding to the real-time image of the meeting member based on the historical width offset corresponding to the previous image of the real-time image of the meeting member; correcting the initial width offset using the reference width offset to obtain the width offset corresponding to the real-time image of the meeting member; calculating the reference height offset corresponding to the real-time image of the meeting member based on the historical height offset corresponding to the previous image of the real-time image of the meeting member; and correcting the initial height offset using the reference height offset to obtain the height offset corresponding to the real-time image of the meeting member.

[0164] The portrait parameters include the portrait mask rectangle parameters in the portrait mask image corresponding to the real-time image of the meeting members. The portrait mask image is the segmentation result obtained by image segmentation technology on the real-time image of the meeting members. The portrait mask image can be used for foreground and background segmentation, thus determining the portrait region based on the portrait mask image. The pixel value of each pixel in the portrait mask image represents the probability that the corresponding pixel belongs to the foreground, i.e., to the portrait pixel. The portrait mask rectangle is the bounding rectangle of the portrait in the portrait mask image, specifically it can be the minimum bounding rectangle. The portrait mask rectangle parameters are the parameters of the portrait mask rectangle, specifically including the portrait width and portrait height of the portrait mask rectangle.

[0165] The preceding frame is the frame that appears before the live feed of the meeting members in chronological order. The historical width offset is the width offset corresponding to the preceding frame, and the reference width offset is a reference offset determined based on the historical width offset. This reference offset is used to correct the initial width offset of the live feed of the meeting members, ensuring that the width offset of the live feed of the meeting members remains consistent with the historical width offset corresponding to the preceding frame within a certain range. In specific implementation, if the preceding frame is the frame preceding the live feed of the meeting members, then the historical width offset is the width offset of the previous frame, and the reference width offset can be directly the width offset of the previous frame. If the preceding frame is the frame preceding the live feed of the meeting members, such as the first eight frames, then the historical width offset is the width offset corresponding to each of the preceding frames, and the reference width offset can be determined based on the average of the width offsets corresponding to each of the preceding frames, such as the arithmetic average or weighted average of the width offsets corresponding to each of the preceding frames. During weighted processing, the width offset of the preceding frame that is strongly related to the current meeting member's live view in time can be assigned a higher weight, while the width offset of the preceding frame that is weakly related to the current meeting member's live view in time can be assigned a lower weight, so that the meeting member's live view can be kept consistent with the preceding frame that is strongly related in time.

[0166] The historical height offset is the height offset corresponding to the previous frame, while the reference height offset is a reference offset determined based on the historical height offset. These are used to correct the initial height offset of the meeting members' live feed, ensuring that the height offset of the meeting members' live feed remains consistent with the historical height offset corresponding to the previous frame within a certain range. In specific implementations, if the previous frame is the frame preceding the meeting members' live feed, the historical height offset is the height offset of that previous frame, and the reference height offset can be directly the height offset of that previous frame. If the previous frame consists of frames preceding the meeting members' live feed, such as the first eight frames, the historical height offset is the height offset corresponding to each of those frames, and the reference height offset can be determined based on the average of the height offsets corresponding to each of those frames. For example, it can be obtained by using the arithmetic average or weighted average of the height offsets corresponding to each of the previous frames. During weighted processing, the height offset of the preceding frame that is strongly related to the current meeting member's live view in time can be assigned a higher weight, while the height offset of the preceding frame that is weakly related to the current meeting member's live view in time can be assigned a lower weight, so that the meeting member's live view can be kept consistent with the preceding frame that is strongly related in time.

[0167] Specifically, the portrait parameters include the portrait mask rectangle parameters in the portrait mask image corresponding to the real-time image of the meeting member. When determining the height and width offsets, the server determines the initial width and initial height offsets based on the portrait mask rectangle parameters and the dimensions of the portrait mask image. For example, the initial width offset can be determined based on the portrait width and the width of the portrait mask image in the portrait mask rectangle parameters; the initial height offset can be determined based on the portrait height and the height of the portrait mask image in the portrait mask rectangle parameters. The server obtains the historical width and historical height offsets corresponding to the previous image of the meeting member's real-time image, and calculates the reference width and reference height offsets based on the historical width and historical height offsets. For example, the reference width and reference height offsets can be calculated by averaging the historical width and historical height offsets. The server corrects the initial width offset using a reference width offset to obtain the width offset corresponding to the real-time image of the meeting members, and corrects the initial height offset using a reference height offset to obtain the height offset corresponding to the real-time image of the meeting members. This ensures that the real-time image of the meeting members corresponds to the image in front, reduces image and face jitter, and thus ensures the accuracy of adaptive adjustment parameters.

[0168] In this embodiment, the initial width offset and initial height offset are determined based on the portrait mask rectangle parameters and the size of the portrait mask image. The initial width offset and initial height offset are then corrected using the historical width offset and historical height offset corresponding to the previous image to obtain the width offset and height offset. By referring to the historical offset corresponding to the previous image to correct the width offset and height offset of the real-time image of the meeting members, it can be ensured that the real-time image of the meeting members corresponds to the previous image, reducing portrait and face jitter, thereby ensuring the accuracy of adaptive adjustment parameters.

[0169] In one embodiment, the reference width offset is the average of the historical width offsets corresponding to the previous frame of the real-time view of the meeting members; the initial width offset is corrected using the reference width offset to obtain the width offset corresponding to the real-time view of the meeting members, including: determining the difference between the initial width offset and the average of the historical width offsets; when the difference is greater than a first preset threshold, the initial width offset is weighted and corrected using the historical width offset corresponding to the previous frame of the real-time view of the meeting members to obtain the corrected width offset; the corrected width offset is then subjected to range limitation processing to obtain the width offset corresponding to the real-time view of the meeting members.

[0170] The reference width offset is the average of the historical width offsets corresponding to the preceding frames of the meeting member's live view. Specifically, it can be the arithmetic mean or weighted average of the historical width offsets corresponding to the preceding frames. During weighted averaging, the width offsets corresponding to preceding frames with strong temporal correlation to the current meeting member's live view can be assigned a higher weight, while the width offsets corresponding to preceding frames with weak temporal correlation can be assigned a lower weight, thus ensuring consistency between the meeting member's live view and the preceding frames with strong temporal correlation. The first preset threshold can be flexibly set according to actual needs to determine whether the width difference between the meeting member's live view and the preceding frame is too large. If the difference is too large, the initial width offset is corrected by weighting it with the historical width offset corresponding to the preceding frame of the meeting member's live view, resulting in a corrected width offset. The corrected width offset is obtained by weighting the initial width offset with the historical width offset corresponding to the previous frame of the meeting members' live view. The weighting distribution can be preset; for example, based on experience, the weight ratio between the historical width offset corresponding to the previous frame of the meeting members' live view and the initial width offset can be set to 9:1. Alternatively, the corrected width offset can also be obtained by weighting the initial width offset with the average of the historical width offsets. Range limitation refers to restricting the value range of the corrected width offset to within a preset value interval.

[0171] Specifically, the reference width offset is the average of the historical width offsets corresponding to the preceding frames of the real-time images of conference members. When correcting the initial width offset of the real-time images of conference members, the server can determine the difference between the initial width offset and the reference width offset, that is, the difference between the initial width offset and the average of the historical width offsets. Specifically, the initial width offset and the reference width offset can be numerically compared to determine the difference between them. The server queries a preset first threshold. If the difference between the determined initial width offset and the reference width offset is greater than the first preset threshold, it indicates that the width of the real-time images of conference members differs significantly from the width of the preceding frames, which may cause jitter and reduce the reliability of the real-time images of conference members. In this case, the server performs weighted correction on the initial width offset using historical width offsets. Specifically, the initial width offset can be weighted and corrected using the historical width offsets corresponding to the preceding frames of the real-time images of conference members to obtain the corrected width offset. The server further performs range limitation processing on the correction width offset to limit the value of the correction width offset to a preset range, thereby obtaining the width offset corresponding to the real-time image of the meeting member. The width offset is used to adaptively adjust the width of the portrait in the real-time image of the meeting member.

[0172] In this embodiment, the reference width offset is the average of the historical width offsets corresponding to the previous frame of the real-time video of the meeting member. When the difference between the initial width offset and the reference width offset is greater than a first preset threshold, indicating a large difference between the width of the real-time video of the meeting member and the width of the previous frame, the initial width offset is weighted and corrected using the historical width offset corresponding to the previous frame of the real-time video of the meeting member. The corrected width offset obtained by the weighted correction is then range-limited to obtain the width offset corresponding to the real-time video of the meeting member. The width of the portrait in the real-time video of the meeting member is adaptively adjusted using the width offset, so that the width of the adaptive portrait matches the width of the corresponding virtual position area. This avoids clutter in the portrait of the meeting member in the real-time video, enhances the realism of the video conference, and thus improves the interactive efficiency of the video conference.

[0173] In one embodiment, the reference height offset is the average of the historical height offsets corresponding to the previous frame of the real-time view of the meeting members; the initial height offset is corrected using the reference height offset to obtain the height offset corresponding to the real-time view of the meeting members, including: determining the difference between the initial height offset and the average of the historical height offsets; when the difference is greater than a second preset threshold, the initial height offset is weighted and corrected using the historical height offset corresponding to the previous frame of the real-time view of the meeting members to obtain the corrected height offset; and the corrected height offset is range-limited to obtain the height offset corresponding to the real-time view of the meeting members.

[0174] The reference height offset is the average of the historical height offsets corresponding to the preceding frames of the meeting members' live feeds. Specifically, it can be the arithmetic mean or weighted average of the historical height offsets corresponding to the preceding frames. During weighted averaging, the height offsets corresponding to preceding frames with strong temporal correlation to the current meeting members' live feeds can be assigned a higher weight, while the height offsets corresponding to preceding frames with weak temporal correlation can be assigned a lower weight. This ensures that the meeting members' live feeds are consistent with the preceding frames with strong temporal correlation. The first preset threshold can be flexibly set according to actual needs to determine whether the height difference between the meeting members' live feeds and the preceding frames is too large. If the difference is too large, the initial height offset is corrected by weighting the historical height offsets, specifically the historical height offsets corresponding to the preceding frames of the meeting members' live feeds, to obtain the corrected height offset. The corrected height offset is obtained by weighting the initial height offset with the historical height offset corresponding to the previous frame of the meeting participants' real-time view. The weighting distribution during the weighting correction can be preset, for example, by setting the weight ratio between the reference height offset and the initial height offset to 9:1 based on experience. Alternatively, the corrected height offset can also be obtained by weighting the initial height offset with the average of the historical height offsets. Range limitation refers to restricting the value range of the corrected height offset to within a preset value interval.

[0175] Specifically, the reference height offset is the average of the historical height offsets corresponding to the previous frame of the meeting member's live view. When correcting the initial height offset of the meeting member's live view, the server can determine the difference between the initial height offset and the reference height offset, that is, the difference between the initial height offset and the average of the historical height offsets. Specifically, the initial height offset and the reference height offset can be numerically compared to determine the difference between them. The server queries a preset first threshold. If the difference between the determined initial height offset and the reference height offset is greater than the first preset threshold, it indicates that the height of the meeting member's live view is significantly different from the height of the previous frame, which may cause jitter and reduce the reliability of the meeting member's live view. In this case, the server uses the historical height offset corresponding to the previous frame of the meeting member's live view to perform weighted correction on the initial height offset to obtain the corrected height offset. The server further performs range restriction processing on the corrected height offset to limit the value of the corrected height offset to a preset value range, thus obtaining the height offset corresponding to the meeting member's live view. The height offset is used to adaptively adjust the height of the person's image in the meeting member's live view.

[0176] In this embodiment, the reference height offset is the average of the historical height offsets corresponding to the previous frame of the real-time view of the meeting member. When the difference between the initial height offset and the reference height offset is greater than a first preset threshold, indicating a large difference between the height of the real-time view of the meeting member and the height of the previous frame, the initial height offset is weighted and corrected by the historical height offset corresponding to the previous frame of the real-time view of the meeting member. The corrected height offset obtained by the weighted correction is then range-limited to obtain the height offset corresponding to the real-time view of the meeting member. The height of the portrait in the real-time view of the meeting member is adaptively adjusted by the height offset, so that the height of the adaptive portrait matches the height of the corresponding virtual location area. This avoids clutter in the portrait of the meeting member in the real-time view, improves the realism of the video conference, and thus improves the interactive efficiency of the video conference.

[0177] In one embodiment, the step of determining the initial width offset includes: determining the image center point of the portrait mask image based on the size of the portrait mask image; determining the portrait center point of the portrait mask rectangle based on the portrait mask rectangle parameters and the size of the portrait mask image; determining the horizontal position difference between the image center point and the portrait center point; and determining the initial width offset based on the horizontal position difference and the image width of the portrait mask image.

[0178] The portrait mask image is a segmentation result obtained by segmenting the real-time image of meeting members using image segmentation technology. The portrait mask image can be used for foreground and background segmentation, thus determining the portrait region. The image center point refers to the center point of the portrait mask image, which can be the geometric center point. The size of the portrait mask image can include its width and height. The portrait mask rectangle is the bounding rectangle of the portrait in the portrait mask image, specifically the minimum bounding rectangle. The portrait mask rectangle parameters are the parameters of the portrait mask rectangle, specifically the portrait width and height. The portrait center point is the center of the portrait mask rectangle, specifically the geometric center point.

[0179] Specifically, the server determines the image region corresponding to the portrait mask image based on its dimensions, such as the image width and height, and further determines the geometric center of this image region as the image center point of the portrait mask image. The server determines the portrait center point of the portrait mask rectangle based on its parameters and the image dimensions; specifically, the portrait center point can be obtained from the geometric center of the portrait mask rectangle. The server calculates the horizontal position difference between the image center point and the portrait center point; specifically, the horizontal position difference can be obtained by subtracting the horizontal positions of the image center point and the portrait center point. Based on the obtained horizontal position difference and the image width of the portrait mask image, the server obtains an initial width offset; specifically, the initial width offset can be obtained by the ratio between the horizontal position difference and the image width of the portrait mask image.

[0180] In this embodiment, an initial width offset is determined based on the horizontal position difference between the center point of the image of the portrait mask image and the center point of the portrait of the portrait mask rectangle, as well as the image width of the portrait mask image. The initial width offset can be used to determine adaptive adjustment parameters, so as to adaptively adjust the portraits in the real-time screen of the meeting members through the adaptive adjustment parameters, and obtain an adaptive portrait for display in a virtual location.

[0181] In one embodiment, the step of determining the initial height offset includes: determining the top and bottom positions of the portrait mask rectangle in the image height direction based on the portrait mask rectangle parameters; determining a first initial height offset based on the top position and the image height of the portrait mask image; determining a second initial height offset based on the bottom position and the image height of the portrait mask image; and obtaining the initial height offset based on the first and second initial height offsets.

[0182] Here, the portrait mask rectangle is the bounding rectangle of the portrait in the portrait mask image, specifically the minimum bounding rectangle. The portrait mask rectangle parameters are the parameters of the portrait mask rectangle, specifically including the portrait width and height. The top position refers to the highest point of the portrait mask rectangle in the image's height direction, and the bottom position refers to the lowest point of the portrait mask rectangle in the image's height direction.

[0183] Specifically, the server determines the top and bottom positions of the portrait mask rectangle in the image height direction based on the portrait mask rectangle parameters. Further, the server determines a first initial height offset based on the top position and the image height of the portrait mask image; specifically, the server can determine the first initial height offset based on the ratio between the top position and the image height of the portrait mask image. The server determines a second initial height offset based on the bottom position and the image height of the portrait mask image; specifically, the server can determine the second initial height offset based on the ratio between the bottom position and the image height of the portrait mask image. The server obtains an initial height offset based on the obtained first and second initial height offsets, which may include both the first and second initial height offsets.

[0184] In this embodiment, a first initial height offset and a second initial height offset are determined based on the top and bottom positions of the portrait mask rectangle in the image height direction, and the image height of the portrait mask image. An initial height offset is then obtained based on these two initial height offsets. The initial height offset can be used to determine adaptive adjustment parameters to adaptively adjust the portraits of meeting members in the real-time frame, resulting in an adaptive portrait for display in a virtual location.

[0185] In one embodiment, the portrait parameters include portrait mask rectangle parameters corresponding to the real-time images of the meeting members; the face parameters include face mask rectangle parameters corresponding to the real-time images of the meeting members; determining a scaling factor based on the portrait parameters, face parameters, and height offset includes: determining a face scaling factor based on the face mask rectangle parameters and height offset; determining a portrait scaling factor based on the portrait mask rectangle parameters and height offset parameters; and fusing the face scaling factor and the portrait scaling factor to obtain a scaling factor.

[0186] The portrait parameters include the portrait mask rectangle parameters corresponding to the real-time images of the meeting members, and the face parameters include the face mask rectangle parameters corresponding to the real-time images of the meeting members. The face scaling factor is the scaling factor required for the face of the portrait during the adaptive adjustment process; the portrait scaling factor is the scaling factor required for the overall portrait during the adaptive adjustment process.

[0187] Specifically, when determining the scaling factor, the server determines the face scaling factor based on the face mask rectangle parameters and height offset, and then scales the faces of the participants in the real-time video feed using this scaling factor. In practice, the face mask rectangle parameters can include the face width and face height. The server can determine the face width scaling factor based on the face width and height offset, and the face height scaling factor based on the face height and height offset. The server then determines the final face scaling factor based on both the face width and height scaling factors; specifically, the smaller value of the face width and height scaling factors can be used as the face scaling factor. Finally, the server determines the overall face scaling factor based on the face mask rectangle parameters and height offset parameters, and then scales the entire face in the real-time video feed using this scaling factor. In practical implementation, the portrait mask rectangle parameters can include the portrait width and portrait height of the portrait mask rectangle. The server can determine the portrait width scaling factor based on the portrait width and height offset, and determine the portrait height scaling factor based on the portrait height and height offset. The server determines the portrait scaling factor based on the portrait width scaling factor and the portrait height scaling factor. Specifically, the average of the portrait width scaling factor and the portrait height scaling factor can be used as the portrait scaling factor.

[0188] After obtaining the face scaling factor and the portrait scaling factor, the server merges them. Specifically, the face scaling factor and the portrait scaling factor can be weighted and merged to obtain the scaling factor. When merging the face scaling factor and the portrait scaling factor, the weights of each factor can be preset, such as a 1:3 ratio.

[0189] In this embodiment, the face scaling factor is determined based on the face mask rectangle parameters and height offset, and the portrait scaling factor is determined based on the portrait mask rectangle parameters and height offset parameters. The scaling factor is obtained by fusing the face scaling factor and the portrait scaling factor. The scaling factor is used to determine the adaptive adjustment parameters, so as to adaptively adjust the portraits in the real-time screen of the meeting members, and obtain the adaptive portraits for display in the virtual location.

[0190] In one embodiment, the portrait parameter includes the number of portrait pixels in the portrait mask image corresponding to the real-time image of the meeting member; the face parameter includes the number of face pixels in the portrait mask image; the method further includes: when it is determined based on the number of portrait pixels or the number of face pixels that the real-time image of the meeting member does not belong to a valid portrait image, obtaining the height offset and width offset of the previous image corresponding to the real-time image of the meeting member; and determining the height offset and width offset of the previous image as the height offset and width offset of the real-time image of the meeting member.

[0191] The image parameters include the number of pixels in the image mask corresponding to the real-time view of the meeting member, and the face parameters include the number of pixels in the face mask image. If the real-time view of the meeting member is a valid image, it means that the image in the real-time view of the meeting member is valid and can be displayed in the virtual frame. If the real-time view of the meeting member is not a valid image, the image in the real-time view of the meeting member is invalid and it is not advisable to display the real-time view of the meeting member in the virtual frame. In this case, the offset parameter of the previous image corresponding to the real-time view of the meeting member is used as the offset parameter of the real-time view of the meeting member, thereby ensuring that the adaptive image can be smoothly displayed in the virtual frame.

[0192] Specifically, the server determines the validity of the real-time images of meeting members based on the number of portrait pixels and face pixels in the corresponding portrait mask image. If, based on the number of portrait pixels or face pixels, the real-time image of a meeting member is determined not to be a valid image, then if the number of portrait pixels is too low or the proportion is too low, the real-time image of the meeting member is determined not to be a valid image; similarly, if the number of face pixels is too low or the proportion is too low, the real-time image of the meeting member is determined not to be a valid image. For real-time images of meeting members that are not valid images, the server determines the previous image corresponding to the real-time image of the meeting member and obtains the height offset and width offset of that previous image. The server directly uses the height offset and width offset of the previous image corresponding to the real-time image of the meeting member as the height offset and width offset of the real-time image of the meeting member, thereby ensuring a smooth change in height and width offsets between different images.

[0193] In this embodiment, for real-time images of meeting members that are determined not to be valid images based on the number of portrait pixels or face pixels, the server uses the height and width offsets of the previous image corresponding to the real-time image of the meeting member as the height and width offsets of the real-time image of the meeting member. This ensures smooth changes in height and width offsets between different images when the real-time image of the meeting member is invalid, so as to achieve an adaptive image that can be smoothly displayed in the virtual frame.

[0194] In one embodiment, such as Figure 14 As shown, the foreground parameters in the real-time view of the meeting members are determined, that is, the portrait parameters and face parameters of the meeting members in the real-time view of the meeting members are determined, including:

[0195] Step 1402: Perform image segmentation processing on the real-time images of the meeting members to obtain the original portrait mask images and original face parameters corresponding to the real-time images of the meeting members.

[0196] Image segmentation refers to the technique and process of dividing an image into several specific regions with unique properties and identifying targets of interest. Specifically, it can include threshold-based segmentation methods, region-based segmentation methods, edge-based segmentation methods, and segmentation methods based on specific theories. Image segmentation is performed on the real-time video feed of meeting participants to determine the foreground region, i.e., the human face region, within the live video feed. The original human face mask image is obtained based on the image segmentation processing of the real-time video feed of meeting participants. The size of the original human face mask image is the same as that of the live video feed of meeting participants. The pixel value of each pixel in the original human face mask image represents the probability that the corresponding pixel belongs to the foreground, i.e., to a human face pixel. The original face parameters are obtained by analyzing the face region in the real-time video feed of meeting participants, and specifically include the number of face pixels in the original image and the original face rectangle.

[0197] Specifically, after the server obtains the real-time images of the conference members captured by each terminal that has joined the video conference, the server performs image segmentation processing on each real-time image of the conference member to obtain the original portrait mask image and original face parameters corresponding to the real-time image of the conference member.

[0198] Step 1404: Perform portrait analysis on the portrait mask image obtained by scaling the original portrait mask image according to the scaling ratio to obtain the portrait parameters corresponding to the portraits in the real-time screen of the meeting members.

[0199] The portrait mask image is a scaled image obtained by scaling the original portrait mask image according to a preset scaling ratio. Scaling the original portrait mask image reduces the amount of data required for portrait analysis and processing, improving efficiency and ensuring the interactive efficiency of video conferencing. Portrait parameters are obtained based on portrait analysis of the portrait mask image and may include parameters such as the portrait mask rectangle parameters and the number of portrait pixels.

[0200] Specifically, the server scales the original portrait mask image according to a preset scaling ratio to obtain a portrait mask image, and performs portrait analysis on the portrait mask image to obtain the portrait parameters corresponding to the portraits in the real-time screen of the meeting members. In practical implementation, the server can sharpen the portrait mask image and determine the portrait parameters based on the sharpened portrait mask image, such as determining the width and height of the portrait mask rectangle in the sharpened portrait mask image, the number of portrait pixels, etc.

[0201] Step 1406: Scale the face parameters of the original image according to the scaling ratio to obtain the face parameters corresponding to the faces in the real-time image of the meeting members.

[0202] The face parameters are obtained by scaling the original image's face parameters according to a predetermined scaling ratio. This scaling ratio is the same as the scaling ratio of the original portrait mask image to the portrait mask image, thus ensuring the correspondence between the face parameters and the portrait parameters. Specifically, the server scales the original image's face parameters according to the scaling ratio of the original portrait mask image to obtain the face parameters corresponding to the faces in the real-time image of the meeting members. The face parameters can include various face-related parameters such as the face mask rectangle parameters and the number of face pixels in the portrait mask image.

[0203] In this embodiment, by performing image segmentation processing on the real-time images of meeting members, the original portrait mask image and the original face parameters are obtained. After scaling the original portrait mask image and the original face parameters according to a preset scaling ratio, the portrait parameters and face parameters are obtained. This can reduce the amount of data for portrait analysis and processing, which is conducive to improving the efficiency of portrait analysis and processing, thereby ensuring the interactive efficiency of video conferencing.

[0204] In one embodiment, adaptive adjustment of the portraits of meeting members in the real-time frame of the meeting members based on adaptive adjustment parameters to obtain corresponding adaptive portraits includes: transforming the adaptive adjustment parameters based on the scaling ratio to obtain transformed adaptive adjustment parameters; performing adaptive portrait adjustment on the real-time frame of the meeting members and the original portrait mask image corresponding to the real-time frame of the meeting members using the transformed adaptive adjustment parameters to obtain an adjusted portrait image and a corresponding adjusted portrait mask image; compositing the adjusted portrait image and the adjusted portrait mask image to obtain a composite image; and scaling the composite image according to the size of the area where the virtual location is located to obtain the adaptive portrait displayed at the virtual location.

[0205] The scaling ratio is the same as the scaling ratio of the original portrait mask image to the portrait mask image. The adaptive adjustment parameters are obtained through portrait analysis based on the scaled portrait mask image. Since the size of the portrait mask image does not correspond to the size of the real-time meeting member image, the obtained adaptive adjustment parameters need to be transformed according to the scaling ratio. This transformed adaptive adjustment parameter is then used to adaptively adjust the portraits of meeting members in the real-time meeting member image, resulting in an adaptive portrait displayed at the virtual location. The portrait adjustment image is the adaptive adjustment result obtained by adaptively adjusting the real-time meeting member image using the transformed adaptive adjustment parameters; the portrait mask adjustment image is the adaptive adjustment result obtained by adaptively adjusting the original portrait mask image using the transformed adaptive adjustment parameters. The composite image is obtained by combining the portrait adjustment image and the portrait mask adjustment image. The composite image is scaled according to the size of the area where the virtual location is located, resulting in an adaptive portrait displayed at the virtual location.

[0206] Specifically, after obtaining the adaptive adjustment parameters, the server transforms the adaptive adjustment parameters according to the scaling ratio to obtain the transformed adaptive adjustment parameters, which are suitable for adaptively adjusting the real-time images of meeting members. The server uses the transformed adaptive adjustment parameters to perform adaptive portrait adjustment on both the real-time images of meeting members and the corresponding original portrait mask images, obtaining an adjusted portrait image and a corresponding adjusted portrait mask image. Specifically, the server can perform adaptive portrait adjustment on the portrait areas in the real-time images of meeting members and the portrait areas in the corresponding original portrait mask images, using parameters such as height offset, width offset, and scaling factors. The server then combines the adjusted portrait image and the corresponding adjusted portrait mask image to obtain a composite image, and scales the composite image according to the size of the corresponding virtual location area to obtain the adaptive portrait displayed at the virtual location.

[0207] In this embodiment, the adaptive adjustment parameters are transformed by scaling the original image mask to the same scaling ratio. The transformed adaptive adjustment parameters are then used to adaptively adjust the images of both the real-time meeting participants and the original image mask. The composite image obtained by combining the adjusted image and the corresponding image mask is then scaled according to the size of the corresponding virtual location's area, resulting in the adaptive image displayed at the virtual location. The adaptive image displayed at the virtual location is obtained by adaptively adjusting the images according to the transformed adaptive adjustment parameters corresponding to the real-time meeting participants' images. The size of the adaptive image matches the size of the corresponding virtual location's area, achieving control over the size of the displayed adaptive image based on the location area of ​​the virtual location in the virtual scene. This avoids cluttered images in the real-time meeting participants' images, enhances the realism of the video conference, and thus improves the interactive efficiency of the video conference.

[0208] In one embodiment, such as Figure 15 As shown, a method for processing video conference images is provided, which can be applied to... Figure 1 Taking the server in the example, the following steps are included:

[0209] Step 1502: Obtain the real-time images of the meeting members captured by each of the multiple terminals that have joined the video conference.

[0210] In this context, "terminal" refers to the user-side device participating in the video conference. "Live feed of conference members" refers to the real-time video captured by the terminal during the video conference, including images of the participants captured by the terminal. Each terminal joining the video conference can activate its camera to capture video and obtain live feeds of the conference members.

[0211] Specifically, the server can obtain real-time images of meeting members captured by each terminal that has joined the video conference. Each terminal captures its corresponding real-time images of meeting members. The server can obtain the corresponding real-time images of meeting members from each terminal and perform adaptive adjustment processing on the real-time images of each meeting member to obtain the corresponding adaptive human image displayed at the virtual position in the virtual frame of the terminal.

[0212] Step 1504: For each meeting member's real-time video feed, perform facial analysis on the meeting member's image in the real-time video feed based on the facial parameters and face parameters corresponding to the image in the real-time video feed, and obtain the adaptive adjustment parameters corresponding to the meeting member's real-time video feed.

[0213] Among them, portrait parameters refer to the parameters of the foreground area in the real-time image of the meeting members. These parameters can be those corresponding to the portraits of the meeting members in the real-time image, specifically including but not limited to the width and height of the portrait rectangle, and the number of pixels in the portrait area. Face parameters refer to the parameters corresponding to the faces of the meeting members in the real-time image, specifically including but not limited to the width and height of the face rectangle, and the number of pixels in the face area. Adaptive adjustment parameters are obtained by performing portrait analysis on the portraits of the meeting members in the real-time image based on the portrait and face parameters, and are used to adaptively adjust the portraits of the meeting members in the real-time image. Adaptive adjustment parameters can specifically include adjustment parameters for adjusting the width of the portraits of the meeting members in the real-time image, adjustment parameters for adjusting the height of the portraits of the meeting members in the real-time image, and scaling parameters for scaling the portraits of the meeting members in the real-time image.

[0214] Specifically, for each meeting member's real-time video feed, the server performs portrait and face segmentation on the live feed to obtain the portrait and face parameters of the meeting member in the live feed. In practical applications, image segmentation technology can be used to segment the live feed image to determine the portrait and face parameters of the meeting member. Based on the obtained portrait and face parameters, the server performs portrait analysis on the meeting member's portrait in the live feed. Specifically, it can analyze the distribution of the meeting member's portrait in the live feed and determine the adjustment parameters that need to be adjusted based on this distribution, thus obtaining the adaptive adjustment parameters corresponding to the meeting member's live feed. The adaptive adjustment parameters include parameters for adjusting the meeting member's portrait in the live feed across multiple dimensions, specifically including adjustment parameters in the width direction, adjustment parameters in the height direction, and scaling parameters.

[0215] Step 1506: Based on the adaptive adjustment parameters, adaptively adjust the portraits of the meeting members in the real-time image of the meeting members to obtain the adaptive portraits corresponding to the portraits in the real-time image of the meeting members.

[0216] Specifically, after obtaining the adaptive adjustment parameters, the server adaptively adjusts the portraits of the meeting members in the real-time screen based on the adaptive adjustment parameters, and obtains the adaptive portraits corresponding to the portraits of the meeting members in the real-time screen. The adaptive portraits are displayed at virtual locations in the virtual background.

[0217] Step 1508: Send the virtual side-by-side images generated based on each adaptive human portrait to each terminal for display on each terminal.

[0218] Among them, virtual co-op screen refers to a virtual screen in which the real-time video screens of each participant in a video conference are displayed in the same frame. Co-op display means that during the video conference, the surrounding environment of multiple participants is replaced with a specified image or video, that is, the participants in the video conference share a virtual background for real-time screen display.

[0219] Specifically, the server adaptively adjusts the real-time images of meeting members captured by each of the multiple terminals joining the video conference. After obtaining adaptive portraits corresponding to the portraits in the real-time images of each meeting member, the server generates a virtual frame-to-frame image based on these adaptive portraits. Specifically, the server can aggregate these adaptive portraits into the virtual frame-to-frame image and send the resulting virtual frame-to-frame image to each terminal for display. When displayed on each terminal, the terminals joining the video conference display the virtual frame-to-frame image. The virtual scene of the virtual frame-to-frame image includes multiple virtual locations to accommodate meeting members. At at least two virtual locations in the virtual scene, adaptive portraits corresponding to the portraits in the real-time images of meeting members captured by at least two of the multiple terminals are displayed. The size of the adaptive portraits displayed at each of the at least two virtual locations matches the size of the corresponding virtual location's area.

[0220] The aforementioned video conferencing image processing method involves taking the real-time images of each participant captured by multiple terminals joining the video conference, performing facial analysis on the participants' images based on their facial and portrait parameters, and then adaptively adjusting these images using the obtained adaptive adjustment parameters. This results in an adaptive portrait, and the virtual frame-to-frame image generated from these adaptive portraits is sent to each terminal for display. In this video conferencing image processing, adaptively adjusting the images of each participant captured by multiple terminals according to the facial and portrait parameters ensures that the size of the adaptive portrait displayed in the virtual frame-to-frame image matches the size of the corresponding virtual location area. This avoids cluttered images in the participants' real-time images, enhances the realism of the video conference, and improves the interactive efficiency.

[0221] This application also provides an application scenario in which the above-described video conferencing interaction method is applied. Specifically, the video conferencing interaction method is applied in this scenario as follows:

[0222] Current online video conferencing systems can aggregate images from multiple devices into a single frame. However, this often results in inconsistent image sizes and an overall disharmonious picture, negatively impacting the realism of the video conference and reducing interaction efficiency. Furthermore, image size is directly related to the distance between the subject and the camera; subjects that are too close or too far away can appear visually "ugly." This reduces the enjoyment of video conferencing, leading most users to opt for audio conferencing instead. The video conferencing interaction method provided in this embodiment, based on an online video conferencing frame-sharing system, standardizes the size of subjects in the frame according to their position. This avoids the negative effect of inconsistent subject sizes due to differences in distance and position from the camera, resulting in an overall disharmonious image. This creates a more realistic atmosphere, bridging the gap between online and in-person meetings, enhancing the online meeting experience, and improving the interaction efficiency of online video conferencing.

[0223] like Figure 16 As shown, in a video conferencing system with a human image co-creation feature that supports co-creation display, after enabling co-creation mode, each terminal collects data and sends it to the server. The server processes the data from each terminal, and after processing by the human image co-creation system, outputs a co-creation image integrating information from each terminal. This processed co-creation image is then returned to each terminal for display. Specifically, the server obtains real-time images of meeting members collected by each terminal joining the video conference, processes them through the human image co-creation system, aggregates the real-time images of meeting members from each terminal into a unified virtual co-creation image, and sends this virtual co-creation image to each terminal for display. Figure 17 As shown, for Figure 16 Each terminal in the system uses a camera to capture images and segment faces, outputting the original portrait, the segmented face mask image, and face parameters, which are then sent to the server. The server performs positional arrangement and image synthesis based on the images input from each terminal, finally returning a virtual frame corresponding to the video conference to the terminal for display. The pixel value of each pixel in the face mask image represents the probability that the corresponding pixel belongs to the foreground, i.e., the pixel value ranges from [0.0, 1.0]. Face parameters include the parameters of the bounding rectangle corresponding to the face region in the image, including the position, length, and height of the face mask rectangle, as well as the number of valid face pixels. By segmenting the captured original image on the terminal, the terminal's computing resources can be fully utilized, ensuring the interactive efficiency of the video conference.

[0224] Furthermore, such as Figure 18As shown, the server obtains real-time images of members captured by the terminal, i.e., the original image, the corresponding original portrait mask image, and the face parameters after scaling the image to 128*128 pixels. Then, the server uniformly adjusts the size of the portraits, adjusting their proportion in the image to a fixed ratio, achieving a harmonious overall effect where individual portraits are uniformly sized in scenes with multiple people in the same frame. Specifically, the server scales the original portrait mask image to 128*128 pixels to obtain the portrait mask image. Based on the portrait mask image, the server determines the portrait parameters. These parameters can be obtained by performing face segmentation on the original image, determining the original face parameters, and then scaling the original face parameters to 128*128 pixels. The portrait parameters include the number of portrait pixels in the portrait mask image (mask_area) and the sharpened portrait mask image (mask_hard). Scaling the original portrait mask image to a 128*128 image size and determining adaptive adjustment parameters based on the scaled portrait mask image can reduce the amount of computation and improve the efficiency of determining adaptive adjustment parameters, thereby improving the interactive efficiency of video conferencing.

[0225] The server determines whether the proportion of the human face in the portrait mask image is too small based on portrait parameters, such as whether it is less than a portrait proportion threshold or a preset portrait proportion ratio; and determines whether the proportion of the face in the portrait mask image is too small based on face parameters, such as whether it is less than a face proportion threshold or a preset face proportion ratio. If the proportion of the human face or the proportion of the face in the portrait mask image is small, the server obtains the adaptive adjustment parameters corresponding to the previous frame's original image and determines the adaptive adjustment parameters corresponding to the portrait mask image based on the adaptive adjustment parameters corresponding to the previous frame's original image. On the other hand, if the proportion of the human face and the proportion of the face in the portrait mask image are both sufficiently large, the server determines the portrait mask rectangle in the portrait mask image and the adaptive adjustment parameters of the original image corresponding to the previous 8 frames. The server determines the adaptive adjustment parameters corresponding to the portrait mask image based on the determined portrait mask rectangle in the portrait mask image and the adaptive adjustment parameters of the original image corresponding to the previous 8 frames' original images. The adaptive adjustment parameters include width offset (shift_on_w), height offset (shift_on_h, crop_down_h), and scaling factor (scale).

[0226] After obtaining the adaptive adjustment parameters corresponding to the portrait mask image, the adaptive adjustment parameters are transformed based on the image size relationship between the portrait mask image and the original image to obtain the transformed adaptive adjustment parameters. The transformed adaptive adjustment parameters are suitable for adaptively adjusting the image size of the original image. Specifically, the transformation formulas may include: real_shift_w = M + shift_on_w * M; real_shift_h = ((1 / scale-1)*N*(1+shift_on_h-crop_down_h)) + (N*shift_on_h); real_crop_down_h = N+N*crop_down_h, where M is the image width of the original image, N is the image height of the original image, and real_shift_w, real_shift_h, real_crop_down_h, and scale constitute the transformed adaptive adjustment parameters. The original image and the original portrait mask image are adaptively adjusted by the transformed adaptive adjustment parameters to obtain the adjusted original image and the adjusted portrait mask image. The adjusted original image and the adjusted portrait mask image are then combined to obtain a composite image. The composite image is scaled according to the size of the corresponding virtual location area to obtain an adaptive portrait. The adaptive portrait is used to display at the corresponding virtual location, and the size of the adaptive portrait matches the size of the corresponding virtual location area.

[0227] Furthermore, such as Figure 19 As shown, when determining the adaptive adjustment parameters, the server scales the original portrait mask image to 128*128, and then performs portrait analysis processing based on the obtained portrait mask image and face parameters through the same function module to obtain the adaptive adjustment parameters. Figure 20As shown, after obtaining the adaptive adjustment parameters, the server adjusts the original image and the original portrait mask image respectively using the transformed adaptive adjustment parameters to obtain the adjusted original image and the adjusted portrait mask image. The adjusted original image and the adjusted portrait mask image are then combined to obtain the composite image. Subsequently, the server scales the composite image according to the size of the corresponding virtual location area to obtain the adaptive image displayed at the virtual location. From the overall server processing perspective, the server obtains the original image, the original portrait mask image, and face parameters. First, it scales the original portrait mask image to 128x128. Then, it calls the "unified function" module to perform portrait analysis to obtain the corresponding adaptive adjustment parameters. Both the original image and the original portrait mask image are adaptively adjusted using these transformed adaptive adjustment parameters to obtain the adjusted original image and the adjusted portrait mask image. Finally, the adjusted original image and the adjusted portrait mask image are combined to obtain the composite image. The composite image can be scaled according to the needs of the virtual location to display the scaled adaptive portrait at the corresponding virtual location.

[0228] For unified function modules, such as Figure 21 As shown, the server obtains the portrait mask image and face parameters. First, it performs calculations on the portrait mask image to obtain the portrait parameters. Then, it uses a threshold to determine whether to directly update the parameter queue or further obtain the portrait mask rectangle. The portrait mask rectangle is the region of the smallest bounding box containing the portrait mask. The offset point is calculated in conjunction with the face parameters. Specifically, in the parameter queue update process, as... Figure 22 As shown, if the number of parameters in the parameter queue is greater than 8, the head member is deleted, that is, the parameter that was first updated in the parameter queue is deleted, and the new parameter is inserted at the tail of the parameter queue to update the parameter queue. The parameter queue maintains a queue with a maximum number of 8 parameters. Each parameter includes 4 members: the horizontal offset of the portrait mask rectangle (shift_on_w), the vertical offsets of the portrait mask rectangle (shift_on_h and crop_down_h), and the overall scaling factor of the portrait mask rectangle (scale). The purpose of the parameter queue is to accumulate parameter information from several frames. When processing the current frame, historical information can be used to perform a smoothing operation on the current frame to obtain the parameters of the current frame. The members stored in each parameter are shift_on_w, shift_on_h, crop_down_h, and scale. The maximum number of parameters in the parameter queue is 8, that is, the adaptive adjustment parameters of the previous 8 frames are accumulated to perform a smoothing operation on the current frame. Figure 23As shown, the dashed box is the portrait mask rectangle. Shift_on_w fills it to a uniform portrait width, and Shift_on_h fills it to a uniform portrait height. Then, its bottom is pressed against the bottom edge of the original image, and its top is filled with crop_down_h to obtain a portrait with appropriate proportions. Finally, it is scaled to the original image size to obtain an image of uniform size.

[0229] Furthermore, the portrait parameters include the sharpened mask image (mask_hard) and the total number of foreground pixels in the portrait mask image (mask_area). For example... Figure 24 As shown, when sharpening the portrait mask image, the pixel value of each pixel in the portrait mask image is determined. If the pixel value is greater than 0.5, the pixel value of the pixel is set to 1; otherwise, the pixel value of the pixel is set to 0. After traversing the portrait mask image, the sharpened mask image mask_hard is obtained, and the total number of foreground pixels mask_area in the portrait mask image is determined based on the sharpened mask image mask_hard.

[0230] When determining the threshold, such as Figure 25 As shown, for portrait and face parameters, it is determined whether the number of portrait pixels (mask_area) is less than 10% of the total number of pixels in the portrait mask image. If so, the portrait area in the portrait mask image is considered to be small, and the frame is discarded. On the other hand, it is determined whether the number of effective face pixels is less than 5% of the number of portrait pixels (mask_area). If so, the face area in the portrait area is considered to be small, and the frame is discarded. After discarding the frame, the parameter queue can be updated based on the adaptive parameters of the previous frame. For example, the adaptive parameters of the previous frame can be used as the adaptive parameters of the current frame. Specifically, the server can copy the queue result of the previous frame and return it as the content of the current queue. This avoids gaps in the queue information of the current frame, which could lead to changes in the calculation logic. It also increases the weight of the information from the previous frame in the overall queue information, making the update information of the next frame smoother than the previous frame. If neither condition is met, the offset point is calculated based on the portrait and face parameters. By determining the proportion of human portrait pixels and the proportion of human face pixels, a frame can be discarded if it is not a human portrait to ensure a more continuous human portrait image.

[0231] When calculating the offset point, it is necessary to calculate the offsets shift_on_h and crop_down_h in the height direction, the offset shift_on_w in the width direction, and the scaling factor scale. For example... Figure 26As shown, after obtaining the initial value of shift_on_w, it is restricted to the range [-0.8, 0.8], and the average value of shift_on_w in the parameter queue is obtained. It is then determined whether the difference between the initial value of shift_on_w and the average value is greater than 0.1. If so, the initial value of shift_on_w is corrected based on the average value, and the corrected shift_on_w is updated in the parameter queue. If the difference between the initial value of shift_on_w and the average value is less than or equal to 0.1, the initial value of shift_on_w is directly updated in the parameter queue. The initial value of shift_on_w can be calculated as: initial value of shift_on_w = 2 × (x-coordinate of the center point of the portrait mask image – x-coordinate of the center point of mask_hard) / width of the portrait mask image. The formula for calculating the difference is: Difference = Initial value of shift_on_w – Average value of shift_on_w in the parameter queue. If the difference is greater than the threshold, it means that the difference between the current frame and the history is too large, and the offset is too large, which manifests as face shaking. Therefore, it is corrected based on the offset of the history frames. The formula for correcting shift_on_w is: shift_on_w = 0.1 × Initial value of shift_on_w + 0.9 × (shift_on_w of the previous frame). After correction, the value of shift_on_w is limited to the range of [-0.8, 0.8], which can be the maximum value, minimum value, or no transformation.

[0232] Furthermore, such as Figure 27As shown, the initial values ​​of `shift_on_h` and `crop_down_h` are obtained. If 1 + initial value of shift_on_h - initial value of crop_down_h >= 0.25, then the average values ​​corresponding to `shift_on_h` and `crop_down_h` in the parameter queue are obtained respectively. The initial value of `shift_on_h` is compared with the average value of `shift_on_h`, and the initial value of `crop_down_h` is compared with the average value of `crop_down_h`. If the difference between the initial value of `shift_on_h` and the average value of `shift_on_h` is greater than 0.1, then the initial value of `shift_on_h` is corrected according to the average value of `shift_on_h`, and the corrected `shift_on_h` is updated to the parameter queue. If the difference between the initial value of `shift_on_h` and the average value of `shift_on_h` is not greater than 0.1, then the initial value of `shift_on_h` is directly updated to the parameter queue. Similarly, if the difference between the initial value of `crop_down_h` and its average value is greater than 0.1, the initial value of `crop_down_h` is corrected based on its average value, and the corrected `crop_down_h` is updated in the parameter queue. If the difference between the initial value of `crop_down_h` and its average value is not greater than 0.1, the initial value of `crop_down_h` is directly updated in the parameter queue. Furthermore, if the condition 1 + initial value of shift_on_h - initial value of `crop_down_h` >= 0.25 is not true, the parameter queue is directly updated based on the initial values ​​of shift_on_h and `crop_down_h`.

[0233] The initial values ​​of shift_on_h and crop_down_h are calculated using the following formulas: shift_on_h = 0.5 – the top y-coordinate of the portrait mask rectangle / the height of the portrait mask image; crop_down_h = 1 – the bottom y-coordinate of the portrait mask rectangle / the height of the portrait mask image. The shift_on_h difference is calculated as: difference = initial shift_on_h value – the average value of shift_on_h in the parameter queue. The formula for correcting shift_on_h is: shift_on_h = 0.1 × initial shift_on_h value + 0.9 × (shift_on_h of the previous frame). The crop_down_h difference is calculated as: difference = initial crop_down_h value – the average value of shift_on_h in the parameter queue; the formula for correcting crop_down_h is: crop_down_h = 0.1 × initial crop_down_h value + 0.9 × (crop_down_h of the previous frame).

[0234] Regarding the determination of the scaling factor, such as Figure 28 As shown, the face scaling factor includes a face width scaling factor and a face height scaling factor. Both the face width scaling factor and the face height scaling factor are initially assigned a value of 1.0. `new_w` and `new_h` are calculated, and `new_w` and `new_h` are used as padding edges. Here, `new_h = h + (shift_on_h - crop_down_h) * h`, `new_w = new_h`, and `h` is the height of the face mask image, which is 128 pixels. If the width of the face mask rectangle, `face_width`, satisfies `face_width < 0.15 × new_w`, then the face width scaling factor `scale_face_w` is determined to be 0.15 * new_w / face_width`; otherwise, the face width scaling factor `scale_face_w` is determined to be 0.3 * new_w / face_width. On the other hand, if the height of the face mask rectangle, `face_height`, satisfies `face_height < 0.2 × new_h`, then the face height scaling factor `scale_face_h` is determined to be 0.2 * new_h / face_height; otherwise, the face height scaling factor `scale_face_h` is determined to be 0.4 * new_h / face_height. The smaller value between the face width scaling factor `scale_face_w` and the face height scaling factor `scale_face_h` is determined as the face scaling factor `scale_face`.

[0235] Furthermore, such as Figure 29 As shown, the scaling factor `scale` is initially assigned a value of 1.0, and `new_w` and `new_h` are determined. If the width of the portrait mask rectangle (`portrait width`) satisfies `portrait width < 0.45 × `new_w`, then the width scaling factor `scale_w` is determined to be 0.45 * `new_w` / `portrait width`; otherwise, the width scaling factor `scale_w` is determined to be 0.8 * `new_w` / `portrait width`. On the other hand, if the height of the portrait mask rectangle (`portrait height`) satisfies `portrait height < 0.65 × `new_h`, then the height scaling factor `scale_h` is determined to be 0.65 * `new_h` / `portrait height`; otherwise, the height scaling factor `scale_h` is determined to be `new_h` / `portrait height`. The portrait scaling factor `scale_box` = (`scale_w` + `scale_h`) / 2, and the scaling factor `scale` = 0.75 × `scale_box` + 0.25 × `scale_face`.

[0236] In the adaptive adjustment processing of the original image and the original portrait mask image based on determined adaptive parameters, when the portrait is too small, such as... Figure 30 As shown, the portrait is cropped from the original image and then enlarged to the original image size. The dashed rectangle represents the target area determined based on the transformed adaptive parameters, and is scaled to the original image size using a scaling factor (scale). When the portrait is too large, such as... Figure 31 As shown, the original image needs to be enclosed in brackets with a fill pixel value of "0". The new image after filling is then scaled down to the original image size using a scaling factor. Figure 32 As shown, after adaptive adjustment processing of the original image and the original portrait mask image, the adjusted original image (in BGR format) and the adjusted portrait mask image (in alpha channel image) are obtained. By merging the adjusted original image and the adjusted portrait mask image, the mask content is interleaved and appended to the BGR image to obtain a composite image in BGR format. The portrait mask image is an alpha channel image obtained based on alpha blending.

[0237] The video conferencing interaction method in this embodiment adaptively scales the portrait at a specified position based on the proportion of the portrait and face in the overall screen, unifying the portrait to a more suitable size, increasing the harmony of the overall screen, thereby creating a more realistic atmosphere of an on-site meeting, bridging the gap between online and on-site meetings, improving the online meeting experience, and increasing the interaction efficiency of online video conferencing.

[0238] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0239] In one embodiment, such as Figure 33 As shown, a video conferencing interaction device 3300 is provided. This device can be a software module, a hardware module, or a combination of both as part of a computer device. Specifically, the device includes: a simultaneous screen display module 3302 and an adaptive human image display module 3304, wherein:

[0240] The co-op display module 3302 is used to respond to the trigger operation of joining a video conference between multiple terminals and display the virtual co-op screen of the video conference. The virtual scene of the virtual co-op screen includes multiple virtual positions to accommodate the meeting members.

[0241] The adaptive human image display module 3304 is used to display adaptive human images corresponding to the human images in the real-time images of meeting members captured by at least two of multiple terminals at at least two virtual locations in a virtual scene.

[0242] At least two virtual locations each display an adaptive human portrait size that matches the size of the corresponding virtual location's location area.

[0243] In one embodiment, the adaptive portrait display module 3304 includes a portrait area size determination module and a target portrait display module; wherein: the portrait area size determination module is used to determine the portrait area size matching the corresponding virtual location based on the size of the location area of ​​the virtual location in the virtual scene and the preset portrait uniform proportion conditions; the target portrait display module is used to display the target adaptive portrait corresponding to the portrait in the real-time screen of the meeting members collected by any one of the multiple terminals at the virtual location in the virtual scene, according to the corresponding portrait area size.

[0244] In one embodiment, the adaptive portrait display module 3304 includes a distribution location determination module, a proportion condition processing module, and a target portrait display module; wherein: the distribution location determination module is used to determine the distribution location of each virtual location in the virtual scene; the proportion condition processing module is used to determine the size of the portrait area matching the corresponding virtual location according to the unified proportion condition of the portrait corresponding to each distribution location; the target portrait display module is used to display the target adaptive portrait corresponding to the portrait in the real-time image of the meeting members collected by any one of the multiple terminals at the virtual location in the virtual scene, according to the corresponding portrait area size.

[0245] In one embodiment, the system further includes a screen sharing area display module and a screen sharing response module; wherein: the screen sharing area display module is used to display a virtual screen sharing area belonging to the virtual scene in the virtual co-creation screen; the screen sharing response module is used to respond to the target terminal among the multiple terminals joining the video conference triggering screen sharing, and to display the screen sharing content of the target terminal in the virtual screen sharing area.

[0246] In one embodiment, the screen sharing area display module is further configured to display a virtual screen sharing area in the background area other than the virtual location in the virtual frame-to-frame image; it also includes a blank screen display module, configured to display a virtual blank screen in the virtual screen sharing area when multiple terminals joining the video conference do not trigger screen sharing; and to display a virtual blank screen in the virtual screen sharing area in response to the target terminal triggering the cancellation of screen sharing.

[0247] In one embodiment, the system further includes a display location movement module, which is used to respond to the first adaptive portrait displayed at a first virtual location in the mobile virtual scene moving to a second virtual location, and at the second virtual location, displaying an adaptive portrait corresponding to the portrait in the real-time image of the meeting member captured by the terminal corresponding to the first adaptive portrait.

[0248] In one embodiment, a network anomaly alert module is also included, which, in response to a target terminal among multiple terminals being in an abnormal network state, displays network anomaly alert information about the meeting member corresponding to the adaptive portrait at the target virtual location of the adaptive portrait corresponding to the portrait in the real-time image of the meeting members collected by the target terminal.

[0249] In one embodiment, an abnormal status prompt module is also included, which is used to respond to the abnormal participation status of a target meeting member in the real-time video of the meeting members collected by the target terminal among multiple terminals, and to display abnormal status prompt information about the target meeting member at the target virtual position of the adaptive human image corresponding to the human image in the real-time video of the meeting members collected by the target terminal.

[0250] In one embodiment, the system further includes a video interface display module, configured to display the video interface of the video conference in response to a trigger operation of joining the video conference; the video interface displays real-time images of meeting members captured by at least one of the multiple terminals; the same-frame display module 3302 is further configured to, in response to triggering the opening of the same-frame mode for the video conference, cancel the display of real-time images of meeting members in the video interface and display a virtual same-frame image of the video conference in the video interface; and includes a same-frame mode end module, configured to, in response to the video conference meeting the conditions for ending the same-frame mode, cancel the display of the virtual same-frame image in the video interface.

[0251] In one embodiment, the system further includes a real-time image acquisition module, a member parameter determination module, a portrait analysis module, and an adaptive adjustment module; wherein: the real-time image acquisition module is used to acquire real-time images of conference members collected by each terminal joining the video conference; the member parameter determination module is used to determine the portrait parameters and face parameters of each conference member in the real-time image; the portrait analysis module is used to perform portrait analysis on the portrait of the conference member in the real-time image based on the portrait parameters and face parameters to obtain the adaptive adjustment parameters corresponding to the real-time image; and the adaptive adjustment module is used to adaptively adjust the portrait of the conference member in the real-time image based on the adaptive adjustment parameters to obtain the corresponding adaptive portrait.

[0252] In one embodiment, the portrait analysis module includes an offset determination module, a scaling factor determination module, and an adjustment parameter acquisition module; wherein: the offset determination module is used to determine the height offset and width offset based on the portrait parameters when the real-time image of the meeting member is determined to be a valid portrait image based on the portrait parameters and face parameters; the scaling factor determination module is used to determine the scaling factor based on the portrait parameters, face parameters, and height offset; and the adjustment parameter acquisition module is used to use the height offset, width offset, and scaling factor as adaptive adjustment parameters corresponding to the real-time image of the meeting member.

[0253] In one embodiment, the portrait parameters include the portrait mask rectangle parameters of the portrait mask rectangle in the portrait mask image corresponding to the real-time image of the meeting member; the offset determination module includes an initial offset determination module, a width offset determination module, and a height offset determination module; wherein: the initial offset determination module is used to determine the initial width offset and the initial height offset based on the portrait mask rectangle parameters and the size of the portrait mask image, respectively; the width offset determination module is used to calculate the reference width offset corresponding to the real-time image of the meeting member based on the historical width offset corresponding to the previous image of the real-time image of the meeting member; and to correct the initial width offset using the reference width offset to obtain the width offset corresponding to the real-time image of the meeting member; the height offset determination module is used to calculate the reference height offset corresponding to the real-time image of the meeting member based on the historical height offset corresponding to the previous image of the real-time image of the meeting member; and to correct the initial height offset using the reference height offset to obtain the height offset corresponding to the real-time image of the meeting member.

[0254] In one embodiment, the reference width offset is the average of the historical width offsets corresponding to the previous frame of the real-time view of the meeting members; the width offset determination module includes a width difference determination module, a width correction module, and a width range limitation module; wherein: the width difference determination module is used to determine the difference between the initial width offset and the average of the historical width offsets; the width correction module is used to perform weighted correction on the initial width offset by using the historical width offset corresponding to the previous frame of the real-time view of the meeting members when the difference is greater than a first preset threshold, to obtain a corrected width offset; the width range limitation module is used to perform range limitation processing on the corrected width offset to obtain the width offset corresponding to the real-time view of the meeting members.

[0255] In one embodiment, the reference height offset is the average of the historical height offsets corresponding to the previous frame of the real-time view of the meeting members; the height offset determination module includes a height difference determination module, a height correction module, and a height range limitation module; wherein: the height difference determination module is used to determine the difference between the initial height offset and the average of the historical height offsets; the height correction module is used to perform weighted correction on the initial height offset by using the historical height offset corresponding to the previous frame of the real-time view of the meeting members when the difference is greater than a second preset threshold, to obtain a corrected height offset; the height range limitation module is used to perform range limitation processing on the corrected height offset to obtain the height offset corresponding to the real-time view of the meeting members.

[0256] In one embodiment, the system further includes an image center point determination module, a portrait center point determination module, a horizontal position difference determination module, and a horizontal position difference processing module; wherein: the image center point determination module is used to determine the image center point of the portrait mask image based on the size of the portrait mask image; the portrait center point determination module is used to determine the portrait center point of the portrait mask rectangle based on the portrait mask rectangle parameters and the size of the portrait mask image; the horizontal position difference determination module is used to determine the horizontal position difference between the image center point and the portrait center point; and the horizontal position difference processing module is used to determine an initial width offset based on the horizontal position difference and the image width of the portrait mask image.

[0257] In one embodiment, the system further includes a portrait position determination module, a portrait position processing module, and an initial height offset determination module; wherein: the portrait position determination module is used to determine the top and bottom positions of the portrait mask rectangle in the image height direction based on the portrait mask rectangle parameters; the portrait position processing module is used to determine a first initial height offset based on the top position and the image height of the portrait mask image; and to determine a second initial height offset based on the bottom position and the image height of the portrait mask image; the initial height offset determination module is used to obtain an initial height offset based on the first initial height offset and the second initial height offset.

[0258] In one embodiment, the portrait parameters include portrait mask rectangle parameters corresponding to the real-time images of the meeting members; the face parameters include face mask rectangle parameters corresponding to the real-time images of the meeting members; the scaling factor determination module includes a face scaling factor module, a portrait scaling factor module, and a factor fusion module; wherein: the face scaling factor module is used to determine the face scaling factor based on the face mask rectangle parameters and the height offset; the portrait scaling factor module is used to determine the portrait scaling factor based on the portrait mask rectangle parameters and the height offset parameters; and the factor fusion module is used to fuse the face scaling factor and the portrait scaling factor to obtain the scaling factor.

[0259] In one embodiment, the portrait parameters include the number of portrait pixels in the portrait mask image corresponding to the real-time image of the meeting member; the face parameters include the number of face pixels in the portrait mask image; and further includes: an invalid image offset determination module, used to obtain the height offset and width offset of the previous image corresponding to the real-time image of the meeting member when it is determined based on the number of portrait pixels or the number of face pixels that the real-time image of the meeting member does not belong to a valid portrait image; and to determine the height offset and width offset of the previous image as the height offset and width offset of the real-time image of the meeting member.

[0260] In one embodiment, the member parameter determination module includes an image segmentation module, a portrait parameter acquisition module, and a face parameter acquisition module; wherein: the image segmentation module is used to perform image segmentation processing on the real-time image of the meeting members to obtain the original portrait mask image and the original face parameters corresponding to the real-time image of the meeting members; the portrait parameter acquisition module is used to perform portrait analysis on the portrait mask image obtained by scaling the original portrait mask image according to the scaling ratio to obtain the portrait parameters corresponding to the portraits in the real-time image of the meeting members; the face parameter acquisition module is used to scale the original face parameters according to the scaling ratio to obtain the face parameters corresponding to the faces in the real-time image of the meeting members.

[0261] In one embodiment, the adaptive adjustment module includes a parameter transformation module, an image adjustment module, an image compositing module, and a scaling module; wherein: the parameter transformation module is used to transform the adaptive adjustment parameters based on the scaling ratio to obtain the transformed adaptive adjustment parameters; the image adjustment module is used to perform adaptive portrait adjustment on the real-time image of the meeting members and the original portrait mask image corresponding to the real-time image of the meeting members using the transformed adaptive adjustment parameters to obtain the portrait adjustment image and the corresponding portrait mask adjustment image; the image compositing module is used to combine the portrait adjustment image and the portrait mask adjustment image to obtain a composite image; and the scaling module is used to scale the composite image according to the size of the location area where the virtual location is located to obtain the adaptive portrait displayed at the virtual location.

[0262] In one embodiment, such as Figure 34 As shown, a video conferencing image processing device 3400 is provided. This device can be a software module, a hardware module, or a combination of both as part of a computer device. Specifically, the device includes: a real-time image acquisition module 3402, an adjustment parameter acquisition module 3404, an adaptive adjustment module 3406, and a simultaneous image distribution module 3408, wherein:

[0263] The real-time video acquisition module 3402 is used to acquire the real-time video of the meeting members captured by multiple terminals that have joined the video conference.

[0264] The parameter acquisition module 3404 is used to perform portrait analysis on the portrait of the meeting member in the real-time image of each meeting member based on the portrait parameters and face parameters corresponding to the portrait in the real-time image of the meeting member, and obtain the adaptive adjustment parameters corresponding to the real-time image of the meeting member.

[0265] The adaptive adjustment module 3406 is used to adaptively adjust the portraits of meeting members in the real-time image of the meeting members based on the adaptive adjustment parameters, so as to obtain the adaptive portraits corresponding to the portraits in the real-time image of the meeting members.

[0266] The on-screen display module 3408 is used to send virtual on-screen displays generated based on each adaptive human image to each terminal for display on each terminal.

[0267] Specific limitations regarding the devices for video conferencing interaction and the devices for processing video conferencing images can be found in the limitations regarding the methods for video conferencing interaction and the methods for processing video conferencing images described above, and will not be repeated here. Each module in the aforementioned devices for video conferencing interaction and video conferencing image processing can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0268] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 35As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The device also includes input / output interfaces (I / O interfaces), which are connection circuits between the processor and external devices for exchanging information. These I / O interfaces are connected to the processor via a bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database stores real-time video data of conference participants. The network interface communicates with external terminals via a network. When the computer program is executed by the processor, it implements a method for processing video conference footage.

[0269] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 36 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The computer device also includes input / output interfaces, which are connection circuits for exchanging information between the processor and external devices. These interfaces are connected to the processor via a bus and are referred to as I / O interfaces. The processor of this computer device provides computing and control capabilities. The memory of the computer device includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The communication interface of the computer device is used for wired or wireless communication with external terminals. Wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a video conferencing interaction method. The display screen of the computer device can be an LCD screen or an e-ink display screen. The input devices of the computer device can be a touch layer covering the display screen, buttons, trackballs, or touchpads located on the computer device casing, or external keyboards, touchpads, or mice, etc.

[0270] Those skilled in the art will understand that Figure 35 and Figure 36 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0271] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0272] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0273] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the steps in the above method embodiments.

[0274] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0275] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0276] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0277] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for video conferencing interaction, characterized in that, The method includes: Acquire real-time video feeds of conference participants captured by each terminal that has joined the video conference; For each meeting member's real-time video feed, determine the image parameters and face parameters of the meeting member in the real-time video feed; When the real-time image of the meeting member is determined to be a valid image based on the portrait parameters and the face parameters, the height offset and width offset are determined according to the portrait parameters. The scaling factor is determined based on the portrait parameters, the face parameters, and the height offset. The height offset, the width offset, and the scaling factor are used as adaptive adjustment parameters corresponding to the real-time screen of the meeting members; Based on the adaptive adjustment parameters, the portraits of the meeting members in the real-time image of the meeting members are adaptively adjusted to obtain the corresponding adaptive portraits; In response to a trigger operation that allows joining a video conference between multiple terminals, a virtual frame of the video conference is displayed, wherein the virtual scene of the virtual frame includes multiple virtual locations for accommodating conference members. At at least two virtual locations in the virtual scene, adaptive human images corresponding to the human images in the real-time images of meeting members collected by at least two of the multiple terminals are displayed; The at least two virtual locations each display an adaptive portrait size that matches the size of the corresponding virtual location's location area.

2. The method according to claim 1, characterized in that, The step of displaying adaptive human images corresponding to the real-time images of meeting members captured by at least two of the multiple terminals at at least two virtual locations in the virtual scene includes: Based on the size of the area where the virtual location is located in the virtual scene, and the preset uniform proportion of human figures, the size of the human figure area matching the corresponding virtual location is determined; At a virtual location in the virtual scene, the target adaptive human image corresponding to the human image in the real-time image of the meeting members collected by any one of the multiple terminals is displayed according to the corresponding human image area size.

3. The method according to claim 1, characterized in that, The step of displaying adaptive human images corresponding to the real-time images of meeting members captured by at least two of the multiple terminals at at least two virtual locations in the virtual scene includes: Determine the distribution of each virtual location within the virtual scene; Based on the uniform proportion of human figures corresponding to each of the aforementioned distribution locations, determine the size of the human figure area that matches the corresponding virtual location; At a virtual location in the virtual scene, the target adaptive human image corresponding to the human image in the real-time image of the meeting members collected by any one of the multiple terminals is displayed according to the corresponding human image area size.

4. The method according to claim 1, characterized in that, The method further includes: In the virtual shared-screen view, a virtual screen sharing area belonging to the virtual scene is displayed; In response to a target terminal among the plurality of terminals joining the video conference triggering screen sharing, the screen sharing content of the target terminal is displayed in the virtual screen sharing area.

5. The method according to claim 4, characterized in that, The provision that displays a virtual screen sharing area belonging to the virtual scene within the virtual shared frame includes: The virtual screen sharing area is displayed in the background area other than the virtual location in the virtual frame-to-frame image. The method further includes: When the multiple terminals joining the video conference do not trigger screen sharing, a virtual blank screen is displayed in the virtual screen sharing area; In response to the target terminal triggering the cancellation of screen sharing, a virtual blank screen is displayed in the virtual screen sharing area.

6. The method according to claim 1, characterized in that, The method further includes: In response to moving the first adaptive portrait displayed at the first virtual location in the virtual scene to the second virtual location, at the second virtual location, the adaptive portrait corresponding to the portrait in the real-time image of the meeting member collected by the terminal corresponding to the first adaptive portrait is displayed.

7. The method according to claim 1, characterized in that, The method further includes: In response to a target meeting member being in an abnormal participation state in the real-time video of the meeting members captured by the target terminal among the multiple terminals, an abnormal state prompt message about the target meeting member is displayed at the target virtual location of the adaptive human image corresponding to the human image in the real-time video of the meeting members captured by the target terminal.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: In response to a triggering operation to join a video conference, the video interface of the video conference is displayed; The video interface displays real-time video feeds of meeting participants captured by at least one of the multiple terminals. The response to a triggering operation in a video conference between multiple terminals, displaying a virtual side-by-side view of the video conference, includes: In response to the video conference triggering the start of the same-frame mode, the real-time screen of the conference members is canceled from being displayed on the video interface, and a virtual same-frame screen of the video conference is displayed on the video interface; The method further includes: In response to the video conference meeting the conditions for ending the same-frame mode, the virtual same-frame screen is no longer displayed in the video interface.

9. The method according to claim 1, characterized in that, The portrait parameters include the portrait mask rectangle parameters in the portrait mask image corresponding to the real-time image of the meeting member; The step of determining the height offset and width offset based on the portrait parameters includes: Based on the portrait mask rectangle parameters and the size of the portrait mask image, the initial width offset and the initial height offset are determined respectively; Calculate the reference width offset corresponding to the real-time screen of the meeting member based on the historical width offset corresponding to the previous screen of the real-time screen of the meeting member; correct the initial width offset using the reference width offset to obtain the width offset corresponding to the real-time screen of the meeting member. Based on the historical height offset corresponding to the previous view of the meeting member's real-time view, a reference height offset corresponding to the meeting member's real-time view is calculated; the initial height offset is corrected using the reference height offset to obtain the height offset corresponding to the meeting member's real-time view.

10. The method according to claim 9, characterized in that, The reference width offset is the average of the historical width offsets corresponding to the previous images of the real-time images of the meeting members. The step of correcting the initial width offset using the reference width offset to obtain the width offset corresponding to the real-time screen of the meeting member includes: Determine the difference between the initial width offset and the average of the historical width offsets; When the difference is greater than the first preset threshold, the initial width offset is weighted and corrected by the historical width offset corresponding to the previous frame of the real-time screen of the meeting member, so as to obtain the corrected width offset. The range of the corrected width offset is limited to obtain the width offset corresponding to the real-time screen of the meeting member.

11. The method according to claim 9, characterized in that, The reference height offset is the average of the historical height offsets corresponding to the previous images of the real-time images of the meeting members. The step of correcting the initial height offset using the reference height offset to obtain the height offset corresponding to the real-time screen of the meeting member includes: Determine the difference between the initial altitude offset and the mean of the historical altitude offsets; When the difference is greater than the second preset threshold, the initial height offset is weighted and corrected by the historical height offset corresponding to the previous frame of the real-time screen of the meeting member, so as to obtain the corrected height offset. The corrected height offset is subjected to range limitation processing to obtain the height offset corresponding to the real-time screen of the meeting member.

12. The method according to claim 9, characterized in that, The steps for determining the initial width offset include: The image center point of the portrait mask image is determined based on its dimensions. The center point of the portrait mask rectangle is determined based on the parameters of the portrait mask rectangle and the size of the portrait mask image. Determine the horizontal positional difference between the center point of the image and the center point of the portrait; The initial width offset is determined based on the horizontal position difference and the image width of the portrait mask image.

13. The method according to claim 9, characterized in that, The steps for determining the initial height offset include: Based on the portrait mask rectangle parameters, determine the top and bottom positions of the portrait mask rectangle in the image height direction; A first initial height offset is determined based on the top position and the image height of the portrait mask image; a second initial height offset is determined based on the bottom position and the image height of the portrait mask image. The initial height offset is obtained based on the first initial height offset and the second initial height offset.

14. The method according to claim 1, characterized in that, The portrait parameters include the portrait mask rectangle parameters corresponding to the real-time image of the meeting member; the face parameters include the face mask rectangle parameters corresponding to the real-time image of the meeting member; determining the scaling factor based on the portrait parameters, the face parameters, and the height offset includes: The face scaling factor is determined based on the face mask rectangle parameters and the height offset. The portrait scaling factor is determined based on the portrait mask rectangle parameters and the height offset. The scaling factor is obtained by fusing the face scaling factor and the portrait scaling factor.

15. The method according to claim 1, characterized in that, The portrait parameter includes the number of portrait pixels in the portrait mask image corresponding to the real-time image of the meeting member; the face parameter includes the number of face pixels in the portrait mask image; the method further includes: When it is determined that the real-time image of the meeting member does not belong to a valid image based on the number of human portrait pixels or the number of face pixels, the height offset and width offset of the previous image corresponding to the real-time image of the meeting member are obtained. The height and width offsets of the previous image are determined as the height and width offsets of the real-time images of the meeting members.

16. The method according to any one of claims 1 to 7 or 9 to 15, characterized in that, The determination of the portrait parameters and facial parameters of the meeting members in the real-time image of the meeting members includes: Image segmentation processing is performed on the real-time images of the meeting members to obtain the original portrait mask image and original face parameters corresponding to the real-time images of the meeting members. Based on the portrait mask image obtained by scaling the original portrait mask image according to the scaling ratio, portrait analysis is performed to obtain the portrait parameters corresponding to the portraits in the real-time screen of the meeting members; The facial parameters of the original image are scaled according to the scaling ratio to obtain the facial parameters corresponding to the faces in the real-time image of the meeting members.

17. The method according to claim 16, characterized in that, The step of adaptively adjusting the portraits of meeting members in the real-time image based on the adaptive adjustment parameters to obtain corresponding adaptive portraits includes: The adaptive adjustment parameters are transformed based on the scaling ratio to obtain the transformed adaptive adjustment parameters. Using the transformed adaptive adjustment parameters, the real-time images of the meeting members and the original portrait mask images corresponding to the real-time images of the meeting members are subjected to adaptive portrait adjustment to obtain the portrait adjustment image and the corresponding portrait mask adjustment image. The portrait adjustment image and the portrait mask adjustment image are combined to obtain a composite image; The synthesized image is scaled according to the size of the area where the virtual location is located, resulting in an adaptive portrait displayed at the virtual location.

18. A device for video conferencing interaction, characterized in that, The device includes: The real-time video acquisition module is used to acquire real-time video feeds of conference members collected by each terminal that has joined the video conference. The member parameter determination module is used to determine the image parameters and face parameters of each meeting member in the real-time image of each meeting member. The portrait analysis module is used to determine a height offset and a width offset based on the portrait parameters and the face parameters when the real-time image of the meeting member is determined to be a valid portrait image; to determine a scaling factor based on the portrait parameters, the face parameters, and the height offset; and to use the height offset, the width offset, and the scaling factor as adaptive adjustment parameters corresponding to the real-time image of the meeting member. An adaptive adjustment module is used to adaptively adjust the portraits of the meeting members in the real-time image of the meeting members based on the adaptive adjustment parameters, so as to obtain the corresponding adaptive portraits. The co-op display module is used to respond to the trigger operation of joining a video conference between multiple terminals and display the virtual co-op screen of the video conference. The virtual scene of the virtual co-op screen includes multiple virtual positions for accommodating the conference members. An adaptive human image display module is used to display adaptive human images corresponding to the human images in the real-time images of meeting members collected by at least two of the multiple terminals at at least two virtual locations in the virtual scene. The at least two virtual locations each display an adaptive portrait size that matches the size of the corresponding virtual location's location area.

19. The video conferencing interaction apparatus according to claim 18, characterized in that, The adaptive portrait display module includes a portrait area size determination module and a target portrait display module; the portrait area size determination module is used to determine the portrait area size that matches the corresponding virtual location based on the size of the location area of ​​the virtual location in the virtual scene and the preset portrait uniform proportion conditions; The target portrait display module is used to display, at a virtual location in the virtual scene, a target adaptive portrait corresponding to the portrait in the real-time image of the meeting members captured by any one of the multiple terminals, according to the corresponding portrait area size.

20. The video conferencing interaction apparatus according to claim 18, characterized in that, The adaptive portrait display module includes a distribution location determination module, a proportion condition processing module, and a target portrait display module; the distribution location determination module is used to determine the distribution location of each virtual location in the virtual scene; the proportion condition processing module is used to determine the size of the portrait area matching the corresponding virtual location according to the unified portrait proportion condition corresponding to each distribution location; The target portrait display module is used to display, at a virtual location in the virtual scene, a target adaptive portrait corresponding to the portrait in the real-time image of the meeting members captured by any one of the multiple terminals, according to the corresponding portrait area size.

21. The video conferencing interaction apparatus according to claim 18, characterized in that, The device further includes a screen sharing area display module and a screen sharing response module; the screen sharing area display module is used to display a virtual screen sharing area belonging to the virtual scene in the virtual frame-to-frame image; the screen sharing response module is used to respond to a target terminal among the multiple terminals joining the video conference triggering screen sharing, and to display the screen sharing content of the target terminal in the virtual screen sharing area.

22. The video conferencing interaction apparatus according to claim 21, characterized in that, The screen sharing area display module is also used to display the virtual screen sharing area in the background area other than the virtual position in the virtual frame image; the device also includes a blank screen display module, used to display a virtual blank screen in the virtual screen sharing area when the multiple terminals joining the video conference do not trigger screen sharing; and to display a virtual blank screen in the virtual screen sharing area in response to the target terminal triggering the cancellation of screen sharing.

23. The video conferencing interaction apparatus according to claim 18, characterized in that, The device further includes a display position movement module, which is used to respond to moving the first adaptive human image displayed at the first virtual position in the virtual scene to the second virtual position, and at the second virtual position, displaying the adaptive human image corresponding to the human image in the real-time picture of the meeting members collected by the terminal corresponding to the first adaptive human image.

24. The video conferencing interaction apparatus according to claim 18, characterized in that, The device also includes an abnormal status prompt module, which is used to respond to the abnormal participation status of a target meeting member in the real-time image of the meeting members collected by the target terminal among the multiple terminals. The module displays abnormal status prompt information about the target meeting member at the target virtual position of the adaptive human image corresponding to the human image in the real-time image of the meeting members collected by the target terminal.

25. The video conferencing interactive apparatus according to any one of claims 18 to 24, characterized in that, The device further includes a video interface display module, used to display the video interface of the video conference in response to a trigger operation of joining the video conference; in the video interface, real-time images of the conference members captured by at least one of the plurality of terminals are displayed; the same-frame display module is further used to cancel the display of the real-time images of the conference members in the video interface and display a virtual same-frame image of the video conference in response to triggering the same-frame mode of the video conference; the device further includes a same-frame mode end module, used to cancel the display of the virtual same-frame image in the video interface in response to the video conference meeting the conditions for ending the same-frame mode.

26. The video conferencing interaction apparatus according to claim 18, characterized in that, The portrait parameters include the portrait mask rectangle parameters in the portrait mask image corresponding to the real-time image of the meeting member; the portrait analysis module includes an initial offset determination module, a width offset determination module, and a height offset determination module; the initial offset determination module is used to determine the initial width offset and the initial height offset based on the portrait mask rectangle parameters and the size of the portrait mask image, respectively; The width offset determination module is used to calculate a reference width offset corresponding to the real-time image of the meeting member based on the historical width offset corresponding to the foreground image of the real-time image of the meeting member; and to correct the initial width offset using the reference width offset to obtain the width offset corresponding to the real-time image of the meeting member. The height offset determination module is used to calculate a reference height offset corresponding to the real-time image of the meeting member based on the historical height offset corresponding to the foreground image of the real-time image of the meeting member; and to correct the initial height offset using the reference height offset to obtain the height offset corresponding to the real-time image of the meeting member.

27. The video conferencing interaction apparatus according to claim 26, characterized in that, The reference width offset is the average of the historical width offsets corresponding to the previous frame of the real-time view of the meeting member; the width offset determination module includes a width difference determination module, a width correction module, and a width range limitation module; the width difference determination module is used to determine the difference between the initial width offset and the average of the historical width offsets; the width correction module is used to perform weighted correction on the initial width offset by using the historical width offset corresponding to the previous frame of the real-time view of the meeting member when the difference is greater than a first preset threshold, to obtain a corrected width offset; the width range limitation module is used to perform range limitation processing on the corrected width offset to obtain the width offset corresponding to the real-time view of the meeting member.

28. The video conferencing interaction apparatus according to claim 26, characterized in that, The reference height offset is the average of the historical height offsets corresponding to the previous frame of the real-time view of the meeting member; the height offset determination module includes a height difference determination module, a height correction module, and a height range limitation module; the height difference determination module is used to determine the difference between the initial height offset and the average of the historical height offsets; the height correction module is used to perform weighted correction on the initial height offset by using the historical height offset corresponding to the previous frame of the real-time view of the meeting member when the difference is greater than a second preset threshold, to obtain a corrected height offset; the height range limitation module is used to perform range limitation processing on the corrected height offset to obtain the height offset corresponding to the real-time view of the meeting member.

29. The video conferencing interaction apparatus according to claim 26, characterized in that, The device further includes an image center point determination module, a portrait center point determination module, a horizontal position difference determination module, and a horizontal position difference processing module; the image center point determination module is used to determine the image center point of the portrait mask image based on the size of the portrait mask image; the portrait center point determination module is used to determine the portrait center point of the portrait mask rectangle based on the portrait mask rectangle parameters and the size of the portrait mask image; the horizontal position difference determination module is used to determine the horizontal position difference between the image center point and the portrait center point; the horizontal position difference processing module is used to determine an initial width offset based on the horizontal position difference and the image width of the portrait mask image.

30. The video conferencing interaction apparatus according to claim 26, characterized in that, The device further includes a portrait position determination module, a portrait position processing module, and an initial height offset determination module; the portrait position determination module is used to determine the top and bottom positions of the portrait mask rectangle in the image height direction based on the portrait mask rectangle parameters; the portrait position processing module is used to determine a first initial height offset based on the top position and the image height of the portrait mask image; and to determine a second initial height offset based on the bottom position and the image height of the portrait mask image; the initial height offset determination module is used to obtain an initial height offset based on the first initial height offset and the second initial height offset.

31. The video conferencing interaction apparatus according to claim 18, characterized in that, The portrait parameters include the portrait mask rectangle parameters corresponding to the real-time images of the meeting members; the face parameters include the face mask rectangle parameters corresponding to the real-time images of the meeting members; the portrait analysis module includes a face scaling factor module, a portrait scaling factor module, and a factor fusion module; the face scaling factor module is used to determine the face scaling factor based on the face mask rectangle parameters and the height offset; The portrait scaling factor module is used to determine the portrait scaling factor based on the portrait mask rectangle parameters and the height offset; The factor fusion module is used to fuse the face scaling factor and the portrait scaling factor to obtain a scaling factor.

32. The video conferencing interaction apparatus according to claim 18, characterized in that, The portrait parameters include the number of portrait pixels in the portrait mask image corresponding to the real-time image of the meeting member; the face parameters include the number of face pixels in the portrait mask image; the device further includes an invalid image offset determination module, used to obtain the height offset and width offset of the previous image corresponding to the real-time image of the meeting member when it is determined based on the number of portrait pixels or the number of face pixels that the real-time image of the meeting member does not belong to a valid portrait image; and to determine the height offset and width offset of the previous image as the height offset and width offset of the real-time image of the meeting member.

33. The apparatus for video conferencing interaction according to any one of claims 18 to 24 or 26 to 32, characterized in that, The member parameter determination module includes an image segmentation module, a portrait parameter acquisition module, and a face parameter acquisition module. The image segmentation module performs image segmentation processing on the real-time image of the meeting members to obtain the original portrait mask image and original face parameters corresponding to the real-time image of the meeting members. The portrait parameter acquisition module performs portrait analysis based on the portrait mask image obtained by scaling the original portrait mask image according to the scaling ratio to obtain the portrait parameters corresponding to the portraits in the real-time image of the meeting members. The face parameter acquisition module scales the original face parameters according to the scaling ratio to obtain the face parameters corresponding to the faces in the real-time image of the meeting members.

34. The video conferencing interaction apparatus according to claim 33, characterized in that, The adaptive adjustment module includes a parameter transformation module, an image adjustment module, an image compositing module, and a scaling module. The parameter transformation module transforms the adaptive adjustment parameters based on the scaling ratio to obtain transformed adaptive adjustment parameters. The image adjustment module performs adaptive portrait adjustment on the real-time image of the meeting member and the corresponding original portrait mask image using the transformed adaptive adjustment parameters to obtain an adjusted portrait image and a corresponding adjusted portrait mask image. The image compositing module combines the adjusted portrait image and the adjusted portrait mask image to obtain a composite image. The scaling module scales the composite image according to the size of the area where the virtual location is located to obtain an adaptive portrait displayed at the virtual location.

35. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 17.

36. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 17.

37. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the method of any one of claims 1 to 17.