Virtual conference processing method, device, equipment and storage medium

By collecting and transmitting facial expression information on the client side during virtual meetings, the problem of poor video synchronization in multi-user online meetings is solved, and efficient facial expression information transmission and immersive experience are achieved.

CN115086594BActive Publication Date: 2025-09-19ALIBABA (CHINA) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202210520743.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-12
Publication Date
2025-09-19
Estimated Expiration
2042-05-12

AI Technical Summary

Technical Problem

In multi-user online meetings, video synchronization is difficult to ensure, especially when the amount of data is large, which is prone to lag, affecting the immersion and efficiency of virtual meetings.

Method used

By collecting the user's facial image on the client, extracting the expression information and sending it to the server for aggregation and synchronization, the client only transmits the expression information to reduce the amount of data, and uses the server to synchronize the expression information of multiple users.

Benefits of technology

It realizes the real-time transmission of expression information, reduces the server processing load, and ensures the transmission timeliness and immersive experience of virtual meetings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115086594B_ABST
    Figure CN115086594B_ABST
Patent Text Reader

Abstract

The present application provides a virtual conference processing method, apparatus, device and storage medium, which is applied to the first client of the first user among multiple users participating in the virtual conference, including: displaying a virtual conference interface including virtual avatars of multiple users, determining the expression information of the first user based on the facial image of the first user; sending the expression information of the first user to the server, so that the server aggregates the expression information of multiple users and synchronizes the expression information of the multiple users to the first client and the second client, the second client corresponding to the second user among the multiple users except the first user. Receive the expression information of the second user sent by the server, locally drive the virtual avatar corresponding to the first user according to the expression information of the first user, and locally drive the virtual avatar corresponding to the second user according to the expression information of the second user. The effect of driving the virtual avatar according to the real expression is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet technology, and in particular to a virtual conference processing method, device, equipment and storage medium. Background Art

[0002] Currently, there are three main methods for online real-time communication: text, audio, and video. Video communication is widely adopted due to its advantages such as real-time, high efficiency, and more realistic experience.

[0003] For example, when multiple users participate in an online meeting, to enhance the immersive experience, each user's client can capture their facial video and send it to a server. The server then synchronizes it with the clients of all other users, thus synchronizing each user's facial video across all participating clients. However, when there are many participants, the amount of facial video required to be transmitted is large, often leading to lags and other issues, making it difficult to ensure facial video synchronization. Summary of the Invention

[0004] The embodiments of the present invention provide a virtual conference processing method, apparatus, device and storage medium, which are used to realize a more realistic virtual conference scene and ensure the timeliness of information transmission of the virtual conference.

[0005] In a first aspect, an embodiment of the present invention provides a virtual conference processing method, which is applied to a first client of any first user among multiple users participating in a virtual conference, the method comprising:

[0006] Displaying a virtual conference interface, wherein the virtual conference interface includes virtual avatars corresponding to the multiple users;

[0007] Acquire the facial image of the first user according to a set sampling frequency;

[0008] determining the facial expression information of the first user based on the facial image of the first user;

[0009] Sending the expression information of the first user to a server, so that the server aggregates the expression information of the multiple users and synchronizes the expression information of the multiple users to the first client and a second client, where the second client corresponds to a second user among the multiple users except the first user;

[0010] Receive the expression information of the second user sent by the server.

[0011] In a second aspect, an embodiment of the present invention provides a virtual conference processing device, which is applied to a first client of any first user among multiple users participating in a virtual conference, and the device includes:

[0012] A display module, configured to display a virtual conference interface, wherein the virtual conference interface includes virtual avatars corresponding to the plurality of users;

[0013] a determination module, configured to obtain a facial image of the first user according to a set sampling frequency; and determine the facial expression information of the first user based on the facial image of the first user;

[0014] a sending module, configured to send the expression information of the first user to a server, so that the server aggregates the expression information of the multiple users and synchronizes the expression information of the multiple users to the first client and a second client, where the second client corresponds to a second user among the multiple users except the first user;

[0015] The receiving module is used to receive the expression information of the user sent by the server.

[0016] In a third aspect, an embodiment of the present invention provides an electronic device comprising: a memory, a processor, a communication interface, and a display; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor can at least implement the virtual conference processing method described in the first aspect.

[0017] In a fourth aspect, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the virtual conference processing method described in the first aspect.

[0018] In a fifth aspect, an embodiment of the present invention provides a virtual conference processing method, which is applied to a first extended reality device of any first user among multiple users participating in the virtual conference, including:

[0019] Displaying a virtual conference interface, wherein the virtual conference interface includes virtual avatars corresponding to the multiple users;

[0020] Acquire the facial image of the first user according to a set sampling frequency;

[0021] determining the facial expression information of the first user based on the facial image of the first user;

[0022] Sending the expression information of the first user to a server, so that the server aggregates the expression information of the multiple users and synchronizes the expression information of the multiple users to the first extended reality device and a second extended reality device, where the second extended reality device corresponds to a second user among the multiple users except the first user;

[0023] Receive the expression information of the second user sent by the server.

[0024] In an embodiment of the present invention, when multiple users enter the same virtual conference through their respective clients, they can select their own virtual avatars to represent themselves. Taking any one of the users (such as the first user) as an example, the first client of the first user can display a virtual conference interface (or virtual conference space, virtual conference room) corresponding to the virtual conference, and the virtual conference interface includes the virtual avatars corresponding to the above-mentioned multiple users. The first client samples the facial image of the first user captured by the camera of the terminal device according to the set sampling frequency, and determines the expression information used to drive the virtual avatar of the first user based on the facial image of the first user, and sends the expression information of the first user to the server. In this way, the server can summarize the expression information of each current user and send the summary result to each client. In this way, the first client can not only obtain the expression information of the first user, but also obtain the expression information of other second users.

[0025] As can be seen, in an embodiment of the present invention, in a virtual conference scenario, multiple users participating in the meeting can select virtual avatars to represent themselves. To obtain the facial expression information used to drive each user's avatar, each client collects the corresponding user's facial image in real time, extracts the facial expression information, and sends the extracted facial expression information to the server, which then aggregates and sends it to each client. Because the transmission bandwidth required for facial expression information is relatively small, facial expression information can be synchronized to each client via the server in a more real-time manner, thereby ensuring the timely transmission of facial expression information and reducing the processing load on the server. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0027] Figure 1 A schematic diagram of a virtual conference system provided by an embodiment of the present invention;

[0028] Figure 2 A flowchart of a virtual conference processing method provided by an embodiment of the present invention;

[0029] Figure 3 A schematic diagram of a virtual conference interface provided by an embodiment of the present invention;

[0030] Figure 4 A flowchart of a method for determining an expression coefficient provided by an embodiment of the present invention;

[0031] Figure 5 A flowchart of another virtual conference processing method provided by an embodiment of the present invention;

[0032] Figure 6 A flowchart of another virtual conference processing method provided by an embodiment of the present invention;

[0033] Figure 7 A schematic diagram of an interactive virtual conference interface provided by an embodiment of the present invention;

[0034] Figure 8 A schematic diagram of an interactive virtual conference interface provided by an embodiment of the present invention;

[0035] Figure 9 A schematic diagram of an interactive virtual conference interface provided by an embodiment of the present invention;

[0036] Figure 10 A schematic diagram of an application of a virtual conference processing method provided by an embodiment of the present invention;

[0037] Figure 11 A schematic diagram of the structure of a virtual conference processing device provided by an embodiment of the present invention;

[0038] Figure 12 A schematic structural diagram of an electronic device provided by an embodiment of the present invention;

[0039] Figure 13 A schematic structural diagram of an extended reality device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0041] The following describes some embodiments of the present invention in detail with reference to the accompanying drawings. The following embodiments and features thereof may be combined with one another unless they conflict with each other. Furthermore, the sequence of steps in the following method embodiments is provided for illustrative purposes only and is not intended to be a strict limitation.

[0042] The purpose of a virtual meeting is to replace face-to-face meetings, that is, to move meetings from face-to-face to virtual mode so that people do not have to go to the physical location of the meeting but can participate at any time and from anywhere.

[0043] In order to enable users to obtain a more immersive experience when using the virtual conference function, an embodiment of the present invention provides a virtual conference solution driven by virtual image expressions, so that when using virtual conferences, users can perceive the expression changes of each user participating in the meeting through the virtual image even if they do not directly capture and display their real avatar images in the virtual conference space.

[0044] The solution provided by the embodiment of the present invention can be executed by a client provided with a virtual conference function.

[0045] With multiple users (such as Figure 1 Take the scenario where N users (shown in FIG) participate in a virtual meeting as an example, Figure 1 The figure illustrates the composition of a virtual conference system: it consists of multiple clients corresponding to multiple users and a server. Each client can see a virtual conference interface (also called a virtual conference space or virtual conference room) corresponding to the virtual conference. The server is used to transmit information with each client.

[0046] Figure 2 A flowchart of a virtual conference processing method provided by an embodiment of the present invention is shown in FIG. Figure 2 As shown, the virtual conference processing method includes the following steps:

[0047] 201. A first client displays a virtual conference interface, where the virtual conference interface includes virtual avatars corresponding to multiple users.

[0048] 202. The first client obtains a facial image of the first user according to a set sampling frequency, and determines expression information of the first user based on the facial image of the first user.

[0049] 203. The first client sends the expression information of the first user to the server, so that the server aggregates the expression information of multiple users and synchronizes the expression information of the multiple users to the first client and a second client, where the second client corresponds to a second user among the multiple users except the first user.

[0050] 204. The first client locally drives a virtual avatar corresponding to the first user according to the expression information of the first user.

[0051] 205. The first client receives the expression information of the second user sent by the server, and locally drives the virtual avatar corresponding to the second user according to the expression information of the second user.

[0052] In this embodiment, assume that multiple users (e.g., N users) participate in the same conference through their respective clients. Any of these users is referred to as the first user, and their corresponding client is referred to as the first client. The other users are referred to as second users, and their clients are referred to as second clients. For ease of description, this embodiment uses the first client as an example to illustrate the execution process for each client. That is, the execution process for the first client is also executed by the other clients.

[0053] It is understandable that in actual applications, a virtual meeting can be organized and created in advance by one of multiple users. When creating the meeting, the user IDs of the users participating in the meeting, the meeting time, the meeting access information and other information will be entered. When the meeting time arrives, multiple users will access the virtual meeting.

[0054] Taking the first user as an example, when the first user accesses the virtual conference through the first client, a virtual conference interface may be displayed, such as Figure 3 As shown, the virtual conference interface can include some environmental objects that simulate a real conference room scene, such as the conference table, display screen, and seats shown in the figure. The number of seats can be set based on the number of participants in the meeting. For example, if the number of participants is set to 8 when creating a meeting, 8 seats will be displayed in the virtual conference interface. These seats can be displayed in corresponding positions according to the set placement rules.

[0055] In addition, Figure 3 As shown in , a virtual avatar can also be displayed on each seat. When the first user accesses the virtual conference through the first client, he can select a seat from multiple seats. At this time, the selected seat can be associated with the user ID of the first user, so that other users entering the virtual conference can see the seat selected by the first user and choose other unselected seats. In addition, as Figure 3 As shown in , optionally, the first user can select a seat by clicking on the virtual avatar displayed on the seat to be selected. At this time, a virtual avatar list can be displayed, which includes multiple selectable virtual avatars. After the first user selects a virtual avatar and confirms it, the virtual avatar will be used as the virtual avatar selected by the first user. At this time, the selected virtual avatar and the seat that triggered the avatar selection both correspond to the first user. That is to say, the first user can trigger the virtual avatar selection operation through the virtual avatar initially displayed on a seat. According to the virtual avatar selection result, the seat and virtual avatar selected by the first user can be known.

[0056] In actual applications, the virtual avatars initially displayed on each seat can be the same or different. Whether different users are allowed to select the same virtual avatar can be set based on the number of virtual avatars included in the virtual avatar list and the total number of participating users.

[0057] It is understandable that after other users enter the virtual conference, they will first perform the selection operation of their corresponding virtual avatars, and based on the selection results, Figure 3 In the virtual conference interface shown, different seats are associated with different user IDs and avatars. Each avatar is displayed in a corresponding frame, such as the rectangular frame shown in the figure. These avatars can be pre-generated 3D avatars, and their expressions are initially set to default.

[0058] During the meeting, the user can also change his or her virtual avatar, that is, reselect his or her virtual avatar. The selection operation is as described above.

[0059] In actual applications, the above-mentioned virtual conference interface can be a virtual conference space environment (as the above-mentioned virtual conference interface) generated by the server based on technologies such as virtual reality (VR) and augmented reality (AR). When the user terminal device installed by the above-mentioned client supports VR and AR technologies, such as a VR helmet or other device, the virtual conference interface containing the above-mentioned three-dimensional objects can be seen through this user terminal device. If the user terminal device does not support the display of three-dimensional images, the three-dimensional images can be converted into two-dimensional images for display.

[0060] After the first user enters the virtual meeting through the first client, they can select their avatar and manually or automatically activate the camera on their terminal device to capture their face. Manually activating the camera function means that a button for activating the camera function can be provided in the virtual meeting interface, and the first user can control the button. Automatically activating the camera function means that after the first user enters the virtual meeting, the first client automatically activates the button.

[0061] After the camera is turned on, it can be configured to continuously collect facial video data and transmit it to the first client. The first client can sample the facial video data at a set sampling frequency to obtain frames of facial images. The sampling interval can be set according to actual needs, and can be set to 50 milliseconds, etc.

[0062] Since the subsequent operations performed on each frame of the face image are the same, the embodiment of the present invention only takes any frame of the face image as an example for description, and for the convenience of description, they are collectively referred to as face images.

[0063] Since the facial image is a real facial image obtained by photographing the first user, expression information reflecting the current expression state of the first user can be extracted from it, so as to drive the first user's virtual avatar in real time based on the expression information, so that the virtual avatar presents an expression state that matches the expression information. In this way, since different users have different current expressions, each user drives his or her own virtual avatar locally according to his or her current expression information through his or her own client, and ultimately each user can see each user's virtual avatar presenting a different expression state through the virtual conference interface, thereby obtaining an immersive conference experience. In summary, the driving of the virtual avatar in the embodiment of the present invention refers to adjusting the expression of the virtual avatar according to the expression information.

[0064] In order to allow other users to see the current expression state of the first user, in an embodiment of the present invention, the first client not only locally drives the first user's virtual avatar based on the first user's expression information, but also sends the first user's expression information to the server. It is understandable that each second client performs the same process and sends the corresponding second user's expression information obtained at this time to the server. The server summarizes the expression information sent by each client and sends the summary result to each client. In this way, each client can obtain the expression information of other users and locally drive the virtual avatars of other users based on the expression information of other users.

[0065] For ease of understanding, for example, suppose user 1, user 2, and user 3 participate in a virtual meeting. At a certain moment, the clients of each of these three users obtain the expression information of the corresponding user and send it to the server respectively. The server can send the expression information of user 2 and user 3 to user 1. Similarly, the expression information of user 1 and user 2 can be sent to user 3. For ease of processing, the server can also aggregate the expression information of user 1, user 2, and user 3 and send it to user 1, user 2, and user 3 respectively. At this time, after user 1's client receives the expression information of the three users, it can drive the virtual avatars of the three users accordingly based on the expression information of the three users locally. It can be seen that this virtual meeting scenario is a many-to-many communication method. Among them, for user 1, its client can download the virtual avatar corresponding to each user from the server in advance, so that the virtual avatar of each user is stored locally.

[0066] In combination with the above examples, it can be seen that the execution timing of the above steps 203 and 204 is not strictly limited. That is to say, the first client can locally drive the virtual avatar of the first user based on the facial expression information of the first user, and at the same time send the facial expression information of the first user to the server, which is synchronized to other second clients by the server. It can also send the facial expression information of the first user to the server first when it is collected, and then after receiving the aggregated facial expression information sent by the server, execute the operation of locally driving the virtual avatar of the first user according to the facial expression information of the first user, and locally driving the virtual avatars of other second users according to the facial expression information of other second users, thereby completing the image processing of rendering and displaying the virtual avatar according to the facial expression information.

[0067] Since only expression information needs to be transmitted between the client and the server during the above interaction process, and the amount of expression information is relatively small, the transmission delay between the client and the server is very small and can be ignored. This allows the expressions to be driven synchronously between the clients, ensuring transmission timeliness.

[0068] It is understandable that during a meeting, the information that needs to be transmitted between the server and the client includes, in addition to the above-mentioned expression information, audio data, that is, the user's speech during the meeting. The transmission of audio data is still synchronously transmitted between the server and the client in the traditional way. I just want to emphasize one point here: assuming that the first user is currently speaking, in order to ensure the consistency of the expression driving effect and pronunciation mouth shape of the first user's virtual avatar, each client needs to ensure the alignment of the audio data and the virtual avatar expression driving effect when driving the first user's virtual avatar.

[0069] The following describes an optional implementation method for determining the first user's facial expression information based on the first user's facial image: extracting multiple facial key points from the first user's facial image and determining the first user's facial expression coefficient based on the multiple facial key points. In other words, based on the plurality of facial key points, corresponding facial expression coefficients can be determined as the first user's facial expression information. In practice, many types of facial expression coefficients (e.g., 52) may be included, with different facial expression coefficients used to adjust facial expressions at different locations on the face.

[0070] The multiple facial key points extracted from the face image actually include key points corresponding to different facial regions, such as forehead, eyebrows, eyes, nose, mouth, cheeks, etc. In the embodiment of the present invention, different methods for determining expression coefficients are provided for key points in different facial regions. Figure 4 .

[0071] Figure 4 A flowchart of a method for determining an expression coefficient provided by an embodiment of the present invention is shown in FIG. Figure 4 As shown, the following steps may be included:

[0072] 401. Input key points of the first facial region into an expression coefficient prediction model to obtain expression coefficients corresponding to the key points of the first facial region.

[0073] 402. Obtain a preset expression coefficient mapping relationship corresponding to the second facial region, where the preset expression coefficient mapping relationship is used to reflect a mapping relationship between the expression coefficient of a target category and the distance between corresponding target key points.

[0074] 403. Determine, based on the key points of the second facial region, a target inter-key point distance value corresponding to the target type expression coefficient.

[0075] 404. Determine expression coefficients corresponding to the key points of the second facial region according to the distance value between the target key points and the preset expression coefficient mapping relationship.

[0076] In this embodiment, two methods are used to jointly complete the determination of the expression coefficient of the first user. One method is to use the expression coefficient prediction model of deep learning, and the other method is to adopt the mapping rule method.

[0077] These two methods are suitable for different facial regions. Generally speaking, facial regions with relatively simple expression changes and high expression accuracy requirements are more suitable for the mapping rule method, such as the upper part of the face, such as the eyes and eyebrows. Facial regions with more complex expression changes and involving more expression coefficients, such as the mouth and cheeks, are more suitable for the prediction model method.

[0078] For example, left and right eyes and left and right eyebrows often only involve blink amplitude (i.e., the size of the eye opening), up and down, left and right movement of the eyeballs, and movement of the eyebrows. The mouth changes state during speech, and even when not speaking, there are many habitual movements, such as pursing lips and yawning.

[0079] Regarding the expression coefficient prediction model: In practical applications, a corresponding expression coefficient prediction model can be trained for each facial area that needs to use the expression coefficient prediction model, such as the expression coefficient prediction model corresponding to the mouth area and the expression coefficient prediction model corresponding to the cheek area. This can more accurately predict the expression coefficient of the corresponding facial area and make model training easier.

[0080] The input of the expression coefficient prediction model is the multiple key points extracted from the corresponding facial area, and the output is the predicted various expression coefficients corresponding to the facial area.

[0081] In practical applications, the structure of the expression coefficient prediction model can be implemented as a neural network model consisting of multiple (e.g., five) fully connected layers. The training samples used in the expression coefficient prediction model training process can be obtained as follows: multiple facial images are acquired, and expression coefficients of target facial regions (e.g., the aforementioned mouth and cheek regions) in the facial images are identified using known expression coefficient recognition software to obtain expression coefficients corresponding to the target facial regions. Furthermore, key point extraction is performed on the facial images to obtain key points corresponding to the target facial regions. In this way, the key points corresponding to the target facial regions serve as training samples, and the expression coefficients corresponding to the target facial regions serve as supervisory information for the training samples, which are then used to train the expression coefficient prediction model corresponding to the target facial regions.

[0082] Regarding the mapping relationship of expression coefficients: Specifically, a certain facial region may correspond to multiple types of expression coefficients, each of which can have a mapping relationship. This mapping relationship can be represented by a mapping function curve, where one axis represents the value of the expression coefficient, and the other axis represents the distance between key points corresponding to the expression coefficient. For example, assuming that one expression coefficient is the degree of left eye blinking, the key point distance information corresponding to this expression coefficient can be the ratio of the left eye height to the eye width. The eye height can be represented by the average distance between the key points corresponding to the upper and lower eye boundaries, and the eye width can be represented by the distance between the key points corresponding to the left and right eye boundaries. A large number of image samples containing eyes in different blinking states can be collected in advance. Expression recognition can be performed on these image samples to obtain the expression coefficient values ​​corresponding to the degree of left eye blinking. The corresponding eye key points can then be extracted and the aforementioned ratios calculated. This results in a large number of coordinate pairs, each consisting of an expression coefficient value and a ratio. Fitting these large number of obtained coordinate pairs can yield an expression coefficient mapping function curve corresponding to the degree of left eye blinking.

[0083] Taking the above-mentioned expression coefficient mapping function curve as an example, multiple key points corresponding to the left eye are extracted from the facial image of the first user, and then the ratio of the eye height to the eye width of the left eye is calculated. The ratio is located in the above-mentioned expression coefficient mapping function curve, and the function value corresponding to the ratio - the expression coefficient value - is determined.

[0084] In summary, the expression coefficients for different facial regions are determined by using mapping rules and deep learning models. Mapping rules provide more accurate results, while deep learning models facilitate predictions. Based on the strengths of both approaches and the varying characteristics of facial expressions, different facial regions are configured and determined using different methods, achieving a balanced balance between accuracy and processing complexity.

[0085] Figure 5 A flowchart of another virtual conference processing method provided by an embodiment of the present invention is shown in FIG. Figure 5 As shown, the method includes the following steps:

[0086] 501. A first client displays a virtual conference interface, where the virtual conference interface includes virtual avatars corresponding to multiple users.

[0087] 502. The first client obtains a facial image of the first user according to a set sampling frequency, extracts a plurality of facial key points from the facial image of the first user, and determines an expression coefficient, head posture information, and head displacement information of the first user based on the plurality of facial key points.

[0088] In this embodiment, in addition to calculating the expression coefficient based on facial key points, the first user's head posture information and / or head displacement information can also be determined based on the facial key points. The head posture information primarily refers to the rotation direction and angle of the first user's head. A rotation matrix can be calculated based on the facial key points, and the posture information can be obtained based on the rotation matrix. For information on calculating the rotation matrix, reference can be made to existing related technologies and will not be elaborated on here.

[0089] Head displacement information refers to the positional movement of the first user's head within the facial image, including the direction and distance of movement. It is understood that the position of the first user's head may be inconsistent between two adjacent facial frames, and this inconsistency is reflected in the positional movement information. In practice, this positional movement information can be determined by comparing the coordinates of key points in the previous facial frame with the coordinates of the corresponding key points in the next facial frame.

[0090] 503. The first client sends the facial expression information, head posture information, and head displacement information of the first user to the server, so that the server aggregates the facial expression information, head posture information, and head displacement information of multiple users, and synchronizes the facial expression information, head posture information, and head displacement information of the multiple users to the first client and the second client, where the second client corresponds to the second user among the multiple users except the first user.

[0091] In this embodiment, in addition to the expression coefficient of the corresponding user determined by each client, the synchronous transmission between the client and the server also includes the head posture information and head displacement information of the corresponding user determined by each client.

[0092] 504. The first client locally drives the virtual avatar corresponding to the first user according to the facial expression information of the first user, adjusts the posture of the virtual avatar of the first user according to the head posture information of the first user, and adjusts the display position of the virtual avatar of the first user in the corresponding display window according to the head displacement information of the first user.

[0093] 505. The first client receives the facial expression information, head posture information, and head displacement information of the second user sent by the server, locally drives the virtual avatar corresponding to the second user according to the facial expression information of the second user, adjusts the posture of the virtual avatar of the second user according to the head posture information of the second user, and adjusts the display position of the virtual avatar of the second user in the corresponding display window according to the head displacement information of the second user.

[0094] Taking the first client as an example, the first client will not only render and display the virtual avatar of each user locally according to the expression coefficient of each user, but will also change the rotation direction and display position of the corresponding user's virtual avatar according to the head posture information and head displacement information of each user, presenting a dynamic update effect in which the virtual avatar of each user changes with the user's real facial expression, posture and position.

[0095] Figure 6 A flowchart of another virtual conference processing method provided by an embodiment of the present invention is shown in FIG. Figure 6 As shown, the method includes the following steps:

[0096] 601. A first client displays a virtual conference interface, where the virtual conference interface includes virtual avatars corresponding to multiple users.

[0097] 602. The first client obtains a facial image of the first user according to a set sampling frequency, and determines expression information of the first user based on the facial image of the first user.

[0098] 603. If the first user is a speaker, the first client analyzes the voice data of the first user to determine the topic type corresponding to the speech content of the first user, and determines whether the expression information of the first user matches the topic type. If so, execute step 604; if not, execute step 605.

[0099] 604. The first client sends the expression information of the first user to the server, so that the server aggregates the expression information of multiple users and synchronizes the expression information of the multiple users to the first client and a second client, where the second client corresponds to a second user among the multiple users except the first user.

[0100] 605. The first client sends the set expression information corresponding to the topic type as the expression information of the first user to the server, so that the server aggregates the expression information of multiple users and synchronizes the expression information of multiple users to the first client and the second client.

[0101] 606. The first client receives the expression information of the second user sent by the server, and locally drives the virtual avatar corresponding to the first user according to the expression information of the first user, and locally drives the virtual avatar corresponding to the second user according to the expression information of the second user.

[0102] In this embodiment, assuming that the first user is currently speaking, after the first user starts speaking, the first client analyzes the voice data sent by the first user to determine the topic type corresponding to the content of the speech.

[0103] Alternatively, the speech data can be converted into text, which is then fed into a pre-set neural network model for predicting topic types. The neural network model then outputs a topic type prediction. In practical applications, different types of topics can be pre-set, and training samples corresponding to each topic can be collected to train the model. For example, topic types could include work reports, free discussions, and other topics, or serious topics, entertainment topics, and so on.

[0104] Optionally, topics can also be identified by keywords. Specifically, common keywords corresponding to different topic types can be pre-set. If the first user's speech content contains keywords corresponding to a certain topic type, the topic type is determined to be the topic type corresponding to the speech content.

[0105] In addition, each topic type can be pre-configured with corresponding expression information, such as setting a value range for each expression coefficient corresponding to each topic type.

[0106] Thus, during the period when the first user is speaking, for each frame of the first user's facial image sampled during the period, after obtaining the first user's expression coefficient according to the scheme introduced in the aforementioned embodiment, the determined expression coefficient of the first user can be compared with the expression coefficient value range corresponding to the topic type currently being spoken by the first user. If it is within the value range, it is considered that the expression coefficient of the first user matches the topic type; otherwise, it does not match.

[0107] If there is a match, the expression coefficient of the first user obtained from the face image will be directly sent to the server; otherwise, a set expression coefficient corresponding to the topic type will be obtained based on the expression coefficient value range corresponding to the topic type (for example, a set of expression coefficients within the value range will be randomly generated), and the set expression coefficient will be sent to the server as the expression information of the first user.

[0108] Based on the solution provided in this embodiment, it is possible to achieve an effect in which the virtual avatar expressions of each user in the virtual conference space match the topic type in the conference, creating an immersive atmosphere in which the content and avatar expressions are more harmonious.

[0109] In addition to the aforementioned topic types, factors that may affect the availability of the first user's facial expression information may also include, for example, the first user's role type. Specifically, correspondences between different role types and expression coefficient value ranges can be pre-set. If the first user belongs to role a, but the currently obtained expression coefficient for the first user does not match the expression coefficient value range corresponding to role a, a set expression coefficient is generated based on the value range to replace the expression coefficient for the first user obtained from the facial image. The user's role can be configured when creating a virtual meeting.

[0110] In each of the above embodiments, the first user's facial expression information (expression coefficient) is obtained based on the first user's facial image. In an optional embodiment, the first user's facial expression information can also be obtained by receiving an expression keyword input by the first user and generating the first user's facial expression information (expression coefficient) based on the expression keyword.

[0111] The corresponding relationship between different expression keywords and expression coefficients can be preset, and the user can be prompted to select the expression keyword to be input for selection as needed.

[0112] In actual applications, there are sometimes situations where the user terminal device does not have a camera. For example, the terminal device used by the first user is a PC without a camera. At this time, the first user can dynamically adjust the expression of his or her virtual avatar in the virtual conference interface by entering expression keywords.

[0113] Regardless of the method used to obtain the first user's facial expression information, the first user may need to adjust the facial expression information. Based on this, in an optional embodiment, the following facial expression information adjustment method is provided:

[0114] First, in the first client, a configuration sub-interface visible only to the first user can be displayed in association with the virtual conference interface, and the configuration sub-interface includes expression configuration items corresponding to the expression information of the first user and the virtual avatar of the first user; thereafter, in response to the first user's configuration adjustment operation on the expression configuration item, the virtual avatar corresponding to the first user is driven in the configuration sub-interface according to the expression information updated after the configuration adjustment operation; thereafter, in response to the first user's confirmation operation on the updated expression information, the virtual avatar corresponding to the first user displayed in the configuration sub-interface is migrated and displayed in the virtual conference interface.

[0115] For ease of understanding, combined Figure 7 Exemplary description. Figure 7In the example, a virtual conference interface 701 is displayed on the first client of the first user, and multiple virtual avatars corresponding to the users are displayed in the virtual conference interface 701, including the virtual avatar A corresponding to the first user. Assume that a set of expression coefficients B1 is extracted based on the face image of the first user. At this time, optionally, as Figure 7 As shown in FIG, the first client may display a configuration sub-interface 702 as shown in the figure. This configuration sub-interface 702 is visible only to the first user, that is, this configuration sub-interface 702 is not synchronized to other clients via the server and is only displayed in the first client. This configuration sub-interface 702 includes a first area for displaying expression configuration items and a second area for displaying the first user's virtual avatar.

[0116] Among them, such as Figure 7 As shown in , a set of expression coefficients B1 consists of several expression coefficients (e.g., expression coefficient 1, expression coefficient 2, etc.). Each expression coefficient corresponds to an adjustment bar and a numerical axis representing the value range of the expression coefficient. This numerical axis and adjustment bar constitute the expression configuration item corresponding to the expression coefficient. Thus, the first user's expression information will correspond to multiple expression configuration items, each used to adjust the various expression coefficients that constitute the expression information.

[0117] The virtual avatar in the second area mentioned above can be copied to the virtual conference interface 701, and the virtual avatar in the second area can be driven based on multiple expression coefficients obtained from the facial image of the first user, so that the virtual avatar presents a corresponding expression. Afterwards, if the first user wants to adjust the expression by watching the driving effect, the adjustment bars corresponding to some expression coefficients can be adjusted in the first area to update the corresponding expression coefficient values. As the expression coefficient values ​​are updated, the expression of the virtual avatar in the second area can be updated. When the user adjusts the expression of the virtual avatar to his satisfaction, he can click the confirmation button set in the first area of ​​the figure to trigger the confirmation operation. Assuming that a new set of expression coefficients B2 is formed at this time, the virtual avatar driven based on the expression coefficient B2 displayed in the configuration sub-interface 702 is copied to the virtual avatar display position corresponding to the first user in the virtual conference interface 701 for replacement.

[0118] The above embodiments introduce some aspects of driving virtual avatar expressions in a virtual conference interface. In addition to performing operations related to the virtual avatar's expressions, other interactive operations can also be performed in the virtual conference interface. This will be explained below with reference to the following embodiments.

[0119] As mentioned above, in addition to the virtual avatar, the virtual conference interface can also include some objects corresponding to the conference scene, such as a virtual display screen, a conference table, etc.

[0120] When the virtual conference interface includes a virtual display screen, in an optional embodiment, information sharing can be implemented in conjunction with the virtual display screen, thereby achieving a simulation effect of projecting shared content onto a real conference terminal screen in a real conference room.

[0121] Specifically, still taking the first client as an example, in response to the information sharing operation triggered by the first user, the shared content is presented on the virtual display screen, and the virtual display screen presenting the shared content is synchronized to the second client through the server.

[0122] For ease of understanding, combined Figure 8 Exemplary description. Figure 8 In the example, the virtual conference interface includes a virtual display screen 801 and an action bar. Based on the various interactive functions provided in the action bar, users can trigger various actions. The action bar includes a share button 802 for triggering information sharing. By triggering this share button 802, a first user can select shared content 803 to be shared with all users. The first client then renders and displays the shared content 803 on the virtual display screen 801. The first user can then view the shared content 803 displayed on the virtual display screen 801 through the first client. To allow other users to also view the shared content 803, the first client takes a screenshot of the virtual display screen 801 containing the shared content 803 and sends it to a server. The server sends the screenshot to each second client, which then renders and displays the screenshot at the location of the virtual display screen in the local display of the virtual conference interface. It will be appreciated that after the first user triggers the information sharing action, the first client can dynamically synchronize the virtual display screen 801 containing the shared content with other clients at a set sampling frequency or based on changes in the content displayed on the virtual display screen 801.

[0123] In addition to the above-mentioned information sharing function, a discussion group function may also be provided in the virtual conference interface to meet the needs of some of the multiple users for group discussion during the conference.

[0124] Still taking the first client as an example, optionally, the method further includes:

[0125] Displaying a discussion group including at least two corresponding users in a virtual conference interface according to the discussion group creation information triggered by the first user, wherein the at least two users include the first user;

[0126] Synchronizing the discussion group creation information to the second client via the server, so that the second client generates the discussion group;

[0127] In response to the operation of switching to the discussion group triggered by the first user, a discussion group conference interface is displayed on the first client, wherein the discussion group conference interface includes virtual avatars of the at least two users migrated from the virtual conference interface.

[0128] For ease of understanding, combined Figure 9 For example, Figure 9 As described in the above, the operation bar in the virtual conference interface 900 may include a button 901 for creating a discussion group. The first user triggers the operation of creating a discussion group through the button and can enter information such as the discussion group name and discussion group members. In this embodiment, it is assumed that the first user chooses to create a discussion group with the second user and the third user, and the name is discussion group 1. After creating discussion group 1, as shown in FIG. Figure 9 As shown in FIG, the first client can display a conference list pop-up box 902 in the local virtual conference interface 900, which displays the various discussion groups currently in the virtual conference. It should be noted that the initial virtual conference participated by all the above multiple users can also be regarded as a special discussion group. The name of this discussion group can be configured by default, such as Figure 9 At the same time, the first client also sends the creation information of the discussion group 1 to the server, which then sends it to each second client. In this way, each second client will also display the conference list pop-up box 902 shown in the figure in the local virtual conference interface.

[0129] The first user can click on the discussion group 1 in the conference list pop-up box, and the first client will replace the originally displayed virtual conference interface 900 with the virtual conference interface 903 corresponding to the discussion group 1. The virtual conference interface 903 includes the virtual avatars of the first user, the second user and the third user. These three virtual avatars are migrated from the previous virtual conference interface. That is to say, after these three users switch to the discussion group 1, the previous virtual conference interface corresponding to the conference hall will no longer contain the virtual avatars of these three users. In addition, Figure 9 As shown in , the virtual conference interface 903 may also include objects such as a virtual display screen, a conference table, and the like.

[0130] Assuming that any one of the first user, the second user and the third user wants to exit discussion group 1 and switch back to the original lobby meeting, they can select the lobby meeting in the above-mentioned meeting list pop-up box 902. At this time, the corresponding client interface will switch to display the virtual meeting interface corresponding to the lobby meeting.

[0131] An embodiment of the present invention also provides a virtual conference processing method executed in the cloud. Several computing nodes (or cloud servers) can be deployed in the cloud, and each computing node has computing, storage and other processing resources. In the cloud, multiple computing nodes can be organized to provide a certain service. Of course, a computing node can also provide one or more services. The cloud can provide the service by providing a service interface to the outside world, and users call the service interface to use the corresponding service. The service interface includes a software development kit (SDK), an application programming interface (API), and the like.

[0132] In accordance with the solution provided by the embodiments of the present invention, a service cluster providing virtual conference service functions can be formed in the cloud. This service cluster can include at least one computing node, namely a cloud server. The service cluster provides a service interface to the outside world. Clients providing virtual conference service functions can call this service interface to interact with the service cluster. Specifically, in the virtual conference processing method provided by the embodiments of the present invention, the service cluster can perform the following steps:

[0133] generating a virtual conference interface corresponding to the virtual conference;

[0134] In response to a request from a client of any one of the plurality of users participating in the virtual conference to access the virtual conference, sending a virtual conference interface to the client of the any one user;

[0135] Receiving user expression information sent by a client of any user, wherein the user expression information is determined by the corresponding client based on a facial image of the corresponding user obtained;

[0136] Aggregate expression information of multiple users;

[0137] Synchronize the expression information of multiple users to the clients of the multiple users.

[0138] For ease of understanding, combined Figure 10 Take the first user among multiple users as an example. The first user's client is installed on Figure 10In the user device E1 shown in the figure, based on the operation of the first user entering the virtual conference, the first client calls the service interface provided by the service cluster E2, and sends a request to join the virtual conference to the service cluster E2 through the user device E1. The service cluster E2 then feeds back the virtual conference interface corresponding to the virtual conference to the user device E1 for display. The camera of the user device E1 is turned on to capture the facial image of the first user. The first client determines the facial expression information of the first user based on the facial image and sends the facial expression information of the first user to the service cluster E2. If other users perform the same process, the service cluster can receive the facial expression information of multiple users sent by their clients respectively, and then send the summary result of the facial expression information of multiple users to the user devices of each user, including the user device E1 of the first user.

[0139] The following describes in detail one or more embodiments of the virtual conference processing device of the present invention. Those skilled in the art will appreciate that these devices can be constructed using commercially available hardware components and configured according to the steps taught in this solution.

[0140] Figure 11 A schematic diagram of the structure of a virtual conference processing device provided by an embodiment of the present invention, wherein the virtual conference processing device is applied to a first client of any first user among multiple users participating in a virtual conference, such as Figure 11 As shown, the device includes: a display module 11, a determination module 12, a sending module 13, and a receiving module 14.

[0141] The display module 11 is configured to display a virtual conference interface, wherein the virtual conference interface includes virtual avatars corresponding to the plurality of users.

[0142] The determination module 12 is configured to obtain the facial image of the first user according to a set sampling frequency; and determine the facial expression information of the first user according to the facial image of the first user.

[0143] The sending module 13 is used to send the expression information of the first user to the server, so that the server can aggregate the expression information of the multiple users and synchronize the expression information of the multiple users to the first client and the second client, where the second client corresponds to the second user among the multiple users except the first user.

[0144] The receiving module 14 is configured to receive the user's expression information sent by the server.

[0145] Optionally, the device further includes: a driving module, configured to locally drive the virtual avatar corresponding to the first user according to the expression information of the first user, and locally drive the virtual avatar corresponding to the second user according to the expression information of the second user.

[0146] Optionally, the determining module 12 is specifically configured to: extract a plurality of facial key points from the facial image of the first user; and determine the expression coefficient of the first user according to the plurality of facial key points.

[0147] Among them, the multiple facial key points include key points corresponding to different facial areas respectively, and the determination module 12 is specifically used to: input the key points of the first facial area into the expression coefficient prediction model to obtain the expression coefficient corresponding to the key points of the first facial area; obtain the preset expression coefficient mapping relationship corresponding to the second facial area; and determine the expression coefficient corresponding to the key points of the second facial area based on the key points of the second facial area and the preset expression coefficient mapping relationship.

[0148] Among them, the preset expression coefficient mapping relationship is used to reflect the mapping relationship between the target type expression coefficient and the distance between the corresponding target key points; the determination module 12 is specifically used to: determine the target key point distance value corresponding to the target type expression coefficient based on the key points of the second facial area; determine the expression coefficient corresponding to the key points of the second facial area based on the target key point distance value and the preset expression coefficient mapping relationship.

[0149] Optionally, the device further comprises a posture processing module configured to determine head posture information of the first user based on the plurality of facial key points. The sending module 13 is further configured to send the head posture information of the first user to the server. The receiving module 14 is further configured to receive the head posture information of the second user sent by the server. The driving module is further configured to adjust the posture of the first user's virtual avatar based on the first user's head posture information and adjust the posture of the second user's virtual avatar based on the second user's head posture information.

[0150] Optionally, the device further includes: a displacement processing module, configured to determine the head displacement information of the first user based on the multiple facial key points, wherein the head displacement information refers to the position movement information of the first user's head in the facial image. The sending module 13 is further configured to send the head displacement information of the first user to the server. The receiving module 14 is further configured to receive the head displacement information of the second user sent by the server. The driving module is further configured to adjust the display position of the first user's virtual avatar in the corresponding display window based on the first user's head displacement information, and adjust the display position of the second user's virtual avatar in the corresponding display window based on the second user's head displacement information.

[0151] Optionally, the apparatus further includes a topic identification module configured to, if the first user is a speaker, analyze the first user's speech data to determine the topic type corresponding to the first user's speech content and to determine whether the first user's expression information matches the topic type. The sending module 13 is further configured to: if a match is found, send the first user's expression information to a server; if a match is found, send the set expression information corresponding to the topic type as the first user's expression information to the server.

[0152] Optionally, the determining module 12 is further configured to: receive an expression keyword input by the first user, and generate expression information of the first user according to the expression keyword.

[0153] Optionally, the display module 11 is further configured to: display, in the first client, a configuration sub-interface visible only to the first user in association with the virtual conference interface, wherein the configuration sub-interface includes expression configuration items corresponding to the expression information of the first user and a virtual avatar of the first user. The driving module is further configured to: in response to the first user performing a configuration adjustment operation on the expression configuration item, drive the virtual avatar corresponding to the first user in the configuration sub-interface according to the expression information updated by the configuration adjustment operation; and in response to the first user performing a confirmation operation on the updated expression information, migrate the virtual avatar corresponding to the first user displayed in the configuration sub-interface to be displayed in the virtual conference interface.

[0154] Optionally, the virtual conference interface includes a virtual display screen; the display module 11 is further configured to present the shared content on the virtual display screen in response to an information sharing operation triggered by the first user; and the sending module 13 is further configured to synchronize the virtual display screen presenting the shared content to the second client via the server.

[0155] Optionally, the display module 11 is further used to: display a discussion group containing at least two corresponding users in the virtual conference interface according to the discussion group creation information triggered by the first user, wherein the at least two users include the first user. The sending module 13 is further used to: synchronize the discussion group creation information to the second client through the server, so that the second client generates the discussion group. The display module 11 is further used to: display a discussion group conference interface on the first client in response to an operation of switching to the discussion group triggered by the first user, wherein the discussion group conference interface includes virtual avatars of the at least two users migrated from the virtual conference interface.

[0156] Figure 11The virtual conference processing device shown can be used to execute the steps in the aforementioned embodiments, and the execution process and effects will not be described in detail here.

[0157] In one possible design, the above Figure 11 The structure of the virtual conference processing device shown can be implemented as an electronic device. The above client is running in the electronic device. Figure 12 As shown, the electronic device may include: a processor 21, a memory 22, a communication interface 23, and a display 24. The memory 22 stores executable code, and when the executable code is executed by the processor 21, the processor 21 can at least implement the virtual conference processing method provided in the above embodiment.

[0158] The electronic device provided in some embodiments of the present invention may be an extended reality device, specifically an external head-mounted display device or an integrated head-mounted display device, etc., which supports XR technology, wherein the external head-mounted display device needs to be used in conjunction with an external processing system (such as a computer processing system).

[0159] Figure 13 A schematic diagram of the internal configuration structure of a head-mounted extended reality device 1300 is shown.

[0160] The display unit 1301 may include a display panel, which is provided on the side surface of the extended reality device 1300 facing the user's face. The display panel may be a single panel, or a left panel and a right panel corresponding to the user's left eye and right eye, respectively. The display panel may be an electroluminescent (EL) element, a liquid crystal display, or a microdisplay with a similar structure, or a retinal direct display or similar laser scanning display. It should be noted that the display unit 1301 should not affect the capture of the user's facial image. For example, the display panel should be able to show the user's eyes and other facial areas.

[0161] The virtual image optical unit 1302 captures the image displayed by the display unit 1301 in a magnified manner and allows the user to observe the displayed image by pressing the magnified virtual image. The display image output to the display unit 1301 can be an image of a virtual scene obtained from a data source such as a content reproduction device (Blu-ray disc or DVD player) or a streaming media server, or an image of a real scene captured using an external camera 1310. In embodiments of the present invention, the image displayed on the display unit 1301 may include a virtual conference interface, etc. In some embodiments, the virtual image optical unit 1302 may include a lens unit, such as a spherical lens, an aspherical lens, a Fresnel lens, etc.

[0162] The input operation unit 1303 includes at least one operation component for performing input operations, such as a key, button, switch or other components with similar functions, receives user instructions through the operation component, and outputs instructions to the control unit 1307.

[0163] The state information acquisition unit 1304 is used to obtain state information of the user using the extended reality device 1300. The state information acquisition unit 1304 may include various types of sensors for detecting state information itself, and may obtain state information from external devices (such as smartphones, watches, and other multi-functional terminals worn by the user) through the communication unit 1305. The state information acquisition unit 1304 may obtain position information and / or posture information of the user's head. The state information acquisition unit 1304 may include one or more of a gyroscope sensor, an acceleration sensor, a global positioning system (GPS) sensor, a geomagnetic sensor, a Doppler effect sensor, an infrared sensor, and a radio frequency field strength sensor.

[0164] The communication unit 1305 performs communication processing with an external device and encoding and decoding processing of communication signals. In addition, the control unit 1307 can send transmission data from the communication unit 1305 to the external device, such as user expression information in an embodiment of the present invention.

[0165] The augmented reality device 1300 may further include a storage unit 1306 , which may store applications or various types of data. For example, content viewed by a user using the augmented reality device 1300 may be stored in the storage unit 1306 , and client programs may be stored in the storage unit 1306 .

[0166] The extended reality device 1300 may further include a control unit 1307, which may include a computer processing unit (CPU) or other device with similar functionality. In some embodiments, the control unit 1307 may be configured to execute an application stored in the storage unit 1306, or the control unit 1307 may be configured to perform the steps disclosed in the embodiments of the present invention.

[0167] The image processing unit 1308 is used to perform signal processing such as image quality correction with respect to the image signal output from the control unit 1307 and conversion of its resolution into a resolution according to the screen resolution of the display unit 1301. Then, the display driving unit 1309 sequentially selects each row of pixels of the display unit 1301 and sequentially scans each row of pixels of the display unit 1301 row by row, thereby providing a pixel signal based on the signal-processed image signal.

[0168] The extended reality device 1300 may further include an external camera 1310. The external camera 1310 may be disposed on the front surface of the main body of the extended reality device 1300. There may be one or more external cameras 1310. In an embodiment of the present invention, the external camera 1310 may be used to capture a face image.

[0169] The augmented reality device 1300 may further include a sound processing unit 1311, which may perform sound quality correction or sound amplification of the sound signal output from the control unit 1307, as well as signal processing of the input sound signal. The sound input / output unit 1312 then outputs the sound after sound processing to the outside and inputs the sound from the microphone.

[0170] It should be noted that Figure 13 The structure or component shown in the dotted box can be independent of the extended reality device 1300, for example, it can be set in an external processing system (such as a computer system) for use with the extended reality device 1300; or, the structure or component shown in the dotted box can be set inside or on the surface of the extended reality device 1300.

[0171] In addition, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the virtual conference processing method provided in the aforementioned embodiment.

[0172] The device embodiments described above are merely illustrative, wherein the network elements described as separate components may or may not be physically separate. Some or all of these modules may be selected based on actual needs to achieve the objectives of this embodiment. Persons of ordinary skill in the art can understand and implement these embodiments without inventive effort.

[0173] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by adding a necessary general hardware platform, and of course can also be implemented by a combination of hardware and software. Based on this understanding, the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A virtual conference processing method, characterized in that: A first client applied to any first user among multiple users participating in a virtual conference includes: Displaying a virtual conference interface, wherein the virtual conference interface includes virtual avatars corresponding to the multiple users; Acquire the facial image of the first user according to a set sampling frequency; determining the facial expression information of the first user based on the facial image of the first user; Sending the expression information of the first user to a server, so that the server aggregates the expression information of the multiple users and synchronizes the expression information of the multiple users to the first client and a second client, where the second client corresponds to a second user among the multiple users except the first user; receiving the expression information of the second user sent by the server; The step of determining the facial expression information of the first user based on the facial image of the first user includes: Extracting a plurality of facial key points from the facial image of the first user, wherein the plurality of facial key points include key points corresponding to different facial regions respectively; Inputting the key points of the first facial region into an expression coefficient prediction model to obtain expression coefficients corresponding to the key points of the first facial region; Obtaining a preset expression coefficient mapping relationship corresponding to the second facial region, the preset expression coefficient mapping relationship being used to reflect a mapping relationship between an expression coefficient of a target category and a distance between corresponding target key points; the mapping relationship being represented by a mapping function curve, wherein one coordinate axis of the mapping function curve represents a value of the expression coefficient, and the other coordinate axis of the mapping function curve represents a distance value between key points corresponding to the expression coefficient; determining, based on the key points of the second facial area, a distance value between target key points corresponding to the expression coefficient of the target category; Determine the expression coefficient corresponding to the key point of the second facial area based on the mapping relationship between the distance value between the target key points and the preset expression coefficient; wherein the expression change corresponding to the first facial area is greater than the expression change corresponding to the second facial area.

2. The method according to claim 1, characterized in that The method further comprises: The virtual avatar corresponding to the first user is locally driven according to the expression information of the first user, and the virtual avatar corresponding to the second user is locally driven according to the expression information of the second user.

3. The method according to claim 2, characterized in that The method further comprises: Determining head posture information of the first user based on the multiple facial key points; Sending the head posture information of the first user to the server; receiving the head posture information of the second user sent by the server; The virtual avatar of the first user is posture-adjusted according to the head posture information of the first user, and the virtual avatar of the second user is posture-adjusted according to the head posture information of the second user.

4. The method according to claim 2, characterized in that The method further comprises: determining head displacement information of the first user based on the multiple facial key points, where the head displacement information refers to position movement information of the first user's head in the facial image; sending the head displacement information of the first user to the server; receiving the head displacement information of the second user sent by the server; The display position of the first user's virtual avatar in the corresponding display window is adjusted according to the first user's head displacement information, and the display position of the second user's virtual avatar in the corresponding display window is adjusted according to the second user's head displacement information.

5. The method according to claim 1, wherein The method further comprises: If the first user is a speaker, analyzing the voice data of the first user to determine a topic type corresponding to the speech content of the first user; The sending the expression information of the first user to the server includes: Determining whether the expression information of the first user matches the topic type; If there is a match, the expression information of the first user is sent to the server; If there is no match, the set expression information corresponding to the topic type is sent to the server as the expression information of the first user.

6. The method according to claim 1, characterized in that The method further comprises: In the first client, a configuration sub-interface visible only to the first user is displayed in association with the virtual conference interface, the configuration sub-interface including expression configuration items corresponding to the expression information of the first user and a virtual avatar of the first user; In response to the first user's configuration adjustment operation on the expression configuration item, driving the virtual avatar corresponding to the first user in the configuration sub-interface according to the expression information updated by the configuration adjustment operation; In response to the first user's confirmation operation on the updated expression information, the virtual avatar corresponding to the first user displayed in the configuration sub-interface is migrated and displayed in the virtual conference interface.

7. The method according to claim 1, characterized in that The virtual conference interface includes a virtual display screen; the method further includes: In response to the information sharing operation triggered by the first user, presenting the shared content on the virtual display screen; The virtual display screen presenting the shared content is synchronized to the second client through the server.

8. The method according to claim 1, characterized in that The method further comprises: Displaying, in the virtual conference interface, a discussion group including at least two corresponding users according to the discussion group creation information triggered by the first user, wherein the at least two users include the first user; Synchronizing the discussion group creation information to the second client through the server, so that the second client generates the discussion group; In response to the operation of switching to the discussion group triggered by the first user, a discussion group conference interface is displayed on the first client, wherein the discussion group conference interface includes virtual avatars of the at least two users migrated from the virtual conference interface.

9. A virtual conference processing device, characterized in that: A first client applied to any first user among multiple users participating in a virtual conference includes: A display module, configured to display a virtual conference interface, wherein the virtual conference interface includes virtual avatars corresponding to the plurality of users; a determination module, configured to obtain a facial image of the first user according to a set sampling frequency; and determine the facial expression information of the first user based on the facial image of the first user; a sending module, configured to send the expression information of the first user to a server, so that the server aggregates the expression information of the multiple users and synchronizes the expression information of the multiple users to the first client and a second client, where the second client corresponds to a second user among the multiple users except the first user; A receiving module, configured to receive the user's expression information sent by the server; wherein the determination module is further configured to: extract a plurality of facial key points from the facial image of the first user, wherein the plurality of facial key points include key points corresponding to different facial regions respectively; input the key points of the first facial region into an expression coefficient prediction model to obtain expression coefficients corresponding to the key points of the first facial region; obtain a preset expression coefficient mapping relationship corresponding to the second facial region, wherein the preset expression coefficient mapping relationship is used to reflect the mapping relationship between the target type expression coefficient and the distance between the corresponding target key points; the mapping relationship is represented by a mapping function curve, wherein one coordinate axis of the mapping function curve represents the value of the expression coefficient, and the other coordinate axis of the mapping function curve represents the distance value between the key points corresponding to the expression coefficient; determine the target distance value between the key points corresponding to the target type expression coefficient based on the key points of the second facial region; determine the expression coefficient corresponding to the key points of the second facial region based on the distance value between the target key points and the preset expression coefficient mapping relationship; wherein the expression change corresponding to the first facial region is greater than the expression change corresponding to the second facial region.

10. An electronic device, characterized in that: include: A memory, a processor, a communication interface, and a display; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor executes the virtual conference processing method according to any one of claims 1 to 8.

11. A non-transitory machine-readable storage medium, characterized in that The non-transitory machine-readable storage medium stores executable code, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the virtual conference processing method according to any one of claims 1 to 8.

12. A virtual conference processing method, characterized in that: A first extended reality device applied to any first user among a plurality of users participating in a virtual meeting includes: Displaying a virtual conference interface, wherein the virtual conference interface includes virtual avatars corresponding to the multiple users; Acquire the facial image of the first user according to a set sampling frequency; determining the facial expression information of the first user based on the facial image of the first user; Sending the expression information of the first user to a server, so that the server aggregates the expression information of the multiple users and synchronizes the expression information of the multiple users to the first extended reality device and a second extended reality device, where the second extended reality device corresponds to a second user among the multiple users except the first user; receiving the expression information of the second user sent by the server; The step of determining the facial expression information of the first user based on the facial image of the first user includes: Extracting a plurality of facial key points from the facial image of the first user, wherein the plurality of facial key points include key points corresponding to different facial regions respectively; Inputting the key points of the first facial region into an expression coefficient prediction model to obtain expression coefficients corresponding to the key points of the first facial region; Obtaining a preset expression coefficient mapping relationship corresponding to the second facial region, the preset expression coefficient mapping relationship being used to reflect a mapping relationship between an expression coefficient of a target category and a distance between corresponding target key points; the mapping relationship being represented by a mapping function curve, wherein one coordinate axis of the mapping function curve represents a value of the expression coefficient, and the other coordinate axis of the mapping function curve represents a distance value between key points corresponding to the expression coefficient; determining, based on the key points of the second facial region, a distance value between target key points corresponding to the expression coefficient of the target category; Determine the expression coefficient corresponding to the key point of the second facial area based on the mapping relationship between the distance value between the target key points and the preset expression coefficient; wherein the expression change corresponding to the first facial area is greater than the expression change corresponding to the second facial area.

Citation Information

Patent Citations

  • Facial expression editing method based on single camera and motion capturing data

    CN103473801A

  • Real-time interactive control method and real-time interactive control device of virtual object

    CN104866101A

  • Virtual scene implementation method and device, terminal and server

    CN108881784A

  • Virtual face generation method

    CN113781610A

  • Network training and video frame processing method and device, equipment and storage medium

    CN114120389A