An interactive processing method and device, a terminal and a medium
By displaying a virtual avatar in the video chat interface and controlling its interaction based on action information, the problem of protecting user privacy in video chats is solved, and data transmission efficiency and user experience are improved.
Patent Information
- Application Number
- CN202110606182.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-31
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2041-05-31
AI Technical Summary
In video conversations, users are unwilling to show their real appearance to protect their privacy, and existing technologies are unable to effectively improve privacy.
By displaying a virtual image of the target conversation object in the video conversation interface and controlling the virtual image to perform interactive actions based on its action information, the transmission of real image images is avoided, and only relevant data is transmitted for rendering.
It achieves the protection of user privacy in video sessions while improving data transmission efficiency and user immersion.
Smart Images

Figure CN113766168B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an interactive processing method, apparatus, terminal and medium. Background Technology
[0002] With the rapid development of science and technology, multiple users in different locations can conduct online conversations via the Internet; for example, multiple users can initiate online video conversations via the Internet. Video conversations are widely used due to their advantages such as convenience, speed, and simplicity.
[0003] In existing video conferencing scenarios, cameras capture and transmit images containing the users' real images so that each user's real image is displayed on their respective terminal screen. In some cases, users may be more concerned about privacy and unwilling to present their real images in video conferencing; therefore, improving the privacy of video conferencing has become a hot research topic. Summary of the Invention
[0004] This application provides an interactive processing method, apparatus, terminal, and medium that enables users to participate in video conversations using virtual avatars instead of their real-life avatars.
[0005] On the one hand, embodiments of this application propose an interactive processing method, which includes:
[0006] During a video session, a video session interface is displayed, which includes a display area for showing the images of the video session objects.
[0007] The target virtual image of the target session object related to the video session is displayed within the image display area;
[0008] Obtain the action information of the target session object, and control the target virtual image displayed in the image display area to perform the target interactive action based on the action information of the target session object.
[0009] On the other hand, embodiments of this application provide an interactive processing apparatus, the apparatus comprising:
[0010] The display unit is used to display a video session interface during a video session, wherein the video session interface includes an image display area for displaying video session objects.
[0011] The processing unit is configured to display a target virtual image of a target session object related to the video session within the image display area;
[0012] The processing unit is further configured to acquire action information of the target session object, and control the target virtual image displayed in the image display area to perform target interactive actions based on the action information of the target session object.
[0013] In one implementation, the processing unit is further used for:
[0014] Display an image selection window, which includes image selection elements;
[0015] In response to a trigger operation on the image selection element, a reference virtual image is displayed in the reference display area, and candidate virtual images are displayed in the image selection window;
[0016] In response to the image selection operation of the candidate virtual image, the reference virtual image is updated and displayed in the reference display area as the target candidate virtual image selected by the image selection operation;
[0017] In response to the virtual avatar confirmation operation, the target candidate virtual avatar is determined as the target virtual avatar.
[0018] In one implementation, the processing unit is further used for:
[0019] The background selection element is displayed in the image selection window;
[0020] In response to a user action on the background selection element, candidate background images are displayed in the image selection window;
[0021] In response to a background selection operation on the candidate background image, the target candidate background image selected by the background selection operation is displayed in the reference display area;
[0022] In response to the background image confirmation operation, the target candidate background image is set as the background image of the image display area.
[0023] In one implementation, the processing unit is further used for:
[0024] The voice selection element is displayed in the image selection window;
[0025] In response to the selection operation of the voice selection element, candidate voice audio processing rules are displayed in the image selection window;
[0026] In response to the confirmation operation of the candidate speech audio processing rule, the candidate speech audio processing rule is determined as the target speech audio processing rule, which is used to simulate the sound signal of the target session object received during the video session.
[0027] In one implementation, the image selection window includes an exit option, and the processing unit is further configured to:
[0028] In response to the selection of the exit option, the video session interface is displayed;
[0029] An environmental image is displayed within the image display area included in the video session interface. The environmental image is obtained by capturing the environment.
[0030] The environmental image is sent to the peer device so that the peer device can display the environmental image. The peer device refers to the device used by other users participating in the video session.
[0031] In one implementation, when the processing unit controls the target virtual avatar displayed in the avatar display area to perform a target interactive action based on the action information of the target session object, it is specifically configured to perform any one or more of the following steps:
[0032] If the action information of the target session object is facial information, then control the target virtual image to perform facial interaction actions;
[0033] If the action information of the target session object is emotional information, then the face of the target virtual avatar is replaced by the target facial resource associated with the emotional information.
[0034] If the action information of the target session object is limb information, then control the target limb of the target virtual image to perform limb actions;
[0035] If the action information of the target session object is location transfer information, then the target virtual image is controlled to perform a location transfer action within the image display area.
[0036] In one implementation, when the processing unit controls the target virtual avatar displayed in the avatar display area to perform a target interactive action based on the action information of the target session object, it is specifically used for:
[0037] Obtain a set of meshes added for the target virtual image. The set of meshes includes multiple meshes and mesh data for each mesh. Each mesh corresponds to an object element. Each mesh consists of at least three mesh vertices. The mesh data for each mesh refers to the state values of each mesh vertex contained in each mesh.
[0038] Based on the action information of the target session object, perform mesh deformation processing on the mesh data of the target mesh in the mesh set;
[0039] Based on the mesh data after mesh deformation processing, the target virtual image that has performed the target interactive action is rendered and displayed. In the rendered and displayed target virtual image that has performed the target interactive action, the position and / or shape of the object elements corresponding to the target mesh change.
[0040] In one implementation, the action information includes facial information, which includes N feature points of the target session object's face, where N is an integer greater than 1; the processing unit is used to perform mesh deformation processing on the mesh data of the target mesh in the mesh set according to the action information of the target session object, specifically for:
[0041] Based on the N feature point information, the updated expression type of the target session object is determined;
[0042] Obtain the expression base coefficients corresponding to the updated expression type;
[0043] Based on the updated expression base coefficients, perform mesh deformation processing on the mesh containing the object elements corresponding to the updated expression type.
[0044] In one implementation, when the processing unit performs mesh deformation processing on the mesh containing the object element corresponding to the updated expression type based on the expression base coefficients corresponding to the updated expression type, it specifically performs the following:
[0045] Obtain the base coefficients of the emoji type before the update;
[0046] The difference between the expression base coefficients corresponding to the updated expression type and the expression base coefficients corresponding to the expression type before the update is calculated to obtain the difference result.
[0047] Based on the difference result, the expression base coefficient of the updated expression state, and the expression base coefficient of the expression state before the update, the mesh of the object element corresponding to the updated expression type is subjected to mesh deformation processing.
[0048] In one implementation, the action information includes limb information. When the processing unit performs mesh deformation processing on the mesh data of the target mesh in the mesh set based on the action information of the target session object, it specifically performs the following:
[0049] Based on the limb information of the target session object, determine the position information of S limb points of the target session object, where S is an integer greater than zero;
[0050] Based on the position information of the S limb points, calculate the angle value of the corresponding torso element;
[0051] The mesh corresponding to the torso element in the target virtual image is subjected to mesh deformation processing based on the angle value.
[0052] In one implementation, the action information includes emotion information. The processing unit, based on the action information of the target session object, controls the target virtual avatar displayed in the avatar display area to perform a target interactive action. Specifically, it is used for:
[0053] The emotional state of the target conversation object is identified based on the emotional information to obtain the current emotional state of the target conversation object;
[0054] Based on the current emotional state, a target facial resource matching the current emotional state is determined;
[0055] The face of the target virtual image is updated using the target facial resources within the image display area to obtain the updated target virtual image.
[0056] In one implementation, the action information includes position transfer information. The processing unit, based on the action information of the target session object, controls the target virtual avatar displayed in the avatar display area to perform a target interactive action. Specifically, it is used for:
[0057] If the position transfer information is the movement information of the target image point, then the target virtual image is controlled to move and be displayed in the image display area according to the movement information of the target image point;
[0058] If the position transfer information is the movement information of the target image region, the size of the display area of the target virtual image in the image display area is adjusted according to the movement information of the target image region.
[0059] In one implementation, the processing unit is further used for:
[0060] When the target terminal initiates a video session, it acquires a target image of the environment.
[0061] The target session object in the target image is subjected to feature recognition processing to obtain the recognition result;
[0062] Assign a virtual avatar that matches the recognition result to the target session object, and identify the virtual avatar that matches the recognition result as the target virtual avatar.
[0063] On the other hand, embodiments of this application provide a terminal, which includes:
[0064] The storage device contains computer programs;
[0065] The processor runs the computer program stored in the storage device to implement the interactive processing method described above.
[0066] On the other hand, embodiments of this application provide a computer-readable storage medium storing a computer application program that, when executed, implements the interactive processing method described above.
[0067] On the other hand, embodiments of this application provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. The processor of a terminal reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the terminal to perform the aforementioned interactive processing method.
[0068] In this embodiment, a target virtual avatar of the target conversation object related to the video conversation can be displayed in the video conversation interface. The target virtual avatar is then driven to perform target interactive actions based on the action information of the target conversation object, allowing the target conversation object to participate in the video conversation through the target virtual avatar. This method of outputting the target virtual avatar in the video conversation interface allows for quick display of the target virtual avatar; furthermore, the target virtual avatar can be used to represent the target conversation object in the video conversation. By using a virtual avatar to simulate real-person interaction, the display of the target conversation object's real image in the video conversation can be avoided, thus protecting the image privacy of the target conversation object. Attached Figure Description
[0069] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0070] Figure 1a A schematic diagram of a social conversation scenario provided by an exemplary embodiment of this application is shown;
[0071] Figure 1b This illustration shows a schematic diagram of data transmission provided by an exemplary embodiment of this application;
[0072] Figure 2 A flowchart illustrating an interactive processing method provided in an exemplary embodiment of this application is shown.
[0073] Figure 3a A schematic diagram of a video session interface provided in an exemplary embodiment of this application is shown;
[0074] Figure 3b This illustration shows a schematic diagram of an environmental image contained in a switching display area provided by an exemplary embodiment of this application;
[0075] Figure 4 This illustration shows a schematic diagram of switching virtual video session modes according to an exemplary embodiment of this application;
[0076] Figure 5 This illustration shows a schematic diagram of selecting a target virtual avatar according to an exemplary embodiment of this application;
[0077] Figure 6 This illustration shows a schematic diagram of a method for displaying a background image when a target virtual image is selected, according to an exemplary embodiment of this application.
[0078] Figure 7 This illustration shows a schematic diagram of selecting a background image within an image selection window, provided by an exemplary embodiment of this application.
[0079] Figure 8 This illustration shows a schematic diagram of selecting speech audio processing rules according to an exemplary embodiment of this application;
[0080] Figure 9 This illustration shows a schematic diagram of exiting a virtual video session mode according to an exemplary embodiment of this application;
[0081] Figure 10 A flowchart illustrating another interactive processing method provided by an exemplary embodiment of this application is shown;
[0082] Figure 11a This illustration shows a schematic diagram of adding a mesh to a target virtual avatar according to an exemplary embodiment of this application;
[0083] Figure 11b This illustration shows a schematic diagram of a mesh deformation process for mesh data provided in an exemplary embodiment of this application;
[0084] Figure 11c This illustration shows a schematic diagram of the mesh data changes of a target mesh during the rendering process of a target virtual image, provided by an exemplary embodiment of this application.
[0085] Figure 11d This illustration shows a schematic diagram of a hierarchically rendered target virtual image provided in an exemplary embodiment of this application;
[0086] Figure 12 This illustration shows a schematic diagram of an exemplary embodiment of the present application, which describes controlling a target virtual avatar to perform target interactive actions based on the facial information of a target session object;
[0087] Figure 13a This illustration shows a schematic diagram of a method for identifying feature points of the face of a target session object, provided by an exemplary embodiment of this application.
[0088] Figure 13b This illustration shows a schematic diagram of an exemplary embodiment of the present application, which describes a method for controlling a virtual image of a target to perform interactive actions based on information from N feature points.
[0089] Figure 13c This illustration shows a schematic diagram of a mouth mesh dynamically changing according to an exemplary embodiment of this application;
[0090] Figure 13d This illustration shows a target conversation object and a target virtual avatar after performing a mouth-opening action, according to an exemplary embodiment of this application.
[0091] Figure 13e This illustration shows a flowchart of an exemplary embodiment of the present application, illustrating how an interpolation algorithm is used to determine the intermediate state of a mouth.
[0092] Figure 14a This illustration shows a schematic diagram of controlling a target virtual avatar to perform a waving action, provided in an exemplary embodiment of this application.
[0093] Figure 14b This illustration shows a schematic diagram of a target virtual image performing target interactive actions by controlling limb points, according to an exemplary embodiment of this application.
[0094] Figure 15a This illustration shows a schematic diagram of replacing the face of a target virtual avatar with a target facial resource, according to an exemplary embodiment of this application.
[0095] Figure 15b This illustration shows a schematic diagram of a preset exaggerated facial expression provided in an exemplary embodiment of this application;
[0096] Figure 15c This illustration shows a flowchart of a process for replacing facial resources provided in an exemplary embodiment of this application;
[0097] Figure 16 This invention provides a schematic diagram of the structure of an interactive processing apparatus according to an exemplary embodiment of the present application.
[0098] Figure 17 The diagram shows a structural schematic of a smart terminal provided in an exemplary embodiment of this application. Detailed Implementation
[0099] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0100] This application relates to virtual avatars, which can refer to a virtual image used by a user to represent themselves. This image can be a fictional model (such as a non-existent cartoon or anime model) or a real model (such as a character model similar to a real person but displayed on the terminal screen). Common virtual avatars include, but are not limited to, virtual character images (such as cartoon characters, anime characters, two-dimensional characters, etc.), virtual animated images (such as cartoon animal images, various object images, etc.). For ease of explanation, the following description will use virtual character images as an example. Using virtual avatars during terminal use can enhance the user's sense of immersion and make the operation more engaging. For example, in video conversation scenarios (such as video calls, video conferences, etc.), using virtual avatars to represent the user in the video conversation, simulating real-person interaction, can enhance the immersion of the video conversation participants. Video conversations can include individual video conversations and group video conversations. In an individual video conversation, the number of users participating is two, and in a group video conversation, the number of users participating is greater than or equal to three. This application does not limit the type of video conversation; this is only a description of the specific type.
[0101] Based on this, this application proposes an interactive processing scheme that can use a target virtual avatar to participate in the video conversation in place of the target conversation object, thereby enabling the target virtual avatar to be quickly displayed in the video conversation interface. Furthermore, it can also obtain the action information of the target conversation object and control the target virtual avatar to flexibly follow the target conversation object and perform target interactive actions based on the action information, thereby enabling participation in the video conversation through the target virtual avatar. This method of using a target virtual avatar to simulate real-person interaction can avoid displaying the real image of the target conversation object in the video conversation, thus protecting the image privacy of the target conversation object.
[0102] The aforementioned interactive processing scheme can be executed by the target terminal. The target terminal here may include, but is not limited to, terminal devices such as smartphones, tablets, laptops, and desktop computers. Applications for executing the interactive processing scheme can be deployed on the target terminal. These applications may include, but are not limited to, IM (Instant Messaging) applications, content interaction applications, etc. Instant messaging applications refer to internet-based applications for instant messaging and social interaction, and may include, but are not limited to, QQ, WeChat, WeChat Work, map applications with social interaction functions, game applications, etc. Content interaction applications refer to applications capable of content interaction, such as online banking, microblogs, personal spaces, news applications, etc. In this way, the target user can open and use the applications deployed on the target terminal to conduct video conversations.
[0103] The following is combined with Figure 1a The present application will provide an exemplary description of video session scenarios involving interactive processing schemes in its embodiments; such as... Figure 1a As shown, the video session in this scenario includes video sessions from instant messaging applications, video conferencing applications, and other related applications that require video communication. The video session participants (i.e., the users involved) include the target user 101 (the user of the target terminal) and other users 102 (such as the target user's friends, colleagues, strangers, etc.). Therefore, the video session interface can be displayed on the screen of the target terminal 1011 used by the target user 101, as well as on other terminals 1021 used by other users 102. The video session interface 103 may include one or a combination of two of the virtual avatars corresponding to the target user 101 and the virtual avatars corresponding to other users 102, such as... Figure 1aThe video session interface shown includes a virtual avatar corresponding to target user 101 and real avatars of other users 102. When the video session interface includes a virtual avatar corresponding to target user 101, the aforementioned target virtual avatar refers to the virtual avatar corresponding to target user 101, and target user 101 refers to the target session object. Similarly, when the video session interface includes both the virtual avatar corresponding to target user 101 and the virtual avatars corresponding to other users 102, ① if the obtained action information is the action information of target user 101, then the target virtual avatar controlling the execution of the target interactive action according to the action information refers to the virtual avatar corresponding to target user 101; conversely, ② if the obtained action information is the action information of other users 102, then the target virtual avatar controlling the execution of the target interactive action according to the action information refers to the virtual avatar corresponding to other users 102. For ease of explanation, this application embodiment uses target user 101 as the target session object and the virtual avatar corresponding to target user 101 as an example to introduce the subsequent related content, which is explained here.
[0104] It is worth noting that this embodiment achieves control over the target virtual image to perform target interactive actions by rendering the target virtual image, without needing to display the real image of the target conversation object within the image display area. Based on this, this embodiment does not need to transmit images of the video conversation object captured by a camera, which includes multiple frames of environmental images captured by the camera. It only needs to transmit the detected data related to the video conversation object (such as facial data, torso data, etc.), and the virtual image performing the target interactive action can be rendered on the other end's device based on this data. Compared to transmitting images, this method of transmitting only the relevant data of the video conversation object reduces the amount of data transmitted and improves data transmission efficiency. The flowchart describing the process of transmitting the relevant data of the video conversation object and rendering the target virtual image on the other end's device based on this data can be found in [reference needed]. Figure 1b .
[0105] This application embodiment can also be combined with blockchain technology. Specifically, the target terminal used to execute the interactive processing scheme can be a node device in a blockchain network. The target terminal can publish the collected action information of the target conversation object (such as body information, facial information, etc.) to the blockchain network, and record the relevant data (type and information of the target interactive action) of the target virtual image controlled by these action information on the blockchain. This can prevent the collected action information of the target conversation object from being tampered with during transmission, and at the same time, it can make each target interactive action executed by the target virtual image effectively traceable. The action information is stored in the form of blocks, which can realize the distributed storage of action information.
[0106] Based on the interactive processing scheme described above, this application proposes a more detailed interactive processing method. The interactive processing method proposed in this application will be described below with reference to the accompanying drawings.
[0107] Please see Figure 2 , Figure 2 This illustration shows a flowchart of an interactive processing method provided by an exemplary embodiment of this application; the interactive processing method can be executed by a target terminal. Figure 2 As shown, the interactive processing method may include, but is not limited to, steps S201-S203.
[0108] S201: During a video session, display the video session interface. When a video session object (i.e., any user participating in the video session) opens and uses the video session function provided by the target application (i.e., any application with video session functionality), the target application can display the video session interface corresponding to the video session. The video session interface can be used to display the images of each video session object participating in the video session, specifically by displaying an environmental image containing the video session object; the environmental image can be obtained by capturing the environment, specifically by using the camera configured on the target terminal to capture the environment in which the target session object is currently located; the process of the camera capturing the environmental image is executed only after the camera is started in the video session scenario. If the camera of a terminal used by a video session object is not turned on in the video session scenario, only the configuration image of that video session object (such as the user's avatar image, blank image, etc.) is displayed in the video session interface. Specifically, the video session object can trigger the camera of the target terminal to be turned on by triggering the camera option included in the video session interface, which will not be elaborated on here. The identification of video conversation objects within environmental images can be based on Artificial Intelligence (AI). AI utilizes digital computers or computers-controlled machines to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results—a theory, method, technology, and application system. Specifically, this can be achieved using Computer Vision (CV) technology under AI, which typically includes common biometric identification technologies such as facial recognition and fingerprint recognition.
[0109] Taking a video session as an example, an exemplary video session interface can be found here. Figure 3a ,like Figure 3aAs shown, the video session interface 301 can be used to display the environmental images corresponding to each user participating in the video session. Specifically, the video session interface 301 includes an image display area for displaying the images of video session objects, with each image display area used to display the environmental image corresponding to one video session object. As described above, the video session object includes each video session object participating in a separate video session, thus... Figure 3a The video session interface 301 shown includes two image display areas, namely image display area 3011 and image display area 3012. These two areas are used to display environmental images corresponding to different video session objects participating in a single video session. The two image display areas can be displayed overlapping in the video session interface, such as... Figure 3a As shown, one image display area is displayed on top of another image display area with a smaller display area; alternatively, two image display areas can be displayed side-by-side (or in parallel, diagonal, or other shapes) in the video session interface. For ease of explanation, the following will use the example of two image display areas overlapping in a single video session and the example of multiple image display areas being displayed side-by-side in a group video session. This is explained here.
[0110] See also Figure 3a This application embodiment supports switching the display of environmental images contained in two image display areas. For example, assuming that image display area 3011 displays the environmental image of a first video session object, and image display area 3012 displays the environmental image of a second video session object, then this application embodiment supports switching the environmental images contained in the two image display areas. That is, the environmental image of the second video session object is displayed in image display area 3011, while the environmental image of the first video session object is displayed in image display area 3012. In specific implementation, it supports video session objects participating in a video session switching the display of environmental images contained in the image display areas on the terminal screens they are using. See also... Figure 3bAssume that the environmental image corresponding to the target session object is displayed in the image display area 3012, while the environmental images corresponding to other video session objects are displayed in the image display area 3011. A switching option 302 is also displayed within the video session images. When the switching option 302 is triggered, it indicates that the target session object wants to switch the display of the environmental images within the image display area; that is, the environmental image corresponding to the target session object is displayed in the image display area 3011, while the environmental images of other video session objects are displayed in the image display area 3012. Of course, besides triggering the switching option to switch the display of environmental images within the image display area, the display of environmental images within the image display area can also be switched using user gestures, such as allowing a video session object to drag the environmental image displayed in one image display area to another, thereby switching the display of environmental images within different image display areas. For ease of explanation, the following description will use the example of the environmental image corresponding to the target session object being displayed in the image display area 3011.
[0111] It is worth noting that the display area, shape, and display region of the image display area on the video conference interface are not fixed. This application embodiment supports the target conference object adjusting all or part of the image display area's attribute information on the target terminal. For example, it supports the target conference object dragging any image display area from a first position to a second position (such as a different position on the video conference interface) on the video conference interface; it also supports the target conference object adjusting the shape of any image display area on the video conference interface from a square to a circle. Specifically, this can be done by selecting any shape from a pre-configured set of shapes, or by the target conference object manually creating a shape; furthermore, it supports the target conference object performing a "pinch" gesture on the video conference interface to enlarge the smaller image display area and shrink the larger image display area; and so on.
[0112] S202: Display the target virtual avatar of the target conversation object related to the video conversation within the avatar display area. This application embodiment supports switching from a real video conversation to a virtual video conversation in a video conversation scenario. A real video conversation refers to a situation where all participants in the video conversation use their real avatars; in this case, the avatar display area displays the real avatar of the video conversation object captured by the camera. A virtual video conversation refers to a situation where all or some participants in the video conversation use virtual avatars to represent themselves in the video conversation. For example, if the target conversation object uses a target virtual avatar to represent itself in the video conversation, then the avatar display area displays the virtual avatar. This method of using virtual avatars to participate in video conversations protects the privacy of the video conversation object's avatar.
[0113] In specific implementation, if the target session object engages in virtual session operations on the video session interface, then the virtual video session mode is activated. In virtual video session mode, the area on the video session interface used to display the environmental image of the target session object will not display the actual image of the target session object, but instead will display a virtual image that substitutes for it. Virtual session operations can include, but are not limited to: operations generated when a virtual session option is triggered; or operations generated when performing shortcut gestures (such as double-click, single-click, drag-and-drop) on the video session interface; and so on. The following mainly uses the example of executing a virtual session operation by triggering a virtual session option. Specifically, a virtual session option is displayed on the video session interface, or a virtual session option exists under any option included in the video session interface; when a virtual session option is triggered, it can be determined that a virtual session operation has been performed. Combined with... Figure 4 Taking the example of a virtual session option appearing under any of the options in the video session interface, the process of executing virtual session operations by triggering the virtual session option will be briefly introduced; for example... Figure 4 As shown, any option 401 (such as animation options, emoticon options, etc.) is displayed in the video session interface. When any option 401 is triggered, an option window 402 is output, which contains virtual session options 4021. If the virtual session option 4021 is triggered, it is determined that the target session object performs a virtual session operation, and then the virtual video session mode is entered. At this time, the target virtual image configured for the target session object, such as the target virtual image 403, can be displayed in the image display area 3011 corresponding to the target session object.
[0114] The methods for determining the target virtual avatar used to represent the target session object in the video session can include various methods, such as the target virtual avatar being configured by the system or selected by the target session object itself. The following is a more detailed explanation of the different methods for determining the target virtual avatar.
[0115] (1) The target virtual avatar is configured by the system. In specific implementation, when the target terminal starts a video session, a target image captured by the environment can be obtained. This target image may include the real image of the target session object (i.e., the target user). Feature recognition processing is performed on the target session object in the target image to obtain the recognition result. A virtual avatar matching the recognition result is assigned to the target session object, and the virtual avatar matching the recognition result is determined as the target virtual avatar. Here, feature recognition can be the recognition of the face of the target session object, which is not limited in this embodiment. In other words, after starting the virtual video session mode, a target image captured by the camera can be obtained, and image recognition processing is performed on the target image to obtain the feature recognition result of the target session object. Then, a virtual avatar matching the feature recognition result is selected from the virtual avatar library according to the feature recognition result, and this virtual avatar is determined as the target virtual avatar. For example: if feature recognition is performed on the face of the target virtual avatar and the facial recognition result indicates that the target session object is a user with a beard, then the selected target virtual avatar matching the facial recognition result can be a virtual avatar containing a beard. For example, if feature recognition is performed on the head of the target virtual avatar and the feature recognition result indicates that the target session object is a user wearing a hat, then the selected target virtual avatar that matches the feature recognition result can be a virtual avatar wearing a hat.
[0116] (2) The target virtual avatar is autonomously selected by the target session object. Specifically, an avatar selection window can be displayed on the terminal screen, including avatar selection elements. In response to a trigger operation on an avatar selection element, a reference virtual avatar is displayed in the reference display area, and candidate virtual avatars are displayed in the avatar selection window. In response to an avatar selection operation on a candidate virtual avatar, the reference virtual avatar in the reference display area is updated to the target candidate virtual avatar selected by the avatar selection operation. In response to a virtual avatar confirmation operation, the target candidate virtual avatar is confirmed as the target virtual avatar. Here, "response" is the opposite of "request." For example, if a trigger operation on an avatar selection element exists in the avatar selection window, and a request to trigger the avatar selection element is generated in the background, this request can be responded to, i.e., the trigger operation on the avatar selection element can be responded to. Furthermore, the target virtual avatar can include anime characters, animal characters, or object characters, thus enriching the selectivity of the video session object.
[0117] The following is combined Figure 4 and Figure 5The implementation method of allowing the target session object to autonomously select a target virtual avatar is described below. First, when the virtual session option 4021 in the option window 402 is triggered, it indicates that the virtual video session mode is enabled, and an avatar selection window 501 can be output. This avatar selection window 501 contains an avatar selection element 5011. The avatar selection window 501 can be the option window 402, but at this time, the information displayed in the window is related to the virtual avatar. Second, if the avatar selection element 5011 is triggered, candidate virtual avatars are displayed in the avatar selection window 501. The number of candidate virtual avatars is greater than or equal to one, such as candidate virtual avatar 50111, etc.; and a reference virtual avatar 50121 is displayed in the reference display area 5012. The reference display area 5012 can refer to the avatar display area 3011, or it can refer to an area specifically used to display the reference virtual avatar. The reference virtual avatar 50121 can be a default virtual avatar, or the virtual avatar used by the target session object before the avatar selection element is triggered. Secondly, in response to the image selection operation of the candidate virtual avatar (such as candidate virtual avatar 50111), the reference virtual avatar 50121 is updated and displayed in the reference display area 5012 to be the target candidate virtual avatar (such as candidate virtual avatar 50111) selected in the image selection operation. In this way, the target session object can preview the appearance and actions of the selected target candidate virtual avatar in the reference display area, thus satisfying the target session object's preview needs for virtual avatars. Finally, if there is a virtual avatar confirmation operation (such as a gesture operation on the reference display area, or a selection operation of the confirmation option, etc.), indicating confirmation that the target candidate virtual avatar displayed in the reference display area is selected as the target virtual avatar, the target candidate virtual avatar is displayed as the target virtual avatar in the image display area of the video session interface. At this time, the target session object can use the target candidate virtual avatar to conduct a session with other video session objects participating in the video session.
[0118] It should be noted that when any option (or element) in the avatar selection window is selected, that selected option can be highlighted to indicate to the target conversation object that the information displayed in the current avatar selection window is related to the selected option. For example, when virtual conversation option 4021 is selected in the avatar selection window, virtual conversation option 4021 can be highlighted in the avatar selection window (e.g., the brightness of virtual conversation option 4021 is higher than the brightness of other options, the transparency of virtual conversation option 4021 is lower than the transparency of other options, etc.) to indicate to the target conversation object that candidate virtual avatars are displayed in the current avatar selection window. Furthermore, since the display area of the terminal screen is limited, some candidate virtual avatars may be hidden. Therefore, the avatar selection window can include a sliding axis, and the hidden candidate virtual avatars can be displayed by sliding the axis. Of course, in addition to sliding the candidate virtual avatars in the avatar selection window by sliding the axis, the candidate virtual avatars can also be displayed by pressing any position in the sliding avatar selection window. This application embodiment does not limit this.
[0119] Understandably, the image displayed in the avatar display area will differ depending on the order in which the avatar selection operation occurs and the camera is turned on. Specifically, in a video conference scenario, if the target user's camera is not on, the camera will automatically turn on after the user performs the avatar selection operation and selects the target virtual avatar. The selected virtual avatar will then be displayed in the avatar display area. If the target user has turned on their camera but has not selected a target virtual avatar and has not enabled virtual video conference mode, the environmental image captured by the camera will be displayed in the avatar display area. After the user performs the avatar selection operation, the environmental image will be replaced with an image containing the target virtual avatar. If the target user has turned on their camera but has not selected a target virtual avatar and has enabled virtual video conference mode, the system-configured virtual avatar will be displayed in the avatar display area. After the user performs the avatar selection operation, the virtual avatar displayed in the avatar display area will be replaced with the target virtual avatar.
[0120] In addition, the avatar display area within the video chat interface also displays a background image. By pairing a background with the target virtual avatar, a more realistic chat environment can be created. Each candidate virtual avatar displayed in the avatar selection window can be paired with a default background image, allowing the candidate virtual avatar to blend well with the background. In specific implementation, as... Figure 5As shown, when the target session object selects any candidate virtual avatar in the avatar selection window, the background image associated with that candidate virtual avatar can also be displayed in the reference display area. An exemplary diagram illustrating the display of a background image when a target virtual avatar is selected can be found here. Figure 6 ,like Figure 6 As shown, when a candidate virtual avatar is selected in the avatar selection window, in addition to displaying the candidate virtual avatar 50111, a background image is also displayed in the reference display area 5012, such as a background image containing multiple fluttering butterflies. In this implementation, the display areas of each candidate virtual avatar in the avatar selection window can also display the background image associated with that candidate virtual avatar, allowing the target audience to conveniently select the target virtual avatar based on both the candidate virtual avatar and the background image, thus enhancing the video conversation experience for the target audience.
[0121] In addition to selecting a background image by choosing a virtual avatar as described above, embodiments of this application also support the target session object freely selecting a background image according to its own preferences; in other words, the target session object can independently select a background image. Specifically, a background selection element is displayed in the avatar selection window; in response to user operation on the background selection element, candidate background images are displayed in the avatar selection window; in response to a background selection operation on a candidate background image, the target candidate background image selected by the background selection operation is displayed in the reference display area; in response to a confirmation operation on the background image, the target candidate background image is set as the background image of the avatar display area. An exemplary schematic diagram of selecting a background image within the avatar selection window can be found... Figure 7 ,like Figure 7 As shown, the image selection window 501 includes a background selection element 5013. When the background selection element 5013 is selected, candidate background images are displayed in the image selection window 501. The number of candidate background images is greater than or equal to one, such as candidate background image 50131. In response to the background selection operation on the candidate background image, the target candidate background image selected by the background selection operation is displayed in the reference display area. It is possible to display only the target candidate background image in the reference display area, or to display both the virtual image and the target candidate background image in the reference display area. This embodiment does not limit this. In response to the background image confirmation operation, indicating confirmation that the candidate background image displayed in the reference display area is used as the background image of the image display area, the video session interface is displayed, and the selected target candidate background image is displayed in the image display area contained in the video session interface.
[0122] It is worth noting that this application embodiment supports selecting the background image only after the target virtual avatar has been selected, which helps the target virtual avatar blend better with the background. In one implementation, if the target session object has not selected the target virtual avatar, the background selection element in the avatar selection window is set to an unselectable state, meaning the background selection element cannot be triggered, thus prompting the target session object to select the target virtual avatar first; conversely, if the target session object has already selected the target virtual avatar, the background selection element is set to an selectable state, meaning the background selection element can be clicked. Of course, if the target session object selects a background image and then selects the target virtual avatar again, the background image corresponding to the target virtual avatar will not be switched when selecting the target virtual avatar; instead, the background image selected by the target session object will be displayed.
[0123] Furthermore, this application embodiment also supports configuring sound effects for the target conversation object to further enrich the gameplay of video conversations. The sound effects configured for the target conversation object may match the target virtual avatar. In this implementation, when the target virtual avatar is selected, a voice and audio processing rule matching the target virtual avatar is used to simulate the sound signal of the target conversation object received during the video conversation, resulting in a sound effect matching the target virtual avatar. This use of a sound effect matching the target virtual avatar for video conversations enhances the privacy of the target conversation object's identity. Alternatively, the target conversation object can select the sound effect. Specifically, a voice selection element can be displayed in the avatar selection window; in response to the selection of the voice selection element, candidate voice and audio processing rules can be displayed in the avatar selection window; in response to the confirmation operation of the candidate voice and audio processing rule, the candidate voice and audio processing rule is determined as the target voice and audio processing rule, which is used to simulate the sound signal of the target conversation object received during the video conversation.
[0124] The implementation method for selecting the above-mentioned speech audio processing rules can be found in [link to relevant documentation]. Figure 8 ,like Figure 8 As shown, the image selection window 501 contains a voice selection element 5014. In response to the selection operation of the voice selection element 5014, one or more candidate voice audio processing rules can be displayed in the image selection window, such as candidate voice audio processing rule 50141, etc. When any candidate voice audio processing rule is selected, the selected candidate voice audio processing rule can be determined as the target voice audio processing rule. In this way, when the sound signal input by the target conversation object is subsequently collected, the target voice audio processing rule is used to simulate the sound signal, so that the output sound effect meets the needs of the target conversation object and improves the experience of the target conversation object participating in the video conversation.
[0125] In summary, the embodiments of this application allow video session participants to freely select background images, virtual avatars, and sound effects, thereby satisfying their personalized needs for virtual avatars, background images, and sound effects and enhancing the interactivity of video session participants.
[0126] S203: Obtain the action information of the target session object, and control the target virtual avatar displayed in the avatar display area to perform the target interactive action based on the action information of the target session object. In virtual video session mode, the target virtual avatar can perform the target interactive action based on the action information of the target session object. This target interactive action corresponds to the action performed by the target session object. Visually, the target virtual avatar follows the target session object in performing the corresponding action. The target session object can conduct video sessions with other session objects through the target virtual avatar.
[0127] In this context, the target interactive action performed by the target virtual avatar is matched with the action information of the target conversation object. This matching can be understood as follows: in one implementation, the target interactive action performed by the target virtual avatar is similar to the action indicated by the action information. For example, if the action information of the target conversation object is facial information, meaning the action information indicates that the target conversation object's face should make an action, then the target virtual avatar can be controlled to perform a facial interactive action; for example, if the facial information indicates that the target conversation object's right eye should blink, then the target virtual avatar's right eye can be controlled to blink. Similarly, if the action information of the target conversation object is limb information, meaning the action information indicates that the target conversation object's limbs should make an action, then the target virtual avatar's limbs can be controlled to perform limb actions; for example, if the limb information indicates that the target conversation object's right arm should be raised, then the target virtual avatar's right arm can be raised. For example, if the action information of the target session object is location transfer information, that is, the action information indicates that the target session object's position has changed, then the target virtual image can be controlled to perform the location transfer action within the image display area; if the action indicated by the location transfer information of the target session object is: the target session object moves to the left along the horizontal direction, then the target virtual image can be controlled to move to the left within the image display area.
[0128] In other implementations, there is a mapping relationship between the target virtual avatar's interactive action and the action indicated by the action information. For example, if the action information of the target conversation object is identified as an emotion, then the face of the target virtual avatar can be replaced with a target facial resource that maps to the emotion indicated by the emotion information. In this case, the appearance of the target virtual avatar (e.g., facial shape) may not be the same as the appearance of the target virtual avatar, but both convey the same emotion. For example, if the target conversation object's emotion is laughter, then a target facial resource that maps to laughter can be used to replace the face of the target virtual avatar, making the displayed expression of the target virtual avatar laugh. Similarly, if the target conversation object's emotion is anger, then a facial resource that maps to anger (e.g., a facial animation resource with eyes spitting fire) can be used to replace the face of the target virtual avatar, making the displayed expression of the target virtual avatar angry. The emotion information corresponding to the action information can be determined through model recognition. By pre-setting the mapping relationship between emotions and facial resources, the aforementioned methods of presenting the face of the target virtual avatar can be implemented to reflect various relatively exaggerated emotions. The implementation process of controlling the target virtual image to perform target interactive actions based on the action information of the target session object will be described in detail in later embodiments, and will not be elaborated here.
[0129] It is worth noting that this application embodiment supports exiting the virtual video session mode, that is, switching from the virtual video session mode to the normal video session mode. In the normal video session mode, the real image of the target session object can be displayed in the image display area included in the video session interface. In specific implementation, an exit option is included in the image selection window, such as... Figure 5 The exit option in the image selection window is the "Clear" option, which is the first option in the image selection window. In response to the selection of the exit option, the video session interface is displayed. An environmental image is displayed in the image display area contained in the video session interface. The environmental image is captured by a camera. The environmental image is sent to the peer device so that the peer device can display the environmental image. The peer device refers to the device used by other users participating in the video session.
[0130] Combined with appendix Figure 9 The following describes the implementation method for exiting the virtual video session mode, such as... Figure 9As shown, the image selection window 501 displays an exit option 901. When the exit option 901 is selected, it indicates that the target session object wants to return from the virtual video session mode to the normal video session mode. At this time, the video session interface is displayed, and the environmental image captured by the camera is displayed in the image display area included in the video session interface. This environmental image can include the actual image of the target session object. Simultaneously, the target terminal can also send the captured environmental image to the peer device, so that the image display area displayed on the peer device also displays the environmental image corresponding to the target session object. Of course, exiting the virtual video session mode is not limited to selecting the exit option, and the exit option does not necessarily have to be displayed in the image selection window. It can also be displayed directly in the video session interface, which allows the target session object to quickly switch video session modes, making the operation simple and convenient.
[0131] In this embodiment, a target virtual avatar can be displayed in the video session interface, and the target virtual avatar can be driven to perform target interactive actions based on the action information of the target session object, allowing the target session object to participate in the video session through the target virtual avatar. This method of outputting the target virtual avatar in the video session interface can quickly display the target virtual avatar; furthermore, the target virtual avatar can be used to represent the target session object in the video session. By using a virtual avatar to simulate real-person interaction, the display of the target session object's real image in the video session can be avoided, thus protecting the image privacy of the target session object.
[0132] Please see Figure 10 , Figure 10 This illustration shows a flowchart of another interactive processing method provided by an exemplary embodiment of this application; this interactive processing method can be executed by a target terminal. Figure 10 As shown, the interactive processing method may include, but is not limited to, steps S1001-S1005.
[0133] S1001: Display the video session interface during a video session.
[0134] S1002: Display the target virtual image of the target session object related to the video session within the image display area.
[0135] It should be noted that the specific implementation methods of steps S1001-S1002 can be found in [reference needed]. Figure 2 The specific implementation details of steps S201-S202 in the illustrated embodiment will not be repeated here.
[0136] S1003: Get the set of meshes added for the target virtual avatar.
[0137] S1004: Obtain the action information of the target session object, and perform mesh deformation processing on the mesh data of the target mesh in the mesh set based on the action information of the target session object.
[0138] S1005: Based on the mesh data after mesh deformation processing, render and display the target virtual image that has performed the target interactive action.
[0139] In steps S1003-S1005, the target virtual image proposed in this application embodiment is a 2D (two-dimensional) virtual image. A 2D virtual image can be called a two-dimensional virtual image, where any point can be represented by the x-axis and y-axis; that is, a two-dimensional virtual image is a planar graphic. This 2D virtual image has low requirements for the configuration capabilities of the target terminal and is easy and quick to move. Its application in video conferencing scenarios can improve the control efficiency of the target virtual image. Furthermore, to make the 2D virtual image perform target interactive actions more realistically and naturally, this application embodiment also employs numerous 3D-to-2D conversion methods. Specifically: ① Rotation can be expressed through the deformation of elements such as the face, head, and hair of the target virtual image, making the target virtual image appear fuller and more three-dimensional within a micro-motion range. ② When the head or torso elements swing, real-world physical elasticity is increased, such as increasing the swinging elasticity of hair and accessory elements, making the target virtual image appear more realistic. ③ When the target virtual image does not perform any interactive actions, you can add the sense of body movement (such as chest movement) when breathing, and randomly make slight body swaying movements to make the target virtual image more realistic.
[0140] To enable the 2D virtual avatar to perform smooth, three-dimensional movements and express emotions in sync with the target conversation object, this embodiment of the application also creates a mesh for each object element contained in the target virtual avatar. All the meshes corresponding to the object elements contained in the target virtual avatar form a mesh set, with one mesh corresponding to one object element. Each mesh consists of at least three mesh vertices. An object element can refer to a single element that makes up the target virtual avatar. For example, the hair of the target virtual avatar can be called a hair element, which is composed of multiple sub-hair elements (such as sideburns, sideburns, etc.); similarly, the arm of the target virtual avatar can be called a limb element, and so on. It should be noted that, to improve the detail and accuracy of the target virtual avatar's execution of target interactive actions, this embodiment of the application supports a mesh corresponding to an object element that can include multiple sub-meshes. This allows for partial changes to the position of the object element corresponding to that mesh to be achieved by controlling a specific sub-mesh. Furthermore, for ease of explanation, this embodiment of the application refers to the mesh corresponding to the object element performing the interactive action as the target mesh, which is explained here.
[0141] An exemplary diagram illustrating the addition of a mesh to a target virtual avatar can be found here. Figure 11a ,like Figure 11a As shown, after adding meshes to the face and hand elements of the original target virtual image shown in the first image, the second image with added meshes is obtained. In the second image, the mesh corresponding to the face element of the target virtual image contains multiple sub-meshes, each composed of multiple connected mesh vertices. For example, the mesh corresponding to the face element includes sub-mesh 1104, sub-mesh 1105, sub-mesh 1106, ..., etc., where sub-mesh 1104 includes three mesh vertices: mesh vertex 11041, mesh vertex 11042, and mesh vertex 11043. It is understandable that... Figure 11a This is merely an exemplary method of adding meshes to the face and hand elements of a target virtual avatar. In actual application scenarios, the number of sub-mesh elements, the number of mesh vertices, and their distribution may vary. This application embodiment does not limit the meshes added to the target virtual avatar, but this is explained here.
[0142] In practice, the driving mechanism for any given grid is achieved by modifying its mesh data, thereby controlling the corresponding object elements within the target virtual object to perform target interactive actions. The set of grids corresponding to the target virtual avatar contains multiple grids and their mesh data. The mesh data for any given grid refers to the state values of its individual vertices. These state values can represent the vertex's position or its positional relationship with other connected vertices. Different state values for each grid vertex result in different interactive actions being performed by the target virtual avatar rendered from the mesh data. This allows the target mesh to deform in real-time using the action information of the target session object, thereby driving the corresponding object elements within the target virtual object to perform target interactive actions. This enables the 2D virtual avatar to execute various 2D patch motion effects, allowing it to flexibly follow the target session object in performing interactive actions.
[0143] Combined with appendix Figure 11b This describes the process of performing mesh deformation processing on the mesh data, such as... Figure 11bAs shown, assuming the target mesh consists of five vertices, namely vertex A, vertex B, vertex C, vertex D, and vertex E, these five vertices are sequentially adjacent. By controlling the position information of all or some of the vertices to change, the mesh data of the target mesh after mesh deformation processing is obtained. Based on the mesh data of the target mesh after mesh deformation processing, the target virtual image is rendered, thereby achieving the effect of controlling the target virtual image to perform interactive actions within the image display area. An exemplary diagram illustrating the changes in the target mesh data during the rendering process of the target virtual image can be found in [reference needed]. Figure 11c ,like Figure 11c As shown, after the mesh corresponding to the left hair element 1107 of the target virtual image is moved to the right along the horizontal direction, the mesh corresponding to the left hair element 1107 can be deformed; then, based on the mesh data of the mesh after the mesh deformation, the target virtual image after the left hair element 1107 has been moved to the left can be rendered.
[0144] It should be noted that in the process of rendering the target virtual image based on grid data, the target virtual image is rendered sequentially according to the hierarchical relationship of the virtual image; for example... Figure 11d As shown, taking the rendering of the face and hair of a target virtual image as an example, the display hierarchy of the face element is lower than that of the hair element. The hair element comprises multiple sub-elements, such as the left hair element and the middle hair element. The display hierarchy of these sub-elements is as follows: the left hair element has a lower display hierarchy than the middle hair element. Based on this, when rendering the target virtual image according to the above display hierarchy from low to high, the rendering order of each element is: face element → left hair element → middle hair element, thus obtaining a target virtual image with a rendered face and hair.
[0145] Based on the above description of the mesh added to the target virtual avatar and the related mesh data, the following section uses facial and limb information as examples of motion information to illustrate... Figure 2 The front-end interface and back-end technology of step S203 in the illustrated embodiment, which controls the target virtual image displayed in the image display area to perform target interactive actions based on the action information of the target session object, are described in detail below. Wherein:
[0146] (1) The action information of the target conversation object is facial information. The actions performed by the target conversation object indicated by the facial information may include: head turning actions (such as turning the head to the side, raising the head, lowering the head, tilting the head, etc.), facial feature changes (such as the mouth being able to reflect the degree of opening and closing of the corners of the mouth upward, downward, or pursed inward, etc.). An exemplary diagram illustrating the control of the target virtual image to perform target interactive actions based on the target conversation object's facial information can be found in [reference needed]. Figure 12 ,like Figure 12 As shown in the first image, if the target session object 1201's facial information indicates that its head is facing the camera, the target virtual avatar's head can be controlled to remain facing forward; if the target session object's facial information indicates that its head is turning to the right, the target virtual avatar's head can be controlled to perform the target interactive action of turning its head to the right; if the target session object's facial information indicates that its head is turning to the left, the target virtual avatar's head can be controlled to perform the target interactive action of turning its head to the left. For example... Figure 12 The second image shows that, based on the degree to which the target conversation object's eyes are open or closed, it represents the target conversation object controlling its eyes to perform different actions, such as blinking, closing its eyes, or staring. Based on the shape of the target conversation object's eyebrows, it represents the target conversation object controlling its eyebrows to perform different actions, such as frowning or raising its eyebrows. In this implementation, if the target conversation object's facial information shows the right eye blinking, then the target virtual character's right eye can be controlled to perform a blinking interaction; if the target conversation object's facial information shows staring, then the target virtual character can be controlled to perform a staring interaction; and if the target conversation object's facial information shows frowning, then the target virtual character's eyebrows can be controlled to perform a frowning interaction.
[0147] In this embodiment, the target virtual image can be controlled to perform target interactive actions by performing mesh deformation processing on the target mesh corresponding to the control object element described above. Specifically, the facial information of the target session object may include N feature points of the target session object's face, where N is an integer greater than 1. An exemplary schematic diagram for identifying the feature points of the target session object's face can be found here. Figure 13a The following is in conjunction with the appendix. Figure 13b This paper introduces the implementation method of controlling the target virtual image to perform target interactive actions based on N feature point information. First, based on the N feature point information, the updated expression type of the target session object is determined (i.e., the expression type of the target session object after performing the action). Second, the expression base coefficient corresponding to the updated expression type is obtained. Then, according to the expression base coefficient corresponding to the updated expression type, the grid where the object element corresponding to the updated expression type is located is subjected to grid deformation processing, which is essentially performing grid deformation processing on the grid data. Finally, based on the grid data after grid deformation processing, the target virtual image performing the target interactive action is rendered and displayed.
[0148] The expression type of the target session object can refer to the type of facial expression on the target session object. Under different expression types, the same object element on the target session object's face will present different appearances. For example, when the expression type is smiling, the appearance of the object element—the mouth—is that the corners of the mouth are turned upwards; similarly, when the expression type is crying, the appearance of the object element—the mouth—is that the corners of the mouth are turned downwards. The expression base coefficient of the expression type can be a coefficient used to represent the appearance of the object element; for example, the expression base coefficient of the mouth can be defined in the range [0,1]. It is assumed that when the expression base coefficient is 0, the mouth is closed; when the expression base coefficient is any value in (0,1) (such as 0.5), the mouth is open; and when the expression base coefficient is 1, the mouth is open to its maximum extent.
[0149] The method for obtaining the updated expression base coefficients corresponding to the expression type described above may include: capturing an image containing the target conversation object using a camera, extracting information from the feature points of the target conversation object's face in the image to obtain N feature point information; and fitting and generating expression base coefficients for each object element of the face under the current expression type based on the N feature point information. Generally, N can be 83, meaning the target conversation object has 83 feature point information, and the number of expression base coefficients for the object elements of the face generated by fitting is 52, meaning the multiple object elements of the target conversation object's face can have a total of 52 expression states.
[0150] It is understandable that during the rendering of the target virtual avatar based on the updated expression base coefficients, the object elements of the target virtual avatar visually exhibit a dynamic change process. In other words, it is also necessary to obtain the expression base coefficients of the object elements in the intermediate state based on the expression base coefficients of the expression types before the update (i.e., before the target session object performs an action) and after the update, so as to render a more continuous dynamic change process of the object elements based on the expression base coefficients before the update, the intermediate state, and the updated state. For example, visually, when the target session object performs the action of opening its mouth, the object element of the target virtual avatar—the mouth—dynamically changes from closed to open, that is, the size of the mouth gradually increases, which is achieved through dynamic changes in the mesh corresponding to the mouth. A schematic diagram of the dynamic changes of the mouth mesh can be found here. Figure 13c In this implementation, a schematic diagram of the target session object and the target virtual avatar after the mouth-opening action can be found here. Figure 13d .
[0151] In practical implementation, the expression base coefficients of the expression type before the update can be obtained. The difference between the expression base coefficients of the updated expression type and those of the original expression type is calculated to obtain the difference result. Then, based on the difference result, the expression base coefficients of the updated expression state, and the expression base coefficients of the original expression state, mesh deformation processing is performed on the mesh containing the object element corresponding to the updated expression type. This yields intermediate mesh data after mesh deformation processing, allowing for a more continuous dynamic change process of the object element to be rendered based on this intermediate mesh data. This process can be simply viewed as an interpolation process. For example, during the process of increasing the size of the mouth, an interpolation algorithm can be used to insert some pixels into the increased area, so that the various feature points on the mouth can be clearly displayed even as the mouth size increases, making the entire dynamic process appear smooth and clear. Interpolation algorithms can include, but are not limited to, linear interpolation algorithms and Bézier curve interpolation algorithms. A flowchart illustrating the process of determining the intermediate state of the mouth using an interpolation algorithm, using the dynamic change of the mouth described above as an example, can be found in [link to flowchart]. Figure 13e .
[0152] (2) The action information of the target session object is limb information. Limb information can be used to reflect the state of the target session object's torso. The torso (or limbs) can include, but is not limited to, the arms (such as the upper arm, forearm, and hand), thighs, calves, feet, etc. For ease of explanation, the arm will be used as an example in the following description. For example, if the limb information of the target session object instructs the target session object to wave its right hand, then the right hand of the target virtual avatar can also be controlled to perform a waving interaction. An exemplary diagram of this implementation method can be found in [reference needed]. Figure 14a ,like Figure 14a As shown, when the target conversation object performs a waving gesture, the target virtual avatar can be flexibly controlled to also perform a waving interaction based on the target conversation object's limb information. Specifically, based on the target conversation object's limb information, the position information of S limb points of the target conversation object can be determined, where S is a positive integer. Then, based on the position information of each of the S limb points, the angle value of the corresponding torso element is calculated. The mesh corresponding to the torso element in the target virtual avatar is then deformed according to the angle value, and the target virtual avatar performing the target interactive action is rendered based on the mesh data after the mesh deformation. In other words, by analyzing the collected limb information of the target conversation object, the position information of each limb point of the torso can be obtained; based on the position information of each limb point, the angle value between each limb point (or torso angle) is calculated, for example, the angle between the elbow and shoulder is calculated to be 60 degrees; then, based on the angle value, the mesh corresponding to the torso element is deformed, thereby controlling the limbs corresponding to the torso element in the target virtual avatar to perform the target interactive action.
[0153] For example, Figure 14b This illustration shows a schematic diagram of a method for controlling a virtual avatar to perform interactive actions using limb positioning, as provided in an embodiment of this application. Figure 14b As shown, the left and right arms of the target virtual character each include three limb points. Taking the left arm as an example, it includes limb point A, limb point B, and limb point C. Based on the position information of limb point A, limb point B, and limb point C, the angle values of the torso elements can be obtained. For example, the angle between the forearm (i.e., the torso element) and the upper arm is valued as 'a', and the angle between the upper arm and the shoulder is valued as 'b'. Then, based on these angle values, the mesh of the corresponding object element of the target virtual character can be deformed, thereby controlling the left arm of the target virtual character to perform the same action as the left arm of the target virtual character. This driving method, which uses limb points to drive the torso's actions, makes the target virtual character more vivid and natural.
[0154] Of course, in addition to the above-described mesh deformation processing of the mesh corresponding to the torso element in the target virtual image based on the angle value, so that the torso element of the target virtual image performs the same angle value as the torso element of the target conversation object; this application embodiment also supports triggering the playback of the torso animation set for the target virtual image after detecting the action type of the target conversation object's limb movement, thereby controlling the target virtual image to perform a torso movement similar to the target conversation object's movement. For example, after detecting that the target conversation object is performing a waving action, the waving action set for the target virtual image can be triggered, at which time the angle value between the upper arm and forearm (or upper arm and shoulder) of the target virtual image when performing the waving action may not match the angle value of the target conversation object. As another example, when detecting that the target conversation object is performing a single-handed heart gesture, a two-handed heart gesture set for the target virtual image can be triggered. As yet another example, when detecting that the target conversation object is performing an "OK" gesture, an "OK" gesture set for the target virtual image can be triggered; and so on. This method, which directly triggers the target virtual avatar to execute a matching torso animation after detecting the target session object's action type, can improve the smoothness and speed of the target virtual avatar's interactive actions to a certain extent.
[0155] In summary, steps S1003-S1005 illustrate an implementation method provided by this application embodiment for controlling a target virtual avatar to perform target interactive actions based on the action information of the target conversation object. Specifically, taking facial information and limb information as examples, the method involves performing mesh deformation processing on the mesh based on the action information. However, it is understood that the implementation method for controlling a target virtual avatar to perform target interactive actions based on the action information of the target conversation object is not necessarily obtained by controlling the mesh to perform mesh deformation processing. Below, taking emotion information and position transfer information as examples, another implementation method for controlling a target virtual avatar to perform target interactive actions, including front-end interface and back-end technology, is given, wherein:
[0156] (1) The action information of the target conversation object is emotional information. Emotional information can be used to indicate the type of emotion of the target conversation object, such as laughing, angry, surprised, crying, etc. It is understood that the above-described method of driving the facial object elements of the target virtual image to perform some actions based on the expression type of the target conversation object can convey the emotion of the target conversation object to a certain extent; however, for the target virtual image in the two-dimensional world, more exaggerated expressions are often used to express emotions. In order to ensure that the target virtual image presents more exaggerated expressions, this application embodiment sets a mapping relationship of exaggerated expressions between the target conversation object and the target virtual image; when the target conversation object is detected to have a preset exaggerated expression, the original facial material of the target virtual image can be replaced with the preset exaggerated expression to change the original facial material of the target virtual image into exaggerated facial material, thereby achieving rapid face changing and improving control efficiency. See also Figure 15a Assuming the target conversation object expresses anger, a preset angry facial resource can be used to replace the original facial resource of the target virtual avatar, resulting in a target virtual avatar conveying the emotion of anger. Similarly, assuming the target conversation object expresses laughter, a preset laughing facial resource can be used to replace the original facial resource of the target virtual avatar, resulting in a target virtual avatar conveying the emotion of laughter. Several exemplary exaggerated expressions are provided in this application's embodiments. Figure 15b .
[0157] In specific implementation, the emotional state of the target conversation object can be identified based on emotional information to obtain the current emotional state of the target conversation object; based on the current emotional state, a target facial resource matching the current emotional state is determined; then, the face of the target virtual image is updated using the target facial resource in the image display area to obtain the updated target virtual image. Optionally, an emotion recognition model can be used to identify the emotional state of the target conversation object. More specifically, deep learning methods can be used to perform emotion recognition and emotion classification on images containing the target conversation object acquired in real time, including but not limited to: surprise, laughter, anger, smile, and other emotions with specific semantics; when the emotion recognition classification corresponding to the target conversation object is identified, the target virtual image can be triggered to execute the corresponding exaggerated expression effect. The exaggerated expression effect of the target virtual image is achieved by: after identifying the emotion classification result of the target conversation object, hiding the original facial resource of the target virtual image and displaying the preset target facial resource on the face, thus achieving the effect of displaying an exaggerated expression. A flowchart of the above-described process of replacing facial resources can be found in [reference needed]. Figure 15c This will not be described in detail here.
[0158] (2) The action information of the target session object is positional transfer information. Positional transfer information can be used to indicate the movement of the target session object in the environment. Specifically, the positional transfer information of a certain object element contained in the target session object can be used to represent the positional transfer information of the target session object. For example, the facial element of the target session object moves to the left horizontally, the face element of the target session object moves upward vertically, the display area of the face element of the target session object shrinks (indicating a change in the distance between the target session object and the terminal screen), and so on. Based on the positional transfer information of the target session object, the target virtual image is driven to perform corresponding positional transfer actions within the image display area, so that the target session object and the target virtual image have a good sense of correspondence.
[0159] In specific implementation, after obtaining the position transfer information of the target session object, if the position transfer information is detected to be the movement information of a target image point, the target virtual image is controlled to move and display within the display area based on the movement information of the target image point; if the position transfer information is the movement information of a target image region, the size of the display area of the target virtual image within the image display area can be adjusted based on the movement information of the target image region. Here, a target image point can refer to a point within the display area where the target virtual image is located, for example, the center point of the target virtual image's face; a target image region can refer to an image region within the display area where the target virtual image is located, such as the face region of the target virtual image. Specifically, to enable the target virtual image to have a three-dimensional perspective effect, this application embodiment supports driving the target virtual image to move and rotate along the x, y, and z axes based on the position transfer information of the target session object, thereby controlling the movement and rotation of the target virtual image within the image display area.
[0160] The implementation logic for driving the target virtual image to move along the x, y, and z axes may include: ① Acquiring multiple consecutive frames of environmental images and identifying the positional shift information of the target image point (such as the center point of the face) of the virtual image in the multiple frames of environmental images, and driving the virtual image to perform a matching movement within the image display area based on the positional shift information. The positional shift information may include any one of the following: horizontal positional shift information along the x-axis, vertical positional shift information along the y-axis, or horizontal positional shift information along the x-axis and vertical positional shift information along the y-axis. ② Acquiring multiple consecutive frames of environmental images and identifying the change information of the display area occupied by the target image region (such as the face region) of the virtual image in the environmental images, and enlarging or shrinking the virtual image within the image display area based on the change information of the display area to achieve control over the dynamic effects of the virtual image's changes in the z-axis direction.
[0161] The implementation logic for driving the target virtual avatar to perform a rotation operation may include: acquiring an environmental image, recognizing the face of the target session object in the environmental image, obtaining the Euler angle of the current face orientation, and then using the Euler angle of the current face orientation to control the mesh corresponding to the face element of the target virtual avatar for mesh deformation processing, thereby achieving the effect of controlling the rotation of the target virtual avatar's face. Of course, if other body parts of the target session object (such as the shoulder) rotate, the above implementation method can be used to control the mesh corresponding to the shoulder element of the target virtual avatar for mesh deformation processing, thereby achieving the effect of controlling the rotation of the target virtual avatar's shoulder; the embodiments of this application do not limit the object element for controlling the rotation of the target virtual avatar.
[0162] In this embodiment, a target virtual avatar can be displayed in the video session interface, and driven to perform target interactive actions based on the action information of the target session object, allowing the target session object to participate in the video session through the target virtual avatar. This method of outputting the target virtual avatar in the video session interface allows for quick display of the virtual avatar; furthermore, the virtual avatar can be used to represent the target session object in the video session, simulating real-person interaction, thus avoiding the display of the target session object's real image in the video session and protecting the target session object's privacy. Additionally, the target virtual avatar is a 2D virtual avatar, and its corresponding grid is driven to flexibly follow the target session object's actions to perform target interactive actions; this innovation in 2D virtual avatars allows these low-cost virtual avatars to achieve a similar level of realism to 3D virtual avatars, thereby reducing the risk of unsuccessful communication with the virtual avatar.
[0163] Figure 16 This illustration shows a schematic diagram of an interactive processing device provided in an embodiment of this application; the interactive processing device is disposed in a target terminal. In some embodiments, the interactive processing device may be a client running on the target terminal (such as an application with video conferencing capabilities); the specific implementation of the units included in the interactive processing device can be referred to the description of the relevant content in the foregoing embodiments. Please refer to... Figure 16 The interactive processing device in this application embodiment includes the following units:
[0164] Display unit 1601 is used to display a video session interface during a video session, the video session interface including an image display area for displaying video session objects;
[0165] Processing unit 1602 is used to display a target virtual image of a target session object related to the video session within the image display area;
[0166] The processing unit 1602 is also used to acquire the action information of the target session object, and control the target virtual image displayed in the image display area to perform the target interactive action based on the action information of the target session object.
[0167] In one implementation, the processing unit 1602 is further configured to:
[0168] Display an image selection window, which includes image selection elements;
[0169] In response to a trigger operation on the image selection element, a reference virtual image is displayed in the reference display area, and candidate virtual images are displayed in the image selection window;
[0170] In response to the image selection operation of the candidate virtual image, the reference virtual image is updated and displayed in the reference display area as the target candidate virtual image selected by the image selection operation;
[0171] In response to the virtual avatar confirmation operation, the target candidate virtual avatar is determined as the target virtual avatar.
[0172] In one implementation, the processing unit 1602 is further configured to:
[0173] The background selection element is displayed in the image selection window;
[0174] In response to a user action on the background selection element, candidate background images are displayed in the image selection window;
[0175] In response to a background selection operation on the candidate background image, the target candidate background image selected by the background selection operation is displayed in the reference display area;
[0176] In response to the background image confirmation operation, the target candidate background image is set as the background image of the image display area.
[0177] In one implementation, the processing unit 1602 is further configured to:
[0178] The voice selection element is displayed in the image selection window;
[0179] In response to the selection operation of the voice selection element, candidate voice audio processing rules are displayed in the image selection window;
[0180] In response to the confirmation operation of the candidate speech audio processing rule, the candidate speech audio processing rule is determined as the target speech audio processing rule, which is used to simulate the sound signal of the target session object received during the video session.
[0181] In one implementation, the image selection window includes an exit option, and the processing unit 1602 is further configured to:
[0182] In response to the selection of the exit option, the video session interface is displayed;
[0183] An environmental image is displayed within the image display area included in the video session interface. The environmental image is obtained by capturing the environment.
[0184] The environmental image is sent to the peer device so that the peer device can display the environmental image. The peer device refers to the device used by other users participating in the video session.
[0185] In one implementation, when the processing unit 1602 controls the target virtual image displayed in the image display area to perform a target interactive action based on the action information of the target session object, it is specifically used to perform any one or more of the following steps:
[0186] If the action information of the target session object is facial information, then control the target virtual image to perform facial interaction actions;
[0187] If the action information of the target session object is emotional information, then the face of the target virtual avatar is replaced by the target facial resource associated with the emotional information.
[0188] If the action information of the target session object is limb information, then control the target limb of the target virtual image to perform limb actions;
[0189] If the action information of the target session object is location transfer information, then the target virtual image is controlled to perform a location transfer action within the image display area.
[0190] In one implementation, when the processing unit 1602 controls the target virtual avatar displayed in the avatar display area to perform a target interactive action based on the action information of the target session object, it is specifically used for:
[0191] Obtain a set of meshes added for the target virtual image. The set of meshes includes multiple meshes and mesh data for each mesh. Each mesh corresponds to an object element. Each mesh consists of at least three mesh vertices. The mesh data for each mesh refers to the state values of each mesh vertex contained in each mesh.
[0192] Based on the action information of the target session object, perform mesh deformation processing on the mesh data of the target mesh in the mesh set;
[0193] Based on the mesh data after mesh deformation processing, the target virtual image that has performed the target interactive action is rendered and displayed. In the rendered and displayed target virtual image that has performed the target interactive action, the position and / or shape of the object elements corresponding to the target mesh change.
[0194] In one implementation, the action information includes facial information, which includes N feature points of the target session object's face, where N is an integer greater than 1; the processing unit 1602 is used to perform mesh deformation processing on the mesh data of the target mesh in the mesh set according to the action information of the target session object, specifically for:
[0195] Based on the N feature point information, the updated expression type of the target session object is determined;
[0196] Obtain the expression base coefficients corresponding to the updated expression type;
[0197] Based on the updated expression base coefficients, perform mesh deformation processing on the mesh containing the object elements corresponding to the updated expression type.
[0198] In one implementation, when processing unit 1602 performs mesh deformation processing on the mesh containing the object element corresponding to the updated expression type based on the expression base coefficients corresponding to the updated expression type, it is specifically used for:
[0199] Obtain the base coefficients of the emoji type before the update;
[0200] The difference between the expression base coefficients corresponding to the updated expression type and the expression base coefficients corresponding to the expression type before the update is calculated to obtain the difference result.
[0201] Based on the difference result, the expression base coefficient of the updated expression state, and the expression base coefficient of the expression state before the update, the mesh of the object element corresponding to the updated expression type is subjected to mesh deformation processing.
[0202] In one implementation, the action information includes limb information. When the processing unit 1602 performs mesh deformation processing on the mesh data of the target mesh in the mesh set based on the action information of the target session object, it specifically performs the following:
[0203] Based on the limb information of the target session object, determine the position information of S limb points of the target session object, where S is an integer greater than zero;
[0204] Based on the position information of the S limb points, calculate the angle value of the corresponding torso element;
[0205] The mesh corresponding to the torso element in the target virtual image is subjected to mesh deformation processing based on the angle value.
[0206] In one implementation, the action information includes emotional information. When the processing unit 1602 controls the target virtual avatar displayed in the avatar display area to perform a target interactive action based on the action information of the target conversation object, it specifically performs the following:
[0207] The emotional state of the target conversation object is identified based on the emotional information to obtain the current emotional state of the target conversation object;
[0208] Based on the current emotional state, a target facial resource matching the current emotional state is determined;
[0209] The face of the target virtual image is updated using the target facial resources within the image display area to obtain the updated target virtual image.
[0210] In one implementation, the action information includes position transfer information. When the processing unit 1602 controls the target virtual avatar displayed in the avatar display area to perform a target interactive action based on the action information of the target session object, it specifically performs the following:
[0211] If the position transfer information is the movement information of the target image point, then the target virtual image is controlled to move and be displayed in the image display area according to the movement information of the target image point;
[0212] If the position transfer information is the movement information of the target image region, the size of the display area of the target virtual image in the image display area is adjusted according to the movement information of the target image region.
[0213] In one implementation, the processing unit 1602 is further configured to:
[0214] When the target terminal initiates a video session, it acquires a target image of the environment.
[0215] The target session object in the target image is subjected to feature recognition processing to obtain the recognition result;
[0216] Assign a virtual avatar that matches the recognition result to the target session object, and identify the virtual avatar that matches the recognition result as the target virtual avatar.
[0217] According to one embodiment of this application, Figure 16 The interactive processing method shown can be constructed by combining each unit into one or more other units, or one or more units can be further divided into multiple functionally smaller units. This can achieve the same operation without affecting the technical effect of the embodiments of this application. The above units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the interactive processing device may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by multiple units working together. According to another embodiment of this application, the interactive processing device can be executed by running on a general-purpose computing device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). Figure 2 and Figure 10The computer program (including program code) for each step involved in the corresponding method shown, to construct such... Figure 16 The interactive processing apparatus shown herein, and the interactive processing method for implementing the embodiments of this application, are described. A computer program may be recorded on, for example, a computer-readable recording medium, loaded onto the aforementioned computing device via the computer-readable recording medium, and executed therein.
[0218] In this embodiment, the display unit 1601 can display the target virtual avatar in the video session interface, and the processing unit 1602 can drive the target virtual avatar to perform target interactive actions based on the action information of the target session object, so that the target session object can participate in the video session through the target virtual avatar. This method of outputting the target virtual avatar in the video session interface can quickly display the target virtual avatar; moreover, the target virtual avatar can be used to represent the target session object in the video session. By using a virtual avatar to simulate real-person interaction, the real image of the target session object can be avoided in the video session, thus protecting the image privacy of the target session object.
[0219] Figure 17 A schematic diagram of the structure of a terminal provided in an exemplary embodiment of this application is shown. Please refer to... Figure 17 The terminal includes a storage device 1701 and a processor 1702. In this embodiment, the terminal also includes a network interface 1703 and a user interface 1704. The terminal can be a smartphone, tablet, smart wearable device, etc., and can access the Internet through the network interface 1703 to communicate and exchange data with servers and other electronic devices. The user interface 1704 can include a touch screen, etc., and can receive user operations and display various interfaces to the user to facilitate receiving user operations.
[0220] Storage device 1701 may include volatile memory, such as random-access memory (RAM); storage device 1701 may also include non-volatile memory, such as flash memory, solid-state drive (SSD), etc.; storage device 1701 may also include combinations of the above types of memory.
[0221] Processor 1702 may be a central processing unit (CPU). Processor 1702 may further include hardware chips. The aforementioned hardware chips may be application-specific integrated circuits (ASICs), programmable logic devices (PLDs), etc. The aforementioned PLDs may be field-programmable gate arrays (FPGAs), generic array logic (GALs), etc.
[0222] The storage device 1701 in this embodiment stores a computer program. The processor 1702 calls the computer program in the storage device. When the computer program is executed, the processor 1702 can be used to implement the above-mentioned functions, such as... Figure 2 as well as Figure 10 The methods described in the corresponding embodiments, etc.
[0223] In one embodiment, the terminal may correspond to the target terminal described above; the storage device 1701 stores a computer program; the processor 1702 loads and executes the computer program to implement the corresponding steps in the above-described interactive processing method embodiment; specifically, the processor 1702 is used to perform the following steps:
[0224] During a video session, a video session interface is displayed, which includes an area for displaying images of the video session objects.
[0225] The target virtual image of the target session object related to the video session is displayed within the image display area;
[0226] Obtain the action information of the target session object, and control the target virtual image displayed in the image display area to perform the target interactive action based on the action information of the target session object.
[0227] In one implementation, the processor 1702 is also used to perform the following steps:
[0228] Display an image selection window, which includes image selection elements;
[0229] In response to a trigger operation on the image selection element, a reference virtual image is displayed in the reference display area, and candidate virtual images are displayed in the image selection window;
[0230] In response to the image selection operation of the candidate virtual image, the reference virtual image is updated and displayed in the reference display area as the target candidate virtual image selected by the image selection operation;
[0231] In response to the virtual avatar confirmation operation, the target candidate virtual avatar is determined as the target virtual avatar.
[0232] In one implementation, the processor 1702 is also used to perform the following steps:
[0233] The background selection element is displayed in the image selection window;
[0234] In response to a user action on the background selection element, candidate background images are displayed in the image selection window;
[0235] In response to a background selection operation on the candidate background image, the target candidate background image selected by the background selection operation is displayed in the reference display area;
[0236] In response to the background image confirmation operation, the target candidate background image is set as the background image of the image display area.
[0237] In one implementation, the processor 1702 is also used to perform the following steps:
[0238] The voice selection element is displayed in the image selection window;
[0239] In response to the selection operation of the voice selection element, candidate voice audio processing rules are displayed in the image selection window;
[0240] In response to the confirmation operation of the candidate speech audio processing rule, the candidate speech audio processing rule is determined as the target speech audio processing rule, which is used to simulate the sound signal of the target session object received during the video session.
[0241] In one implementation, the image selection window includes an exit option, and the processor 1702 is further configured to perform the following steps:
[0242] In response to the selection of the exit option, the video session interface is displayed;
[0243] An environmental image is displayed within the image display area included in the video session interface. The environmental image is obtained by capturing the environment.
[0244] The environmental image is sent to the peer device so that the peer device can display the environmental image. The peer device refers to the device used by other users participating in the video session.
[0245] In one implementation, when the processor 1702 controls the target virtual avatar displayed in the avatar display area to perform a target interactive action based on the action information of the target session object, it is specifically used to perform any one or more of the following steps:
[0246] If the action information of the target session object is facial information, then control the target virtual image to perform facial interaction actions;
[0247] If the action information of the target session object is emotional information, then the face of the target virtual avatar is replaced by the target facial resource associated with the emotional information.
[0248] If the action information of the target session object is limb information, then control the target limb of the target virtual image to perform limb actions;
[0249] If the action information of the target session object is location transfer information, then the target virtual image is controlled to perform a location transfer action within the image display area.
[0250] In one implementation, when the processor 1702 controls the target virtual avatar displayed in the avatar display area to perform a target interactive action based on the action information of the target session object, it specifically performs the following steps:
[0251] Obtain a set of meshes added for the target virtual image. The set of meshes includes multiple meshes and mesh data for each mesh. Each mesh corresponds to an object element. Each mesh consists of at least three mesh vertices. The mesh data for each mesh refers to the state values of each mesh vertex contained in each mesh.
[0252] Based on the action information of the target session object, perform mesh deformation processing on the mesh data of the target mesh in the mesh set;
[0253] Based on the mesh data after mesh deformation processing, the target virtual image that has performed the target interactive action is rendered and displayed. In the rendered and displayed target virtual image that has performed the target interactive action, the position and / or shape of the object elements corresponding to the target mesh change.
[0254] In one implementation, the action information includes facial information, which includes N feature points of the face of the target session object, where N is an integer greater than 1. When the processor 1702 performs mesh deformation processing on the mesh data of the target mesh in the mesh set based on the action information of the target session object, it specifically executes the following steps:
[0255] Based on the N feature point information, the updated expression type of the target session object is determined;
[0256] Obtain the expression base coefficients corresponding to the updated expression type;
[0257] Based on the updated expression base coefficients, perform mesh deformation processing on the mesh containing the object elements corresponding to the updated expression type.
[0258] In one implementation, when the processor 1702 performs mesh deformation processing on the mesh containing the object element corresponding to the updated expression type based on the expression base coefficients corresponding to the updated expression type, it specifically performs the following steps:
[0259] Obtain the base coefficients of the emoji type before the update;
[0260] The difference between the expression base coefficients corresponding to the updated expression type and the expression base coefficients corresponding to the expression type before the update is calculated to obtain the difference result.
[0261] Based on the difference result, the expression base coefficient of the updated expression state, and the expression base coefficient of the expression state before the update, the mesh of the object element corresponding to the updated expression type is subjected to mesh deformation processing.
[0262] In one implementation, the action information includes limb information. When the processor 1702 performs mesh deformation processing on the mesh data of the target mesh in the mesh set based on the action information of the target session object, it specifically performs the following steps:
[0263] Based on the limb information of the target session object, determine the position information of S limb points of the target session object, where S is an integer greater than zero;
[0264] Based on the position information of the S limb points, calculate the angle value of the corresponding torso element;
[0265] The mesh corresponding to the torso element in the target virtual image is subjected to mesh deformation processing based on the angle value.
[0266] In one implementation, the action information includes emotion information. When the processor 1702 controls the target virtual avatar displayed in the avatar display area to perform a target interactive action based on the action information of the target session object, it specifically performs the following steps:
[0267] The emotional state of the target conversation object is identified based on the emotional information to obtain the current emotional state of the target conversation object;
[0268] Based on the current emotional state, a target facial resource matching the current emotional state is determined;
[0269] The face of the target virtual image is updated using the target facial resources within the image display area to obtain the updated target virtual image.
[0270] In one implementation, the action information includes position transfer information. When the processor 1702 controls the target virtual image displayed in the image display area to perform a target interactive action based on the action information of the target session object, it specifically performs the following steps:
[0271] If the position transfer information is the movement information of the target image point, then the target virtual image is controlled to move and be displayed in the image display area according to the movement information of the target image point;
[0272] If the position transfer information is the movement information of the target image region, the size of the display area of the target virtual image in the image display area is adjusted according to the movement information of the target image region.
[0273] In one implementation, the processor 1702 is also used to perform the following steps:
[0274] When the target terminal initiates a video session, it acquires a target image of the environment.
[0275] The target session object in the target image is subjected to feature recognition processing to obtain the recognition result;
[0276] Assign a virtual avatar that matches the recognition result to the target session object, and identify the virtual avatar that matches the recognition result as the target virtual avatar.
[0277] In this embodiment, the processor 1702 can display a target virtual avatar in the video session interface and drive the target virtual avatar to perform target interactive actions based on the action information of the target session object, enabling the target session object to participate in the video session through the target virtual avatar. This method of outputting the target virtual avatar in the video session interface can quickly display the target virtual avatar; furthermore, the target virtual avatar can be used to represent the target session object in the video session. By using a virtual avatar to simulate real-person interaction, the display of the target session object's real image in the video session can be avoided, thus protecting the image privacy of the target session object.
[0278] This application embodiment also provides a computer-readable storage medium (Memory), which is a memory device in an electronic device used to store programs and data. It is understood that the computer-readable storage medium here can include both built-in storage media in the electronic device and extended storage media supported by the electronic device. The computer-readable storage medium provides storage space that stores the processing system of the electronic device. Furthermore, the storage space also stores computer programs (including program code) suitable for loading and execution by the processor 1702. It should be noted that the computer-readable storage medium here can be high-speed RAM memory or non-volatile memory, such as at least one disk storage device; optionally, it can also be at least one computer-readable storage medium located remotely from the aforementioned processor.
[0279] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A terminal's processor reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the terminal to perform the interactive processing methods provided in the various alternative embodiments described above.
[0280] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0281] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0282] The above description discloses only preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.
Claims
1. An interactive processing method, characterized in that, include: During a video session, a video session interface is displayed, which includes an area for displaying the visual representation of the video session object. The target virtual image of the target session object related to the video session is displayed in the image display area; the target virtual image is a 2D virtual image, and each object element contained in the 2D virtual image is made into a grid, with one grid corresponding to one object element; The grid corresponds to grid data, which refers to the state values of each grid vertex contained in the grid. The state values include: the position information of the grid vertex, or the positional relationship between the grid vertex and other connected grid vertices. Changes in the grid data drive the grid, and the grid drive is used to control the object elements corresponding to the grid in the 2D virtual image to perform target interactive operations. Obtain the action information of the target session object, and perform mesh deformation processing on the mesh data of the target mesh in the mesh set according to the action information; Based on the mesh data after mesh deformation processing, a target virtual image that has performed the target interactive action is rendered and displayed; wherein, in the rendered and displayed target virtual image that has performed the target interactive action, the position and / or shape of the object elements corresponding to the target mesh changes; the rendering method of the 2D virtual image includes: rendering layer by layer according to the display level of each object element included in the 2D virtual image. Among them, the multiple video session objects participating in the video session transmit motion information about the video session objects. The motion information is obtained by detecting images of the video session objects captured by the camera. The motion information includes one or more of the following: facial information, emotional information, body information, and positional transfer information of the video session objects.
2. The method as described in claim 1, characterized in that, The method further includes: Display an image selection window, which includes image selection elements; In response to a trigger operation on the image selection element, a reference virtual image is displayed in the reference display area, and candidate virtual images are displayed in the image selection window; In response to the image selection operation of the candidate virtual image, the reference virtual image is updated and displayed in the reference display area as the target candidate virtual image selected by the image selection operation; In response to the virtual avatar confirmation operation, the target candidate virtual avatar is determined as the target virtual avatar.
3. The method as described in claim 2, characterized in that, The method further includes: The background selection element is displayed in the image selection window; In response to a user action on the background selection element, candidate background images are displayed in the image selection window; In response to a background selection operation on the candidate background image, the target candidate background image selected by the background selection operation is displayed in the reference display area; In response to the background image confirmation operation, the target candidate background image is set as the background image of the image display area.
4. The method as described in claim 2, characterized in that, The method further includes: The voice selection element is displayed in the image selection window; In response to the selection operation of the voice selection element, candidate voice audio processing rules are displayed in the image selection window; In response to the confirmation operation of the candidate speech audio processing rule, the candidate speech audio processing rule is determined as the target speech audio processing rule, which is used to simulate the sound signal of the target session object received during the video session.
5. The method according to any one of claims 2-4, characterized in that, The image selection window includes an exit option, and the method further includes: In response to the selection of the exit option, the video session interface is displayed; An environmental image is displayed within the image display area included in the video session interface. The environmental image is obtained by capturing the environment. The environmental image is sent to the peer device so that the peer device can display the environmental image. The peer device refers to the device used by other users participating in the video session.
6. The method as described in claim 1, characterized in that, The target interactive action performed by the target virtual avatar is matched with the collected action information of the target session object; the matching includes any one or more of the following steps: If the action information of the target session object is facial information, then control the target virtual image to perform facial interaction actions; If the action information of the target session object is emotional information, then the face of the target virtual avatar is replaced by the target facial resource associated with the emotional information. If the action information of the target session object is limb information, then control the target limb of the target virtual image to perform limb actions; If the action information of the target session object is location transfer information, then the target virtual image is controlled to perform a location transfer action within the image display area.
7. The method as described in claim 1, characterized in that, The set of meshes added to the target virtual image includes multiple meshes and mesh data for each mesh; any mesh is composed of at least three mesh vertices, and the mesh data for any mesh refers to the state values of each mesh vertex contained in any mesh.
8. The method as described in claim 7, characterized in that, The action information is the facial information, which includes N feature points of the target session object's face, where N is an integer greater than 1; the step of performing mesh deformation processing on the mesh data of the target mesh in the mesh set based on the action information includes: Based on the N feature point information, the updated expression type of the target session object is determined; Obtain the expression base coefficients corresponding to the updated expression type; Based on the updated expression base coefficients, perform mesh deformation processing on the mesh containing the object elements corresponding to the updated expression type.
9. The method as described in claim 8, characterized in that, The step of performing mesh deformation processing on the mesh containing the object element corresponding to the updated expression type based on the expression base coefficients includes: Obtain the base coefficients of the emoji type before the update; The difference between the expression base coefficients corresponding to the updated expression type and the expression base coefficients corresponding to the expression type before the update is calculated to obtain the difference result. Based on the difference result, the expression base coefficient of the updated expression state, and the expression base coefficient of the expression state before the update, the mesh of the object element corresponding to the updated expression type is subjected to mesh deformation processing.
10. The method as described in claim 7, characterized in that, The action information is the limb information, and the step of performing mesh deformation processing on the mesh data of the target mesh in the mesh set based on the action information includes: Based on the limb information of the target session object, determine the position information of S limb points of the target session object, where S is an integer greater than zero; Based on the position information of the S limb points, calculate the angle value of the corresponding torso element; The mesh corresponding to the torso element in the target virtual image is subjected to mesh deformation processing based on the angle value.
11. The method as described in claim 1, characterized in that, The action information is the emotion information, and the method further includes: The emotional state of the target conversation object is identified based on the emotional information to obtain the current emotional state of the target conversation object; Based on the current emotional state, a target facial resource matching the current emotional state is determined; The face of the target virtual image is updated using the target facial resources within the image display area to obtain the updated target virtual image.
12. The method as described in claim 1, characterized in that, The action information is the location transfer information, and the method further includes: If the position transfer information is the movement information of the target image point, then the target virtual image is controlled to move and be displayed in the image display area according to the movement information of the target image point; If the position transfer information is the movement information of the target image region, the size of the display area of the target virtual image in the image display area is adjusted according to the movement information of the target image region.
13. The method as described in claim 1, characterized in that, The method further includes: When the target terminal starts a video session, acquire the target image obtained from the environment; The target session object in the target image is subjected to feature recognition processing to obtain the recognition result; Assign a virtual avatar that matches the recognition result to the target session object, and determine the virtual avatar that matches the recognition result as the target virtual avatar.
14. An interactive processing device, characterized in that, include: The display unit is used to display a video session interface during a video session, wherein the video session interface includes an image display area for displaying video session objects. The processing unit is configured to display a target virtual image of a target session object related to a video session within the image display area; the target virtual image is a 2D virtual image, and each object element contained in the 2D virtual image is made into a grid, with one grid corresponding to one object element; The grid corresponds to grid data, which refers to the state values of each grid vertex contained in the grid. The state values include: the position information of the grid vertex, or the positional relationship between the grid vertex and other connected grid vertices. Changes in the grid data drive the grid, and the grid drive is used to control the object elements corresponding to the grid in the 2D virtual image to perform target interactive operations. The processing unit is further configured to acquire the action information of the target session object, perform mesh deformation processing on the mesh data of the target mesh in the mesh set according to the action information, render and display the target virtual image that has performed the target interactive action based on the mesh data after mesh deformation processing, wherein the position and / or shape of the object element corresponding to the target mesh changes in the rendered and displayed target virtual image that has performed the target interactive action; the rendering method of the 2D virtual image includes: rendering layer by layer according to the display level of each object element included in the 2D virtual image; wherein the multiple video session objects participating in the video session transmit action information about the video session object, the action information being obtained by detecting the image of the video session object captured by the camera; the action information includes one or more of the following: facial information, emotional information, limb information, and position transfer information of the video session object.
15. A smart terminal, characterized in that, include: Storage devices and processors; The storage device stores a computer program; The processor executes a computer program stored in the storage device to implement the interactive processing method as described in any one of claims 1-13.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer application program, which, when executed, performs the interactive processing method as described in any one of claims 1-13.
17. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium, and the processor executes the computer instructions to implement the interactive processing method as described in any one of claims 1-13.
Citation Information
Patent Citations
KR20200109634A