Digital meeting interaction method, device, system and related device
By using a cloud server-edge node separation rendering method, the problems of high hardware resource consumption, high network bandwidth consumption, and poor privacy in digital virtual conference systems are solved, achieving high-efficiency rendering performance and real-time interaction, and enhancing user data security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-12
- Publication Date
- 2026-04-07
AI Technical Summary
Existing digital virtual conferencing systems have high requirements for user terminal hardware, consume a lot of computing and storage resources, consume a lot of network bandwidth, have poor privacy, and have poor real-time user interaction.
The cloud server-edge node separation rendering method is adopted. By analyzing user semantics and intent through edge nodes, a private semantic library is built, and private semantic scene construction and rendering are performed on edge nodes, reducing the consumption of hardware resources. At the same time, the cloud server performs cloud rendering and compositing of public semantic scenes, realizing cloud-edge collaboration.
It improves rendering performance, reduces hardware resource consumption, enhances real-time interaction, and eliminates the need to upload user audio and video data to the cloud, thus improving user data security.
Smart Images

Figure CN116527839B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of communication technology, and in particular to a digital conference interaction method, an edge node device, a cloud server device, a digital conference interaction system, a computer-readable storage medium, and an electronic device. Background Technology
[0002] Digital humans are created through 3D graphical character modeling and combined with artificial intelligence technology to visualize and virtually simulate the human body. They are generally divided into virtual avatar digital humans and intelligent service digital humans. Among them, virtual avatar digital humans serve as an idealized version of oneself in the virtual world, representing the user's personal representation in the virtual world, and are often used in scenarios such as virtual reality conferences and metaverse conferences.
[0003] Existing digital virtual conferencing systems have the following drawbacks:
[0004] 1. It has high requirements for user terminal hardware. User identification analysis and 3D digital human rendering require a lot of computing resources and storage resources to save data, which will generally reduce the rendering quality and user experience.
[0005] 2. It causes high network bandwidth consumption;
[0006] 3. Users' private data, such as audio and video, needs to be uploaded to the cloud for identification, resulting in poor privacy. 4. User interaction lacks real-time responsiveness, impacting user experience.
[0007] Therefore, overcoming the shortcomings of existing digital conferencing interaction technologies, such as limited hardware performance, high network bandwidth consumption, poor privacy, and poor interactivity, is a technical problem that urgently needs to be solved by those skilled in the art.
[0008] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0009] The purpose of this disclosure is to provide a digital conference interaction method, edge node device, cloud server device, digital conference interaction system, computer-readable storage medium, and electronic device, so as to at least solve the technical problems of limited hardware performance, high network bandwidth consumption, poor privacy, and poor interactivity in related technologies.
[0010] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.
[0011] The technical solution disclosed herein is as follows:
[0012] According to one aspect of this disclosure, a digital conferencing interaction method is provided, comprising: a user terminal accessing a conferencing system via a web page and sending user data to an edge node; the edge node analyzing user semantics and intent based on the user data and requesting a public semantic rendering stream from a cloud server; the cloud server requesting 3D model data corresponding to the user from cloud storage based on information stored in a public semantic library, and constructing and rendering a public semantic scene; the cloud server returning the cloud rendering stream to the edge node; the edge node constructing a private semantic library based on the user data and requesting data from cloud storage, and constructing and rendering a private semantic scene; and the edge node rendering and compositing the cloud rendering stream and the edge rendering stream; and the edge node returning the composite stream to the user terminal, enabling the user terminal to receive, in real time on a web page, a view of the conferencing scene of interest and a digital human user, the view including the digital human user's real-time expressions, actions, and position status.
[0013] In some embodiments of this disclosure, prior to the step of the edge node returning the synthesized stream to the user terminal, the method further includes: if the user's semantics involve interaction with other users, the edge node needs to request the edge rendering stream of the interacting user from other edge nodes.
[0014] In some embodiments of this disclosure, a public semantic library stores 3D scene information, including basic information of the 3D model, position data, rotation data, and scaling data. The basic information of the 3D model is used to represent a unique 3D scene; the position data, rotation data, and scaling data are used to construct the 3D scene.
[0015] In some embodiments of this disclosure, a private semantic library stores 3D scene information, including user basic information, expression-driven, head-driven, action-driven, semantic tags, and basic information of the 3D model. The user basic information is used to represent a unique and specific user; expression-driven, head-driven, and action-driven are used to drive the digital human; and semantic tags are used to represent the results of the edge nodes analyzing the user's semantics and intentions, pointing to specific user basic information.
[0016] In some embodiments of this disclosure, the method further includes: obtaining expression-driven parameters from the edge nodes based on the user's face blending deformation parameters, wherein the user's face blending deformation parameters are obtained based on expression recognition of the user; calculating head-driven parameters from the edge nodes based on the user's face blending deformation parameters; obtaining user action tags from the edge nodes based on a combination of action recognition, sentiment analysis, and semantic analysis of the user; and matching pre-made animations from an action library with the action tags as action-driven parameters.
[0017] In some embodiments of this disclosure, the step of rendering and compositing the cloud rendering stream and the edge rendering stream by the edge node includes: the edge node occupies the pixels in each frame of the edge rendering stream in the cloud rendering keyframe by using placeholders, and replaces the pixel values in the cloud rendering stream with the pixel values in the edge rendering stream.
[0018] In some embodiments of this disclosure, the step of the edge node analyzing user semantics and intent based on user data and requesting a public semantic rendering stream from the cloud server further includes: the cloud server training an artificial intelligence model for analyzing user semantics and intent; the cloud server distributing the trained artificial intelligence model to the edge node and cloud storage; the edge node generating user semantic tags based on the artificial intelligence model to identify user data; and the edge node sending the user's semantic tags to the cloud server to request a public semantic rendering stream.
[0019] According to another aspect of this disclosure, an edge node device is provided, comprising: a user data analysis module, an edge rendering stream generation module, a rendering compositing module, and a private semantic library; the user data analysis module is used to acquire user data from a user terminal, analyze user semantics and intent based on the user data, and request a public semantic rendering stream from a cloud server, wherein the user data is sent by the user terminal through a web page; the edge rendering stream generation module is used to construct a private semantic library based on the user data, request data from cloud storage, and perform private semantic scene construction and edge rendering; the rendering compositing module is used to perform rendering compositing of the cloud rendering stream and the edge rendering stream; and return the compositing stream to the user terminal, enabling the user terminal to receive, in real time on a web page, the scene of interest and the user's digital human image, including the digital human's real-time expressions, actions, and position status; wherein, the cloud rendering stream is generated by the cloud server based on information stored in the public semantic library, requesting 3D model data corresponding to the user from cloud storage, constructing a public semantic scene, generating cloud rendering, and returning it to the edge node device.
[0020] According to another aspect of this disclosure, a cloud server device includes: a cloud rendering stream generation module and a public semantic library; the cloud rendering stream generation module is used to respond to edge nodes' requests for a public semantic rendering stream of user semantics and intent, request corresponding 3D model data from cloud storage based on information stored in the public semantic library, construct and cloud render a public semantic scene; and return the cloud rendering stream to the edge nodes, so that the edge nodes can construct a private semantic library based on user data, request data from cloud storage, construct a private semantic scene and perform edge rendering, and render and synthesize the cloud rendering stream and the edge rendering stream, and return the synthesized stream to the user terminal, so that the user terminal can receive the meeting scene of interest and the user's digital human image in real time on the web page, the image including the digital human's real-time expressions, actions, and position status.
[0021] According to another aspect of this disclosure, a digital conference interaction system includes: an edge node, a cloud server, and cloud storage. The edge node includes: a user data analysis module, an edge rendering stream generation module, a rendering compositing module, and a private semantic library. The cloud server includes: a cloud rendering stream generation module and a public semantic library. The user data analysis module is used to analyze user semantics and intentions based on user data and request a public semantic rendering stream from the cloud server. This user data is sent by the user terminal accessing the conference system via a web page. The cloud rendering stream generation module is used to request 3D model data corresponding to the user from the cloud storage based on information stored in the public semantic library, construct a public semantic scene, and perform cloud rendering; and return the cloud rendering stream to the edge node. The edge rendering stream generation module is used to construct a private semantic library based on user data and request data from the cloud storage to construct a private semantic scene and perform edge rendering. The rendering compositing module is used to render and composite the cloud rendering stream and the edge rendering stream; and return the composite stream to the user terminal, enabling the user terminal to receive, in real-time, the conference scene of interest and the user's digital human image on a web page. This image includes the digital human's real-time expressions, actions, and position status.
[0022] According to another aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the above-described digital conferencing interaction method by executing the executable instructions.
[0023] According to another aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described digital conferencing interaction method.
[0024] The digital human conference interaction method proposed in this embodiment of the present disclosure, based on cloud server-edge node separate rendering, splits computing resources to achieve separate rendering of 3D scenes and digital humans, reducing the limitation of single hardware; and uses cloud storage to store 3D models, animations and other data, further reducing the consumption of hardware resources, thereby improving the overall rendering performance.
[0025] Furthermore, the cloud server loads and renders public 3D scenes based on a public semantic library, and proposes a rendering compositing algorithm. Based on keyframes and placeholders, it synthesizes rendering streams at edge nodes, realizing cloud-edge collaboration, reducing bandwidth resource requirements, and improving the real-time performance of interactions.
[0026] Furthermore, the edge server builds a private semantic library based on multimodal recognition, and realizes a user semantically driven digital human based on the edge-to-edge collaboration strategy. This eliminates the need to upload users' private data such as audio and video to the cloud for recognition and storage, thereby improving the security of user data.
[0027] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0028] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0029] Figure 1 A flowchart illustrating a digital conferencing interaction method according to an embodiment of this disclosure is shown.
[0030] Figure 2 This diagram illustrates a step in a digital conferencing interaction method according to an embodiment of the present disclosure, in which an edge node analyzes user semantics and intent based on user data and requests a rendering stream of public semantics from a cloud server.
[0031] Figure 3 This diagram illustrates a flowchart of a digital human-driven strategy in a digital conference interaction method according to an embodiment of the present disclosure.
[0032] Figure 4 This diagram illustrates the overall architecture of a digital conference interaction system according to an embodiment of the present disclosure.
[0033] Figure 5 A schematic diagram of an edge node device according to an embodiment of this disclosure is shown.
[0034] Figure 6 A schematic diagram of a cloud server device according to an embodiment of this disclosure is shown.
[0035] Figure 7 A schematic block diagram of an electronic device based on a digital conferencing interaction method is shown in an embodiment of this disclosure. Detailed Implementation
[0036] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0037] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0038] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this disclosure, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0039] In view of the technical problems existing in the above-mentioned related technologies, the present disclosure provides a digital conference interaction method to solve at least one or all of the above-mentioned technical problems.
[0040] It should be noted that the nouns or terms used in the embodiments of this application can be referenced from each other and will not be repeated here.
[0041] The following will describe in more detail each step of the routing node quality monitoring method in this exemplary embodiment, with reference to the accompanying drawings and embodiments.
[0042] Figure 1 A flowchart illustrating a digital conferencing interaction method according to an embodiment of this disclosure is shown. Figure 1 As shown, method 100 may include the following steps:
[0043] In step S110, the user terminal accesses the conference system via a web page and sends user data to the edge node.
[0044] User data includes user interactions, audio and video data, etc.
[0045] Each user terminal corresponds to a unique edge node.
[0046] In step S120, the edge node analyzes the user's semantics and intent based on the user data and requests the public semantic rendering stream from the cloud server.
[0047] In step S130, the cloud server requests the corresponding 3D model data from the cloud storage based on the information stored in the public semantic library, and performs public semantic scene construction and cloud rendering.
[0048] Among them, the public semantic library is used to render public scenes.
[0049] Among them, the 3D model data in cloud storage is a public 3D scene.
[0050] In step S140, the cloud server returns the cloud rendering stream to the edge node.
[0051] In step S150, the edge nodes construct a private semantic library based on user data and request data from cloud storage to construct a private semantic scene and perform edge rendering.
[0052] Each edge node has a private semantic library used to render the real-time expressions, actions, and location of the user corresponding to that edge node.
[0053] Among them, the data requested by the edge node from the cloud storage based on private semantics can be a pre-made 3D animation.
[0054] In step S160, the edge nodes perform rendering compositing on the cloud rendering stream and the edge rendering stream.
[0055] In step S170, the edge node returns the synthesized stream to the user terminal, enabling the user terminal to receive the meeting scene of interest and the image of the user's digital human in real time on the web page. The image includes the digital human's real-time expressions, actions, and location status.
[0056] The digital human conference interaction method proposed in this embodiment of the present disclosure, based on cloud server-edge node separate rendering, splits computing resources to achieve separate rendering of 3D scenes and digital humans, reducing the limitation of single hardware; and uses cloud storage to store 3D models, animations and other data, further reducing the consumption of hardware resources, thereby improving the overall rendering performance.
[0057] Furthermore, the cloud server loads and renders public 3D scenes based on a public semantic library, and proposes a rendering compositing algorithm. Based on keyframes and placeholders, it synthesizes rendering streams at edge nodes, realizing cloud-edge collaboration, reducing bandwidth resource requirements, and improving the real-time performance of interactions.
[0058] Furthermore, the edge server builds a private semantic library based on multimodal recognition, and realizes a user semantically driven digital human based on the edge-to-edge collaboration strategy. This eliminates the need to upload users' private data such as audio and video to the cloud for recognition and storage, thereby improving the security of user data.
[0059] In some embodiments of this disclosure, prior to step S170, the method may further include: if the user's semantics involve interaction with other users, then the edge node needs to request the edge rendering stream of the interacting user from other edge nodes.
[0060] In other words, the digital interactive conference screen in this disclosed method may include one or more users. If there are multiple users, there are multiple edge nodes corresponding to each user. In this case, when the edge node synthesizes the rendering stream, it needs to obtain not only the cloud rendering stream from the cloud server but also the edge rendering streams of other users from other edge nodes. Then, the edge node synthesizes its own edge rendering stream, other edge rendering streams, and cloud rendering streams together into one screen.
[0061] Specifically, the steps performed by other edge nodes may include: acquiring user data based on the user terminal corresponding to the edge node; analyzing user semantics and intent based on the user data; building a private semantic library based on the user data and requesting data from cloud storage to perform private semantic scene construction and edge rendering; and then sending the edge rendering stream to the edge node that performs rendering and compositing.
[0062] The method of this disclosure can enable users to receive images of meeting scenes of interest and digital human figures of other users in real time on a web page, thereby achieving real-time intelligent interaction in digital meetings.
[0063] Furthermore, the edge node sends a request to other edge nodes where the interactive user is located. The other edge nodes perform edge rendering based on the private semantic library on that node and return the rendering stream to the edge node, thus realizing edge-to-edge collaboration.
[0064] In some embodiments of this disclosure, the steps in step S120 may further include, for example... Figure 2 The method steps shown are as follows: Figure 2 As shown:
[0065] In step S210, an artificial intelligence model for analyzing user semantics and intent is trained by the cloud server.
[0066] In step S220, the cloud server distributes the trained artificial intelligence model to the edge nodes and cloud storage.
[0067] In step S230, the edge nodes generate semantic tags for users based on user data identified by an artificial intelligence model.
[0068] The semantic tag (SemanticID) corresponds one-to-one with the user ID.
[0069] In step S240, the edge node sends the user's semantic tags to the cloud server to request the rendering stream of public semantics.
[0070] Among them, the public semantic rendering stream is obtained by the edge nodes requesting the corresponding 3D scene model from the cloud storage based on the semantic tags, and then the edge nodes load, drive and render the 3D model.
[0071] Each semantic tag points to a specific 3D scene model that corresponds to it.
[0072] By training and analyzing models of user data semantics and intent on cloud servers, and then distributing them to edge nodes and cloud storage, user data analysis is completed to form a user intelligent perception system supported by edge-cloud collaboration. This embodiment of the present disclosure alleviates the inaccuracy of user data analysis caused by the weak data processing capabilities and poor data correlation of current edge node devices, while meeting real-time requirements.
[0073] In some embodiments of this disclosure, the three-dimensional scene information stored in the public semantic library includes basic information of the three-dimensional model, position data, rotation data, and scaling data. The basic information of the three-dimensional model is used to represent a unique and specific three-dimensional scene; the position data, rotation data, and scaling data are used to construct the three-dimensional scene.
[0074] Among them, the basic information of a 3D model is the information that represents the unique identity of the 3D model, such as an identifier (ID).
[0075] Specifically, the public semantic library can be represented as: [ModelID, Position, Rotation, Scale], where ModelID represents the ID of the 3D model in the scene, and Position, Rotation, and Scale represent the position, rotation, and scaling data of the model, respectively, used to construct the 3D scene.
[0076] The cloud rendering stream rendered based on 3D scene information in a public semantic library, using the method of this disclosure, only reflects key information of the public scene, has low real-time requirements, and thus saves bandwidth usage.
[0077] In some embodiments of this disclosure, a private semantic library stores 3D scene information, including user basic information, expression-driven, head-driven, action-driven, semantic tags, and basic information of the 3D model. The user basic information is used to represent a unique and specific user; expression-driven, head-driven, and action-driven are used to drive the digital human; and semantic tags are used to represent the results of edge nodes analyzing user semantics and intentions, pointing to specific user basic information.
[0078] Specifically, the private semantic library can be represented as: [UserID, Url, ModelID, Face, Head, Action, Transform, SemanticID], where UserID uniquely identifies the user, Url represents the user's stream address, ModelID represents the digital human model ID corresponding to the user, Face, Head, and Action represent the user's facial expression data, head transformation, and action labels, respectively, used to drive the digital human, Transform represents the current 3D transformation of the user's digital human, and SemanticID represents the user's semantic label obtained from the above artificial intelligence analysis, pointing to a certain user ID.
[0079] The method of this disclosure allows an edge server to build a private semantic library based on multimodal recognition, and to realize a user-semantic driven digital human based on an edge-to-edge collaboration strategy. Furthermore, it eliminates the need to upload users' private data such as audio and video to the cloud for recognition and storage, thereby improving the security of user data.
[0080] In some embodiments of this disclosure, the methods for driving the digital human's facial expression, head movement, and motion can be, for example... Figure 3 One example of a digital human-driven method is shown, such as Figure 3 As shown, method 300 may include the following steps:
[0081] In step S310, the edge nodes obtain the expression driver based on the user's face blending deformation parameters, which are obtained by recognizing the user's expression.
[0082] In step S320, the head drive is calculated by the edge nodes based on the user's face blending deformation parameters.
[0083] In step S330, the edge nodes obtain the user's action tags based on a combination of action recognition, sentiment analysis, and semantic analysis.
[0084] In step S340, the edge nodes match the pre-made animations in the action library according to the action tags as the action drivers.
[0085] Specifically, the expression-driven Face can be represented as [UserID, BlendShapes], where the face blending deformation parameters (BlendShapes) are extracted based on the expression recognition algorithm, and the head motion parameters are calculated through the face parameters to achieve head-driven Head: [UserID, BlendShapes]. The action-driven Action can be represented as [UserID, Hand, Emotion, Semantic], which is the user action tag obtained based on action recognition, sentiment analysis, and semantic analysis. Then, the action tag is matched with pre-made animations in the action library to drive the digital human's actions.
[0086] The digital human driving method disclosed in this embodiment can reflect the user's state in the real physical world in real time, thereby improving the user's experience and feelings.
[0087] In some embodiments of this disclosure, the rendering stream synthesis method in step S160 may further include: using placeholders to place pixels in each frame of the edge rendering stream in the cloud rendering keyframe by the edge node, and replacing the pixel values in the cloud rendering stream with the pixel values in the edge rendering stream.
[0088] The cloud-edge separation rendering architecture of this disclosure can separate computing resources and storage resources. The cloud is responsible for rendering relatively static and complex public scenes, while the edge is responsible for rendering digital humans that change in real time, thereby improving the overall rendering performance.
[0089] Furthermore, by synthesizing rendering streams at edge nodes based on keyframes and placeholders, cloud-edge collaboration is achieved, reducing bandwidth resource requirements and improving the real-time performance of interactions.
[0090] In some embodiments of this disclosure, a digital conferencing interaction system is also proposed, such as... Figure 4 As shown, the digital conference interaction 400 may include: edge node 410, cloud server 420, and cloud storage 430. The edge node 410 includes: user data analysis module 418, edge rendering stream generation module 414, rendering compositing module 412, and private semantic library 416. The cloud server 420 includes: cloud rendering stream generation module 422 and public semantic library 424.
[0091] The user data analysis module 418 can be used to analyze user semantics and intent based on user data and request a public semantic rendering stream from the cloud server 420. This user data is sent by the user terminal through a web-based access to the conference system.
[0092] The cloud rendering stream generation module 422 is used to request the corresponding 3D model data from the cloud storage 430 based on the information stored in the public semantic library 424, to construct the public semantic scene and perform cloud rendering; and to return the cloud rendering stream to the edge node 410.
[0093] The edge rendering stream generation module 414 is used to build a private semantic library 416 based on user data, and request data from cloud storage 430 to perform private semantic scene construction and edge rendering; and
[0094] The rendering and compositing module 412 is used to render and composite the cloud rendering stream and the edge rendering stream; and return the composite stream to the user terminal, so that the user terminal can receive the meeting scene of interest and the image of the user's digital human in real time on the web page. The image includes the digital human's real-time expressions, actions and position status.
[0095] The digital human conference interaction system based on cloud-edge separation rendering proposed in this disclosure splits computing resources to achieve separate rendering of 3D scenes and digital humans, reducing the limitation on single hardware; and uses cloud storage to store 3D models, animations and other data, further reducing the consumption of hardware resources, thereby improving the overall rendering performance.
[0096] Furthermore, the cloud server loads and renders public 3D scenes based on a public semantic library, and proposes a rendering compositing algorithm. Based on keyframes and placeholders, it synthesizes rendering streams at edge nodes, realizing cloud-edge collaboration, reducing bandwidth resource requirements, and improving the real-time performance of interactions.
[0097] Furthermore, the edge server builds a private semantic library based on multimodal recognition, and realizes a user semantically driven digital human based on the edge-to-edge collaboration strategy. This eliminates the need to upload users' private data such as audio and video to the cloud for recognition and storage, thereby improving the security of user data.
[0098] In some embodiments of this disclosure, if the user's semantics involve interaction with other users, the edge node 410 is also used to request the edge rendering stream of the interacting user from other edge nodes.
[0099] Specifically, the system may also include one or more other edge nodes N 440, which may include: a user data analysis module 448, an edge rendering stream generation module 444, a rendering compositing module 442, and a private semantic library 446. Specifically, the user data analysis module 448 obtains user data based on the user terminal corresponding to the edge node N, analyzes user semantics and intent based on the user data, constructs a private semantic library 416 based on the user data, requests data from cloud storage 430, performs private semantic scene construction and edge rendering, and then sends the edge rendering stream to the edge node 410 for compositing.
[0100] In some embodiments of this disclosure, the public semantic library 424 is used to store three-dimensional scene information, including basic information of the three-dimensional model, position data, rotation data, and scaling data. The basic information of the three-dimensional model is used to represent a unique and specific three-dimensional scene; the position data, rotation data, and scaling data are used to construct the three-dimensional scene.
[0101] In some embodiments of this disclosure, a private semantic library 446 is used to store three-dimensional scene information, including user basic information, expression-driven, head-driven, action-driven, semantic tags, and basic information of the three-dimensional model. The user basic information is used to represent a unique and specific user; expression-driven, head-driven, and action-driven are used to drive the digital human; and semantic tags are used to represent the results of edge nodes analyzing user semantics and intentions, pointing to specific user basic information.
[0102] In some embodiments of this disclosure, the edge node 410 may further include: a digital human driving module, used to obtain expression driving based on the user's face blending deformation parameters obtained by performing expression recognition on the user; to calculate head driving based on the user's face blending deformation parameters; to obtain the user's action tags based on the user's action recognition, sentiment analysis, and semantic analysis; and to match pre-made animations in the action library as action driving based on the action tags.
[0103] In some embodiments of this disclosure, the edge node 410 can also be used to place pixels in each frame of the edge rendering stream in the cloud rendering keyframe using placeholders, and replace the pixel values in the cloud rendering stream with the pixel values in the edge rendering stream.
[0104] In some embodiments of this disclosure, the cloud server 420 can also be used to analyze the artificial intelligence model of user semantics and intent; distribute the trained artificial intelligence model to edge nodes and cloud storage; the edge node 410 can also be used to generate user semantic tags based on user data identified by the artificial intelligence model; and the edge node 410 can send the user's semantic tags to the cloud server 420 to request a public semantic rendering stream.
[0105] In some embodiments of this disclosure, an edge node device is also proposed, such as Figure 5As shown, the edge node device 500 may include: a user data analysis module 510, an edge rendering stream generation module 520, a rendering compositing module 530, and a private semantic library 540. The user data analysis module 510 is used to obtain user data from the user terminal, analyze user semantics and intent based on the user data, and request a public semantic rendering stream from the cloud server. This user data is sent by the user terminal through a web page. The edge rendering stream generation module 520 is used to build a private semantic library based on the user data, request data from cloud storage, and perform private semantic scene construction and edge rendering. The rendering compositing module 530 is used to perform rendering compositing of the cloud rendering stream and the edge rendering stream, and return the composite stream to the user terminal, so that the user terminal can receive the meeting scene of interest and the image of the user's digital human in real time on the web page. This image includes the digital human's real-time expressions, actions, and position status. The cloud rendering stream is generated by the cloud server based on information stored in the public semantic library, requesting 3D model data corresponding to the user from the cloud storage, constructing a public semantic scene, generating cloud rendering, and returning it to the edge node device.
[0106] In some embodiments of this disclosure, a cloud server device is also proposed, such as... Figure 6 As shown, the cloud server device 600 may include: a cloud rendering stream generation module 610 and a public semantic library 620. The cloud rendering stream generation module 610 is used to respond to the edge node's request for a public semantic rendering stream of user semantics and intent. Based on the information stored in the public semantic library 620, it requests the corresponding 3D model data from the cloud storage to construct and render the public semantic scene. The cloud rendering stream is then returned to the edge node, which constructs a private semantic library based on the user data, requests data from the cloud storage, constructs a private semantic scene, and performs edge rendering. The cloud rendering stream and the edge rendering stream are then rendered and synthesized, and the synthesized stream is returned to the user terminal, so that the user terminal can receive the meeting scene of interest and the user's digital human in real time on the web page. The screen includes the digital human's real-time expressions, actions, and position status.
[0107] Regarding the digital conference interaction system 400, edge node device 500, and cloud server device 600 in the above embodiments, the specific methods by which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0108] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."
[0109] The following reference Figure 7 To describe an electronic device 700 according to such an embodiment of the present disclosure. Figure 7 The electronic device 700 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.
[0110] like Figure 7 As shown, the electronic device 700 is manifested in the form of a general-purpose computing device. The components of the electronic device 700 may include, but are not limited to: at least one processing unit 710, at least one storage unit 720, and a bus 730 connecting different system components (including storage unit 720 and processing unit 710).
[0111] The storage unit stores program code that can be executed by the processing unit 710, causing the processing unit 710 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processing unit 710 can perform actions such as... Figure 1 In step S110, the user terminal accesses the conference system via a web page and sends user data to the edge node; in step S120, the edge node analyzes the user's semantics and intent based on the user data and requests a public semantic rendering stream from the cloud server; in step S130, the cloud server requests 3D model data corresponding to the user from the cloud storage based on the information stored in the public semantic library, and performs public semantic scene construction and cloud rendering; in step S140, the cloud server returns the cloud rendering stream to the edge node; in step S150, the edge node constructs a private semantic library based on the user data and requests data from the cloud storage, and performs private semantic scene construction and edge rendering; in step S160, the edge node performs rendering synthesis of the cloud rendering stream and the edge rendering stream; and in step S170, the edge node returns the synthesized stream to the user terminal, so that the user terminal can receive the conference scene of interest and the image of the user's digital human in real time on the web page, which includes the digital human's real-time expressions, actions, and position status.
[0112] Storage unit 720 may include readable media in the form of volatile storage units, such as random access memory (RAM) 721 and / or cache memory 722, and may further include read-only memory (ROM) 723.
[0113] The storage unit 720 may also include a program / utility 724 having a set (at least one) of program modules 725, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0114] Bus 730 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0115] Electronic device 700 can also communicate with one or more external devices (e.g., keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with electronic device 700, and / or any device that enables electronic device 700 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 750. Furthermore, electronic device 700 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 760. As shown, network adapter 760 communicates with other modules of electronic device 700 via bus 730. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0116] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible implementations, various aspects of this disclosure may also be implemented as a program product including program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of this disclosure described in the "Exemplary Methods" section above.
[0117] The program product for implementing the above-described method according to embodiments of the present disclosure may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used or used in conjunction with an instruction execution system, server, terminal, or device.
[0118] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, server, terminal, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0119] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, server, terminal, or device.
[0120] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0121] Program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0122] According to one aspect of this disclosure, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional implementations of the above embodiments.
[0123] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0124] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0125] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0126] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.
Claims
1. A digital conferencing interaction method, characterized in that, The method includes: User terminals access the conferencing system via a web interface and send user data to edge nodes; The edge nodes analyze user semantics and intent based on the user data and request a public semantic rendering stream from the cloud server; The cloud server requests 3D model data corresponding to the user from the cloud storage based on information stored in the public semantic library, and performs public semantic scene construction and cloud rendering. The cloud server returns the cloud rendering stream to the edge node; The edge nodes construct a private semantic library based on user data and request data from cloud storage to perform private semantic scene construction and edge rendering; and The edge nodes render and synthesize the cloud rendering stream and the edge rendering stream. The edge node returns the synthesized stream to the user terminal, enabling the user terminal to receive the meeting scene and the user's digital human image of interest in real time on the web page. The image includes the digital human's real-time expressions, movements, and location status.
2. The digital conference interaction method according to claim 1, characterized in that, Prior to the step of returning the synthesized stream to the user terminal by the edge node, the method further includes: If the semantics of the user involve interaction with other users, then the edge node needs to request the edge rendering stream of the interacting user from other edge nodes.
3. The digital conference interaction method according to claim 1, characterized in that, The public semantic library stores 3D scene information, including basic information of the 3D model, position data, rotation data, and scaling data. The basic information of the 3D model is used to represent a unique 3D scene; the position data, rotation data, and scaling data are used to construct the 3D scene.
4. The digital conference interaction method according to claim 1, characterized in that, The The private semantic library stores 3D scene information, including user basic information, expression-driven, head-driven, action-driven, semantic tags, and basic information of the 3D model. Among them, user basic information is used to represent a unique user; expression-driven, head-driven, and action-driven are used to drive the digital human; semantic tags are used to represent the results of the edge nodes analyzing the user's semantics and intentions, pointing to specific user basic information.
5. The digital conferencing interaction method according to claim 4, characterized in that, The method further includes: The expression driver is obtained from the edge nodes based on the user's face blending deformation parameters, which are obtained by performing expression recognition on the user. The head drive is calculated by the edge nodes based on the user's face blending deformation parameters; The edge nodes obtain the user's action tags based on action recognition, sentiment analysis, and semantic analysis. The edge nodes match pre-made animations from the action library based on the action tags, which serve as the action drivers.
6. The digital conferencing interaction method according to claim 1 or 2, characterized in that, The steps for rendering and compositing the cloud rendering stream and the edge rendering stream by the edge nodes include: The edge node places the pixels in each frame of the edge rendering stream in the cloud rendering keyframe using placeholders, and replaces the pixel values in the cloud rendering stream with the pixel values in the edge rendering stream.
7. The digital conference interaction method according to claim 1, characterized in that, The step of the edge node analyzing user semantics and intent based on the user data and requesting a public semantic rendering stream from the cloud server further includes: An artificial intelligence model is trained by the cloud server to analyze user semantics and intent; The cloud server distributes the trained artificial intelligence model to the edge nodes and the cloud storage. The edge nodes generate semantic tags for the user based on the user data identified by the artificial intelligence model; and The edge node sends the user's semantic tags to the cloud server to request a public semantic rendering stream.
8. An edge node device, characterized in that, The device includes: a user data analysis module, an edge rendering stream generation module, a rendering compositing module, and a private semantic library; The user data analysis module is used to obtain user data from the user terminal, analyze user semantics and intent based on the user data, and request a public semantic rendering stream from the cloud server. The user data is sent by the user terminal through a web page. The edge rendering stream generation module is used to build a private semantic library based on user data, request data from cloud storage, and perform private semantic scene construction and edge rendering. The rendering and compositing module is used to render and composite the cloud rendering stream and the edge rendering stream; and return the composite stream to the user terminal, so that the user terminal can receive the meeting scene of interest and the image of the user's digital human in real time on the web page. The image includes the digital human's real-time expressions, actions and position status. The cloud rendering stream is generated by the cloud server based on information stored in the public semantic library, requesting 3D model data corresponding to the user from the cloud storage, constructing the public semantic scene and generating cloud rendering, and returning it to the edge node device.
9. A cloud server device, characterized in that, The cloud server device includes: a cloud rendering stream generation module and a public semantic library; The cloud rendering stream generation module is used to respond to edge nodes' requests for rendering streams of public semantics based on user semantics and intents. Based on information stored in the public semantic library, it requests corresponding 3D model data from cloud storage to construct and render the public semantic scene. The cloud rendering stream is then returned to the edge nodes, which construct private semantic libraries based on user data, request data from cloud storage, construct private semantic scenes, and perform edge rendering. The module also renders and synthesizes the cloud rendering stream and the edge rendering stream, returning the synthesized stream to the user terminal. This allows the user terminal to receive real-time images of the meeting scene and the user's digital human on a webpage, including the digital human's real-time expressions, actions, and location status.
10. A digital conference interaction system, characterized in that, The system includes: edge nodes, cloud servers, and cloud storage. The edge nodes include: a user data analysis module, an edge rendering stream generation module, a rendering compositing module, and a private semantic library. The cloud servers include: a cloud rendering stream generation module and a public semantic library. The user data analysis module is used to analyze user semantics and intent based on the user data and request a public semantic rendering stream from the cloud server. The user data is sent by the user terminal through a web page accessing the conference system. The cloud rendering stream generation module is used to request 3D model data corresponding to the user from the cloud storage based on the information stored in the public semantic library, to construct and render the public semantic scene, and to return the cloud rendering stream to the edge node. The edge rendering stream generation module is used to build a private semantic library based on user data, request data from cloud storage, and perform private semantic scene construction and edge rendering; and The rendering and compositing module is used to render and composite the cloud rendering stream and the edge rendering stream; and return the composite stream to the user terminal, so that the user terminal can receive the meeting scene and the user's digital human image of interest in real time on the web page, the image including the digital human's real-time expression, movement and position status.
11. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the digital conference interaction method according to any one of claims 1 to 7 by executing the executable instructions.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the digital conference interaction method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Semantic fidelity-oriented virtual conference method and three-dimensional virtual conference system
CN114363557A
Multi-user cooperation method and system for mobile virtual reality based on edge calculation
CN115865916A