Video conference interaction method and device based on artificial intelligence, server and medium
By obtaining video streaming data and character information, generating conference scene pages, configuring virtual interactive props and identifying action display effects, the problem of insufficient interaction convenience in video conferences is solved, and a more intuitive and intelligent interactive experience is achieved.
Patent Information
- Application Number
- CN202510704436.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-25
AI Technical Summary
The interaction of participants in existing video conferences is poor, and they mainly rely on chat boxes to achieve it, lacking intuition and interactivity.
By obtaining the video stream data and role information of participants, a meeting scene page is generated, and virtual interactive props are configured based on the role information, and the human action recognition model is used to identify the action, and the corresponding special effects are displayed.
It improves the interactive convenience and intelligence of video conferencing, allowing participants to interact through intuitive virtual interactive props and special effects to enhance the meeting effect.
Smart Images

Figure CN120378570A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a video conferencing interaction method, device, server, and storage medium based on artificial intelligence. Background Art
[0002] Video conferencing is a form of holding meetings using cameras and networks, which has high flexibility and convenience and is widely used in various industries. For example, managers and salespersons in insurance companies discuss marketing strategies through video conferencing. Another example is that purchasers of drugs or medical devices in hospitals communicate with representatives of pharmaceutical companies about the purchase requirements of drugs or medical devices through video conferencing. Currently, during video conferencing, whether it is a birthday party or a company regular meeting, the interaction between participants is mainly achieved through a chat box, and the convenience is poor. Therefore, how to improve the convenience of video conferencing interaction is an urgent problem to be solved currently. Summary of the Invention
[0003] Embodiments of this application provide a video conferencing interaction method, device, server, and storage medium based on artificial intelligence, aiming at the convenience of video conferencing interaction.
[0004] In a first aspect, embodiments of this application provide a video conferencing interaction method based on artificial intelligence, including:
[0005] Obtain video stream data and role information of each participant in the video conference;
[0006] Extract the real portrait of each participant from the video stream data of each participant, and generate a conference scene page according to the real portrait and role information of each participant;
[0007] Configure corresponding virtual interaction props for each participant according to the role information of each participant;
[0008] Identify the human body movements of each participant through a preset human body movement recognition model to obtain the human body movements of each participant;
[0009] In response to the human body movements of at least one participant being the same as the human body movements corresponding to the virtual interaction props, display special effects corresponding to at least one of the virtual interaction props on the conference scene page.
[0010] In a second aspect, embodiments of this application further provide a video conferencing interaction device, where the video conferencing interaction device includes:
[0011] An obtaining module, configured to obtain video stream data and role information of each participant in the video conference;
[0012] The obtaining module is further configured to extract the real portrait of each participant from the video stream data of each participant;
[0013] The generating module is configured to generate a conference scene page according to the real portrait and role information of each participant;
[0014] The configuration module is configured to configure corresponding virtual interaction props for each participant according to the role information of each participant;
[0015] The action recognition module is configured to recognize the human body actions of each participant through a preset human body action recognition model to obtain the human body actions of each participant;
[0016] The interaction module is configured to, in response to the human body actions of at least one participant being the same as the human body actions corresponding to the virtual interaction props, display special effects corresponding to at least one of the virtual interaction props in the conference scene page.
[0017] In a third aspect, an embodiment of the present application further provides a server, which includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the video conference interaction method described in the first aspect are implemented.
[0018] In a fourth aspect, an embodiment of the present application further provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the video conference interaction method described in the first aspect are implemented.
[0019] Based on the role information of each participant, the embodiment of the present application configures corresponding virtual interaction props for each participant, and recognizes the human body actions of each participant through a human body action recognition model to obtain the human body actions of each participant. Then, when the human body actions of at least one participant are the same as the human body actions corresponding to the virtual interaction props, special effects corresponding to at least one of the virtual interaction props are displayed in the conference scene page, thereby effectively improving the convenience of video conference interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1It is a schematic diagram of the application environment of the video conferencing interaction method provided by an embodiment of the present application;
[0022] Figure 2 It is a schematic flowchart of a video conferencing interaction method provided by an embodiment of the present application;
[0023] Figure 3 It is a schematic block diagram of a video conferencing interaction device provided by an embodiment of the present application;
[0024] Figure 4 is Figure 3 a schematic block diagram of a sub-module of the video conferencing interaction device in
[0025] Figure 5 It is a schematic block diagram of the structure of a server provided by an embodiment of the present application.
[0026] The realization, functional features, and advantages of the purpose of the present application will be further described with reference to the accompanying drawings in conjunction with the embodiments. Specific Embodiments
[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0028] The flowchart shown in the accompanying drawings is only an example illustration, and does not necessarily include all the contents and operations / steps, nor does it necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may change according to the actual situation.
[0029] The video conferencing interaction method based on artificial intelligence provided by the embodiments of the present application can be applied in, for example, Figure 1In the application environment, the terminal device 100 communicates with the server 200 through a network. The participants use the terminal device 100 to access the server 200 to enter a video conference, and the terminal device 100 sends the video stream data of the participants to the server. The server 200 obtains the video stream data and role information of each participant participating in the video conference; extracts the real portrait of each participant from the video stream data of each participant, and generates a conference scene page according to the real portrait and role information of each participant (the server 200 can send the conference scene page to the terminal device 100 logged in by each participant for the terminal device 100 to display the conference scene page); configures corresponding virtual interaction props for each participant according to the role information of each participant; identifies the human actions of each participant through a preset human action recognition model to obtain the human actions of each participant; in response to the human actions of at least one participant being the same as the human actions corresponding to the virtual interaction props, display special effects corresponding to at least one virtual interaction prop on the conference scene page.
[0030] Among them, the terminal device can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can include a single server or a server cluster composed of multiple servers. The server can be an independent physical server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0031] The following will describe some embodiments of the present application in detail with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0032] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a video conference interaction method based on artificial intelligence provided by an embodiment of the present application.
[0033] As Figure 2 shown, the video conference interaction method includes steps S101 to S105.
[0034] Step S101, obtain the video stream data and role information of each participant participating in the video conference.
[0035] In this embodiment, the video stream data of each participant in the video conference is uploaded to the server by the terminal device with a camera after shooting the participants. The role information of the participants may include the birthday protagonist, the audience, the speaker, or the host, etc. For example, the manager of an insurance company reserves a video conference and sends the video conference number or the video conference link to the insurance salesperson. The insurance salesperson enters the video conference number or clicks on the video conference link through the terminal device to join the reserved video conference. Another example is that the purchaser of drugs or medical devices in a hospital reserves a video conference and sends the video conference number or the video conference link to the pharmaceutical company representative. The pharmaceutical company representative enters the video conference number or clicks on the video conference link through the terminal device to join the reserved video conference.
[0036] In some embodiments, when a participant enters the video conference, the identification information of the participant is obtained, and based on the identification information of the participant, role information is assigned to the participant. For example, if the identification information of the participant is found in the preset mapping relationship table, the role information corresponding to the identification information of the participant in the mapping relationship table is assigned to the participant. If the identification information of the participant is not found in the preset mapping relationship table, the role information of the audience is assigned to the participant. Among them, the mapping relationship between the identification information and the role information in the preset mapping relationship table can be set based on the actual situation, and the embodiments of the present application do not make specific limitations on this.
[0037] Step S102: Extract the real portrait of each participant from the video stream data of each participant, and generate a conference scene page according to the real portrait and role information of each participant.
[0038] In this embodiment, the preset portrait detection model can be used to perform portrait detection on the video stream data of each participant to obtain the portrait detection result, and according to this portrait detection result, the real portrait of each participant is extracted from the video stream data of each participant. Among them, the preset portrait detection model is obtained by pre-training the target detection model according to a plurality of first training samples. The first training samples include training images and the annotated portrait detection results.
[0039] In some embodiments, the preset portrait detection model is obtained by pre-training the target detection model according to a plurality of first training samples and then performing model pruning and quantization processing. Among them, the target detection model includes the YOLOv7 model or the Faster R-CNN model. In this embodiment, the model pruning and quantization processing are performed on the trained target detection model, which can reduce the amount of calculation, thereby realizing millisecond-level multi-person portrait detection and improving the efficiency of portrait detection.
[0040] In some embodiments, the size information of the human portrait detection frame adopted by the human portrait detection model when detecting the video stream data of the participants is related to the size information of the images included in the video stream data. Among them, if the size information of the images included in the video stream data is different, the size information of the human portrait detection frame adopted by the human portrait detection model when detecting the video stream data of the participants is different. The size information of the human portrait detection frame adopted in this embodiment is related to the size information of the images included in the video stream data, which can meet the requirements of the human portrait detection accuracy for images with different size information and improve the accuracy of human portrait detection.
[0041] In some embodiments, the size information of the human portrait detection frame includes the aspect ratio and / or the size. For example, the aspect ratio of the human portrait detection frame adopted by the human portrait detection model when detecting the video stream data of the participants is the same as the aspect ratio of the images included in the video stream data and / or the size of the human portrait detection frame adopted by the human portrait detection model when detecting the video stream data of the participants is positively correlated with the size of the images included in the video stream data.
[0042] In some embodiments, generating a meeting scene page according to the real human portraits and role information of each participant may include: obtaining a meeting scene map matching the type of the video conference; determining a target meeting layout through a preset meeting layout determination model based on the meeting scene map, the role information of each participant, and the total number of participants, where the meeting layout determination model is obtained by pre-training a neural network model according to a plurality of second training samples; and arranging the real human portraits of each participant in the meeting scene map according to the target meeting layout and the role information of each participant to obtain the meeting scene page. This embodiment realizes the automatic configuration of the meeting scene page, improves the intelligence of the video conference, and the configured meeting scene page is more in line with the real meeting scene, enabling the participants to be more focused during the video conference, thereby improving the meeting effect.
[0043] In some embodiments, the meeting layout determination model is obtained by pre-fine-tuning a BERT (Bidirectional Encoder Representations from Transformers) model according to a plurality of second training samples, and the second training samples include meeting scene map samples, role information samples, total number of participants samples, and labeled meeting layouts.
[0044] In some embodiments, obtaining a meeting scene graph matching the type of video conference may include: querying a mapping relationship table between types and meeting scene graphs stored in advance to obtain a meeting scene graph matching the type of video conference. Among them, different types match different meeting scene graphs, and the pre-stored mapping relationship table between types and meeting scene graphs is set based on actual situations, and the embodiments of the present application do not make specific limitations in this regard. For example, the types of video conferences include birthday parties, report meetings, job fairs, defense meetings, or lecture meetings, etc. The mapping relationship table includes meeting scene graphs corresponding to birthday parties, meeting scene graphs corresponding to report meetings, meeting scene graphs corresponding to job fairs, meeting scene graphs corresponding to defense meetings, or meeting scene graphs corresponding to lecture meetings, and these meeting scene graphs are all different.
[0045] In some embodiments, the target meeting layout includes the positions of each participant in the meeting scene graph and the role information corresponding to these positions. According to the target meeting layout and the role information of each participant, the real portraits of each participant are laid out in the meeting scene graph to obtain a meeting scene page, including: determining the target positions of the real portraits of each participant in the meeting scene graph according to the positions of each participant in the meeting scene graph, the role information corresponding to these positions, and the role information of each participant; laying out the real portraits of each participant at the corresponding target positions in the meeting scene graph to obtain a meeting scene page. For example, the real portraits of participants with role information as the birthday protagonist or the host are displayed at the central position of the meeting scene graph, and the real portraits of participants with role information as the audience are displayed at the edge positions of the meeting scene graph.
[0046] In some embodiments, incremental learning techniques can be used to update and optimize the meeting layout determination model. For example, when updating the meeting layout determination model, the Elastic Weight Consolidation (EWC) algorithm can be used to update the parameters of the meeting layout determination model to avoid catastrophic forgetting. Another example is to update the meeting layout determination model by means of a dynamic replay buffer, such as storing a small number of representative samples of old meeting scenes, mixing them with the data of new meeting scenes, and then training the meeting layout determination model to maintain the memory of the meeting layout determination model for historical scenes. Another example is to use a lightweight network branch (such as an adapter module) to process new meeting scenes, keeping the parameters of the meeting layout determination model fixed and only updating the parameters of the lightweight network branch to achieve rapid deployment.
[0047] Step S103: Configure corresponding virtual interaction props for each participant according to the role information of each participant.
[0048] In this embodiment, the virtual interactive prop is a virtual interactive prop, which may include a virtual cheering stick, a virtual cheering sign, a virtual clapping board, a virtual like sign, a virtual blowing kiss sign or a virtual love sign, etc. Among them, different role information corresponds to different virtual interactive props.
[0049] For example, the role information may include the birthday protagonist, the audience, the speaker or the host, etc. For the participants with the role information of the audience, virtual interactive props such as virtual cheering sticks, virtual cheering signs, virtual clapping boards, virtual like signs, virtual blowing kiss signs or virtual love signs can be configured for them. For the participants with the role information of the birthday protagonist, the speaker or the host, virtual interactive props such as virtual blowing kiss signs or virtual love signs can be configured for them.
[0050] In some embodiments, configuring corresponding virtual interactive props for each participant according to the role information of each participant may include: for each participant, querying the mapping relationship table between the role information and the virtual interactive prop information to obtain the virtual interactive prop information corresponding to the role information of the participant, and configuring the corresponding virtual interactive prop for the participant according to the virtual interactive prop information. Among them, the mapping relationship table between the role information and the virtual interactive prop information can be set based on the actual situation, and the embodiments of the present application do not make specific limitations on this.
[0051] In some embodiments, the method further includes: displaying virtual decoration props on the real portrait of the target participant, where the role information of the target participant is the same as the role information corresponding to the type of the video conference. Among them, the virtual decoration props move along with the movement of the real portrait of the target participant, and the virtual decoration props include a virtual birthday crown or a virtual highlighter. For example, a virtual birthday crown is displayed on the top of the real portrait of the birthday protagonist. For another example, a virtual highlighter is displayed in the hand of the real portrait of the lecturer. In this embodiment, by displaying virtual decoration props on the real portrait of the target participant, it is convenient to distinguish the participants of different roles, so that the participants can pay more attention to the target participant with virtual decoration props displayed, so as to improve the meeting effect.
[0052] Step S104: Recognize the body movements of each participant through a preset body movement recognition model to obtain the body movements of each participant.
[0053] In this embodiment, the body movement recognition model may be a real-time body movement recognition model, and the body movement recognition model is obtained by training a neural network model in advance according to a plurality of third training samples, and the third training samples include sample portraits and labeled body movements. Among them, the neural network model may include an OpenPose model, an AlphaPose model or a PoseNet model, etc.
[0054] In some embodiments, the human motion recognition model is a neural network model after inference optimization by TensorRT. The human motion recognition model in this embodiment is a neural network model after inference optimization by TensorRT, which can improve the running speed of the human motion recognition model on a graphics processing unit (GPU).
[0055] Step S105: In response to the human body movement of at least one participant being the same as the human body movement corresponding to the virtual interactive prop, displaying a special effect corresponding to at least one virtual interactive prop in the conference scene page.
[0056] This embodiment configures corresponding virtual interactive props for each participant based on the role information of each participant, and recognizes the human motion of each participant through a human motion recognition model to obtain the human motion of each participant. Then, when the human motion of at least one participant is the same as the human motion corresponding to the virtual interactive prop, the special effects corresponding to at least one virtual interactive prop are displayed in the conference scene page, thereby effectively improving the convenience of video conference interaction.
[0057] In some embodiments, different virtual interactive props correspond to different special effects, and different virtual interactive props correspond to different human movements. For example, the human movement corresponding to the virtual cheering stick is the movement of the participant raising one hand and waving it left and right, and the corresponding special effect is the cheering stick special effect, the human movement corresponding to the virtual cheering board is the movement of the participant raising both hands and waving it left and right, and the corresponding special effect is the cheering board special effect, the human movement corresponding to the virtual applause is the movement of clapping with both hands, and the corresponding special effect is the applause special effect, the human movement corresponding to the virtual like board is the movement of like with one hand, and the corresponding special effect is the like special effect, and the human movement corresponding to the virtual kiss board is the movement of kissing, and the corresponding special effect is the kissing special effect.
[0058] In some embodiments, after step S102, it also includes: obtaining voice data of the target participant; identifying page adjustment instructions on the voice data through a preset page adjustment instruction recognition model to obtain a page adjustment instruction recognition result; in response to the presence of a page adjustment instruction in the page adjustment instruction recognition result, adjusting the conference scene page, wherein the conference process link corresponding to the conference scene page before the adjustment is different from the conference process link corresponding to the conference scene page after the adjustment. The page adjustment instruction recognition model is obtained by pre-training the BERT model or the GPT model based on multiple fourth training samples, and the fourth training samples include voice samples and annotated page adjustment instructions. This embodiment improves the intelligence of the video conference by identifying the voice of the target participant, such as the host, and thereby automatically adjusting the conference scene page to enter different links of the conference process.
[0059] For example, the host's voice is "Now, let me introduce the birthday star to you all". At this time, the page adjustment instruction recognition model recognizes the page adjustment instruction for the voice "Now, let me introduce the birthday star to you all", and obtains a page adjustment instruction for adjusting the real portrait of the birthday star to the main display area of the conference scene page. In response to this page adjustment instruction, the content displayed in the main display area of the conference scene page is switched to the real portrait of the birthday star. When the host's voice is "Please wish the birthday star", the page adjustment instruction recognition model recognizes the page adjustment instruction for the voice "Please wish the birthday star", and obtains a page adjustment instruction for adjusting the real portraits of the current birthday-wishing people and the birthday star to the main display area (central area) of the conference scene page. In response to this page adjustment instruction, the content displayed in the main display area of the conference scene page is switched to the real portraits of the current birthday-wishing people and the birthday star.
[0060] In some embodiments, the conference processes of different types of video conferences are different. For example, the conference process of a birthday party includes a birthday star introduction session, a birthday star blessing session, a birthday song singing session, a wishing session, and a group photo session, etc. In the birthday star introduction session, the real portrait of the birthday star is displayed in the main display area (central area) of the conference scene page. In the birthday star blessing session, the real portraits of the current birthday-wishing people and the birthday star are displayed side by side in the main display area of the conference scene page. In the birthday song singing session, the real portrait of the birthday star and the birthday song lyrics are displayed in the main display area of the conference scene page. In the wishing session, a virtual birthday cake is displayed in the main display area of the conference scene page. In the group photo session, the real portraits of all people are displayed on the conference scene page.
[0061] Another example is that the conference process of a training conference includes an instructor introduction session, a training session, a Q&A session, and a group photo session, etc. In the instructor introduction session, the real portrait of the instructor is displayed in the main display area (central area) of the conference scene page. In the training session, the real portrait of the instructor and the training materials are displayed in the main display area of the conference scene page. In the Q&A session, the real portraits of the current questioners and the instructor are displayed side by side in the main display area of the conference scene page. In the group photo session, the real portraits of all people are displayed on the conference scene page. Still another example is that the conference process of a discussion conference includes a theme introduction session, a discussion session, and a group photo session, etc. The conference process of an award ceremony includes a guest introduction session, an award presentation session, a recipient speech session, and a group photo session, etc.
[0062] In some embodiments, when training the page adjustment instruction recognition model, sequence modeling technology, such as recurrent neural network (RNN) or Transformer, is introduced to convert the instruction recognition task into a sequence labeling problem to better capture the sequential relationship and instruction semantics in the process, thereby achieving more accurate and intelligent conference process control. For example, the host's instructions (such as "start speech → play slides → enter the question and answer session") are converted into sequence labeling tasks, and Bi-LSTM or Transformer encoders are used to capture temporal dependencies. The self-attention mechanism is introduced in the Transformer to identify key actions in the instructions (such as "play slides") and ignore redundant information.
[0063] In some embodiments, the human body motion of a participant whose role information is an audience is recognized by a preset human body motion recognition model; in response to the human body motion of a participant whose role information is an audience being a preset questioning motion, the real portrait of the audience and the real portrait of the lecturer are displayed in the main display area of the conference scene page. Among them, the preset questioning motion can be set based on actual conditions, and the embodiments of the present application do not make specific limitations on this. For example, the preset questioning motion is a hand-raising motion. This embodiment can automatically switch the real portrait of the questioner in the main display area, thereby improving the intelligence of the video conference.
[0064] In some embodiments, a virtual birthday cake is displayed on the conference scene page, and after step S102, the process further includes: recognizing the human body motion of the participant whose role information is the birthday protagonist through a preset human body motion recognition model; in response to the human body motion of the participant whose role information is the birthday protagonist being the action of blowing out candles, controlling the virtual candles on the virtual birthday cake to be extinguished. This embodiment improves the fun of holding birthday parties through video conferencing.
[0065] In some embodiments, after step S102, it also includes: in response to a group photo start instruction triggered by a target participant, displaying a group photo countdown in the conference scene page; extracting a group photo portrait of each participant from the video stream data when the group photo countdown returns to zero through a preset target detection and instance segmentation model; generating a group photo based on the group photo portrait of each participant and the role information of each participant. Among them, the preset target detection and instance segmentation model may include a Mask R-CNN model. This embodiment can ensure that each person participating in the video conference can be photographed, thereby improving the integrity of the people in the group photo.
[0066] In some embodiments, the video conferencing method further includes: recognizing the voice data of the target participant through a preset voice recognition model to obtain text information, and triggering a group photo start instruction in response to the presence of a keyword for triggering the group photo start instruction in the text information. Alternatively, recognizing the gestures of the target participant through a preset gesture recognition model to obtain the current gestures of the target participant; triggering a group photo start instruction in response to the current gestures of the target participant being the same as the preset gestures for triggering the group photo start instruction. Alternatively, the conference scene page includes a group photo button, and the terminal device triggers a group photo start instruction in response to the click operation of the target participant on the group photo button, and sends the group photo start instruction to the server.
[0067] In some embodiments, generating a group photo according to the group photo portraits of each participant and the role information of each participant may include: obtaining a group photo style template corresponding to the type of the video conference; determining the target positions of the group photo portraits of each participant in the group photo style template according to the role information of each participant; filling the group photo portraits of each participant into the corresponding target positions in the group photo style template to obtain a group photo. In this embodiment, a group photo of a corresponding style can be generated through the group photo style template, and the quality of the group photo is better.
[0068] In some embodiments, after generating a group photo according to the group photo portraits of each participant and the role information of each participant, it further includes: sending the group photo to the email of each participant. In this embodiment, by sending the group photo to the email of each participant, it is not necessary for the participants to manually capture the group photo during the video conference, which is convenient for each participant to obtain the group photo after the video conference ends, improving the user experience.
[0069] Please refer to Figure 3 , Figure 3 which is a schematic block diagram of an AI-based video conferencing interaction device provided by an embodiment of the present application.
[0070] As Figure 3 shown, the video conferencing interaction device 300 includes:
[0071] An obtaining module 310, configured to obtain the video stream data and role information of each participant participating in the video conference;
[0072] The obtaining module 310 is further configured to extract the real portrait of each participant from the video stream data of each participant;
[0073] A generating module 320, configured to generate a conference scene page according to the real portrait and role information of each participant;
[0074] A configuration module 330, configured to configure corresponding virtual interaction props for each of the participants according to the role information of each participant.
[0075] An action recognition module 340, configured to recognize the human actions of each participant through a preset human action recognition model to obtain the human actions of each participant.
[0076] An interaction module 350, configured to, in response to the human actions of at least one participant being the same as the human actions corresponding to the virtual interaction props, display special effects corresponding to at least one of the virtual interaction props on the conference scene page.
[0077] In some embodiments, the interaction module 350 is further configured to display virtual decoration props on the real portrait of the target participant, where the role information of the target participant is the same as the role information corresponding to the type of the video conference.
[0078] In some embodiments, as Figure 4 shown, the generation module 320 includes:
[0079] An acquisition sub-module 321, configured to acquire a conference scene graph matching the type of the video conference.
[0080] A determination sub-module 322, configured to determine a target conference layout based on the conference scene graph, the role information of each participant, and the total number of participants through a preset conference layout determination model, where the conference layout determination model is obtained by training a neural network model according to a plurality of training samples.
[0081] A layout sub-module 323, configured to layout the real portraits of each participant in the conference scene graph according to the target conference layout and the role information of each participant to obtain a conference scene page.
[0082] In some embodiments, the video conference interaction device 300 further includes:
[0083] The acquisition module 310 is further configured to acquire the voice data of the target participant.
[0084] An instruction recognition module, configured to recognize page adjustment instructions from the voice data through a preset page adjustment instruction recognition model to obtain a page adjustment instruction recognition result.
[0085] An adjustment module, configured to, in response to the existence of a page adjustment instruction in the page adjustment instruction recognition result, adjust the conference scene page, where the conference process link corresponding to the conference scene page before adjustment is different from the conference process link corresponding to the conference scene page after adjustment.
[0086] In some embodiments, a virtual birthday cake is displayed on the conference scene page, and the video conferencing interaction device 300 further includes:
[0087] An action recognition module 340, configured to recognize the human body actions of the participants whose role information is the birthday protagonist through a preset human body action recognition model;
[0088] A control module, configured to control the virtual candles on the virtual birthday cake to go out in response to the human body action of the participant whose role information is the birthday protagonist being a candle-blowing action.
[0089] In some embodiments, the video conferencing interaction device 300 further includes:
[0090] A countdown module, configured to display a group photo countdown on the conference scene page in response to a group photo start instruction triggered by a target participant;
[0091] A portrait extraction module, configured to extract the group photo portraits of each participant from the video stream data when the group photo countdown reaches zero through a preset target detection and instance segmentation model;
[0092] A group photo module, configured to generate a group photo according to the group photo portraits of each participant and the role information of each participant.
[0093] In some embodiments, the video conferencing interaction device 300 further includes:
[0094] A sending module, configured to send the group photo to the email of each participant.
[0095] It should be noted that those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described device and each module and unit can refer to the corresponding processes in the foregoing embodiments of the video conferencing interaction method, and will not be described herein again.
[0096] The device provided in the above embodiment can be implemented in the form of a computer program, and the computer program can run on a server as shown in Figure 5 shown.
[0097] Please refer to Figure 5 , Figure 5 which is a schematic block diagram of the structure of a server provided in an embodiment of the present application.
[0098] As shown in Figure 5 shown, the server includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory may include a storage medium and an internal memory.
[0099] The storage medium can store an operating system and a computer program. The computer program includes program instructions which, when executed, can cause the processor to execute any one of the video conferencing interaction methods.
[0100] The processor is used to provide computing and control capabilities to support the operation of the entire server.
[0101] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 5 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the server to which the solution of this application is applied. The specific server may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0102] It should be understood that the processor can be a Central Processing Unit (CPU), and the processor can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0103] Among them, in one embodiment, the processor is used to run the computer program stored in the memory to implement the following steps:
[0104] Obtain the video stream data and role information of each participant in the video conference;
[0105] Extract the real portrait of each participant from the video stream data of each participant, and generate a conference scene page according to the real portrait and role information of each participant;
[0106] Configure corresponding virtual interaction props for each participant according to the role information of each participant;
[0107] Identify the human actions of each participant through a preset human action recognition model to obtain the human actions of each participant;
[0108] In response to the body movements of at least one of the participating persons being the same as the body movements corresponding to the virtual interaction prop, display special effects corresponding to at least one of the virtual interaction props within the conference scene page.
[0109] In some embodiments, the processor is further configured to implement:
[0110] Display virtual decoration props on the real portrait of the target participating person, where the role information of the target participating person is the same as the role information corresponding to the type of the video conference.
[0111] In some embodiments, when the processor generates a conference scene page according to the real portraits and role information of each of the participating persons, it is configured to implement:
[0112] Obtain a conference scene map that matches the type of the video conference;
[0113] Determine a target conference layout through a preset conference layout determination model based on the conference scene map, the role information of each of the participating persons, and the total number of participating persons, where the conference layout determination model is obtained by training a neural network model in advance according to a plurality of training samples;
[0114] Layout the real portraits of each of the participating persons in the conference scene map according to the target conference layout and the role information of each of the participating persons to obtain a conference scene page.
[0115] In some embodiments, after the processor generates a conference scene page according to the real portraits and role information of each of the participating persons, it is further configured to implement:
[0116] Obtain the voice data of the target participating person;
[0117] Identify page adjustment instructions for the voice data through a preset page adjustment instruction recognition model to obtain a page adjustment instruction recognition result;
[0118] In response to the existence of page adjustment instructions in the page adjustment instruction recognition result, adjust the conference scene page, where the conference process link corresponding to the conference scene page before adjustment is different from the conference process link corresponding to the conference scene page after adjustment.
[0119] In some embodiments, a virtual birthday cake is displayed on the conference scene page. After the processor generates a conference scene page according to the real portraits and role information of each of the participating persons, it is further configured to implement:
[0120] Identify the body movements of the participating person whose role information is the birthday protagonist through a preset body movement recognition model;
[0121] In response to the human body movement of the participant whose role information is the birthday protagonist being the action of blowing out the candles, control the virtual candles on the virtual birthday cake to go out.
[0122] In some embodiments, after the processor implements generating a conference scene page according to the real portraits and role information of each participant, it is further configured to implement:
[0123] In response to the group photo start instruction triggered by the target participant, display a group photo countdown on the conference scene page;
[0124] Extract the group photo portraits of each participant from the video stream data when the group photo countdown reaches zero through a preset target detection and instance segmentation model;
[0125] Generate a group photo according to the group photo portraits of each participant and the role information of each participant.
[0126] In some embodiments, after the processor implements generating a group photo according to the group photo portraits of each participant and the role information of each participant, it is further configured to implement:
[0127] Send the group photo to the email of each participant.
[0128] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the server described above can refer to the corresponding process in the foregoing embodiments of the video conferencing interaction method, and will not be elaborated here.
[0129] From the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a server (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0130] The embodiments of this application further provide a storage medium, on which a computer program is stored. The computer program includes program instructions, and the method implemented when the program instructions are executed can refer to various embodiments of the video conferencing interaction method of this application.
[0131] Among them, the storage medium can be volatile or non-volatile. The storage medium can be the internal storage unit of the server described in the foregoing embodiments, such as the hard disk or memory of the server. The storage medium can also be an external storage device of the server, such as a plug-in hard disk equipped on the server, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0132] Further, the storage medium may mainly include a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.; the storage data area may store data created according to the use of the blockchain node, etc.
[0133] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, an application service layer, etc.
[0134] It should be understood that the non-company software tools or components appearing in the embodiments of this application are only introduced by way of example and do not represent actual use. The terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0135] It should also be understood that the term "and / or" used in this application specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. It should be noted that in this article, the term "comprises", "comprising" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or system comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or system. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or system comprising the element.
[0136] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments. As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A video conferencing interaction method based on artificial intelligence, characterized in that, Including: Obtaining video stream data and role information of each participant in the video conference; Extracting the real portrait of each participant from the video stream data of each participant, and generating a conference scene page according to the real portrait and role information of each participant; Configuring corresponding virtual interaction props for each participant according to the role information of each participant; Identifying the human body movements of each participant through a preset human body movement recognition model to obtain the human body movements of each participant; In response to the human body movements of at least one participant being the same as the human body movements corresponding to the virtual interaction props, displaying special effects corresponding to at least one of the virtual interaction props in the conference scene page.
2. The video conferencing interaction method according to claim 1, wherein The method further includes: Displaying virtual decoration props on the real portrait of the target participant, where the role information of the target participant is the same as the role information corresponding to the type of the video conference.
3. The video conferencing interaction method according to claim 1, wherein The generating a conference scene page according to the real portrait and role information of each participant includes: Obtaining a conference scene map matching the type of the video conference; Determining a target conference layout through a preset conference layout determination model based on the conference scene map, the role information of each participant, and the total number of participants, where the conference layout determination model is obtained by training a neural network model according to multiple training samples in advance; Laying out the real portraits of each participant in the conference scene map according to the target conference layout and the role information of each participant to obtain a conference scene page.
4. The video conferencing interaction method according to claim 1, wherein After generating the conference scene page according to the real portrait and role information of each participant, it further includes: Obtaining the voice data of the target participant; Identifying a page adjustment instruction from the voice data through a preset page adjustment instruction recognition model to obtain a page adjustment instruction recognition result; In response to the existence of a page adjustment instruction in the page adjustment instruction recognition result, adjusting the conference scene page, where the conference process link corresponding to the conference scene page before adjustment is different from the conference process link corresponding to the conference scene page after adjustment.
5. The video conferencing interaction method according to claim 1, wherein The conference scene page displays a virtual birthday cake. After generating the conference scene page according to the real portrait and role information of each participant, it further includes: Identifying the human body movements of the participant with the role information of the birthday protagonist through a preset human body movement recognition model; In response to the human body movement of the participant with the role information of the birthday protagonist being a candle-blowing action, controlling the virtual candles on the virtual birthday cake to go out.
6. The video conferencing interaction method according to any one of claims 1-5, characterized in that, After generating the conference scene page according to the real portrait and role information of each participant, it further includes: In response to a group photo start instruction triggered by the target participant, displaying a group photo countdown in the conference scene page; Extracting the group photo portraits of each participant from the video stream data when the group photo countdown reaches zero through a preset target detection and instance segmentation model. Generate a group photo based on the group photo portrait of each said participant and the role information of each said participant.
7. The video conferencing interaction method according to claim 6, wherein After generating the group photo based on the group photo portrait of each said participant and the role information of each said participant, it further includes: Send the group photo to the email of each said participant.
8. An artificial intelligence-based video conferencing interactive device, characterized in that, The video conferencing interaction device includes: An acquisition module for acquiring the video stream data and role information of each participant participating in the video conference; The acquisition module is further configured to extract the real portrait of each said participant from the video stream data of each said participant; A generation module for generating a conference scene page based on the real portrait and role information of each said participant; A configuration module for configuring corresponding virtual interaction props for each said participant according to the role information of each said participant; An action recognition module for recognizing the human actions of each said participant through a preset human action recognition model to obtain the human actions of each said participant; An interaction module for, in response to the human actions of at least one said participant being the same as the human actions corresponding to the virtual interaction props, displaying the special effects corresponding to at least one said virtual interaction prop within the conference scene page.
9. A server, characterized in that, The server includes a processor, a memory, and a computer program stored on the memory and executable by the processor, wherein when the computer program is executed by the processor, the steps of the video conferencing interaction method according to any one of claims 1 to 7 are implemented.
10. A storage medium for computer-readable storage, characterized in that, A computer program is stored on the storage medium, wherein when the computer program is executed by a processor, the steps of the video conferencing interaction method according to any one of claims 1 to 7 are implemented.