Method and device for processing video conference based on semantic information, equipment and medium
By obtaining and processing the reference portrait change feature information of the video conferencing terminal, slicing the audio data and synthesizing the video stream, the problem of audio and video out-synchronization during video stream synthesis is solved, real-time synthesis and synchronization of high-definition video streams is realized, and the quality of video conferencing is improved.
Patent Information
- Application Number
- CN202311844754.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-25
AI Technical Summary
In video conferencing, due to the performance differences and bandwidth fluctuations of conference terminals, the video streams uploaded by some terminals are not high-definition and screen stuttering occur, and the existing technology is difficult to realize real-time synthesis and synchronization of high-definition video streams.
By obtaining the reference portrait change feature information of the target conference terminal, processing audio data slices and feature information, synthesize digital portrait video streams and sending them to other terminals, reducing bandwidth requirements and realizing the synthesis of high-definition video streams.
It solves the problem of audio and video out-synchronization during video stream synthesis, improves the quality of video streams, ensures that other terminals display high-definition video conferencing images, and reduces bandwidth requirements.
Smart Images

Figure CN120378569A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and in particular, to a method, apparatus, device, and medium for processing video conferences based on semantic information. Background Art
[0002] With the development of communication technologies, users have increasingly higher requirements for the quality and efficiency of communication, and their demands have become more diverse and differentiated. In the video conferencing scenario, users are not satisfied with merely being able to see real-time video images, and their demand for high-definition, high-quality, and highly stable video conferencing services is becoming increasingly strong.
[0003] In a video conference, the video stream captured by the conference terminal is uploaded to the server, and then distributed by the server to the display devices at each venue for display. However, due to the different performances of the conference terminals, some conference terminals may upload high-definition video streams, while some may not. If all conference terminals are replaced with high-definition cameras, video conferencing images may freeze during video stream transmission due to large bandwidth fluctuations and other reasons. Summary of the Invention
[0004] In view of the above problems, a method, apparatus, device, and medium for processing video conferences based on semantic information are provided to overcome or at least partially solve the above problems, including:
[0005] A method for processing video conferences based on semantic information, the method including:
[0006] Obtain the reference portrait change feature information associated with the target conference terminal;
[0007] When a trigger event for the target conference terminal is detected, obtain a data stream collected by the target conference terminal that includes at least audio data;
[0008] Slice the audio data in the data stream to obtain a plurality of audio segments;
[0009] According to the plurality of audio segments and the reference portrait change feature information, perform video stream synthesis to obtain a video stream for presenting a digital portrait, and send the video stream to other conference terminals.
[0010] Optionally, the obtaining the reference portrait change feature information associated with the target conference terminal includes:
[0011] Obtain the audio material and the corresponding portrait material associated with the target conference terminal;
[0012] Generate the reference portrait change feature information associated with the target conference terminal according to the audio material and the corresponding portrait material.
[0013] Optionally, performing video stream synthesis based on the multiple audio segments and the reference portrait variation feature information to obtain a video stream for presenting a digital portrait includes:
[0014] Generating target portrait variation feature information corresponding to the multiple audio segments according to the multiple audio segments and the reference portrait variation feature information;
[0015] Performing video stream synthesis according to the multiple audio segments and the target portrait variation feature information to obtain a video stream for presenting a digital portrait.
[0016] Optionally, performing video stream synthesis according to the multiple audio segments and the target portrait variation feature information to obtain a video stream for presenting a digital portrait includes:
[0017] Obtaining target image data associated with the target conference terminal;
[0018] Performing video stream synthesis according to the multiple audio segments, the target portrait variation feature information, and the target image data to obtain a video stream for presenting a digital portrait.
[0019] Optionally, the data stream further includes video feature attribute information generated according to the video stream collected by the target conference terminal. Performing video stream synthesis according to the multiple audio segments and the reference portrait variation feature information to obtain a video stream for presenting a digital portrait includes:
[0020] Obtaining target image data associated with the target conference terminal;
[0021] Performing video stream synthesis according to the multiple audio segments, the target portrait variation feature information, the video feature attribute information, and the target image data to obtain a video stream for presenting a digital portrait.
[0022] Optionally, performing slicing processing on the audio data in the data stream to obtain multiple audio segments includes:
[0023] Determining the audio variation feature information of the audio data in the data stream;
[0024] Performing slicing processing on the audio data in the data stream according to the audio variation feature information to obtain multiple audio segments.
[0025] Optionally, the data stream containing at least audio data is the data stream corresponding to the video key area in the video stream sent by the target conference terminal, and the synthesized video stream is used to replace the video key area.
[0026] An apparatus for processing video conferences based on semantic information, the apparatus comprising:
[0027] A reference portrait change feature information acquisition module, configured to acquire reference portrait change feature information associated with the target conference terminal;
[0028] A data stream acquisition module, configured to acquire a data stream including at least audio data collected by the target conference terminal when a trigger event for the target conference terminal is detected;
[0029] An audio slicing module, configured to slice the audio data in the data stream to obtain a plurality of audio segments;
[0030] A video stream synthesis module, configured to perform video stream synthesis according to the plurality of audio segments and the reference portrait change feature information to obtain a video stream for presenting a digital portrait, and send the video stream to other conference terminals.
[0031] An electronic device, comprising a processor, a memory, and a computer program stored on the memory and capable of running on the processor, where when the computer program is executed by the processor, the method for processing video conferences based on semantic information as described above is implemented.
[0032] A computer-readable storage medium, on which a computer program is stored, where when the computer program is executed by a processor, the method for processing video conferences based on semantic information as described above is implemented.
[0033] The embodiments of the present invention have the following advantages:
[0034] In the embodiments of the present invention, by acquiring reference portrait change feature information associated with the target conference terminal, when a trigger event for the target conference terminal is detected, acquiring a data stream including at least audio data collected by the target conference terminal, slicing the audio data in the data stream to obtain a plurality of audio segments, performing video stream synthesis according to the plurality of audio segments and the reference portrait change feature information to obtain a video stream for presenting a digital portrait, and sending the video stream to other conference terminals, video stream synthesis is achieved according to portrait change features, thereby being able to solve the problem of audio-visual asynchronization when synthesizing video streams and improving the quality of video streams.
[0035] Moreover, by synthesizing a video stream according to the audio data uploaded by the conference terminal, the requirements for aspects such as bandwidth for the data uploaded by the conference terminal are reduced, data can be uploaded with a lower bandwidth and high-definition video conference images can be displayed on other terminals, thereby being able to make up for the situation where the video stream is difficult to meet the requirements due to reasons such as non-high-definition shooting and large bandwidth fluctuations. Description of the Drawings
[0036] To more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0037] Figure 1 is a flowchart of the steps of a method for processing video conferences based on semantic information provided by some embodiments of the present invention;
[0038] Figure 2a is a schematic diagram of a system architecture provided by some embodiments of the present invention;
[0039] Figure 2b is a schematic diagram of another system architecture provided by some embodiments of the present invention;
[0040] Figure 3 is a flowchart of the steps of another method for processing video conferences based on semantic information provided by some embodiments of the present invention;
[0041] Figure 4 is a flowchart of the steps of another method for processing video conferences based on semantic information provided by some embodiments of the present invention;
[0042] Figure 5 is a flowchart of the steps of another method for processing video conferences based on semantic information provided by some embodiments of the present invention;
[0043] Figure 6 is a flowchart of the steps of another method for processing video conferences based on semantic information provided by some embodiments of the present invention;
[0044] Figure 7 is a block diagram of the structure of a device for processing video conferences based on semantic information provided by some embodiments of the present invention. Detailed implementation manners
[0045] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0046] In the related art, semantic prediction is to identify natural language after inputting a whole passage, and output the identified result for semantic feature prediction. However, the result obtained by this method is not real-time, and the execution of the result is also unpredictable. Therefore, in the process of real-time ultra-clear portrait synthesis, the semantic prediction technology cannot be incorporated, and the simulation of ultra-clear portraits cannot be made more accurate and vivid.
[0047] In the embodiments of the present invention, by combining the voice materials and human pixel materials input by each participant before the video conference, the predicted portrait change feature information is generated. During the video conference, the conference server switches to receive the audio stream or the audio stream and video feature attribute information, slices the real-time voice, and inputs the portrait change feature information, the received audio stream or the audio stream and video feature attribute information into the portrait synthesis model to generate a high-definition portrait video stream. The generated ultra-clear portrait can continuously fit the actions of the portrait speaking in combination with the semantic prediction of the context without affecting the progress of the video conference, thereby solving the problem of real-time semantic and facial expression synchronization of the characters in the picture during the high-definition video conference.
[0048] The present invention will be further described below with reference to the accompanying drawings:
[0049] Refer to Figure 1 , which shows a flowchart of the steps of a method for processing a video conference based on semantic information provided by some embodiments of the present invention. This method can be applied to a server in a video conference.
[0050] In some examples, the server can be a single server or multiple servers. For example Figure 2a , the server can include a conference server (such as an all-in-one machine, XMCU, etc.) and a video post-processing server (which can also be an AI server). The conference server and the video post-processing server are connected to each other. This method can be partially executed on the conference server and partially executed on the video post-processing server. For example, the relevant content of video stream synthesis can be executed by the video post-processing server. The relevant content of video stream synthesis can include AI processing (such as determining digital portraits, super-resolution reconstruction, etc.) and streaming media processing (such as audio encoding and decoding, video encoding and decoding), and other content can be executed by the conference server.
[0051] Among them, the video post-processing server (AI server) is a data server that can provide artificial intelligence. It can be used to support local applications and web pages, and can also provide complex AI models and services for cloud and local servers.
[0052] In some examples, the video post - processing server can be set with an AI conference management module, a user management module, and a proxy service module. Among them, the AI conference management module can be used to manage the establishment of a communication channel with the conference server, calculate and classify the received data stream, and at the same time encode and decode the audio stream and the video stream, and synthesize a high - definition digital portrait video stream. The user management module can be used to manage the user information in the video conference. When a conference terminal needs to create a high - definition digital portrait video stream, the user management module will create the user information of a certain conference terminal in this video conference. The user name, portrait data, etc. used in this user information can be synchronized by the conference server to the AI server when the conference terminal starts the conference, mainly including the conference ID, the user ID when logging in to the conference terminal, the image ID, the identifier of the conference terminal, etc. The proxy module can be used to establish a communication channel with the conference server, and then can implement functions such as file upload and file download. After the proxy module establishes a communication channel with the conference server, it can forward files.
[0053] Such as Figure 2b , the AI server can generate personalized label information of the participants (i.e., reference portrait change feature information) according to the materials obtained before the start of the conference and in combination with the portrait image data and conference scene image data obtained in advance. Then, during the video conference, it can obtain real - time data input, and then output a high - definition digital personnel conference video stream by combining digital humans, audio streams, super - resolution reconstruction, and the personalized label information of the participants.
[0054] Specifically, it can include the following steps:
[0055] Step 101, obtain the reference portrait change feature information associated with the target conference terminal.
[0056] Before the start of the video conference, the reference portrait change feature information can be generated according to the collected materials. In some examples, the reference portrait change feature information of different participants is different, that is, the reference portrait change feature information is the personalized information of the participants and can be stored for the participants, and then can be used in subsequent other conferences.
[0057] In some examples, the reference portrait change feature information can include the corresponding relationship between the semantic type and the portrait expression. For example, when person A reads "I am very happy", the displacement of the facial features of the portrait picture of this person is calculated through the portrait change model, and the offset of the center point of the eyes after change from the center point of the eyes before change is recorded, and this offset is used to generate the reference portrait change feature information.
[0058] For another example, for multiple individuals (such as Individual A, Individual B, and Individual C), the change characteristics of their expressions under different material texts. For Individual A, "happy, eyebrow change offset (-2, -1)"; for Individual A, "sad, eyebrow change offset (+3, +5)". Record this offset to generate reference portrait change characteristic information.
[0059] In some embodiments of the present invention, obtaining the reference portrait change characteristic information associated with the target conference terminal includes:
[0060] Obtain the audio material and corresponding portrait material associated with the target conference terminal; based on the audio material and the corresponding portrait material, generate the reference portrait change characteristic information associated with the target conference terminal.
[0061] In practical applications, by providing material texts to the participants, the audio material and the corresponding portrait material (the portrait material can be in the form of an image or a video) collected when the participants read the material texts can be recorded. Then, through the analysis of the audio material and the corresponding portrait material, the reference portrait change characteristic information can be generated and stored for subsequent use.
[0062] In some examples, each person records for common phrases (such as "Hello", "Goodbye", "The weather is nice today") and common modal particles (such as "Not bad", "Happy", "Angry") according to the same material text. The material text is randomly extracted from a modal particle library, and the modal particle library is updated regularly. When recording the voice material, the portrait pictures of the person when reading the material text are also captured. Multiple expression portrait pictures are captured, and then the language material and the portrait pictures are input into the portrait change model. Then, the characteristic information of the portrait change under the personality characteristics can be measured according to the expression changes of each person.
[0063] In some examples, an interaction interface can be provided to the user, and the material text is presented to the user in the interaction interface. The user can read in accordance with the material text, and the material information when the user reads is collected.
[0064] Step 102, in the case of detecting a trigger event for the target conference terminal, obtain a data stream collected by the target conference terminal that at least includes audio data.
[0065] Among them, the target conference terminal and other conference terminals are participating terminals in the same video conference.
[0066] In a video conference, the target conference terminal can collect a video stream and upload it to the server. Then, the server sends the video stream to other conference terminals and displays it through a display device.
[0067] In some scenarios, the video stream uploaded by the target conference terminal to the server may not meet the requirements. For example, the camera connected to the target conference terminal can only capture a video stream with low clarity and cannot capture a high-definition video stream. Another example is that due to large fluctuations in bandwidth, the video stream uploaded by the target conference terminal may experience stuttering in the video conference screen. In this case, when a trigger event is detected, the video stream of the video conference can be synthesized to ensure that the video stream of the video conference can meet the requirements.
[0068] In some examples, when a trigger event is detected, the system can switch from the general processing mode (i.e., directly uploading the video stream captured by the target conference terminal to the server in the above text, and then the server sending it to other conference terminals) to the AI processing mode, and then synthesize the video stream of the video conference in the AI processing mode (i.e., switch to collecting at least the data stream including audio data by the target conference terminal and subsequent processing).
[0069] In some embodiments of the present invention, the trigger event for the target conference terminal may include any one of the following: receiving a video synthesis request sent by the target conference terminal, and the bandwidth fluctuation amplitude of the channel connected to the target conference terminal being greater than a threshold value.
[0070] In some implementation manners, an interactive interface can be provided to the user on the target conference terminal. The user can perform manual operations through the interactive interface, and then generate and send a video synthesis request to the server through the target conference terminal. When the server receives the video synthesis request, the trigger event is detected. For example, when the camera connected to the target conference terminal can only capture a video stream with low clarity and cannot capture a high-definition video stream, the user can control the target conference terminal to generate and send a video synthesis request through the interactive interface; another example is that due to large fluctuations in bandwidth, the video stream uploaded by the target conference terminal experiences stuttering in the video conference screen, and the user can control the target conference terminal to generate and send a video synthesis request through the interactive interface.
[0071] In some other implementation manners, the server can establish communication channels with each participating terminal. The server can be provided with a bandwidth fluctuation monitoring module, and the bandwidth fluctuation monitoring module can detect the bandwidth fluctuation conditions of each communication channel. Each channel can be set with a certain bandwidth threshold. After the video conference is started, the server can detect whether the bandwidth fluctuation amplitude of the channel connected to the target conference terminal is greater than the threshold value (such as 100K) through the bandwidth fluctuation monitoring module. When it is detected that the bandwidth fluctuation amplitude of the channel connected to the target conference terminal is greater than the threshold value, the trigger event is detected.
[0072] After detecting a trigger event, the server can switch from obtaining the video stream collected by the target conference terminal (collected by the microphone and camera) to obtaining the data stream containing at least audio data (collected by the microphone) collected by the target conference terminal.
[0073] In some embodiments of the present invention, before obtaining the data stream containing at least audio data collected by the target conference terminal, it further includes:
[0074] Sending a first message to the target conference terminal; wherein, the first message is used to control the target conference terminal to stop sending the video stream and send the data stream containing at least audio data.
[0075] After detecting a trigger event, the server can send a first message to the target conference terminal. After receiving the first message, the target conference terminal can stop sending the video stream (it can also stop collecting the video stream), and then can send the data stream containing at least audio data (the data stream containing at least audio data does not contain video data) collected to the server.
[0076] In some examples, such as Figure 2a , it can be detected by the conference server whether a trigger event occurs. When a trigger event is detected, the conference server sends a first message to the target conference terminal, establishes a communication connection between the conference server and the video post-processing server, sends the data stream containing at least audio data sent by the target conference terminal to the video post-processing server, and the video post-processing server performs related operations of video synthesis according to the data stream containing at least audio data.
[0077] In some examples, a scheduling platform can be set up. The scheduling platform can be used to control the number of cameras and microphones allowed for users to turn on and off, and can also control whether to allow users to turn on the AI function. In some examples, affected by the computing power of the background GPU cluster, it may not be allowed for each user to randomly turn on and off the AI enhancement and synthesis functions.
[0078] In some examples, the server can be initialized first. Then, the scheduling platform, the conference server, and the video post-processing server can establish a TCP Socket connection. The participating terminals can perform user login. After the user login is successful, the conference terminal in charge of hosting can start a meeting through the scheduling platform. The scheduling platform can then control the server to open the meeting room and feedback the meeting information to the conference terminal in charge of hosting. Other participating terminals can join the meeting room through the meeting information, and the scheduling platform can control the participating terminals to turn on the camera and microphone, and can also control the opening of the video synthesis function.
[0079] Step 103: Slice the audio data in the data stream to obtain multiple audio segments.
[0080] For the data stream collected during a video conference, the audio data contained in the data stream can be sliced, that is, an entire audio data is sliced into multiple audio segments, which facilitates the analysis of each audio segment, so that the audio segment corresponds to the corresponding portrait expression.
[0081] In some embodiments of the present invention, the slicing process of the audio data in the data stream to obtain multiple audio segments includes:
[0082] Determine the audio change feature information of the audio data in the data stream; according to the audio change feature information, slice the audio data in the data stream to obtain multiple audio segments.
[0083] In practical applications, the audio change feature information can be obtained through the analysis of the audio data, and then the audio data can be sliced into multiple audio segments according to the audio change feature information.
[0084] In some examples, the audio change feature information can be the pause duration and sound intensity of the sound, and slicing can be performed according to the pause duration and sound intensity of the sound. For example, if the pause duration of the sound exceeds 3s, slicing is performed. Another example is that the audio with the sound intensity within a preset range is sliced into one audio segment.
[0085] Step 104, according to the multiple audio segments and the reference portrait change feature information, perform video stream synthesis to obtain a video stream for presenting a digital portrait, and send the video stream to other conference terminals.
[0086] Among them, a digital human is a humanoid image made through computer technology or a product made through computer software. They have the appearance or behavior pattern of a human, but they are not a video of a certain person in the real world. They can run and exist independently. The ontology of a digital human exists in a computing device (such as a computer, mobile phone, VR headset, etc.) and is presented through a display device so that humans can see it with their eyes or can interact through voice. They have an independent personality setting, a specific identity, an independent name, an independent character image, and an independent knowledge base to answer specific questions.
[0087] After obtaining a data stream containing at least audio data, that is, when the target conference terminal no longer sends a video stream to the server, the server can use the data stream containing at least audio data for video stream synthesis. Specifically, a video stream for presenting a digital portrait can be synthesized by combining multiple audio segments in the data stream and the reference portrait change feature information, and the synthesized video stream can be sent to other conference terminals to ensure the high definition of the video stream in the video conference and reduce the occupancy of bandwidth.
[0088] In some examples, the video stream synthesis process is a super-resolution reconstruction process. Super-resolution reconstruction is to improve the resolution of the original image through hardware or software methods. The process of obtaining a high-resolution image from a series of low-resolution images is super-resolution reconstruction. Specifically, when there is a lag in a video conference or when the clarity of the video stream captured by a conference terminal is low, it is switched to a high-definition video stream to improve the quality of the played video and ensure the normal progress of the video conference.
[0089] In some examples, the target conference terminal can encapsulate the data stream containing at least audio data obtained by collection in a signaling message, send it to the conference server through the signaling message, the conference server forwards it to the video post-processing server, the video post-processing server calls the streaming media library for decoding processing, sends the decoded data stream to the AI module, performs AI synthesis processing in the AI module, and then can send the synthesized video stream to other conference terminals to display the synthesized video stream on other conference terminals.
[0090] In some examples, the target conference terminal needs to be authorized by the scheduling platform before sending the data stream to the server. This platform can be set to automatically trigger the authorization process when certain conditions are met, or it can be defaulted that the scheduling platform has been authorized. For example, when the conference terminal sends a request authorization signaling message to the conference server, the scheduling platform needs to confirm the authorization. When the conference server sends an authorization notification signaling message to the conference terminal, the scheduling platform can default to authorization.
[0091] In some embodiments of the present invention, performing video stream synthesis according to the multiple audio segments and the reference portrait change feature information to obtain a video stream for presenting a digital portrait includes:
[0092] Generating target portrait change feature information corresponding to the multiple audio segments according to the multiple audio segments and the reference portrait change feature information; performing video stream synthesis according to the multiple audio segments and the target portrait change feature information to obtain a video stream for presenting a digital portrait.
[0093] After the audio data is segmented into multiple audio segments, for each audio segment, the corresponding reference portrait change feature information is determined, and then the corresponding target portrait change feature information can be generated. Then, video stream synthesis is performed by combining the multiple audio segments and the target portrait change feature information.
[0094] In some embodiments of the present invention, performing video stream synthesis according to the multiple audio segments and the target portrait change feature information to obtain a video stream for presenting a digital portrait includes:
[0095] Obtain the target image data associated with the target conference terminal; perform video stream synthesis based on the multiple audio segments, the target portrait change feature information, and the target image data to obtain a video stream for presenting a digital portrait.
[0096] In practical applications, after joining a video conference, the target conference terminal can upload the target image data to the server. The server can find the target image data synchronized by the user when creating the conference from the uploaded images (if no image data is synchronized during creation, an image data can be randomly matched from the in-memory photos in the user management module). For example, the image data can include portrait image data and conference scene image data, and the server can receive a data stream containing audio data, decode it by the media library and convert it into text information. Then, the target image data, the text information, and the target portrait change feature information can be input into a super-resolution reconstruction model to generate a video stream for presenting a digital portrait, that is, output a video stream of the portrait speaking.
[0097] Among them, the super-resolution reconstruction model is a semi-supervised model (a model trained by inputting portrait data and semantic information and intervened manually, which is more accurate than a fully supervised model and a supervised model). The video stream of the portrait speaking generated will be synthesized with the audio stream to generate an audio-video stream of a high-definition digital portrait, and at the same time, encoding is performed.
[0098] In some embodiments of the present invention, the data stream collected by the target conference terminal and at least containing audio data may further contain video feature attribute information generated according to the video stream collected by the target conference terminal. The video feature attribute information can be generated by the target conference terminal according to its own collected video stream and is sent to the server together with the video data in the data stream.
[0099] In some examples, the video feature attribute information may include real-time facial micro-expression information for motion compensation and reconstruction, important area texture and light and shadow information, hand motion capture data information, etc. In some examples, audio feature attribute information can also be generated according to the audio data, which can further enable better integration of audio and video, including one-dimensional temporal information for multi-modal alignment and video generation, audio compression / restoration inference compensation information, and keyword information for predicting and perceiving key speech time points, etc.
[0100] In some embodiments of the present invention, when the target conference terminal is in the first shooting mode, the data stream collected by the target conference terminal and uploaded to the server may include audio data, that is, only audio data is sent (video collection is turned off, and only audio collection is turned on). For example, the real speaker's image is sitting and only the part above the shoulders is exposed, and there are basically no body movements, that is, only the part above the user's shoulders is shot (half-body shooting mode). In this shooting mode, the synthesized digital image also only has the upper body part, so only audio data needs to be sent.
[0101] In some embodiments of the present invention, when the target conference terminal is in the second shooting mode, the data stream collected by the target conference terminal and uploaded to the server may include audio data and video feature attribute information (audio collection and video collection are turned on, but the collected video is only used to generate video feature attribute information, and the video stream is not uploaded). For example, the real speaker's image is standing and has more body movements, that is, the user's body movements need to be shot (full-body shooting mode). In this shooting mode, the synthesized digital image can be the full body of the person, that is, audio data and video feature attribute information need to be sent.
[0102] In some embodiments of the present invention, before the meeting starts, the conference terminal responsible for hosting can manually input: half-body or full-body shooting, or it can be determined by the conference server by obtaining the video conference screen as half-body or full-body shooting.
[0103] In some embodiments of the present invention, the video stream synthesis is performed according to the multiple audio segments and the reference portrait change feature information to obtain a video stream for presenting a digital portrait, including:
[0104] Obtain the target image data associated with the target conference terminal; perform video stream synthesis according to the multiple audio segments, the target portrait change feature information, the video feature attribute information, and the target image data to obtain a video stream for presenting a digital portrait.
[0105] When the data stream sent by the target conference terminal includes audio data and video feature attribute information, the image data synchronized when the user created the meeting can be searched (if no image data is synchronized during creation, an image data is randomly matched from the memory photos of the user management module). For example, the image data may include portrait image data and conference scene image data, and the data stream containing audio data received can be decoded by the media library and converted into text information. Then, the target image data, text information, video feature attribute information, and target portrait change feature information are input into the super-resolution reconstruction model to generate a video stream for presenting a digital portrait, that is, an output video stream of the person speaking.
[0106] In some embodiments of the present invention, before sending the video stream to other conference terminals, it further includes: creating a simulated user for the target conference terminal and sending a second message to the other conference terminals; wherein, the second message is used to control the other conference terminals to modify the subscription to the original user of the target conference terminal to a subscription to the simulated user.
[0107] In some embodiments of the present invention, the sending of the video stream to other conference terminals includes:
[0108] Sending the video stream to the other conference terminals through the simulated user.
[0109] In the case where no trigger event is detected, the server sends the video stream collected by the target conference terminal to other conference terminals, and the other conference terminals subscribe to the data of the corresponding user of the target conference terminal. After detecting the trigger event, since the target conference terminal no longer sends the video stream, a simulated user for the target conference terminal can be created and a second message (broadcast) can be sent to other conference terminals. The second message may include information about the simulated user. After receiving the second message, the other conference terminals can modify the subscription to the original user of the target conference terminal to a subscription to the simulated user.
[0110] After synthesizing the video stream according to the data stream sent by the target conference terminal, the server can publish the video stream in the name of the simulated user, and other conference terminals subscribing to the simulated user can then receive the video stream sent by the simulated user through the server.
[0111] In some embodiments of the present invention, the data stream containing at least audio data is the data stream corresponding to the video key area in the video stream sent by the target conference terminal, and the synthesized video stream is used to replace the video key area.
[0112] In a video conference, some conference terminals may use non-high-definition shooting, or the video stream may be difficult to meet the requirements due to large fluctuations in bandwidth during upload, such as low video clarity. Then, when a trigger event is detected, the video key area in the video stream can be determined for optimization.
[0113] In some examples, the video key area may be a video frame in the video stream with a clarity less than a threshold (such as a resolution less than 1024), that is, a low-definition area. For example, the video key area may include the portrait area of the speaker and the environmental area around the portrait area of the speaker (i.e., the venue environment).
[0114] In practical applications, the server can monitor the video stream sent by the target conference terminal, such as detecting which video frames have a clarity less than the threshold, and then the video key area can be determined from the video stream. In some examples, the conference server can forward the video stream to the video post-processing server, and then the video post-processing server monitors the video key area in the video stream.
[0115] In some examples, the triggering event for the target conference terminal includes: the video key area in the video stream sent by the target conference terminal has a clarity less than the threshold.
[0116] In some embodiments, the server can monitor the video stream sent by the target conference terminal in real time, such as detecting which video frames have a clarity less than the threshold. When it is detected that the video stream has a video key area with a clarity less than the threshold, the triggering event is detected. In some examples, the conference server can forward the video stream to the video post-processing server, and then the video post-processing server monitors the video key area in the video stream, and then can notify the conference server to perform subsequent processing.
[0117] After determining the video key area, video stream synthesis can be performed on the video key area, and the synthesized video stream is used to replace the video key area in the original video stream, and the new video stream is sent to other conference terminals, so as to ensure that the video stream can meet the requirements.
[0118] In some examples, the target image data may include venue image data and portrait image data of the participants corresponding to the target conference terminal. The portrait image data may be a real portrait taken before the meeting starts. Then the video stream for replacing the video key area may include the following replacement content:
[0119] The real portrait captured in real time in the original video stream → replaced with → the real portrait taken before the meeting starts → the video stream for replacing the video key area
[0120] The meeting scene captured in real time in the original video stream → replaced with → the meeting scene taken before the meeting starts → the video stream for replacing the video key area
[0121] In an embodiment of the present invention, by obtaining the reference portrait change feature information associated with the target conference terminal, when a trigger event for the target conference terminal is detected, a data stream collected by the target conference terminal and containing at least audio data is obtained. The audio data in the data stream is sliced to obtain a plurality of audio segments. According to the plurality of audio segments and the reference portrait change feature information, a video stream synthesis is performed to obtain a video stream for presenting a digital portrait, and the video stream is sent to other conference terminals, realizing video stream synthesis according to the portrait change features, and further being able to solve the problem of audio-visual asynchrony when synthesizing a video stream, and improving the quality of the video stream.
[0122] Moreover, by synthesizing a video stream according to the audio data uploaded by the conference terminal, the requirements for bandwidth and other aspects when the conference terminal uploads data are reduced, and data can be uploaded with a lower bandwidth and high-definition video conference images can be displayed on other terminals, thereby being able to make up for the situation where the video stream is difficult to meet the requirements due to reasons such as non-high-definition shooting and large bandwidth fluctuations.
[0123] Referring to Figure 3 , a flowchart of steps of another method for processing a video conference based on semantic information provided by some embodiments of the present invention is shown, and specifically may include the following steps:
[0124] Step 301, obtain the audio material and the corresponding portrait material associated with the target conference terminal.
[0125] Step 302, generate the reference portrait change feature information associated with the target conference terminal according to the audio material and the corresponding portrait material.
[0126] Step 303, when a trigger event for the target conference terminal is detected, obtain a data stream collected by the target conference terminal and containing at least audio data.
[0127] Step 304, slice the audio data in the data stream to obtain a plurality of audio segments.
[0128] Step 305, perform video stream synthesis according to the plurality of audio segments and the reference portrait change feature information to obtain a video stream for presenting a digital portrait, and send the video stream to other conference terminals.
[0129] Referring to Figure 4 , a flowchart of steps of another method for processing a video conference based on semantic information provided by some embodiments of the present invention is shown, and specifically may include the following steps:
[0130] Step 401, obtain the reference portrait change feature information associated with the target conference terminal.
[0131] Step 402, when a trigger event for the target conference terminal is detected, obtain a data stream collected by the target conference terminal and containing at least audio data.
[0132] Step 403, slice the audio data in the data stream to obtain a plurality of audio segments;
[0133] Step 404, generate target portrait change feature information corresponding to the plurality of audio segments according to the plurality of audio segments and the reference portrait change feature information.
[0134] Step 405, perform video stream synthesis according to the plurality of audio segments and the target portrait change feature information to obtain a video stream for presenting a digital portrait, and send the video stream to other conference terminals.
[0135] Refer to Figure 5 , which shows a flowchart of steps of another method for processing a video conference based on semantic information provided by some embodiments of the present invention, and may specifically include the following steps:
[0136] Step 501, obtain the reference portrait change feature information associated with the target conference terminal.
[0137] Step 502, when a trigger event for the target conference terminal is detected, obtain a data stream collected by the target conference terminal and containing at least audio data.
[0138] Step 503, slice the audio data in the data stream to obtain a plurality of audio segments;
[0139] Step 504, generate target portrait change feature information corresponding to the plurality of audio segments according to the plurality of audio segments and the reference portrait change feature information.
[0140] Step 505, obtain the target image data associated with the target conference terminal.
[0141] Step 506, perform video stream synthesis according to the plurality of audio segments, the target portrait change feature information, and the target image data to obtain a video stream for presenting a digital portrait, and send the video stream to other conference terminals.
[0142] Refer to Figure 6 , which shows a flowchart of steps of another method for processing a video conference based on semantic information provided by some embodiments of the present invention, and may specifically include the following steps:
[0143] Step 601, obtain the reference portrait change feature information associated with the target conference terminal.
[0144] Step 602, when a trigger event for the target conference terminal is detected, obtain a data stream collected by the target conference terminal and containing at least audio data; wherein, the data stream further contains video feature attribute information generated according to the video stream collected by the target conference terminal.
[0145] Step 603, slice the audio data in the data stream to obtain a plurality of audio segments;
[0146] Step 604, generate target portrait change feature information corresponding to the plurality of audio segments according to the plurality of audio segments and the reference portrait change feature information.
[0147] Step 605, obtain target image data associated with the target conference terminal;
[0148] Step 606, perform video stream synthesis according to the plurality of audio segments, the target portrait change feature information, the video feature attribute information, and the target image data to obtain a video stream for presenting a digital portrait, and send the video stream to other conference terminals.
[0149] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.
[0150] Refer to Figure 7 , which shows a schematic structural diagram of a device for processing video conferences based on semantic information provided by some embodiments of the present invention, and specifically may include the following modules:
[0151] A reference portrait change feature information acquisition module 701, configured to acquire reference portrait change feature information associated with the target conference terminal;
[0152] A data stream acquisition module 702, configured to obtain a data stream collected by the target conference terminal and containing at least audio data when a trigger event for the target conference terminal is detected;
[0153] An audio slicing module 703, configured to slice the audio data in the data stream to obtain a plurality of audio segments;
[0154] The video stream synthesis module 704 is configured to perform video stream synthesis based on the multiple audio segments and the reference portrait change feature information, obtain a video stream for presenting a digital portrait, and send the video stream to other conference terminals.
[0155] In some embodiments of the present invention, obtaining the reference portrait change feature information associated with the target conference terminal includes:
[0156] Obtaining the audio material and the corresponding portrait material associated with the target conference terminal;
[0157] Generating the reference portrait change feature information associated with the target conference terminal according to the audio material and the corresponding portrait material.
[0158] In some embodiments of the present invention, performing video stream synthesis based on the multiple audio segments and the reference portrait change feature information to obtain a video stream for presenting a digital portrait includes:
[0159] Generating the target portrait change feature information corresponding to the multiple audio segments according to the multiple audio segments and the reference portrait change feature information;
[0160] Performing video stream synthesis according to the multiple audio segments and the target portrait change feature information to obtain a video stream for presenting a digital portrait.
[0161] In some embodiments of the present invention, performing video stream synthesis based on the multiple audio segments and the target portrait change feature information to obtain a video stream for presenting a digital portrait includes:
[0162] Obtaining the target image data associated with the target conference terminal;
[0163] Performing video stream synthesis according to the multiple audio segments, the target portrait change feature information, and the target image data to obtain a video stream for presenting a digital portrait.
[0164] In some embodiments of the present invention, the data stream further includes video feature attribute information generated according to the video stream collected by the target conference terminal. Performing video stream synthesis based on the multiple audio segments and the reference portrait change feature information to obtain a video stream for presenting a digital portrait includes:
[0165] Obtaining the target image data associated with the target conference terminal;
[0166] Performing video stream synthesis according to the multiple audio segments, the target portrait change feature information, the video feature attribute information, and the target image data to obtain a video stream for presenting a digital portrait.
[0167] In some embodiments of the present invention, the slicing process of the audio data in the data stream to obtain multiple audio segments includes:
[0168] Determine the audio change feature information of the audio data in the data stream;
[0169] According to the audio change feature information, perform a slicing process on the audio data in the data stream to obtain multiple audio segments.
[0170] In some embodiments of the present invention, the data stream containing at least audio data is the data stream corresponding to the video key area in the video stream sent by the target conference terminal, and the synthesized video stream is used to replace the video key area.
[0171] In the embodiments of the present invention, by obtaining the reference portrait change feature information associated with the target conference terminal, in the case of detecting a trigger event for the target conference terminal, obtaining a data stream containing at least audio data collected by the target conference terminal, performing a slicing process on the audio data in the data stream to obtain multiple audio segments, and performing video stream synthesis based on the multiple audio segments and the reference portrait change feature information to obtain a video stream for presenting a digital portrait, and sending the video stream to other conference terminals, it realizes video stream synthesis according to the portrait change feature, and thus can solve the problem of audio-visual asynchrony when synthesizing the video stream, and improves the quality of the video stream.
[0172] Moreover, by synthesizing the video stream according to the audio data uploaded by the conference terminal, the requirements for bandwidth and other aspects of the data uploaded by the conference terminal are reduced, and data can be uploaded with a lower bandwidth and high-definition video conference images can be displayed on other terminals, thus being able to make up for the situation where the video stream is difficult to meet the requirements due to reasons such as non-high-definition shooting and large bandwidth fluctuations.
[0173] Some embodiments of the present invention also provide an electronic device, which may include a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the method for processing video conferences based on semantic information as described above is implemented.
[0174] Some embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, the method for processing video conferences based on semantic information as described above is implemented.
[0175] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, refer to the partial description of the method embodiments.
[0176] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0177] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0178] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0179] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0180] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0181] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide for implementing the process Figure 1 in one process or multiple processes and / or blocks Figure 1 steps for the functions specified in one block or multiple blocks.
[0182] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
[0183] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article or terminal device comprising the above element.
[0184] The above provides a detailed introduction to the method, apparatus, device and medium for processing video conferencing based on semantic information. In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for processing video conferences based on semantic information, characterized in that, The method includes: Obtaining reference portrait change feature information associated with the target conference terminal; When a trigger event for the target conference terminal is detected, obtaining a data stream collected by the target conference terminal that at least includes audio data; Performing slicing processing on the audio data in the data stream to obtain a plurality of audio segments; Performing video stream synthesis based on the plurality of audio segments and the reference portrait change feature information to obtain a video stream for presenting a digital portrait, and sending the video stream to other conference terminals.
2. The method according to claim 1, characterized in that The obtaining of the reference portrait change feature information associated with the target conference terminal includes: Obtaining audio materials and corresponding portrait materials associated with the target conference terminal; Generating the reference portrait change feature information associated with the target conference terminal according to the audio materials and the corresponding portrait materials.
3. The method according to claim 1 or 2, characterized in that, The performing of video stream synthesis based on the plurality of audio segments and the reference portrait change feature information to obtain a video stream for presenting a digital portrait includes: Generating target portrait change feature information corresponding to the plurality of audio segments according to the plurality of audio segments and the reference portrait change feature information; Performing video stream synthesis based on the plurality of audio segments and the target portrait change feature information to obtain a video stream for presenting a digital portrait.
4. The method according to claim 3, characterized in that The performing of video stream synthesis based on the plurality of audio segments and the target portrait change feature information to obtain a video stream for presenting a digital portrait includes: Obtaining target image data associated with the target conference terminal; Performing video stream synthesis based on the plurality of audio segments, the target portrait change feature information, and the target image data to obtain a video stream for presenting a digital portrait.
5. The method according to claim 3, characterized in that, The data stream further includes video feature attribute information generated according to the video stream collected by the target conference terminal. The performing of video stream synthesis based on the plurality of audio segments and the reference portrait change feature information to obtain a video stream for presenting a digital portrait includes: Obtaining target image data associated with the target conference terminal; Performing video stream synthesis based on the plurality of audio segments, the target portrait change feature information, the video feature attribute information, and the target image data to obtain a video stream for presenting a digital portrait.
6. The method according to claim 1, wherein The performing of slicing processing on the audio data in the data stream to obtain a plurality of audio segments includes: Determining audio change feature information of the audio data in the data stream; Performing slicing processing on the audio data in the data stream according to the audio change feature information to obtain a plurality of audio segments.
7. The method according to claim 1, wherein The data stream that at least includes audio data is the data stream corresponding to the video key area in the video stream sent by the target conference terminal, and the synthesized video stream is used to replace the video key area.
8. A device for processing video conferences based on semantic information, characterized in that, The device includes: A reference portrait change feature information acquisition module, configured to obtain reference portrait change feature information associated with the target conference terminal; A data stream acquisition module, configured to obtain a data stream collected by the target conference terminal that at least includes audio data when a trigger event for the target conference terminal is detected; An audio slicing module, configured to slice the audio data in the data stream to obtain a plurality of audio segments; A video stream synthesis module, configured to perform video stream synthesis according to the plurality of audio segments and the reference portrait change feature information to obtain a video stream for presenting a digital portrait, and send the video stream to other conference terminals.
9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, it implements the method for processing a video conference based on semantic information according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the method for processing a video conference based on semantic information according to any one of claims 1 to 7.