Video conference data transmission method and video conference system
By using a distributed video conferencing system and transmitting background data and historical facial expression data index information via the access network, the problem of video stuttering was solved, and the stability and flexibility of high-quality video transmission were achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-22
- Publication Date
- 2026-03-27
AI Technical Summary
The problem of video stuttering in existing video conferencing systems is difficult to solve effectively. Existing solutions such as increasing bandwidth or compressing audio and video have not significantly improved the situation. In particular, video stuttering still occurs frequently when backbone network bandwidth is uncontrollable and terminal computing power is limited.
The system adopts a distributed video conferencing architecture, connecting the conferencing terminals and edge servers through the access network. The target background data is transmitted only in the first stage of the video conferencing, and real-time audio and historical facial expression data index information are transmitted in the second stage. Non-speaking servers combine and generate the target video data and send it, reducing the data volume and bandwidth requirements of the backbone network.
Without increasing backbone network bandwidth or compressing audio and video, it significantly reduces the probability of video stuttering, improves video transmission quality and stability, and adapts to data transmission needs under different user model conditions.
Smart Images

Figure CN115941881B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, and in particular to a video conference data transmission method and a video conference system. BACKGROUND
[0002] A video conference system refers to a system in which individuals or groups in two or more different places communicate with each other in real time and interactively by transmitting voice, images, and file data through transmission lines and multimedia devices. In current video conference systems, each conference terminal directly interacts through a wide area network. During the interaction, the video data requires high bandwidth and high terminal codec. However, during the interaction, the wide area network bandwidth often fluctuates, and the audio and video codec fluctuates, which causes the video to freeze or drop frames during the conference.
[0003] To alleviate this phenomenon, the current solutions mainly include increasing bandwidth or compressing audio and video. For the solution of increasing bandwidth, there are two defects: first, when a user increases the bandwidth, only the egress bandwidth of the wide area network can be increased, but the backbone network bandwidth cannot be increased, and the backbone network bandwidth is not determined by the user; second, even if the backbone network bandwidth is expanded, situations such as burst traffic causing congestion of a link of the backbone network still occur, which is a temporary solution. For the solution of compressing audio and video, there are also two defects: first, the compression ratio of the video is not high, and the video still requires a large bandwidth compared with the audio; second, the higher the compression ratio, the higher the requirement for the computing power of the terminal codec, and when the terminal computing fluctuates, the video still freezes.
[0004] In summary, for the video freezing problem in the video conference system, the current solutions of increasing bandwidth or compressing audio and video cannot achieve good results. SUMMARY
[0005] Embodiments of the present application provide a video conference data transmission method, device, system, electronic device, and storage medium to alleviate the technical problem of video freezing in the current video conference data transmission process.
[0006] To solve the above technical problem, embodiments of the present application provide the following technical solutions:
[0007] The present application provides a video conference data transmission method applicable to a video conference system. The video conference system includes at least two conference terminals and at least two conference servers. Each conference server is connected through a backbone network, and each conference server is connected to a corresponding conference terminal through an access network. The conference terminal includes a speaking conference terminal and a non-speaking conference terminal, and the conference server includes a speaking conference server and a non-speaking conference server. The video conference data transmission method is applied to the speaking conference server, and the method includes:
[0008] In the first stage of the video conference, target background data of the speaking user is received from the speaking conference terminal, and the target background data is sent to the non-speaking conference server;
[0009] In the second stage of the video conference, real-time audio data and real-time expression data of the speaking user are received from the speaking conference terminal, target historical expression data index information of the speaking user is obtained according to the real-time audio data, the real-time expression data and a user model of the speaking user;
[0010] The real-time audio data and the target historical expression data index information are sent to the non-speaking conference server, so that the non-speaking conference server acquires target historical expression data from a historical expression data set of the speaking user according to the target historical expression data index information, combines the target historical expression data and the target background data to obtain target video data, and sends the target video data and the real-time audio data to a corresponding non-speaking conference terminal.
[0011] Meanwhile, the embodiment of the application further provides a video conference data transmission method, which is suitable for a video conference system, the video conference system comprising at least two conference terminals and at least two conference servers, each conference server being connected through a backbone network, each conference server being connected with a corresponding conference terminal through an access network, the conference terminals comprising speaking conference terminals and non-speaking conference terminals, the conference servers comprising speaking conference servers and non-speaking conference servers, the video conference data transmission method being applied to the non-speaking conference server, and the method comprising:
[0012] In the first stage of the video conference, target background data of the speaking user is received from the speaking conference terminal, and the target background data is sent to the non-speaking conference server;
[0013] In the second stage of the video conference, real-time audio data and target historical expression data index information of the speaking user are received from the speaking conference server, the target historical expression data index information being obtained by the speaking conference server according to the real-time audio data, real-time expression data of the speaking user and a user model, the real-time audio data and the real-time expression data being sent to the speaking conference server by the speaking conference terminal;
[0014] Target historical expression data is acquired from a historical expression data set of the speaking user according to the target historical expression data index information, the target historical expression data and the target background data are combined to obtain target video data, and the target video data and the real-time audio data are sent to a corresponding non-speaking conference terminal.
[0015] The application also provides a video conference data transmission device, which is suitable for a video conference system, the video conference system comprising at least two conference terminals and at least two conference servers, each conference server being connected through a backbone network, each conference server being connected with a corresponding conference terminal through an access network, the conference terminals comprising speaking conference terminals and non-speaking conference terminals, the conference servers comprising speaking conference servers and non-speaking conference servers, the video conference data transmission device being applied to the speaking conference server, and the device comprising:
[0016] a first sending module, configured to receive target background data of a speaking user sent by the speaking conference terminal in a first stage of a video conference, and send the target background data to the non-speaking conference server;
[0017] a first receiving module, configured to receive real-time audio data and real-time expression data of the speaking user sent by the speaking conference terminal in a second stage of the video conference, and obtain target historical expression data index information of the speaking user according to the real-time audio data, the real-time expression data and a user model of the speaking user;
[0018] a second sending module, configured to send the real-time audio data and the target historical expression data index information to the non-speaking conference server, so that the non-speaking conference server acquires target historical expression data from a historical expression data set of the speaking user according to the target historical expression data index information, combines the target historical expression data and the target background data to obtain target video data, and sends the target video data and the real-time audio data to a corresponding non-speaking conference terminal.
[0019] The application also provides a video conference data transmission device, which is suitable for a video conference system, the video conference system comprising at least two conference terminals and at least two conference servers, each conference server being connected through a backbone network, each conference server being connected with a corresponding conference terminal through an access network, the conference terminals comprising speaking conference terminals and non-speaking conference terminals, the conference servers comprising speaking conference servers and non-speaking conference servers, the video conference data transmission device being applied to the non-speaking conference server, and the device comprising:
[0020] a second receiving module, configured to receive target background data of a speaking user sent by the speaking conference server in a first stage of a video conference, the target background data being sent by the speaking conference terminal to the speaking conference server;
[0021] a third receiving module, configured to receive, in the second phase of the video conference, real-time audio data of the speaking user and target historical expression data index information sent by the speaking conference server, the target historical expression data index information being obtained by the speaking conference server according to the real-time audio data, real-time expression data of the speaking user and a user model, the real-time audio data and the real-time expression data being sent by the speaking conference terminal to the speaking conference server;
[0022] a third sending module, configured to obtain target historical expression data from a historical expression data set of the speaking user according to the target historical expression data index information, combine the target historical expression data and the target background data to obtain target video data, and send the target video data and the real-time audio data to a corresponding non-speaking conference terminal.
[0023] The application further provides a video conference system, comprising at least two conference terminals and at least two conference servers, each conference server being connected through a backbone network, each conference server being connected with a corresponding conference terminal through an access network, the conference terminals comprising speaking conference terminals and non-speaking conference terminals, and the conference servers comprising speaking conference servers and non-speaking conference servers, wherein:
[0024] the speaking conference server is configured to, in a first phase of a video conference, receive target background data of a speaking user sent by the speaking conference terminal, and send the target background data to the non-speaking conference server;
[0025] the speaking conference server is configured to, in a second phase of the video conference, receive real-time audio data and real-time expression data of the speaking user sent by the speaking conference terminal, obtain target historical expression data index information of the speaking user according to the real-time audio data, the real-time expression data and a user model of the speaking user, and send the real-time audio data and the target historical expression data index information to the non-speaking conference server;
[0026] the non-speaking conference server is configured to, in the second phase of the video conference, obtain target historical expression data from a historical expression data set of the speaking user according to the target historical expression data index information, combine the target historical expression data and the target background data to obtain target video data, and send the target video data and the real-time audio data to a corresponding non-speaking conference terminal.
[0027] The application further provides an electronic device comprising a memory and a processor, wherein the memory stores an application program, and the processor is configured to run the application program in the memory to execute the steps in the video conference data transmission method according to any one of the preceding embodiments.
[0028] The embodiment of the present application provides a computer readable storage medium, which stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the video conference data transmission method.
[0029] Beneficial effects: the present application provides a video conference data transmission method, device, system, electronic equipment and storage medium, the video conference data transmission method is applied to a video conference system, a distributed architecture is adopted in the video conference system, at least two conference terminals and at least two conference servers are arranged, each conference server is connected through a backbone network, and each conference server is connected with a corresponding conference terminal through an access network. In the first stage of the video conference, the speaking conference terminal collects target background data of the speaking user and sends the target background data to the speaking conference server, and the speaking conference server synchronizes the received target background data to each non-speaking conference server. In the second stage of the video conference, the speaking conference terminal collects real-time expression data and real-time audio data of the speaking user and sends the real-time expression data and the real-time audio data to the speaking conference server, the speaking conference server obtains target historical expression data index information according to the real-time expression data, the real-time audio data and the user model of the speaking user, and then sends the real-time audio data and the target historical expression data index information to the non-speaking conference server, the non-speaking conference server finds corresponding target historical expression data from a historical expression data set of the speaking user according to the index information, and combines the target historical expression data with the target background data received in the first stage to form complete target video data, and finally sends the target video data and the real-time audio data to the corresponding non-speaking conference terminal for display. That is, the present application only transmits complete audio and video between the conference terminal and the conference server connected through the access network, and only transmits target background data in the first stage and real-time audio data and target historical expression data index information in the second stage between the conference servers connected through the backbone network, and then calculates and combines to obtain complete audio and video. Since the bandwidth of the access network is controllable and is exclusively enjoyed by the user, the video transmission between the conference terminal and the conference server can be realized without bandwidth jitter by adjusting the bandwidth of the access network, and although the bandwidth of the backbone network is uncontrollable, the amount of data transmitted through the backbone network in the same period is reduced, which reduces the requirement for the bandwidth of the backbone network, so that the scheme in the present application can significantly reduce the probability of video lag without increasing the bandwidth of the backbone network or compressing the audio and video. BRIEF DESCRIPTION OF DRAWINGS
[0030] The technical scheme and other beneficial effects of the present application will be apparent from the following detailed description of the specific embodiments of the present application combined with the accompanying drawings.
[0031] Figure 1 FIG. 1 is a structural schematic diagram of a video conference system in the prior art.
[0032] Figure 2 The working schematic diagram of the video conference system provided by the embodiment of the present application in the initial use stage.
[0033] Figure 3 The working schematic diagram of the video conference system provided by the embodiment of the present application in the first stage of the video conference.
[0034] Figure 4 The working schematic diagram of the video conference system provided by the embodiment of the present application in the second stage of the video conference.
[0035] Figure 5 The first flow schematic diagram of the video conference data transmission method provided by the embodiment of the present application.
[0036] Figure 6 The second flow schematic diagram of the video conference data transmission method provided by the embodiment of the present application.
[0037] Figure 7 The third flow schematic diagram of the video conference data transmission method provided by the embodiment of the present application.
[0038] Figure 8 The fourth flow schematic diagram of the video conference data transmission method provided by the embodiment of the present application.
[0039] Figure 9 The first structural schematic diagram of the video conference data transmission device provided by the embodiment of the present application.
[0040] Figure 10 The second structural schematic diagram of the video conference data transmission device provided by the embodiment of the present application.
[0041] Figure 11 The structural schematic diagram of the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0042] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0043] The embodiments of the present application provide a video conference data transmission method, device, system, electronic device and storage medium, to alleviate the technical problem of video lag in the current video conference data transmission process.
[0044] As Figure 1As shown in the structural schematic diagram of a video conference system in the prior art, the video conference system includes a conference server and at least two conference terminals. The conference server is centrally deployed in an IDC (Internet Data Center) machine room. Each conference terminal accesses a backbone network of a wide area network through a BRAS (Broadband Remote Access Server) and transmits audio and video data through the backbone network.
[0045] During a video conference, one of the conference terminals is a speaking conference terminal, i.e., a conference terminal used by a speaker, and the others are non-speaking conference terminals, i.e., conference terminals used by other participants. For ease of description, the conference terminal 1 represents the speaking conference terminal and the conference terminal 2 represents the non-speaking conference terminal in the embodiments of the present application. During the video conference, the audio and video data of the speaking user need to be synchronized to all the non-speaking users. Therefore, the conference terminal 1 needs to collect the audio and video data of the speaker and send the data to the conference terminal 2. In the existing architecture, the conference terminal directly sends complete audio and video data to the conference server through the backbone network, and the conference server distributes the data to the conference terminal 2 through the backbone network and displays the data to the non-speaking user by the conference terminal 2.
[0046] From the above content, it can be known that the conference server in the video conference system in the prior art adopts a centralized deployment architecture. Therefore, complete audio and video data need to pass through the backbone network during transmission. Since the bandwidth of the backbone network is not determined by the user and the video data has a high demand for bandwidth, when the bandwidth of the backbone network is dithered or burst traffic occurs, the video data will have packet loss or transmission delay during transmission through the wide area network, which easily leads to video freezing or frame loss of the conference terminal 2. In addition, the prior art also adopts a scheme of compressing and transmitting the audio and video data and decoding the data after receiving. However, the compression ratio is not high, and the terminal has a high requirement for the coding and decoding capability. When the terminal computing is dithered, the video will still be frozen.
[0047] Based on the above defects, the present application provides a distributed video conference system which can significantly reduce the probability of video freezing without increasing the bandwidth of the backbone network or compressing the audio and video. The structural diagram of the video conference system provided by the embodiments of the present application is as shown in Figures 2 to 4As shown, the video conference system includes at least two conference terminals and at least two conference servers, each conference server is connected through a backbone network, each conference server is connected with a corresponding conference terminal through an access network, all conference terminals can include a speaking conference terminal and a non-speaking conference terminal, and all conference servers can include a speaking conference server and a non-speaking conference server. For ease of illustration, in this embodiment and the following embodiments, the conference terminal 1 represents the speaking conference terminal, the conference server 1 represents the speaking conference server, the conference terminal 2 represents the non-speaking conference terminal, and the conference server 2 represents the non-speaking conference server.
[0048] The video conference system adopts a distributed architecture, and there are two or more conference terminals and two or more conference servers. Each conference server is deployed in an edge access machine room, each conference terminal is connected with the conference server closest in physical distance through an access network, each conference server is connected to the backbone network through a BRAS, and the conference servers are connected through the backbone network. Since the access network can be controlled by the user, the bandwidth can be set according to the needs, so it can be considered that there is no technical defect of video lag caused by bandwidth jitter between the conference terminal and the conference server even if the video data is directly transmitted.
[0049] In the video conference, the identity of each conference terminal is determined according to the speaking state of each participant. When a participant is currently a speaking user, the conference terminal used by the participant is a speaking conference terminal, and other participants are non-speaking users, and the conference terminals used by the participants are non-speaking conference terminals. At the same time in the video conference, there can be only one speaking conference terminal, and there can be multiple non-speaking conference terminals. It should be noted that in the video conference, the speaking user is not necessarily a fixed participant. For example, in a video conference with multiple participants, each participant takes turns to speak, and the identity of each conference terminal also changes. That is, the speaking conference terminal and the non-speaking conference terminal referred to in the embodiments of the present application are not specific to one or a few fixed conference terminals in the video conference system, but are determined according to the real-time speaking situation of the actual conference scene. Any conference terminal in the video conference system can be a speaking conference terminal, and can also be a non-speaking conference terminal.
[0050] Similarly, when the identity of each conference terminal changes, the identity of the corresponding connected conference server also changes, when a conference terminal is a speaking conference terminal, the corresponding connected conference server is a speaking conference server, and other conference terminals are non-speaking conference terminals, and the corresponding connected conference server is a non-speaking conference server. That is, the speaking conference server and the non-speaking conference server referred to in the embodiments of the present application do not specifically refer to a certain or several fixed conference servers in the video conference system, but are determined according to the real-time speaking situation of the actual conference scene. Any conference server in the video conference system can be a speaking conference server, and can also be a non-speaking conference server.
[0051] It should be noted that, Figures 2 to 4 The system architecture diagram shown is only an example, and the servers and scenarios described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, as the system evolves and new business scenarios appear, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems. It should be noted that the order of description of the following embodiments is not a limitation on the preferred order of the embodiments.
[0052] In addition, it should be noted that the video conference system of the present application can be used as a separate system, or can be combined with an existing video conference system as a whole system. When it is used as a separate system, it can only include two or more distributed conference servers and corresponding conference terminals, at this time, all data transmission paths are speaking conference terminal-speaking conference server-non-speaking conference server-non-speaking conference terminal. When it is a whole, it can include both distributed conference servers in the edge machine room and central conference servers in the IDC machine room, and the data transmission between each conference terminal can be determined according to whether there is a distributed conference server nearby that can be accessed. If there is a distributed conference server nearby, the conference data transmission path is speaking conference terminal-speaking conference server-non-speaking conference server-non-speaking conference terminal, if not, the conference data transmission path is still speaking conference terminal-central server-non-speaking conference terminal. That is, the video conference system in the present application can be improved on the basis of the original video conference system, and the improvement process does not affect the work of the existing video conference system. After the improvement is completed, part or all of the business can be migrated according to the needs to realize the rapid switching of the business process.
[0053] The video conference data transmission method of the present application involves three time periods, namely the initial stage of use of the video conference system, the stage before each video conference (hereinafter referred to as the first stage of the video conference), and the stage during each video conference (hereinafter referred to as the second stage of the video conference). In the video conference system provided by the present application, the speaking conference server is configured to, in the first stage of the video conference, receive target background data of a speaking user sent by a speaking conference terminal, and send the target background data to the non-speaking conference server; the speaking conference server is configured to, in the second stage of the video conference, receive real-time audio data and real-time expression data of the speaking user sent by the speaking conference terminal, obtain target historical expression data index information of the speaking user according to the real-time audio data, the real-time expression data and a user model of the speaking user, and send the real-time audio data and the target historical expression data index information to the non-speaking conference server; and the non-speaking conference server is configured to, in the second stage of the video conference, obtain target historical expression data from a historical expression data set of the speaking user according to the target historical expression data index information, combine the target historical expression data and the target background data to obtain target video data, and send the target video data and the real-time audio data to a corresponding non-speaking conference terminal.
[0054] Since the video conference data transmission method of the present application involves the interaction between the speaking conference server and the non-speaking conference server in the video conference system, in the following embodiments of the present application, the video conference system and the video conference data transmission method based on the system will be described from the perspective of the speaking conference server and the non-speaking conference server respectively.
[0055] As shown in Figure 5 , it is the first flowchart of the video conference data transmission method provided by the embodiment of the present application, which corresponds to the initial stage of use of the video conference system, as shown in Figure 6 , it is the second flowchart of the video conference data transmission method provided by the embodiment of the present application, which corresponds to the first stage and the second stage of the video conference. Figure 5 and Figure 6 are described from the perspective of the speaking conference server.
[0056] Please refer to Figure 5 , in the initial stage of use of the video conference system, the method specifically comprises:
[0057] S11: obtaining historical expression data and historical audio data of at least one registered user.
[0058] The user can use the video conference system only after registration. In the initial stage of use of the video conference system, the speaking conference server and the non-speaking conference server can be used for registration of a new user. The embodiment is described with respect to the speaking conference server. Any speaking conference server can obtain historical expression data and historical audio data of at least one registered user. The historical expression data includes a plurality of historical expression frames. The historical expression frames are used to record facial expression information of the registered user. The historical audio data includes a plurality of historical audio frames. The historical audio frames are used to record voice information of the registered user.
[0059] S12: Obtain a user model and a historical expression data set of each registered user according to the historical expression data and the historical audio data of each registered user.
[0060] For each registered user, each historical expression frame of the registered user is stored to obtain the historical expression data set. When the number of historical expression frames in the historical expression data set is sufficient (for example, more than 100,000), a user model of the registered user is obtained according to the historical expression data and the historical audio data of the registered user. The user model is trained by using the historical expression frames and the corresponding historical audio frames of the registered user as training data. The historical audio frame corresponding to the historical expression frame can be a historical audio frame at the same time or an audio segment composed of a plurality of historical audio frames in the same time period. The specific selection can be made according to the training needs. After the training is completed, the combination of any expression frame and the corresponding audio frame of the registered user can be identified with high accuracy. Since the audio data can reflect the emotions of the user such as joy, anger, sorrow, and the like, the combination of the expression frame and the audio frame is used as the training data to establish the user model. It is equivalent to using only the expression frame. The identification accuracy of the trained user model is higher, and the identification effect is better.
[0061] In one embodiment, the step of obtaining a user model and a historical expression data set of each registered user according to historical expression data and historical audio data of each registered user includes: generating a historical expression data set of each registered user according to all historical expression frames of each registered user, and generating historical expression frame index information of the historical expression frames in each historical expression data set; obtaining an initial user model of each registered user, and taking the historical expression frames and the historical audio frames of each registered user as training input data and the historical expression frame index information as training output data, training each initial user model to obtain the user model of each registered user.
[0062] After obtaining the historical expression data set of a certain registered user, the corresponding historical expression frame index information can be generated for each historical expression frame in the historical expression data set. At the same time, the initial user model of the registered user is obtained, and the initial user model of the registered user is a neural network model, which can be ResNet or other types of neural network models. After training with training data of different registered users respectively, the user models of different registered users are obtained. Specifically, for the initial user model, the binary information of each historical expression frame and the corresponding historical audio frame of a certain registered user is combined into new picture information, and N historical expression frames can be combined to obtain N groups of training input data. At the same time, the index information of each historical expression frame of the registered user forms N groups of training output data respectively. The initial user model is trained based on the N groups of training input data and training output data, so that the model can output the historical expression frame index information corresponding to the real-time expression frame of the registered user after feature extraction and classification, and the similarity between the historical expression frame corresponding to the historical expression frame index information and the real-time expression frame meets the expectation. After training, the user model of the registered user is obtained. The above operation is performed once for all registered users, and the user models of the registered users can be obtained in the speech conference server.
[0063] S13: Synchronize the user models and historical expression data sets of the registered users to other conference servers.
[0064] After the user model of each registered user is established, it can be synchronized to other conference servers, so that other conference servers have the user model of the registered user. The model establishment and synchronization of each registered user is a dynamic process, that is, the speech conference server will continuously obtain the historical expression data and historical speech data of each registered user. When the amount of these data of a certain registered user reaches the expectation, the model training and synchronization of the registered user will be performed, and other registered users are always in the data collection stage until the condition for starting training is reached.
[0065] The above entire process can be referred to Figure 2 In this stage, the conference server 1 can obtain the user models and historical expression data sets of the registered users 1, and synchronize them to the conference server 2 through a wide area network.
[0066] In the early stage of use, the video conference system still adopts the architecture shown in Figure 1 , and the workflow is the same as Figure 1The historical expression data and the historical audio data can be collected when the user submits the registration information and during the subsequent video conference. When the number of historical expression frames in the historical expression data does not meet the expectation, the current data transmission process is still maintained and the data collection is continuously performed. When the historical expression frame data is large enough, the establishment and synchronization of the user model are automatically performed. This process is a self-learning process without awareness. After the model is established and synchronized, seamless switching from the current video conference data transmission method to the video conference data transmission method in the application can be realized.
[0067] As shown in FIG. 1, in the first phase and the second phase of the video conference, the method specifically includes: Figure 6
[0068] S21: In the first phase of the video conference, target background data of a speaking user sent by a speaking conference terminal is received, and the target background data is sent to a non-speaking conference server.
[0069] The first phase of the video conference is a phase in which each participant user has not started speaking after the video conference starts. In this phase, if it is clear in the video conference that only one or several users can speak in this conference, these users are speaking users, and other users are non-speaking users. If it is not clear in the video conference, since each participant user may speak subsequently, each participant user can be regarded as a speaking user in this phase, and other participant users are regarded as non-speaking users when the user is a speaking user. Each speaking conference terminal collects and calculates the target background data of the speaking user of each speaking conference terminal, that is, all video information of each participant user except the expression, and sends the target background data to the speaking conference server closest to itself through the access network. The speaking conference server synchronizes the target background data to other non-speaking conference servers, so that each participant conference server has the target background data of all participant users. The specific transmission process of the data in the above process can be referred to in the description of FIG. 2. Figure 3 .
[0070] In this step, the target background data is transmitted between the speaking conference terminal and the speaking conference server through the access network. Since the access network can be controlled by the user himself, the bandwidth can be set according to the need, so it can be considered that the transmission process of the target background data will not cause the problem of video freezing caused by bandwidth jitter. The target background data is transmitted between the speaking conference server and the non-speaking conference server through the backbone network. Although it is also video data, on the one hand, it does not include the expression information of the participant user, so the data amount is reduced compared with the complete video data, and thus the requirement for the bandwidth is also reduced. On the other hand, no user is speaking in this phase, and the video of the speaking user does not need to be displayed, so the requirement for the transmission time is relatively sufficient. In combination of the two, high-quality target background data can be obtained in each conference server.
[0071] S22: In the second stage of the video conference, real-time audio data and real-time expression data of the speaking user sent by the speaking conference terminal are received, and target historical expression data index information of the speaking user is obtained according to the real-time audio data, the real-time expression data and a user model of the speaking user.
[0072] The second stage of the video conference is a stage in which the video conference is formally conducted and the user speaks, and in this stage, there is a clear speaking user and non-speaking user at each time. When a speaking user starts to speak, the speaking conference terminal where the speaking user is located collects and calculates real-time audio data and real-time expression data of the speaking user, the real-time audio data includes a plurality of real-time audio frames, and the real-time expression data includes a plurality of real-time expression frames. The real-time expression data only includes expression information of the speaking user and does not include other background information. The speaking conference server calls a user model of the speaking user to perform feature recognition and classification on the combination of the real-time audio data and the real-time expression data, determines the most similar historical expression frame of each real-time expression frame in a historical expression data set of the speaking user, and obtains target historical expression frame index information of each historical expression frame in the historical expression data set of the speaking user.
[0073] In an embodiment, S22 specifically includes: judging whether there is a user model of the speaking user; if there is, inputting each real-time expression frame and a corresponding real-time audio frame into the user model of the speaking user to obtain target historical expression frame index information corresponding to each real-time expression frame.
[0074] In the above embodiments, it is mentioned that the user model of each registered user needs a certain amount of data to be trained and synchronized, and therefore, there can be a situation that the user model of a speaking user has not been trained and synchronized. Therefore, after receiving the real-time video data and extracting each real-time expression frame, the speaking conference server first judges whether it has the user model of the speaking user, if it exists, the user model is called to perform feature recognition and classification, and the target historical expression frame index information corresponding to each real-time expression frame is output. If it does not exist, the following methods can be used by considering the terminal computing capacity and data transmission amount and other factors: one is that the speaking conference terminal directly collects real-time expression data and real-time audio data and sends them to the speaking conference server, and then the speaking conference server sends them to other non-speaking conference servers. The non-speaking conference server will combine the real-time expression data with the target background data received in the first phase of the video conference to obtain complete target video data, and finally send the target video data and real-time audio data to the corresponding non-speaking conference terminal. The other is that the speaking conference terminal directly collects real-time audio data and real-time video data (including real-time background data and real-time expression data) and sends them to the speaking conference server, and then the speaking conference server sends them to other non-speaking conference servers, and finally sends the real-time video data and real-time audio data to the corresponding non-speaking conference terminal. That is, the data transmission mechanism of the present application can flexibly adapt to the data transmission requirements in the two states of no registered user model and registered user model.
[0075] S23: Send the real-time audio data and the target historical expression data index information to the non-speaking conference server, so that the non-speaking conference server acquires the target historical expression data from the historical expression data set of the speaking user according to the target historical expression data index information, combines the target historical expression data and the target background data to obtain target video data, and sends the target video data and the real-time audio data to the corresponding non-speaking conference terminal.
[0076] The speaking conference server sends the target historical expression data index information and the real-time audio data of the speaking user to the non-speaking conference server. Since the non-speaking conference server stores the historical expression data set of all registered users, after obtaining the target historical expression data index information, it can acquire the target historical expression data corresponding to the target historical expression data index information from the historical expression data set of the speaking user, that is, the most similar target historical expression frame of each real-time expression frame, and then combine each target historical expression frame and the target background data to obtain multiple complete target video frames. The target video frame has the background information and expression information of the speaking user. Finally, the target video data and the real-time audio data are sent to the corresponding non-speaking conference terminal, so that each non-speaking user can obtain the complete audio and video of the speaking user. The specific data transmission process in the above process can be referred to Figure 4 .
[0077] In the embodiment of the present application, the speaking conference server only sends the target historical expression index information to the non-speaking conference server through the backbone network, instead of directly sending the video, so that the amount of data transmitted in the backbone network is reduced, and the requirement for bandwidth is also reduced. In addition, by establishing the user model of each registered user, the target historical expression index information is obtained based on the user model, and finally the target historical expression frames are obtained. These target historical expression frames have high similarity with the real-time expression frames, so even if the expression of the speaking user is not transmitted in real time, the complete target video data obtained by finally combining also has high quality and restoration degree, and the experience of the non-speaking user is not affected.
[0078] In an embodiment, after S23, it further includes: updating the historical expression data set and the corresponding historical expression frame index information of the speaking user based on the real-time expression frame; and updating and training the user model of the speaking user based on the real-time expression frame, the real-time audio frame and the historical expression frame index information.
[0079] For the user model of each registered user, the more the amount of training sample data is, the better the recognition effect is. In each video conference, the speaking conference server can obtain new real-time expression frames of the speaking user, and then these real-time expression frames can be put into the historical expression data set of the speaking user as new historical expression frames, and new historical expression frame index information is added. Then, the new real-time expression frame and the corresponding real-time audio frame are used as training input data, and the new historical expression frame index information is used as training output data, so that the user model of the speaking user is updated and trained. After the update and training reaches a predetermined number of times, or the sample of the update and training reaches an expectation, the updated user model and the historical expression data set of the speaking user can be synchronized to other conference servers again.
[0080] Through the above-mentioned manner, the user model is continuously accumulated and continuously learned, so that if the same speaking user appears in the subsequent video conference, the recognition effect of the expression and audio of the speaking user will be improved.
[0081] In an embodiment, the method further includes: in the second phase of the video conference, receiving the real-time audio data and the real-time video data of the speaking user sent by the speaking conference terminal; and sending the real-time audio data and the real-time video data to the non-speaking conference server, so that the non-speaking conference server sends the target video data and the real-time audio data to the corresponding non-speaking conference terminal.
[0082] In the above embodiments, the target video data is obtained based on the target background data before the meeting and the real-time facial expression data during the meeting. This method is suitable for situations where the background of the speaking user remains unchanged or remains essentially unchanged during the meeting. If the background of the speaking user changes at a certain moment in the meeting, for example, from a user speaking scenario to a scenario of displaying meeting documents or images, the speaking conference terminal will directly collect real-time video data and real-time audio data and send them to the speaking conference server. The speaking conference server will also directly send real-time video data and real-time audio data to the non-speaking conference server. That is, the data transmission method of this application can flexibly apply to both situations where the background remains unchanged and where the background changes, achieving a balance between ensuring video quality and reducing the amount of data transmitted.
[0083] like Figure 7 The diagram shown is a third flowchart of the video conferencing data transmission method provided in this application embodiment, corresponding to the initial stage of use of the video conferencing system, such as... Figure 8 The diagram shown is a fourth flowchart of the video conferencing data transmission method provided in this application embodiment, which corresponds to the first and second stages of the video conferencing. Figure 7 and Figure 8 All explanations are from the perspective of the non-speaking conference server.
[0084] Please see Figure 7 In the early stages of using video conferencing systems, this method specifically included:
[0085] S31: Obtain historical facial expression data and historical audio data of at least one registered user.
[0086] Similarly, in the initial stages of using the video conferencing system, any non-speaking conferencing server can also obtain the historical facial expression data and historical audio data of at least one registered user. Users can only use the video conferencing system after registration. The historical facial expression data includes multiple historical facial expression frames, which are used to record the registered user's facial expression information. The historical audio data includes multiple historical audio frames, which are used to record the registered user's voice information.
[0087] S32: Based on the historical facial expression data and historical audio data of each registered user, obtain the user model and historical facial expression dataset for each registered user.
[0088] The historical expression data set of each registered user is obtained by storing each historical expression frame of the registered user. When the number of historical expression frames in the historical expression data set is large enough (for example, more than 100,000), the user model of the registered user can be modeled according to the historical expression data and the historical audio data of the registered user. The user model is trained by using the historical expression frames and the corresponding historical audio frames of the registered user as training data. The historical audio frame corresponding to the historical expression frame can be a historical audio frame at the same time, or an audio segment composed of multiple historical audio frames in the same time period. The specific selection can be made according to the training needs. After the training is completed, the feature recognition of any combination of the expression frame and the corresponding audio frame of the registered user can be performed, and the recognition accuracy is high. Since the audio data can reflect the emotions of the user such as joy, anger, sorrow, and the like, the combination of the expression frame and the audio frame is used as the training data to establish the user model, which is equivalent to using only the expression frame. The recognition accuracy of the trained user model is higher, and the recognition effect is better.
[0089] In an embodiment, the step of obtaining the user model of each registered user and the historical expression data set according to the historical expression data and the historical audio data of each registered user includes: generating the historical expression data set of each registered user according to all historical expression frames of each registered user, and generating historical expression frame index information of the historical expression frames in each historical expression data set; obtaining an initial user model of each registered user, and respectively taking the historical expression frames and the historical audio frames of each registered user as training input data, and taking the historical expression frame index information as training output data, to train each initial user model to obtain the user model of each registered user.
[0090] After obtaining the historical expression data set of a certain registered user, the corresponding historical expression frame index information can be generated for each historical expression frame in the historical expression data set. At the same time, the initial user model of the registered user is obtained, and the initial user model of the registered user is a neural network model, which can be ResNet or other types of neural network models. After training with training data of different registered users, the user model of each different registered user is obtained. Specifically, for the initial user model, the binary information of each historical expression frame and the corresponding historical audio frame of a certain registered user is combined into new picture information, and N historical expression frames can be combined to obtain N groups of training input data. At the same time, the index information of each historical expression frame of the registered user forms N groups of training output data, respectively. The initial user model is trained based on the N groups of training input data and training output data, so that the model can output the historical expression frame index information corresponding to the real-time expression frame of the registered user after feature extraction and classification, and the similarity between the historical expression frame corresponding to the historical expression frame index information and the real-time expression frame meets the expectation. After training, the user model of the registered user is obtained. The above operation is performed once for all registered users, and the user model of each registered user can be obtained in the non-speaking conference server.
[0091] S33: Synchronize the user model and the historical expression data set of each registered user to other conference servers.
[0092] After the user model of each registered user is built, it can be synchronized to other conference servers, so that other conference servers have the user model of the registered user. The model building and synchronization of each registered user is a dynamic process, that is, the non-speaking conference server will continuously obtain the historical expression data and historical speech data of each registered user. When the amount of these data of a certain registered user reaches the expectation, the model training and synchronization of the registered user will be performed, and other registered users are always in the data collection stage until the condition for starting training is reached.
[0093] The entire process described above can be referred to Figure 2 In this stage, the conference server 2 can obtain the user model and the historical expression data set of each registered user 2, and synchronize them to the conference server 1 through a wide area network.
[0094] In the early stage of use, the video conference system still uses the architecture shown in Figure 1 , and the workflow is the same as Figure 1The historical expression data and the historical audio data can be collected by the non-speaking conference server when the user submits the registration information and during the subsequent video conference. When the number of historical expression frames in the historical expression data does not meet the expectation, the current data transmission process is still maintained and the data collection is continuously performed. When the historical expression frame data is large enough, the user model is automatically established and synchronized. This process is a self-learning process without awareness. After the model is established and synchronized, seamless switching from the current video conference data transmission method to the video conference data transmission method in the application can be realized.
[0095] As shown in FIG. 1, in the first phase and the second phase of the video conference, the method specifically includes: Figure 8
[0096] S41: In the first phase of the video conference, target background data of a speaking user is received, which is sent by a speaking conference server. The target background data is sent by a speaking conference terminal to the speaking conference server.
[0097] The first phase of the video conference is a phase in which each participant has not started speaking after the video conference starts. In this phase, if it is clear in the video conference that only one or several users can speak in this conference, these users are speaking users, and other users are non-speaking users. If it is not clear in the video conference, since each participant may speak later, each participant can be regarded as a speaking user in this phase, and other participants are regarded as non-speaking users when the participant is a speaking user. Each speaking conference terminal collects and calculates the target background data of the speaking user, that is, all video information of each participant except the expression, and sends the target background data to the nearest speaking conference server through the access network. The speaking conference server synchronizes the target background data to other non-speaking conference servers, so that each non-speaking conference server can obtain and save the target background data. The specific transmission process of the data in the above process can be referred to in the description of FIG. 2. Figure 3 .
[0098] In this step, the target background data is transmitted between the speaking conference terminal and the speaking conference server through the access network. Since the access network can be controlled by the user himself, the bandwidth can be set according to the need, so it can be considered that the transmission process of the target background data will not cause the problem of video freezing caused by bandwidth jitter. The target background data is transmitted between the speaking conference server and the non-speaking conference server through the backbone network. Although it is also video data, on the one hand, it does not include the expression information of the participant, so the data amount is reduced compared with the complete video data, and thus the requirement for bandwidth is also reduced. On the other hand, no user is speaking in this stage, and the video of the speaking user does not need to be displayed, so the requirement for transmission time is relatively sufficient. In combination of the two, high-quality target background data can be obtained in each conference server.
[0099] S42: In the second stage of the video conference, the real-time audio data of the speaking user and the target historical expression data index information sent by the speaking conference server are received, the target historical expression data index information is obtained by the speaking conference server according to the real-time audio data, the real-time expression data of the speaking user and the user model, and the real-time audio data and the real-time expression data are sent by the speaking conference terminal to the speaking conference server.
[0100] The second stage of the video conference is also the stage in which the video conference is formally carried out and the user speaks, and in this stage, there is a clear speaking user and a non-speaking user at each time. When a speaking user starts to speak, the speaking conference terminal where the speaking user is located collects and calculates the real-time audio data and the real-time expression data of the speaking user, the real-time audio data includes a plurality of real-time audio frames, and the real-time expression data includes a plurality of real-time expression frames, the real-time expression data only includes the expression information of the speaking user and does not include other background information. The speaking conference server calls the user model of the speaking user to perform feature recognition and classification on the combination of the real-time audio data and the real-time expression data, determines the most similar historical expression frame of each real-time expression frame in the historical expression data set of the speaking user, and obtains the target historical expression frame index information of each historical expression frame in the historical expression data set of the speaking user. Finally, the speaking conference server sends the target historical expression data index information of the speaking user and the real-time audio data to the non-speaking conference server, and the non-speaking conference server receives these data.
[0101] The target historical expression data index information obtained by the speaking conference server is obtained through user model processing, and the user model of each registered user needs a certain amount of data to be trained and synchronized. Therefore, the speaking conference server may not have the user model of the speaking user after receiving the real-time video data and extracting each real-time expression frame. If the user model exists, the user model is called to perform feature recognition and classification, and the target historical expression frame index information corresponding to each real-time expression frame is output. Then, the target historical expression frame index information is sent to the non-speaking conference server, and the non-speaking conference server performs subsequent processing. If the user model does not exist, the terminal computing capacity and data transmission amount and other factors are considered, and any of the following methods is used: one is that the speaking conference terminal directly collects real-time expression data and real-time audio data and sends them to the speaking conference server, and then the speaking conference server sends them to other non-speaking conference servers. The non-speaking conference server combines the real-time expression data with the target background data received in the first phase of the video conference to obtain complete target video data, and finally sends the target video data and the real-time audio data to the corresponding non-speaking conference terminal. The other is that the speaking conference terminal directly collects real-time audio data and real-time video data (including real-time background data and real-time expression data) and sends them to the speaking conference server, and then the speaking conference server sends them to other non-speaking conference servers, and finally sends the real-time video data and the real-time audio data to the corresponding non-speaking conference terminal. That is, the data transmission mechanism of the present application can flexibly adapt to the data transmission requirements in the two states of no registered user model and registered user model.
[0102] S43: According to the target historical expression data index information, the target historical expression data is obtained from the historical expression data set of the speaking user, the target historical expression data and the target background data are combined to obtain the target video data, and the target video data and the real-time audio data are sent to the corresponding non-speaking conference terminal.
[0103] Since the non-speaking conference server stores the historical expression data set of all registered users, after obtaining the target historical expression data index information, the target historical expression data corresponding to the target historical expression data index information can be obtained from the historical expression data set of the speaking user, that is, the most similar target historical expression frame of each real-time expression frame. Then, each target historical expression frame and the target background data are combined to obtain a plurality of complete target video frames, which have the background information and expression information of the speaking user. Finally, the target video data and the real-time audio data are sent to the corresponding non-speaking conference terminal, so that each non-speaking user can obtain the complete audio and video of the speaking user. The specific data transmission process in the above process can be referred to in Figure 4 .
[0104] In the embodiments of the present application, the non-speaking conference server receives only the target historical expression index information instead of the complete video from the backbone network, so that the amount of data transmitted in the backbone network is reduced, and the requirement for bandwidth is also reduced. In addition, the target historical expression index information is obtained based on the user model of each registered user, and finally the target historical expression frames are obtained. These target historical expression frames have high similarity with the real-time expression frames, so that even if the expression of the speaking user is not transmitted in real time, the complete target video data obtained by finally combining has high quality and restoration degree, and the experience of the non-speaking user is not affected.
[0105] In an embodiment, the method further comprises: in the second phase of the video conference, receiving the real-time audio data and the real-time video data of the speaking user sent by the speaking conference server; and sending the real-time audio data and the real-time video data to the corresponding non-speaking conference terminal
[0106] In the above embodiments, the target video data is obtained based on the target background data before the conference and the real-time expression data in the conference. This method is suitable for the case where the background of the speaking user does not change or basically does not change during the conference. If the background of the speaking user changes at a certain moment of the conference, for example, the scene of the user speaking is converted into the scene of displaying conference documents or pictures, then the speaking conference terminal directly collects the real-time video data and the real-time audio data and sends them to the speaking conference server, and the speaking conference server directly sends the real-time video data and the real-time audio data to the non-speaking conference server. That is, the data transmission method of the present application can be flexibly applied to the case where the background remains unchanged and the case where the background changes, and the balance between ensuring the video quality and reducing the amount of transmitted data is achieved.
[0107] In combination with Figures 2 to 8 It can be seen that in the video conference data transmission method and the video conference system of the present application, only audio and video are transmitted between the conference terminal 1 and the conference server 1 connected through the access network, and between the conference terminal 2 and the conference server 2, in the first phase, only target background data is transmitted, in the second phase, only real-time audio data and target historical expression data index information are transmitted, and then complete audio and video are calculated and combined. Since the bandwidth of the access network is controllable and is exclusively used by the user, the video transmission between the conference terminal and the conference server can be realized without jitter by adjusting the bandwidth of the access network. Although the bandwidth of the backbone network is uncontrollable, the amount of data transmitted through the backbone network in the same period is reduced, and these data reduce the requirement for the bandwidth of the backbone network. Therefore, by comprehensively considering the above two aspects, the scheme in the present application can significantly reduce the probability of video freezing without increasing the bandwidth of the backbone network or compressing the audio and video.
[0108] On the basis of the method described in the above embodiment, the present embodiment will be further described from the perspective of a video conference data transmission device, please refer to Figure 9 The video conference data transmission device is applied to a speaking conference server, and specifically can include:
[0109] A first sending module 10 is configured to receive target background data of a speaking user sent by a speaking conference terminal in a first stage of a video conference, and send the target background data to a non-speaking conference server.
[0110] A first receiving module 20 is configured to receive real-time audio data and real-time expression data of the speaking user sent by the speaking conference terminal in a second stage of the video conference, and obtain target historical expression data index information of the speaking user according to the real-time audio data, the real-time expression data and a user model of the speaking user.
[0111] A second sending module 30 is configured to send the real-time audio data and the target historical expression data index information to the non-speaking conference server, so that the non-speaking conference server acquires target historical expression data from a historical expression data set of the speaking user according to the target historical expression data index information, combines the target historical expression data and the target background data to obtain target video data, and sends the target video data and the real-time audio data to a corresponding non-speaking conference terminal.
[0112] Correspondingly, the present application further provides a video conference data transmission device, please refer to Figure 10 The video conference data transmission device is applied to a non-speaking conference server, and specifically can include:
[0113] A second receiving module 40 is configured to receive target background data of a speaking user sent by the speaking conference server in a first stage of a video conference, the target background data being sent by a speaking conference terminal to the speaking conference server.
[0114] A third receiving module 50 is configured to receive real-time audio data and target historical expression data index information of the speaking user sent by the speaking conference server in a second stage of the video conference, the target historical expression data index information being obtained by the speaking conference server according to real-time expression data of the speaking user and a user model, and the real-time audio data and the real-time expression data being sent by the speaking conference terminal to the speaking conference server.
[0115] The third sending module 60 is used to obtain target historical expression data from the historical expression dataset of the speaking user according to the target historical expression data index information, combine the target historical expression data and the target background data to obtain target video data, and send the target video data and the real-time audio data to the corresponding non-speaking conference terminal.
[0116] Unlike existing technologies, the video conferencing data transmission device provided in this application transmits complete audio and video only between the conference terminal and the conference server connected via the access network. Between the conference servers connected via the backbone network, only target background data is transmitted in the first stage, and only real-time audio data and target historical facial expression data index information are transmitted in the second stage. The complete audio and video are then calculated and combined. Since the bandwidth of the access network is controllable and exclusively used by the user, video transmission between the conference terminal and the conference server can be made free of bandwidth jitter by adjusting the bandwidth of the access network. Although the bandwidth of the backbone network is uncontrollable, the reduced amount of data transmitted through the backbone network at the same time reduces the bandwidth requirements of the backbone network. Therefore, combining these two factors, the solution in this application can significantly reduce the probability of video stuttering without increasing the backbone network bandwidth or compressing the audio and video.
[0117] Accordingly, embodiments of this application also provide an electronic device, such as... Figure 11 As shown, the electronic device may include a radio frequency (RF) circuit 101, a memory 102 including one or more computer-readable storage media, an input unit 103, a display unit 104, a sensor 105, an audio circuit 106, a WiFi module 107, a processor 108 including one or more processing cores, and a power supply 109, among other components. Those skilled in the art will understand that... Figure 11 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0118] The radio frequency circuit 101 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and hands it over to one or more processors 108 for processing; additionally, it transmits uplink data to the base station. The memory 102 can be used to store software programs and modules. The processor 108 executes various functional applications and video conferencing data transmission by running the software programs and modules stored in the memory 102. The input unit 103 can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical, or trackball signal inputs related to customer settings and function control.
[0119] The display unit 104 can be used to display information input by the client or provided to the client and various graphical client interfaces of the server, which can be composed of graphics, text, icons, video and any combination thereof.
[0120] The electronic device can also include at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. The audio circuit 106 includes a speaker, which can provide an audio interface between the client and the electronic device.
[0121] WiFi belongs to short-range wireless transmission technology, and the electronic device can help the client send and receive emails, browse web pages, and follow streaming media through the WiFi module 107, which provides the client with wireless broadband Internet access. Although Figure 11 The WiFi module 107 is shown, but it is understood that it is not a necessary component of the electronic device and can be omitted as needed without changing the essence of the application.
[0122] The processor 108 is the control center of the electronic device, which connects all parts of the phone through various interfaces and lines, executes various functions of the electronic device and processes data by running or executing software programs and / or modules stored in the memory 102 and calling data stored in the memory 102, thereby monitoring the entire phone.
[0123] The electronic device also includes a power supply 109 (such as a battery) for powering the components. Preferably, the power supply can be logically connected to the processor 108 through a power management system, so that the power management system can realize functions such as charge management, discharge management, and power consumption management.
[0124] Although not shown, the electronic device can also include a camera, a Bluetooth module, etc., which will not be described here. In this embodiment, the processor 108 in the server will load one or more executable files corresponding to the processes of one or more application programs into the memory 102 according to the following instructions, and run the application programs stored in the memory 102 by the processor 108, thereby realizing the following functions:
[0125] In the first stage of the video conference, the target background data of the speaking user sent by the speaking conference terminal is received, and the target background data is sent to the non-speaking conference server;
[0126] In the second stage of the video conference, the real-time audio data and real-time expression data of the speaking user sent by the speaking conference terminal are received, and the target historical expression data index information of the speaking user is obtained according to the real-time audio data, the real-time expression data and the user model of the speaking user;
[0127] send the real-time audio data and the target historical expression data index information to the non-speaking conference server, so that the non-speaking conference server acquires target historical expression data from the historical expression data set of the speaking user according to the target historical expression data index information, combines the target historical expression data and the target background data to obtain target video data, and sends the target video data and the real-time audio data to the corresponding non-speaking conference terminal.
[0128] Or implement the following functions:
[0129] In the first phase of the video conference, target background data of a speaking user is received, which is sent by a speaking conference terminal to a speaking conference server;
[0130] In the second phase of the video conference, real-time audio data and target historical expression data index information of the speaking user are received, which are sent by the speaking conference server, the target historical expression data index information is obtained by the speaking conference server according to the real-time audio data, real-time expression data of the speaking user and a user model, and the real-time audio data and the real-time expression data are sent by the speaking conference terminal to the speaking conference server;
[0131] According to the target historical expression data index information, target historical expression data is acquired from the historical expression data set of the speaking user, the target historical expression data and the target background data are combined to obtain target video data, and the target video data and the real-time audio data are sent to the corresponding non-speaking conference terminal.
[0132] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the detailed description above, which will not be repeated here.
[0133] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by instructions controlling related hardware, which can be stored in a computer readable storage medium and loaded and executed by a processor.
[0134] To this end, an embodiment of the present application provides a computer readable storage medium, which stores a plurality of instructions capable of being loaded by a processor to implement the following functions:
[0135] In the first phase of the video conference, target background data of a speaking user is received, which is sent by a speaking conference terminal to a speaking conference server;
[0136] In the second stage of the video conference, the real-time audio data and real-time expression data of the speaking user sent by the speaking conference terminal are received, and target historical expression data index information of the speaking user is obtained according to the real-time audio data, the real-time expression data and a user model of the speaking user;
[0137] The real-time audio data and the target historical expression data index information are sent to the non-speaking conference server, so that the non-speaking conference server obtains target historical expression data from a historical expression data set of the speaking user according to the target historical expression data index information, combines the target historical expression data and target background data to obtain target video data, and sends the target video data and the real-time audio data to a corresponding non-speaking conference terminal.
[0138] Or the following functions are implemented:
[0139] In the first stage of the video conference, target background data of a speaking user sent by the speaking conference server is received, and the target background data is sent by the speaking conference terminal to the speaking conference server.
[0140] In the second stage of the video conference, real-time audio data and target historical expression data index information of the speaking user sent by the speaking conference server are received, the target historical expression data index information is obtained by the speaking conference server according to the real-time audio data, real-time expression data of the speaking user and a user model, and the real-time audio data and the real-time expression data are sent by the speaking conference terminal to the speaking conference server.
[0141] Target historical expression data is obtained from a historical expression data set of the speaking user according to the target historical expression data index information, the target historical expression data and the target background data are combined to obtain target video data, and the target video data and the real-time audio data are sent to a corresponding non-speaking conference terminal.
[0142] The above describes in detail a video conference data transmission method and a video conference system, an electronic device and a computer readable storage medium provided by an embodiment of the application. The principle and implementation manner of the application are described by applying specific examples. The above embodiment descriptions are only used to help understand the technical solutions and core ideas of the application. Those skilled in the art should understand that the technical solutions recorded in the above embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the application.
Claims
1. A method of video conference data transmission, characterized by, The application is suitable for a video conference system, which comprises at least two conference terminals and at least two conference servers, each conference server is connected through a backbone network, each conference server is connected with a corresponding conference terminal through an access network, the conference terminal comprises a speaking conference terminal and a non-speaking conference terminal, the conference server comprises a speaking conference server and a non-speaking conference server, the video conference data transmission method is applied to the speaking conference server, and the method comprises the following steps: In the first stage of the video conference, target background data of a speaking user sent by the speaking conference terminal is received, and the target background data is sent to the non-speaking conference server; In the second stage of the video conference, real-time audio data and real-time expression data of the speaking user sent by the speaking conference terminal are received, target historical expression data index information of the speaking user is obtained according to the real-time audio data, the real-time expression data and a user model of the speaking user; The real-time audio data and the target historical expression data index information are sent to the non-speaking conference server, so that the non-speaking conference server acquires target historical expression data from a historical expression data set of the speaking user according to the target historical expression data index information, combines the target historical expression data and the target background data to obtain target video data, and sends the target video data and the real-time audio data to a corresponding non-speaking conference terminal.
2. The video conference data transmission method of claim 1, wherein, Before the step of receiving the target background data of the speaking user sent by the speaking conference terminal, the method further comprises the following steps: Obtaining historical expression data and historical audio data of at least one registered user; Obtaining a user model and a historical expression data set of each registered user according to the historical expression data and the historical audio data of each registered user; Synchronizing the user model and the historical expression data set of each registered user to other conference servers.
3. The video conference data transmission method of claim 2, wherein, The step of obtaining the user model and the historical expression data set of each registered user according to the historical expression data and the historical audio data of each registered user comprises the following steps: Generating a historical expression data set of each registered user according to all historical expression frames of each registered user, and generating historical expression frame index information of the historical expression frames in each historical expression data set; Obtaining an initial user model of each registered user, taking the historical expression frames and the historical audio frames of each registered user as training input data and the historical expression frame index information as training output data, and training the initial user model to obtain the user model of each registered user.
4. The video conference data transmission method of claim 3, wherein, The step of obtaining the target historical expression data index information of the speaking user according to the real-time audio data, the real-time expression data and the user model of the speaking user comprises the following steps: Judging whether the user model of the speaking user exists or not; If the user model of the speaking user exists, inputting each real-time expression frame and a corresponding real-time audio frame into the user model of the speaking user to obtain target historical expression frame index information corresponding to each real-time expression frame.
5. The video conference data transmission method of claim 4, wherein, After the step of sending the target video data and the real-time audio data to the corresponding non-speaking conference terminal, the method further comprises the following steps: updating a historical expression dataset and corresponding historical expression frame index information of the speaking user based on the real-time expression frame; updating training a user model of the speaking user based on the real-time expression frame, the real-time audio frame and the historical expression frame index information.
6. A method of video conference data transmission, characterized by, The video conference data transmission method is suitable for a video conference system, the video conference system comprising at least two conference terminals and at least two conference servers, each conference server being connected through a backbone network, and each conference server being connected with a corresponding conference terminal through an access network, the conference terminals comprising speaking conference terminals and non-speaking conference terminals, the conference servers comprising speaking conference servers and non-speaking conference servers, the video conference data transmission method being applied to the non-speaking conference server, and the method comprising: in a first phase of the video conference, receiving target background data of a speaking user sent by a speaking conference server, the target background data being sent by the speaking conference terminal to the speaking conference server; in a second phase of the video conference, receiving real-time audio data and target historical expression data index information of the speaking user sent by the speaking conference server, the target historical expression data index information being obtained by the speaking conference server according to the real-time audio data, real-time expression data of the speaking user and a user model, the real-time audio data and the real-time expression data being sent by the speaking conference terminal to the speaking conference server; obtaining target historical expression data from a historical expression dataset of the speaking user according to the target historical expression data index information, combining the target historical expression data and the target background data to obtain target video data, and sending the target video data and the real-time audio data to a corresponding non-speaking conference terminal.
7. The video conference data transmission method of claim 6, wherein, Before the step of receiving the target background data of the speaking user sent by the speaking conference server, the method further comprises: obtaining historical expression data and historical audio data of at least one registered user; obtaining a user model and a historical expression dataset of each registered user according to the historical expression data and the historical audio data of each registered user; synchronizing the user model and the historical expression dataset of each registered user to other conference servers.
8. The video conference data transmission method of claim 7, wherein, The step of obtaining a user model and a historical expression dataset of each registered user according to the historical expression data and the historical audio data of each registered user comprises: generating the historical expression dataset of each registered user according to all historical expression frames of each registered user, and generating historical expression frame index information of the historical expression frames in each historical expression dataset; obtaining an initial user model of each registered user, and training each initial user model by taking the historical expression frames and the historical audio frames of each registered user as training input data and taking the historical expression frame index information as training output data, to obtain the user model of each registered user.
9. The video conference data transmission method of claim 8, wherein, The step of obtaining target historical expression data from a historical expression dataset of the speaking user according to the target historical expression data index information, and combining the target historical expression data and the target background data to obtain target video data comprises: According to the target historical expression data index information, a target historical expression frame corresponding to each real-time expression frame is obtained from a historical expression data set of the speaking user; Each target historical expression frame is combined with the target background data to obtain each target video frame.
10. A video conferencing system characterized by The video conference system includes at least two conference terminals and at least two conference servers, each conference server is connected through a backbone network, and each conference server is connected with a corresponding conference terminal through an access network. The conference terminal includes a speaking conference terminal and a non-speaking conference terminal, and the conference server includes a speaking conference server and a non-speaking conference server. Wherein: The speaking conference server is configured to, in a first stage of the video conference, receive target background data of a speaking user sent by the speaking conference terminal, and send the target background data to the non-speaking conference server; The speaking conference server is configured to, in a second stage of the video conference, receive real-time audio data and real-time expression data of the speaking user sent by the speaking conference terminal, obtain target historical expression data index information of the speaking user according to the real-time audio data, the real-time expression data and a user model of the speaking user, and send the real-time audio data and the target historical expression data index information to the non-speaking conference server; The non-speaking conference server is configured to, in the second stage of the video conference, obtain target historical expression data from a historical expression data set of the speaking user according to the target historical expression data index information, combine the target historical expression data and the target background data to obtain target video data, and send the target video data and the real-time audio data to a corresponding non-speaking conference terminal.
Citation Information
Patent Citations
Extensible framework for executable annotations in electronic content
CN113287138A
Method for controlling flux of audio and video flow transferred in IP video meeting system
CN1540954A