Video conference method, system and device and storage medium
By using pre-trained parameters to adjust the model in real-time prediction of video code rate and adjusting video parameters in the video conferencing system, the problems of data transmission delay and unstable video quality in the video conferencing system are solved, and a more stable and smooth video conferencing experience is achieved.
Patent Information
- Application Number
- CN202510244040.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-11-29
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-03
AI Technical Summary
Currently, video conferencing systems are facing the problems of data transmission delay and unstable video quality, especially when network conditions fluctuate, which may lead to audio and video out of synchronization, picture stuttering or distortion, affecting the real-time and effectiveness of the meeting.
The model predicts the best video bit rate in real time through pre-trained parameters, adjusts the video parameters to adapt to changes in network and user behavior, and reduces the video picture quality to reduce network transmission pressure when the search service is called.
It realizes the stability of video quality under different network environments, reduces lag and distortion, and improves the fluency of meetings and the integrity of information transmission.
Smart Images

Figure CN120050384A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network communication technology, and in particular, to a method, system, device, and storage medium for video conferencing. Background Art
[0002] With the development of Internet technology, video conferencing has become an important assistant for remote communication and is widely used in fields such as business, education, and healthcare. However, current video conferencing systems still face many challenges. First, data transmission latency is a prominent problem. Fluctuations in network conditions can lead to audio-video desynchronization, affecting the real-time nature of the conference. Especially in scenarios such as legal or medical consultations, this latency may lead to misunderstandings or decision-making mistakes. Second, video quality is unstable in different network environments, often resulting in frozen or distorted images due to bandwidth limitations, making it difficult to effectively convey the conference content. Summary of the Invention
[0003] The purpose of the present invention is to provide a method, system, device, and storage medium for video conferencing to improve the above problems. To achieve the above purpose, the technical solutions adopted by the present invention are as follows:
[0004] In a first aspect, the present application provides a method for video conferencing, including:
[0005] Obtain network performance information, user behavior information, video quality information, and retrieval service usage information, where the retrieval service usage information includes the call status of the retrieval service;
[0006] Using the network performance information, user behavior information, and video quality information as inputs, obtain a predicted bitrate through a pre-constructed first parameter adjustment model;
[0007] According to the call status of the retrieval service, determine whether the retrieval service is called. If not, adjust the video parameters according to the predicted bitrate; if so, obtain a preset resolution and a preset frame rate, calculate the video bitrate to obtain a target bitrate, and adjust the video parameters according to the target bitrate.
[0008] In a second aspect, the present application further provides a system for video conferencing, including:
[0009] An acquisition module, configured to obtain network performance information, user behavior information, video quality information, and retrieval service usage information, where the retrieval service usage information includes the call status of the retrieval service;
[0010] A first processing module, configured to use the network performance information, user behavior information, and video quality information as inputs, and obtain a predicted bitrate through a pre-constructed first parameter adjustment model;
[0011] A calculation module, configured to determine whether the retrieval service is invoked according to the invocation status of the retrieval service. If not, the video parameters are adjusted according to the predicted bit rate. If so, a preset resolution and a preset frame rate are obtained, and the video bit rate is calculated to obtain a target bit rate, and the video parameters are adjusted according to the target bit rate.
[0012] In a third aspect, the present application further provides a device for a video conference, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above video conference method are implemented.
[0013] In a fourth aspect, the present application further provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above video conference method are implemented.
[0014] The beneficial effects of the present invention are as follows:
[0015] The present invention predicts the optimal video bit rate in real time during the conference through a pre-trained first parameter adjustment model, reduces stuttering or distortion, ensures video quality, and at the same time adjusts the video data transmission parameters according to the user's retrieval behavior during the conference, improves the response speed of the retrieval service, and also ensures the smoothness of the conference video, ensuring the integrity of the information obtained by the user.
[0016] Other features and advantages of the present invention will be described in the subsequent description, and some of them will become obvious from the description, or be understood by implementing the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a flowchart of the method for a video conference in an embodiment of the present application;
[0019] Figure 2 It is a schematic structural diagram of the system for a video conference in an embodiment of the present application;
[0020] Figure 3 It is a schematic structural diagram of the device for a video conference in an embodiment of the present application.
[0021] Markings in the figure: 100 - acquisition module; 200 - first processing module; 300 - calculation module; 400 - second processing module; 500 - third processing module; 600 - retrieval module; 800 - device for video conferencing; 810 - processor; 820 - memory; 803 - multimedia component; 804 - I / O interface; 805 - communication component. Detailed implementation manners
[0022] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention generally described and illustrated in the figures herein can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but is merely representative of selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] It should be noted that like reference numerals and letters denote like items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0024] Embodiment 1
[0025] Refer to Figure 1 , this embodiment provides a method for video conferencing, including steps S100, S200 and S300; in addition, this method requires the use of a pre-trained first parameter adjustment model, and the construction steps of the first parameter adjustment model are as follows:
[0026] Step S001, data collection and preparation;
[0027] Collect the following types of data for model training:
[0028] Network performance data: bandwidth (which must change frequently to reflect the actual network condition), latency, packet loss rate, etc.;
[0029] User behavior data: such as the user is "watching", the on / off states of the user's camera and microphone, etc.;
[0030] Video quality data: video resolution (in the viewing state), frame rate, video bit rate (the optimal video transmission bit rate under various network conditions);
[0031] Timestamp data: record the timestamps corresponding to the collected data.
[0032] After collecting the above data records from historical meetings, a time series data set is formed. Then, the data is integrated and unified in format to ensure data integrity, remove duplicate values and outliers, and form a structured data.
[0033] Step S002, data preprocessing;
[0034] Normalize feature data such as bandwidth, latency, and packet loss rate to the range of 0 - 1, and use one-hot encoding to process categorical feature data such as video resolution and frame rate, and convert it into numerical feature data.
[0035] Generate feature vectors based on the processed features, and combine the required information into the input format:
[0036] pythonfeatures = [bandwidth_normalized, latency_normalized, packet_loss_normalized, resolution_one_hot, frame_rate_one_hot, user_state_encoded], and the output value is the required optimal video bitrate.
[0037] Divide the integrated data set into a training set (70%), a validation set (15%), and a test set (15%). Ensure random division to ensure that the distribution of each subset is as similar as possible.
[0038] Step S003, model construction and training;
[0039] Use the LSTM model for time series prediction.
[0040] Input layer: Accept the above processed feature vectors.
[0041] LSTM layer: Establish a deep LSTM network to process time series data. Multiple LSTM layers can be selected to improve the model's prediction ability.
[0042] Output layer: One node, output the predicted value of the video bitrate.
[0043] Compile the model: Select appropriate loss functions and optimizers. Example code: pythonmodel.compile(loss='mean_squared_error', optimizer='adam').
[0044] Training the model: Input the training set into the model for training, and at the same time use the validation set to evaluate the model performance. Set parameters: the batch size of each training (e.g., 32), the number of training epochs (e.g., 50). Example code: python model.fit(X_train, y_train, validation_data=(X_val, y_val), epochs=50, batch_size=32).
[0045] Model evaluation: After training is completed, use the test set to evaluate the model performance and calculate the prediction accuracy. The metrics to focus on include: mean squared error (MSE), root mean squared error (RMSE).
[0046] Step S004: Model optimization;
[0047] Adopt early stopping method (EarlyStopping), stop training when the validation loss no longer decreases to prevent overfitting. Conduct hyperparameter tuning, optimize hyperparameters such as learning rate, number of LSTM layers, etc. through grid search (GridSearch) or random search (RandomSearch), and finally obtain the first parameter-adjusted model.
[0048] Step S100: Obtain network performance information, user behavior information, video quality information, and retrieval service usage information. The retrieval service usage information includes the call status of the retrieval service;
[0049] The network performance information includes the current bandwidth, latency, packet loss rate, etc., which can be collected regularly using monitoring tools; the user behavior information includes the behavior of the user "watching", detecting the opening and closing status of the user's camera and microphone, and other operations of the user on the interface, etc.;
[0050] The video quality information includes the resolution, frame rate, etc. of the current conference video;
[0051] The retrieval service usage information includes the call status of the retrieval service, that is, whether it is called, and also includes the response status of each retrieval result. The retrieval results include text data, image data, and video data.
[0052] Step S200: Use the network performance information, user behavior information, and video quality information as inputs, and obtain the predicted bitrate through the pre-built first parameter-adjusted model.
[0053] After processing the acquired data according to the above step S002, the data is input into the first parameter adjustment model to obtain the predicted bitrate. This step performs real-time prediction of the video bitrate based on the current network, user behavior, and video quality conditions, thereby determining the optimal encoding and decoding parameters, which can effectively improve the data transmission efficiency of video conferencing, reduce video frame freezes or distortions, and ensure video quality.
[0054] When disputes or conflicts occur in data transmission, traditional systems usually lack an intelligent mediation mechanism, the processing process is slow, and manual intervention is required, which affects decision-making efficiency. However, the first parameter adjustment model of this application can achieve intelligent mediation and dynamically adjust encoding according to the predicted optimal bitrate.
[0055] The method of this application is implemented based on a video conferencing program with a retrieval function. The retrieval function enables users to retrieve historical information of the meeting, and the historical information at least includes the text record and video of the meeting. Therefore, this method may further include the steps of:
[0056] Acquire the audio data of the video conference, perform speech recognition processing on the audio data, and obtain and store the text data after conversion of the audio data;
[0057] Perform semantic recognition processing on the text data to obtain segmented text data divided by content; in this step, the text can be pre-corrected through a backend text proofreading algorithm to ensure the accuracy of the text record.
[0058] Receive a retrieval request and obtain a retrieval keyword, and perform a retrieval in the segmented text data based on the retrieval keyword to obtain the first target text data. For example, when a user searches for "project progress", using it as a keyword, search the database for segmented text data containing "project progress", and present the retrieved text and the corresponding timestamp to the user. In addition, a link relationship can be established between the text and the corresponding video segment according to the timestamp, enabling the user to directly jump to the corresponding audio and video playback position by clicking on the text segment, enhancing the efficiency of information acquisition.
[0059] Step S300: According to the call status of the retrieval service, determine whether the retrieval service is called. If not, adjust the video parameters according to the predicted bitrate; if so, obtain the preset resolution and preset frame rate, calculate the video bitrate to obtain the target bitrate, and adjust the video parameters according to the target bitrate;
[0060] The retrieval service is a function already developed in existing video conferencing software, that is, the server stores the data of this video conference. While the user is having a video conference, the user can retrieve the information before this video conference through keywords, and obtain historical text and video from the server to review the meeting content.
[0061] When the client performs a search, the current conference video needs to keep playing, and historical data also needs to be obtained from the server, which will further increase the pressure on network transmission. In fact, when the user has the behaviors of searching and reviewing, the picture of the current conference video is not the focus of the user's attention, and even the user hardly pays attention to the current video picture. Therefore, the picture quality of the conference video can be reasonably reduced at this time to reduce the pressure on network transmission and improve the response speed of the search service.
[0062] The reasonable reduction of the picture quality of the conference video can be achieved by presetting a lower resolution and frame rate. For example, the resolution can be fixed at 480p or 360p. Then, the required bit rate can be calculated directly according to the selected resolution and frame rate, which can be achieved through the bit rate formula: [Bit rate = Resolution × Frame rate × Color depth].
[0063] In order to more flexibly adjust the picture quality and optimize the conference experience, the picture quality can also be adjusted according to the user's search behavior. As an optional implementation method, the obtaining of the preset resolution and preset frame rate, and the calculation of the video bit rate to obtain the target bit rate include:
[0064] When the search service is in the called state, obtain the response search results in the response state according to the search service usage information, that is, the content that the user needs to review, such as text, audio, or video, etc.
[0065] Obtain the corresponding preset resolution and preset frame rate according to the response search results; the transmission of different search results requires different amounts of bandwidth. For the transmission of complex data, the current picture quality needs to be adjusted to be smaller accordingly. Therefore, when reviewing historical videos, the resolution and frame rate of the current conference video should be adjusted to be smaller.
[0066] Judge whether there are two or more response search results at the same time. If so, select the preset resolution and preset frame rate corresponding to the response search result with the highest priority in the order of decreasing priority of video data, image data, and text data. If not, obtain the preset resolution and preset frame rate corresponding to the only response search result.
[0067] Different types of search results correspond to preset different resolutions and frame rates. When multiple search results are responded at the same time, the resolution and frame rate corresponding to the video data shall prevail first, and then the image data (or audio data) shall prevail; that is, the target resolution and target frame rate are obtained.
[0068] According to the target resolution and target frame rate, calculate the video bit rate through the bit rate formula to obtain the target bit rate.
[0069] As an alternative implementation, based on the above embodiments, the method further includes:
[0070] Classify the text data by content using semantic recognition technology, and at the same time generate the category features corresponding to the text data, and associate the category features with the corresponding video data according to the timestamp; for example, the content of the text data can be classified as discussion, process demonstration, decision-making, introduction, etc., and different contents correspond to different video requirements. For example, for discussion content, the audio information should be presented emphatically, and the resolution and frame rate can be reduced. For process demonstration content, the picture information should be presented emphatically, and it is best to ensure a certain resolution and frame rate at the same time; therefore, the present application takes the content type as one of the factors for regulating the picture quality.
[0071] When a video data acquisition request is received, search for the category features of the video data to be responded to;
[0072] Based on the retrieval service usage information, obtain the historical response times of the video data to be responded to for all users to obtain the importance feature of the video data to be responded to, and obtain the historical response times of the video data to be responded to for all users to obtain the response time feature of the video data to be responded to; there are multiple user terminals in the same video conference, and the retrieval behaviors of these users may have a high degree of repetition. Therefore, the picture quality to be output currently can be determined according to the attention of other users to this historical video in the past time, optimizing the user experience.
[0073] Using the network performance information, user behavior information, video quality information, category features, importance features, and response time features as inputs, obtain the historical video resolution and historical video frame rate through a pre-constructed second parameter adjustment model; the training method of the second parameter adjustment model is similar to that of the first parameter adjustment model;
[0074] Calculate the historical video bit rate according to the historical video resolution and historical video frame rate through the bit rate formula;
[0075] According to the historical video bit rate and bandwidth data, calculate the maximum bit rate of the conference video, that is, the sum of the two video bit rates cannot exceed the bandwidth limit, otherwise the video will freeze;
[0076] According to the maximum bit rate of the conference video, the preset resolution, and the current frame rate, calculate the predicted frame rate of the conference video; the current frame rate is the currently more appropriate frame rate automatically adjusted by the first parameter adjustment model and can be used as a reference for the future frame rate. That is, under the condition of not exceeding the maximum bit rate, it is preferably to maintain the current frame rate, that is, use the current frame rate as the preset frame rate. If the current frame rate cannot be maintained due to the limitation of the maximum bit rate, the preset frame rate is calculated and determined according to the maximum bit rate.
[0077] After the parameters of the conference video and the historical video are determined, the data of the two videos are encoded simultaneously with different parameter settings to improve the encoding and decoding speed and optimize the transmission efficiency.
[0078] When the user conducts a retrieval review, they may miss the content of the current video conference. To avoid missing information transmission, as an optional implementation method, this method further includes step S400:
[0079] While obtaining the audio data of the video conference, record the time stamps corresponding to the audio data, and store the time stamps and the text data after audio conversion in a corresponding manner to obtain a first data set;
[0080] According to the call status of the retrieval service, obtain the time stamp corresponding to the retrieval request and the time stamp when the retrieval service call ends to obtain a start time stamp and an end time stamp;
[0081] Based on the start time stamp and the end time stamp, obtain a query time range, and find the target time stamps falling within the query time range in the first data set in a traversing or filtering manner, and obtain the second target text data corresponding to the target time stamps;
[0082] Generate a text push request based on the second target text data to push the found text to the user; Pseudo-code example: # Get the content missed by the user missed_content = find_missed_contents(user_state, video_text_records) # Push processing if missed_content: video_player.show_overlay("Content you missed during retrieval:", missed_content).
[0083] This step determines a time range according to the user's retrieval start time and exit time, searches for text records within this time range, and automatically pushes the content that the user may have missed after the user's retrieval to avoid information omission.
[0084] As an optional implementation method, in order to further reduce the pressure of data transmission, this method further includes step S500:
[0085] Obtain the video data of the video conference and continuously transmit the video data;
[0086] Based on the video data, determine whether the images of the video conference simultaneously meet a first preset condition and a second preset condition. The first preset condition is that the image is recognized as a presentation, and the second preset condition is that each frame of the image remains unchanged within a preset time range;
[0087] If the video conference image simultaneously meets the first preset condition and the second preset condition, obtain the image data of any frame within a preset time, and replace the video data with the image data for transmission;
[0088] Specifically, it is possible to identify whether the image is a presentation by the program status, video screen, and audio content of sharing the screen. If the current meeting is in the presentation mode and the page remains unchanged for a long time, this static image data of this frame can be temporarily used to replace the video data, that is, changing the video stream to an image stream to reduce the data transmission load and optimize the viewing experience.
[0089] If the video conference image does not simultaneously meet the first preset condition and the second preset condition, that is, the video image has changed, stop the transmission of the image data, and at the same time resume the transmission of the video stream to ensure that the user can immediately see the new content.
[0090] To improve the security of data transmission in the meeting, this application adds Perfect Forward Secrecy (PFS) to the Real-Time Communication (RTC) transmission protocol to ensure that the key exchange mechanism can guarantee that even in the case of long-term key leakage, the past communication content cannot be decrypted. In this embodiment, the Diffie-Hellman key exchange algorithm that supports PFS is selected and meets the low-latency requirements of RTC. The specific operations are as follows:
[0091] When establishing a session, use the Diffie-Hellman algorithm to generate a temporary key pair. The server and the client each keep the private key and exchange the public key through the signaling protocol.
[0092] Both parties calculate the shared key using the exchanged public key and their respective private keys of the hash values, and confirm that the shared keys calculated by both parties are the same by verifying the shared key.
[0093] The server uses the shared key and a random number to generate a session key and sends it to the client for encrypting and decrypting the media stream.
[0094] The server generates a session ID or a session ticket and updates it regularly, and sends the session ID or the session ticket to the terminal connected to it, that is, the client, so that the session ID or the session ticket is regularly updated in the client; at the same time, the expired session ID should be cleared in time to ensure the uniqueness and unpredictability of the session ID to prevent session hijacking.
[0095] The server stores the session key and the session status information in the database. The session status information includes encryption parameters, session ID or session ticket; the session ticket should be encrypted and stored, and signed with the private key of the server to ensure that only the authenticated client can restore the session.
[0096] Obtain the connection status information of the connected terminal, determine whether the connection is disconnected according to the connection status information. If so, receive the latest session ID or session ticket sent by the connected terminal, and verify the validity of the session ID or session ticket. If valid, extract the saved session key from the database; reconnect with the connected terminal based on the extracted session key and return the encrypted session data, avoiding regenerating the session key. By reusing the saved session key, the server and the client avoid re-performing the entire handshake process (including public key exchange and a brand-new key negotiation), which greatly reduces the latency of connection recovery and improves the user experience.
[0097] The present invention improves the data transmission efficiency of video conferencing through the above method, can adapt to different network environments, and at the same time ensures the video quality, providing a stable and smooth meeting experience even during network fluctuations; the optimization of the retrieval service and the push of meeting content during retrieval can effectively improve the efficiency of users obtaining information and the integrity of information, avoid information transmission omissions, and improve the efficiency of remote collaboration.
[0098] Embodiment 2
[0099] See Figure 2 , this application also provides a video conferencing system, including:
[0100] An acquisition module 100, configured to acquire network performance information, user behavior information, video quality information, and retrieval service usage information, where the retrieval service usage information includes the call status of the retrieval service;
[0101] A first processing module 200, configured to use the network performance information, user behavior information, and video quality information as inputs, and obtain a predicted bitrate through a pre-constructed first parameter adjustment model;
[0102] A calculation module 300, configured to determine whether the retrieval service is called according to the call status of the retrieval service. If not, adjust the video parameters according to the predicted bitrate; if so, obtain a preset resolution and a preset frame rate, calculate the video bitrate to obtain a target bitrate, and adjust the video parameters according to the target bitrate.
[0103] As an optional implementation manner, the video conferencing system further includes:
[0104] A second processing module 400, configured to acquire the audio data of the video conference, process the audio data using speech recognition technology, obtain the text data after conversion of the audio data, and store it;
[0105] A third processing module 500, configured to process the text data using semantic recognition technology to obtain segmented text data divided by content;
[0106] A retrieval module 600, configured to receive a retrieval request and obtain retrieval keywords, and retrieve in the segmented text data based on the retrieval keywords to obtain first target text data.
[0107] Embodiment 3
[0108] Corresponding to the above method embodiment, in this embodiment, a video conferencing device is further provided. A video conferencing device described below can be correspondingly referred to the video conferencing method described above.
[0109] Figure 3 It is a block diagram of a video conferencing device 800 shown according to an exemplary embodiment. As Figure 3 shown, the video conferencing device 800 includes a processor 801 and a memory 802. The video conferencing device 800 may further include one or more of a multimedia component 803, an input / output (I / O) interface 804, and a communication component 805. Among them, the processor 801 is used to control the overall operation of the video conferencing device 800 to complete all or part of the steps in the above video conferencing method. The memory 802 is used to store various types of data to support the operation of the video conferencing device 800. These data may include, for example, commands for any application or method operating on the video conferencing device 800, and application-related data, such as contact data, sent and received messages, pictures, audio, video, and so on. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0110] The multimedia component 803 may include a screen and an audio component. Among them, the screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone, and the microphone is used to receive external audio signals.
[0111] The received audio signal can be further stored in the memory 802 or sent via the communication component 805. The audio component also includes at least one speaker for outputting the audio signal. The I / O interface 804 provides an interface between the processor 801 and other interface modules, and the other interface modules may be a keyboard, a mouse, buttons, etc. These buttons can be virtual buttons or physical buttons. The communication component 805 is used for the device 800 of the video conference to communicate with other devices in a wired or wireless manner. The wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G or 4G, or a combination of one or more of them. Accordingly, the communication component 805 may include: a Wi-Fi module, a Bluetooth module, an NFC module.
[0112] In an exemplary embodiment, the device 800 for mutual signature and verification of digital files may be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for performing the above-described method of the video conference.
[0113] In another exemplary embodiment, there is also provided a computer-readable storage medium including program commands, and when the program commands are executed by a processor, the steps of the above-described method of the video conference are implemented. For example, the computer-readable storage medium may be the above-mentioned memory 802 including program commands, and the above program commands may be executed by the processor 801 of the device 800 of the video conference to complete the above-described method of the video conference.
[0114] Embodiment 4
[0115] Corresponding to the above method embodiment of the video conference, in this embodiment, there is also provided a readable storage medium, and a readable storage medium described below can be correspondingly referred to the method of the video conference described above.
[0116] A readable storage medium has a computer program stored thereon, and when the computer program is executed by a processor, the steps of the above-described method embodiment of the video conference are implemented.
[0117] Specifically, the readable storage medium may be various readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc., which can store program codes.
[0118] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0119] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A video conferencing method, characterized in that: include: Acquiring network performance information, user behavior information, video quality information, and retrieval service usage information, wherein the retrieval service usage information includes a call status of the retrieval service; Taking the network performance information, user behavior information and video quality information as input, a predicted bit rate is obtained through a pre-constructed first parameter adjustment model; According to the calling status of the retrieval service, it is determined whether the retrieval service is called. If not, the video parameters are adjusted according to the predicted bit rate. If so, the preset resolution and the preset frame rate are obtained, and the video bit rate is calculated to obtain the target bit rate, and the video parameters are adjusted according to the target bit rate.
2. The video conferencing method according to claim 1, characterized in that: The method further comprises: Acquire audio data of the video conference, perform speech recognition processing on the audio data, obtain text data converted from the audio data and store it; Performing semantic recognition processing on the text data to obtain segmented text data divided by content; A search request is received and a search keyword is obtained, and a search is performed in the segmented text data based on the search keyword to obtain first target text data.
3. The video conferencing method according to claim 2, characterized in that: The method further comprises: While obtaining the audio data of the video conference, recording the timestamp corresponding to the audio data, and storing the timestamp and the converted text data in correspondence to obtain a first data set; According to the calling status of the retrieval service, obtain the timestamp corresponding to the retrieval request and the timestamp when the retrieval service call ends, and obtain the start timestamp and the end timestamp; Obtaining a query time range based on the start timestamp and the end timestamp, searching the first data set for a target timestamp that falls within the query time range, and obtaining second target text data corresponding to the target timestamp; A text push request is generated based on the second target text data.
4. The video conferencing method according to claim 1, characterized in that: The method further comprises: Acquire video data of a video conference and continuously transmit the video data; Based on the video data, determining whether the image of the video conference satisfies a first preset condition and a second preset condition at the same time, wherein the first preset condition is that the image is identified as a presentation, and the second preset condition is that each frame of the image remains unchanged within a preset time range; If the video conference image satisfies both the first preset condition and the second preset condition, then obtaining image data of any frame of image within a preset time, and replacing the video data with the image data for transmission; If the video conference image does not satisfy the first preset condition and the second preset condition at the same time, the transmission of the image data is stopped and the transmission of the video data is resumed.
5. The video conferencing method according to claim 1, characterized in that: The retrieval service usage information also includes a response status of each retrieval result, wherein the retrieval result includes text data, image data, and video data; the obtaining of a preset resolution and a preset frame rate, and calculating a video bit rate to obtain a target bit rate includes: When the search service is in a called state, obtaining a response search result in a responding state according to the search service usage information; Obtain the corresponding preset resolution and preset frame rate according to the response search result; Determine whether there are two or more response search results at the same time. If so, select the preset resolution and preset frame rate corresponding to the response search result with the highest priority in the order of decreasing priority of video data, image data and text data. If not, obtain the preset resolution and preset frame rate corresponding to the only response search result; obtain the target resolution and target frame rate; According to the target resolution and target frame rate, the video bit rate is calculated using the bit rate formula to obtain the target bit rate.
6. The video conferencing method according to claim 5, characterized in that: The network performance information includes bandwidth data, and the video quality information includes a current frame rate; the method further includes: Classify the text data by content through semantic recognition processing, generate category features corresponding to the text data, and associate the category features with the corresponding video data according to the timestamp; When receiving a request for obtaining video data, searching for category features of the video data to be responded to; Based on the retrieval service usage information, the historical response times of the video data to be responded to by all users are obtained to obtain the importance characteristics of the video data to be responded to, and the historical response times of the video data to be responded to by all users are obtained to obtain the response time characteristics of the video data to be responded to; Taking the network performance information, user behavior information, video quality information, category features, importance features, and response time features as input, a historical video resolution and a historical video frame rate are obtained through a pre-constructed second parameter adjustment model; Obtain the historical video bit rate according to the historical video resolution and the historical video frame rate; Calculate the maximum bit rate of the conference video based on historical video bit rate and bandwidth data; According to the highest bit rate, preset resolution and current frame rate of the conference video, the predicted frame rate of the conference video is calculated and used as the preset frame rate.
7. A video conferencing system, characterized in that: include: An acquisition module, used to acquire network performance information, user behavior information, video quality information and retrieval service usage information, wherein the retrieval service usage information includes a call status of the retrieval service; A first processing module, configured to use the network performance information, user behavior information and video quality information as inputs and obtain a predicted bit rate through a pre-built first parameter adjustment model; The calculation module is used to determine whether the retrieval service is called according to the calling status of the retrieval service. If not, the video parameters are adjusted according to the predicted bit rate; if so, the preset resolution and the preset frame rate are obtained, and the video bit rate is calculated to obtain the target bit rate, and the video parameters are adjusted according to the target bit rate.
8. A video conferencing system according to claim 7, characterized in that: Also includes: The second processing module is used to obtain audio data of the video conference, process the audio data using speech recognition technology, obtain text data converted from the audio data and store it; A third processing module is used to process the text data using semantic recognition technology to obtain segmented text data divided by content; The retrieval module is used to receive a retrieval request and obtain a retrieval keyword, and perform a search in the segmented text data based on the retrieval keyword to obtain the first target text data.
9. A video conferencing device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the video conferencing method according to any one of claims 1 to 6 are implemented.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the video conferencing method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Video code rate self-adaptive adjustment method and device and electronic equipment
CN109982118A
Method and device for dynamically regulating and controlling video playing by applying deep search
CN110087110A
Video code rate adjustment method and device, server and storage medium
CN110290402A
Video data transmission method, device and equipment
CN116095324A
Video decision code rate determination method and device, storage medium and electronic device
CN117640920A
Cited By
Network adaptive video conference transmission optimization method and system
CN120358347A
Network-adaptive video conferencing transmission optimization method and system
CN120358347B
Optimized cross-media heterogeneous feature fusion indexing method
CN121117243A