A method, system, device and storage medium of video conferencing

By adjusting the model with pre-trained parameters in real time, the video bitrate and parameters of video conferencing are optimized, which solves the problems of data transmission latency and unstable quality in video conferencing, and achieves more stable video quality and more efficient information transmission.

CN120050384BActive Publication Date: 2025-11-04SICHUAN XINYUNDIAO TECHNOLOGY SERVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510244040.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2024-11-29
Filing Date
2025-03-03
Publication Date
2025-11-04
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

Current video conferencing systems face challenges in terms of data transmission latency and unstable video quality, especially when network conditions fluctuate, affecting real-time performance and the integrity of information transmission.

Method used

The model is adjusted in real time by pre-trained parameters to predict the optimal video bitrate, and video parameters are dynamically adjusted based on user behavior and retrieval service status to optimize data transmission.

Benefits of technology

It improves the smoothness and information integrity of video conferencing, reduces screen stuttering and distortion, and enhances user experience and retrieval service response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050384B_ABST
    Figure CN120050384B_ABST
Patent Text Reader

Abstract

The application provides a video conference method, system, device and storage medium, and relates to the technical field of network communication.The method comprises the following steps: acquiring network performance information, user behavior information, video quality information and search service usage information, wherein the search service usage information comprises a calling state of the search service; taking the network performance information, the user behavior information and the video quality information as inputs, and obtaining a predicted code rate through a pre-constructed first parameter adjustment model; judging whether the search service is called according to the calling state of the search service; if not, adjusting the video parameters according to the predicted code rate; if yes, acquiring a preset resolution and a preset frame rate, calculating a video code rate, obtaining a target code rate, and adjusting the video parameters according to the target code rate.The application predicts the optimal video code rate in the conference process through the pre-trained first parameter adjustment model, and adjusts the video data transmission parameters according to the search behavior of the user, thereby improving the information acquisition efficiency of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network communication technology, and more specifically, to a method, system, device, and storage medium for video conferencing. Background Technology

[0002] With the development of internet technology, video conferencing has become an important tool for remote communication, widely used in business, education, and healthcare. However, current video conferencing systems still face many challenges. First, data transmission latency is a prominent issue. Fluctuations in network conditions can cause audio and video to become out of sync, affecting the real-time nature of the meeting. This latency is particularly problematic in scenarios such as legal or medical consultations, where it can lead to misunderstandings or flawed decision-making. Second, video quality is unstable under different network environments, often resulting in stuttering or distortion due to bandwidth limitations, making it difficult to effectively convey meeting content. Summary of the Invention

[0003] The purpose of this invention is to provide a video conferencing method, system, device, and storage medium to improve the aforementioned problems. To achieve the above objective, the technical solution adopted by this invention is as follows:

[0004] In a first aspect, this application provides a method for video conferencing, comprising:

[0005] Obtain network performance information, user behavior information, video quality information, and retrieval service usage information, wherein the retrieval service usage information includes the retrieval service invocation status;

[0006] Using the network performance information, user behavior information, and video quality information as input, the predicted bitrate is obtained by adjusting the model through a pre-constructed first parameter.

[0007] Based on the call status of the search service, determine whether the search service has been called. If not, adjust the video parameters according to the predicted bitrate. If yes, obtain the preset resolution and preset frame rate, calculate the video bitrate, obtain the target bitrate, and adjust the video parameters according to the target bitrate.

[0008] Secondly, this application also provides a video conferencing system, comprising:

[0009] The acquisition module is used to acquire network performance information, user behavior information, video quality information, and retrieval service usage information, wherein the retrieval service usage information includes the retrieval service invocation status;

[0010] The first processing module is used to take the network performance information, user behavior information and video quality information as input, and adjust the model through the pre-constructed first parameter to obtain the predicted bitrate.

[0011] The calculation module is used to determine whether the retrieval service has been invoked based on the invocation status of the retrieval service. If not, the video parameters are adjusted according to the predicted bitrate. If so, the preset resolution and preset frame rate are obtained, the video bitrate is calculated, the target bitrate is obtained, and the video parameters are adjusted according to the target bitrate.

[0012] Thirdly, this application also provides a video conferencing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the video conferencing method described above.

[0013] Fourthly, this application also provides a readable storage medium storing a computer program that, when executed by a processor, implements the steps of the video conferencing method described above.

[0014] The beneficial effects of this invention are as follows:

[0015] This invention uses a pre-trained model with adjusted first parameters to predict the optimal video bitrate in real time during a meeting, reducing stuttering or distortion and ensuring video quality. It also adjusts video data transmission parameters based on user search behavior during the meeting, improving the response speed of the search service and ensuring the smoothness of the meeting video, thus guaranteeing the completeness of information obtained by the user.

[0016] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a video conferencing method in an embodiment of this application;

[0019] Figure 2 This is a schematic diagram of the system structure of video conferencing in an embodiment of this application;

[0020] Figure 3 This is a schematic diagram of the device structure for video conferencing in an embodiment of this application.

[0021] The diagram is labeled as follows: 100 - Acquisition module; 200 - First processing module; 300 - Calculation module; 400 - Second processing module; 500 - Third processing module; 600 - Retrieval module; 800 - Video conferencing equipment; 810 - Processor; 820 - Memory; 803 - Multimedia component; 804 - I / O interface; 805 - Communication component. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0023] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0024] Example 1

[0025] See Figure 1 This embodiment provides a video conferencing method, including steps S100, S200, and S300. Additionally, this method requires a pre-trained first parameter adjustment model, the construction steps of which are as follows:

[0026] Step S001: Data collection and preparation;

[0027] Collect the following types of data for model training:

[0028] Network performance data: bandwidth (must be changed frequently to reflect the actual network conditions), latency, packet loss rate, etc.

[0029] User behavior data: such as whether a user is "watching", the on / off status of the user's camera and microphone, etc.;

[0030] Video quality data: video resolution (in viewing mode), frame rate, video bitrate (optimal video transmission bitrate under various network conditions);

[0031] Timestamp data: Records the timestamps corresponding to the collected data.

[0032] After collecting and recording the above data from historical meetings, a time-series dataset was created. The data was then integrated and formatted to ensure integrity, removing duplicates and outliers to form a structured dataset.

[0033] Step S002: Data preprocessing;

[0034] Feature data such as bandwidth, latency, and packet loss rate are normalized to the range of 0-1. One-hot encoding is used to process categorical feature data such as video resolution and frame rate, transforming them into numerical feature data.

[0035] Generate feature vectors based on the processed features, and combine the required information into the input format:

[0036] The pythonfeatures = [bandwidth_normalized, latency_normalized, packet_loss_normalized, resolution_one_hot, frame_rate_one_hot, user_state_encoded] outputs the optimal video bitrate.

[0037] The integrated dataset was divided into a training set (70%), a validation set (15%), and a test set (15%). Random partitioning was ensured to guarantee that the distribution of each subset was as similar as possible.

[0038] Step S003: Model building and training;

[0039] Using LSTM models for time series forecasting.

[0040] Input layer: Receives the feature vectors after the above processing.

[0041] LSTM layer: A deep LSTM network is built to process time series data. Multiple layers of LSTM can be selected to improve the model's predictive ability.

[0042] Output layer: One node that outputs the predicted video bitrate.

[0043] Compile the model: Choose an appropriate loss function and optimizer. Example code: pythonmodel.compile(loss='mean_squared_error', optimizer='adam').

[0044] Training the model: Input the training set into the model for training, and use the validation set to evaluate the model's performance. Parameter settings: batch size (e.g., 32) for each training session, number of training epochs (e.g., 50). Example code: `python model.fit(X_train, y_train, validation_data = (X_val, y_val), epochs = 50, batch_size = 32)`.

[0045] Model evaluation: After training, the model performance is evaluated using a test set, and the accuracy of the predictions is calculated. The metrics of interest include: mean squared error (MSE) and root mean square error (RMSE).

[0046] Step S004: Model optimization;

[0047] Early stopping is employed to stop training when the validation loss no longer decreases, preventing overfitting. Hyperparameter tuning is then performed using grid search or random search to optimize hyperparameters such as the learning rate and the number of LSTM layers, ultimately yielding the first parameter-adjusted model.

[0048] Step S100: Obtain network performance information, user behavior information, video quality information, and retrieval service usage information, wherein the retrieval service usage information includes the retrieval service invocation status;

[0049] Network performance information includes current bandwidth, latency, and packet loss rate, which can be collected periodically using monitoring tools; user behavior information includes the user's "viewing" behavior, the on / off status of the user's camera and microphone, and other user operations on the interface.

[0050] Video quality information includes the resolution and frame rate of the current meeting video;

[0051] The information used by the retrieval service includes the call status of the retrieval service, i.e. whether it has been called, and the response status of each type of retrieval result, which includes text data, image data and video data.

[0052] Step S200: Using the network performance information, user behavior information, and video quality information as input, the predicted bitrate is obtained by adjusting the model through the pre-constructed first parameter.

[0053] After processing the acquired data according to step S002 above, the data is input into the first parameter adjustment model to obtain the predicted bitrate. This step predicts the video bitrate in real time based on the current network, user behavior and video quality, thereby determining the optimal encoding and decoding parameters, which can effectively improve the data transmission efficiency of video conferencing, reduce screen stuttering or distortion, and ensure video quality.

[0054] When disputes or conflicts arise in data transmission, traditional systems typically lack intelligent mediation mechanisms, resulting in slow processing and reliance on manual intervention, which impacts decision-making efficiency. However, the first parameter adjustment model of this application can achieve intelligent mediation by dynamically adjusting the encoding based on the predicted optimal bit rate.

[0055] The method in this application is based on a video conferencing program with a retrieval function. This retrieval function allows users to search for historical information about the meeting, which includes at least text recordings and video of the meeting. Therefore, this method may also include the following steps:

[0056] Acquire audio data from video conferences, perform speech recognition processing on the audio data, obtain the converted text data from the audio data, and store it.

[0057] The text data is subjected to semantic recognition processing to obtain segmented text data divided by content; in this step, the text can be pre-corrected by a backend text proofreading algorithm to ensure the accuracy of the text record.

[0058] The system receives a search request and obtains search keywords. Based on these keywords, it performs a search within the segmented text data to obtain the first target text data. For example, when a user searches for "project progress," the system uses this as a keyword to search the database for segmented text data containing "project progress," presenting the retrieved text and its corresponding timestamp to the user. Furthermore, links can be established between the text and corresponding video clips based on the timestamps, allowing users to directly jump to the corresponding audio / video playback location by clicking on a text clip, thus enhancing the efficiency of information retrieval.

[0059] Step S300: Based on the call status of the search service, determine whether the search service has been called. If not, adjust the video parameters according to the predicted bitrate. If yes, obtain the preset resolution and preset frame rate, calculate the video bitrate, obtain the target bitrate, and adjust the video parameters according to the target bitrate.

[0060] The retrieval service is a feature already developed in existing video conferencing software. The server stores the data of the current video conference, and while the user is participating in the video conference, they can use keywords to retrieve information from previous video conferences, obtain historical text and video from the server, and thus review the conference content.

[0061] When a user performs a search, the current meeting video needs to continue playing, and historical data also needs to be retrieved from the server, which further increases the pressure on network transmission. In reality, when users are searching and reviewing, the current meeting video is not their focus, and they may not even pay attention to it. Therefore, the video quality can be appropriately reduced at this time to reduce the pressure on network transmission and improve the response speed of the search service.

[0062] The quality of the meeting video can be reasonably reduced by setting a lower resolution and frame rate, such as setting the resolution to 480p or 360p. The required bitrate can then be calculated directly based on the selected resolution and frame rate using the bitrate formula: [Bitrate = Resolution × Frame Rate × Color Depth].

[0063] To more flexibly adjust image quality and optimize the meeting experience, image quality can also be adjusted based on user search behavior. As an optional implementation method, obtaining a preset resolution and preset frame rate, and calculating the video bitrate to obtain the target bitrate includes:

[0064] When the retrieval service is in the invoked state, the response retrieval results in the response state are obtained based on the retrieval service usage information, that is, the content that the user needs to review, such as text, audio or video.

[0065] The preset resolution and preset frame rate are obtained based on the response search results. The transmission of different search results requires different amounts of bandwidth. For the transmission of complex data, the current picture quality needs to be adjusted to a lower level. Therefore, when reviewing historical videos, the resolution and frame rate of the current meeting video should be adjusted to a lower level.

[0066] Determine whether there are two or more response search results at the same time. If so, select the preset resolution and preset frame rate corresponding to the response search result with the highest priority in descending order of priority, such as video data, image data, and text data. If not, obtain the preset resolution and preset frame rate corresponding to the unique response search result.

[0067] Different types of search results correspond to different preset resolutions and frame rates. When multiple search results respond simultaneously, the resolution and frame rate corresponding to the video data are given priority, followed by the image data (or audio data); that is, the target resolution and target frame rate are obtained.

[0068] The target bitrate is obtained by calculating the video bitrate using the bitrate formula based on the target resolution and target frame rate.

[0069] As an optional implementation, based on the above embodiments, the method further includes:

[0070] This application employs semantic recognition technology to classify text data according to its content, generates category features corresponding to the text data, and associates these category features with the corresponding video data based on timestamps. For example, the content of the text data can be classified as discussion, process demonstration, decision-making, introduction, etc. Different content corresponds to different video requirements. For example, for discussion content, the focus should be on presenting audio information, and the resolution and frame rate can be reduced. For process demonstration content, the focus should be on presenting visual information, and it is best to maintain a certain resolution and frame rate at the same time. Therefore, this application considers content type as one of the factors for controlling picture quality.

[0071] When a request to retrieve video data is received, the category characteristics of the video data to be responded to are searched.

[0072] Based on the information obtained from the retrieval service, the historical number of responses to the video data to be responded to by all users is obtained, thus obtaining the importance characteristics of the video data to be responded to. The historical response time of the video data to be responded to by all users is also obtained, thus obtaining the response time characteristics of the video data to be responded to. There are multiple user terminals in the same video conference, and the retrieval behavior of these users may have a high degree of repetition. Therefore, the current output picture quality can be determined based on the attention of other users to this historical video in the past time, thereby optimizing the user experience.

[0073] Using the network performance information, user behavior information, video quality information, category features, importance features, and response time features as input, the historical video resolution and historical video frame rate are obtained through a pre-constructed second parameter adjustment model; the training method of the second parameter adjustment model is similar to that of the first parameter adjustment model.

[0074] The historical video bitrate is calculated using the bitrate formula based on the historical video resolution and historical video frame rate.

[0075] Based on historical video bitrate and bandwidth data, calculate the maximum bitrate of the meeting video. That is, the sum of the two video bitrates cannot exceed the bandwidth limit, otherwise the video will stutter.

[0076] The predicted frame rate of the conference video is calculated based on the highest bitrate, preset resolution, and current frame rate. The current frame rate is a suitable frame rate that is automatically adjusted by the first parameter adjustment model and can be used as a reference for the future frame rate. That is, under the condition of not exceeding the highest bitrate, it is preferable to maintain the current frame rate, that is, the current frame rate can be used as the preset frame rate. If the current frame rate cannot be maintained due to the limitation of the highest bitrate, the preset frame rate is calculated and determined based on the highest bitrate.

[0077] Once the parameters for both the conference video and the historical video are determined, the data from both videos are encoded simultaneously with different parameter settings to improve encoding and decoding speed and optimize transmission efficiency.

[0078] When users are reviewing a video conference, they may miss the content of the current video conference. To avoid missing information, as an optional implementation method, this method also includes step S400:

[0079] While acquiring the audio data from the video conference, the timestamps corresponding to the audio data are recorded, and the timestamps are stored in correspondence with the text data converted from the audio data to obtain the first dataset;

[0080] Based on the call status of the search service, obtain the timestamp corresponding to the search request and the timestamp at which the search service call ended, and get the start timestamp and end timestamp.

[0081] Based on the start and end timestamps, the query time range is obtained. The target timestamp falling within the query time range is searched in the first dataset by traversal or filtering, and the second target text data corresponding to the target timestamp is obtained.

[0082] A text push request is generated based on the second target text data to push the searched text to the user; Pseudocode example: # Get the content missed by the user missed_content = find_missed_contents(user_state, video_text_records) # Push processing if missed_content: video_player.show_overlay("The content you missed during the search:", missed_content).

[0083] This step determines a time range based on the user's search start and exit times, searches for text records within that time range, and automatically pushes content that the user might have missed after the search, thus avoiding information omissions.

[0084] As an optional implementation, to further reduce the pressure on data transmission, this method also includes step S500:

[0085] Acquire video data from the video conference and continuously transmit the video data;

[0086] Based on the video data, it is determined whether the images of the video conference simultaneously meet the first preset condition and the second preset condition. The first preset condition is that the image is recognized as a presentation, and the second preset condition is that each frame of the image remains unchanged within a preset time range.

[0087] If the video conference image simultaneously meets the first preset condition and the second preset condition, then the image data of any frame within a preset time period is obtained, and the video data is replaced with the image data for transmission.

[0088] Specifically, the program status, video feed, and audio content of the shared screen can be used to identify whether an image is a presentation. If the current meeting is in presentation mode and the page remains unchanged for a long time, the video data can be temporarily replaced with this frame of static image data, that is, the video stream can be converted into an image stream to reduce the data transmission load and optimize the viewing experience.

[0089] If the video conference image does not simultaneously meet the first and second preset conditions, i.e., the video image has changed, then image data transmission is stopped, while video stream transmission is resumed to ensure that users can immediately see the new content.

[0090] To enhance the security of data transmission during meetings, this application incorporates Perfect ForwardSecrecy (PFS) into the Real-Time Communication (RTC) transmission protocol. This ensures that the key exchange mechanism guarantees that past communications cannot be decrypted even in the event of long-term key leakage. This embodiment selects the Diffie-Hellman key exchange algorithm, which supports PFS and meets the low-latency requirements of RTC. The specific operation is as follows:

[0091] When establishing a session, a temporary key pair is generated using the Diffie-Hellma algorithm. The server and client each retain their private keys and exchange public keys through a signaling protocol.

[0092] Both parties use the exchanged public key and the private key of each hash value to calculate the shared key. By verifying the shared key, they confirm that the shared key calculated by both parties is the same.

[0093] The server uses a shared key and a random number to generate a session key and sends it to the client for encrypting and decrypting media streams.

[0094] The server generates session IDs or session tickets and updates them periodically, sending the session IDs or session tickets to the connected terminals, i.e., clients, so that the session IDs or session tickets are updated periodically in the clients; at the same time, expired session IDs should be cleaned up in a timely manner to ensure the uniqueness and unpredictability of session IDs and prevent session hijacking.

[0095] The server stores the session key and session state information in the database. The session state information includes encryption parameters, session ID, or session ticket. The session ticket should be stored in an encrypted manner and signed using the server's private key to ensure that only authenticated clients can resume the session.

[0096] The system acquires connection status information from the connected terminal and determines whether the connection is broken based on this information. If so, it receives the latest session ID or session ticket from the connected terminal and verifies the validity of the session ID or session ticket. If valid, it retrieves the saved session key from the database. Based on the retrieved session key, it reconnects with the connected terminal and returns encrypted session data, avoiding the need to regenerate the session key. By reusing the saved session key, the server and the client avoid re-performing the entire handshake process (including public key exchange and new key negotiation), which greatly reduces connection recovery latency and improves user experience.

[0097] This invention improves the data transmission efficiency of video conferencing through the above-described method, adapts to different network environments, and ensures video quality, providing a stable and smooth conferencing experience even during network fluctuations. The optimization of the retrieval service and the push of meeting content during retrieval can effectively improve the efficiency and completeness of information acquisition for users, avoid information omissions, and improve the efficiency of remote collaboration.

[0098] Example 2

[0099] See Figure 2 This application also provides a video conferencing system, comprising:

[0100] The acquisition module 100 is used to acquire network performance information, user behavior information, video quality information, and retrieval service usage information, wherein the retrieval service usage information includes the retrieval service call status;

[0101] The first processing module 200 is used to take the network performance information, user behavior information and video quality information as input, and adjust the model through a pre-constructed first parameter to obtain the predicted bitrate.

[0102] The calculation module 300 is used to determine whether the retrieval service has been called based on the call status of the retrieval service. If not, the video parameters are adjusted according to the predicted bitrate. If so, the preset resolution and preset frame rate are obtained, the video bitrate is calculated, the target bitrate is obtained, and the video parameters are adjusted according to the target bitrate.

[0103] As an optional implementation, the video conferencing system further includes:

[0104] The second processing module 400 is used to acquire audio data from the video conference, process the audio data using speech recognition technology, obtain text data after audio data conversion, and store it.

[0105] The third processing module 500 is used to process the text data using semantic recognition technology to obtain segmented text data divided by content.

[0106] The retrieval module 600 is used to receive retrieval requests and obtain retrieval keywords, and to perform retrieval in the segmented text data based on the retrieval keywords to obtain the first target text data.

[0107] Example 3

[0108] Corresponding to the above method embodiments, this embodiment also provides a video conferencing device. The video conferencing device described below can be referred to in correspondence with the video conferencing method described above.

[0109] Figure 3 This is a block diagram illustrating a video conferencing device 800 according to an exemplary embodiment. Figure 3 As shown, the video conferencing device 800 includes a processor 801 and a memory 802. The video conferencing device 800 may also include one or more of a multimedia component 803, an input / output (I / O) interface 804, and a communication component 805. The processor 801 controls the overall operation of the video conferencing device 800 to complete all or part of the steps in the video conferencing method described above. The memory 802 stores various types of data to support the operation of the video conferencing device 800. This data may include, for example, commands for any application or method operating on the video conferencing device 800, and application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0110] Multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals.

[0111] The received audio signal can be further stored in memory 802 or transmitted via communication component 805. The audio component also includes at least one speaker for outputting audio signals. I / O interface 804 provides an interface between processor 801 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons can be virtual or physical. Communication component 805 is used for wired or wireless communication between the video conferencing device 800 and other devices. Wireless communication includes, for example, Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof; therefore, the corresponding communication component 805 may include a Wi-Fi module, a Bluetooth module, or an NFC module.

[0112] In an exemplary embodiment, the digital document mutual signing and verification device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the video conferencing method described above.

[0113] In another exemplary embodiment, a computer-readable storage medium including program commands is also provided, which, when executed by a processor, implement the steps of the video conferencing method described above. For example, the computer-readable storage medium may be the memory 802 including program commands described above, which may be executed by the processor 801 of the video conferencing device 800 to complete the video conferencing method described above.

[0114] Example 4

[0115] Corresponding to the above video conferencing method embodiments, this embodiment also provides a readable storage medium. The readable storage medium described below can be referred to in conjunction with the video conferencing method described above.

[0116] A readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the video conferencing method embodiments described above.

[0117] The readable storage medium can specifically be a USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, or any other readable storage medium capable of storing program code.

[0118] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0119] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for video conferencing, characterized in that, include: Obtain network performance information, user behavior information, video quality information, and retrieval service usage information, wherein the retrieval service usage information includes the retrieval service invocation status; Using the network performance information, user behavior information, and video quality information as input, the predicted bitrate is obtained by adjusting the model through a pre-constructed first parameter. Based on the call status of the search service, determine whether the search service has been called. If not, adjust the video parameters according to the predicted bitrate. If yes, obtain the preset resolution and preset frame rate, calculate the video bitrate, obtain the target bitrate, and adjust the video parameters according to the target bitrate. The retrieval service usage information also includes the response status of each retrieval result, which includes text data, image data, and video data; the process of obtaining the preset resolution and preset frame rate, and calculating the video bitrate to obtain the target bitrate includes: When the search service is in the invoked state, retrieve the response search results in the response state based on the search service usage information; Based on the response search results, obtain the corresponding preset resolution and preset frame rate; Determine whether two or more response search results exist simultaneously. If so, select the preset resolution and preset frame rate corresponding to the response search result with the highest priority in descending order of priority: video data, image data, and text data. If not, obtain the preset resolution and preset frame rate corresponding to the unique response search result; and obtain the target resolution and target frame rate. The target bitrate is obtained by calculating the video bitrate using the bitrate formula based on the target resolution and target frame rate.

2. The video conferencing method according to claim 1, characterized in that, The method further includes: Acquire audio data from video conferences, perform speech recognition processing on the audio data, obtain the converted text data from the audio data, and store it. The text data is subjected to semantic recognition processing to obtain segmented text data divided by content; The system receives a search request and obtains search keywords. Based on the search keywords, it performs a search on the segmented text data to obtain the first target text data.

3. The video conferencing method according to claim 2, characterized in that, The method further includes: While acquiring the audio data from the video conference, the timestamps corresponding to the audio data are recorded, and the timestamps are stored in correspondence with the converted text data to obtain the first dataset; Based on the call status of the search service, obtain the timestamp corresponding to the search request and the timestamp at which the search service call ended, and get the start timestamp and end timestamp. Based on the start and end timestamps, the query time range is obtained. The target timestamp falling within the query time range is found in the first dataset, and the second target text data corresponding to the target timestamp is obtained. A text push request is generated based on the second target text data.

4. The video conferencing method according to claim 1, characterized in that, The method further includes: Acquire video data from the video conference and continuously transmit the video data; Based on the video data, it is determined whether the images of the video conference simultaneously meet the first preset condition and the second preset condition. The first preset condition is that the image is recognized as a presentation, and the second preset condition is that each frame of the image remains unchanged within a preset time range. If the video conference image simultaneously meets the first preset condition and the second preset condition, then the image data of any frame within a preset time period is obtained, and the video data is replaced with the image data for transmission. If the video conference image does not simultaneously meet the first and second preset conditions, image data transmission will be stopped, while video data transmission will be resumed.

5. The video conferencing method according to claim 1, characterized in that, The network performance information includes bandwidth data, and the video quality information includes the current frame rate; the method further includes: The text data is classified according to its content through semantic recognition processing, generating category features corresponding to the text data, and then associating the category features with the corresponding video data based on the timestamp. When a request to retrieve video data is received, the category characteristics of the video data to be responded to are searched. Based on the information used by the retrieval service, the historical number of responses to the video data to be responded to by all users is obtained, and the importance feature of the video data to be responded to is obtained. The historical response time of the video data to be responded to by all users is obtained, and the response time feature of the video data to be responded to is obtained. Using the network performance information, user behavior information, video quality information, category features, importance features, and response time features as inputs, the historical video resolution and historical video frame rate are obtained by adjusting the pre-built second parameter model. The historical video bitrate is obtained based on the historical video resolution and historical video frame rate. Calculate the highest bitrate of the conference video based on historical video bitrate and bandwidth data; The predicted frame rate of the conference video is calculated based on the highest bitrate, preset resolution, and current frame rate, and is used as the preset frame rate.

6. A video conferencing system, characterized in that, include: The acquisition module is used to acquire network performance information, user behavior information, video quality information, and retrieval service usage information, wherein the retrieval service usage information includes the retrieval service invocation status; The first processing module is used to take the network performance information, user behavior information and video quality information as input, and adjust the model through the pre-constructed first parameter to obtain the predicted bitrate. The calculation module is used to determine whether the retrieval service has been invoked based on its invocation status. If not, it adjusts the video parameters according to the predicted bitrate; if so, it obtains the preset resolution and preset frame rate, calculates the video bitrate to obtain the target bitrate, and adjusts the video parameters according to the target bitrate. The retrieval service usage information also includes the response status of each retrieval result, and the retrieval results include text data, image data, and video data. Obtaining the preset resolution and preset frame rate, and calculating the video bitrate to obtain the target bitrate includes: When the search service is in the invoked state, retrieve the response search results in the response state based on the search service usage information; Based on the response search results, obtain the corresponding preset resolution and preset frame rate; Determine whether two or more response search results exist simultaneously. If so, select the preset resolution and preset frame rate corresponding to the response search result with the highest priority in descending order of priority: video data, image data, and text data. If not, obtain the preset resolution and preset frame rate corresponding to the unique response search result; and obtain the target resolution and target frame rate. The target bitrate is obtained by calculating the video bitrate using the bitrate formula based on the target resolution and target frame rate.

7. A video conferencing system according to claim 6, characterized in that, Also includes: The second processing module is used to acquire audio data from the video conference, process the audio data using speech recognition technology, obtain the converted text data from the audio data, and store it. The third processing module is used to process the text data using semantic recognition technology to obtain segmented text data divided by content. The retrieval module is used to receive retrieval requests and obtain retrieval keywords, and to perform retrieval in the segmented text data based on the retrieval keywords to obtain the first target text data.

8. A video conferencing device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the video conferencing method as described in any one of claims 1 to 5.

9. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the video conferencing method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Video transmission method and device, equipment and storage medium

    CN118175356A