Method, device and medium for locating network video conference lagging reason
By adding echo codes during the audio and video data encoding process and transmitting the decoding results back, the latency and accuracy issues in locating the cause of video conferencing lag were resolved. This enabled rapid and accurate identification of the cause of lag, improving meeting efficiency and stability.
Patent Information
- Application Number
- CN202410785256.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-17
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-06-17
AI Technical Summary
Existing technologies suffer from delays in locating the cause of network video conferencing lag, inability to accurately pinpoint the nature of network problems, and time wasted by the host in confirming network conditions, which also affects the meeting atmosphere.
Echo codes are added during the audio and video data encoding process to generate audio and video files. After decoding by the user client, the echo codes are sent back to the host client. The host client determines the network status based on the echo codes and quickly locates the cause of the stuttering.
It enables quick and accurate identification of the cause of buffering in online video conferencing, reducing wasted time, ensuring the normal progress of the meeting, and improving meeting efficiency and stability.
Smart Images

Figure CN118827916B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of data processing, and particularly relates to a method and device for positioning the cause of network video conference lag, equipment and medium. BACKGROUND
[0002] In the process of network video conference or video live broadcast, video lag or audio lag problems are difficult to avoid due to network fluctuations or different devices. When such a situation is encountered, the video conference host needs to ask the participants from time to time whether they can hear the sound or see the video image to determine which side the network fluctuation occurs and make timely adjustments to ensure the quality of the video conference. Therefore, it is very important to determine where the network video conference lag problem occurs.
[0003] In the existing method for positioning the cause of network video conference lag, the host needs to interrupt the conference to ask the participants for feedback information, which not only wastes time but also affects the atmosphere and progress of the entire conference to some extent. Therefore, a method is needed that can enable the host to discover network lag problems and accurately locate the cause in the process of network video conference. SUMMARY
[0004] The embodiments of the present application provide a method and device for positioning the cause of network video conference lag, which can enable the host to discover network lag problems and accurately locate the cause in the process of network video conference.
[0005] In a first aspect, the embodiments of the present application provide a method for positioning the cause of network video conference lag, which is applied to a host client. The method can include:
[0006] Adding echo codes every preset addition time in the encoding process of audio and video data to generate an audio and video file, wherein the audio and video file includes first video echo codes and first audio echo codes;
[0007] Continuously sending the audio and video file to at least one user client, so that the at least one user client sends second video echo codes and second audio echo codes to the host client after playing the audio and video file, wherein the second video echo codes are video echo codes generated by the user client after decoding the audio and video file, and the second audio echo codes are audio echo codes generated by the user client after decoding the audio and video file;
[0008] Determining the network status of each user client in a first preset time period according to the first video echo codes and first audio echo codes sent to each user client in the first preset time period, and the second video echo codes and second audio echo codes sent back by each user client.
[0009] According to the network status of each user client in the first preset time period, the reason for the freezing of the network video conference is determined.
[0010] In a second aspect, an embodiment of the present application provides a method for positioning a reason for freezing of a network video conference, applied to a user client, which can include:
[0011] receiving an audio and video file sent by a host client, the audio and video file including a first video echo code and a first audio echo code;
[0012] decoding the audio and video file to obtain the first video echo code and the first audio echo code;
[0013] generating a second video echo code and a second audio echo code according to the first video echo code and the first audio echo code;
[0014] sending the second video echo code and the second audio echo code to the host client.
[0015] In a third aspect, an embodiment of the present application provides a positioning device for a reason for freezing of a network video conference, applied to a host client, which can include:
[0016] a generating module, configured to add an echo code in an encoding process of audio and video data every preset adding time to generate an audio and video file, the audio and video file including a first video echo code and a first audio echo code;
[0017] a first sending module, configured to continuously send the audio and video file to at least one user client, so that the at least one user client sends a second video echo code and a second audio echo code to the host client after playing the audio and video file, the second video echo code being a video echo code generated by the user client after decoding the audio and video file, and the second audio echo code being an audio echo code generated by the user client after decoding the audio and video file;
[0018] a first determining module, configured to determine network status of each user client in a first preset time period according to the first video echo code and the first audio echo code sent to each user client in the first preset time period, and the second video echo code and the second audio echo code sent back by each user client;
[0019] a second determining module, configured to determine a reason for freezing of the network video conference according to the network status of each user client in the first preset time period.
[0020] In a fourth aspect, an embodiment of the present application provides a positioning device for a reason for freezing of a network video conference, applied to a user client, which can include:
[0021] receive a video file and an audio file sent by a host client, the video file comprising a first video echo code and the audio file comprising a first audio echo code;
[0022] decode the video file and the audio file to obtain the first video echo code and the first audio echo code;
[0023] generate a second video echo code and a second audio echo code according to the first video echo code and the first audio echo code;
[0024] send the second video echo code and the second audio echo code to the host client.
[0025] In a fifth aspect, an electronic device is provided, and the device comprises:
[0026] a processor;
[0027] a memory for storing processor-executable instructions;
[0028] The processor is configured to execute the instructions to implement the method for locating the cause of the network video conference lag as shown in any one of the embodiments of the first aspect and the second aspect.
[0029] In a sixth aspect, an electronic device is provided, and the device comprises:
[0030] In a seventh aspect, a computer program product is provided, and the computer program product comprises a computer program stored in a readable storage medium, and at least one processor of the device reads and executes the computer program from the storage medium, so that the device implements the method for locating the cause of the network video conference lag as shown in any one of the embodiments of the first aspect and the second aspect.
[0031] The embodiments of the present application provide a method, device, and medium for locating the cause of the network video conference lag, and the present application has the following beneficial effects compared with the prior art:
[0032] The network video conference lag reason positioning method, device, equipment and medium provided by the embodiment of the application, the host client adds echo code in the encoding process of audio and video data every preset adding time, generates an audio and video file, and continuously sends the audio and video file to at least one user client. After the first video echo code and the first audio echo code in the audio and video file are decoded at the user client, the second video echo code and the second audio echo code are generated and returned to the host client. The host client can determine the network state of each user client by matching the first video echo code and the first audio echo code with the second video echo code and the second audio echo code, and finally quickly determine the lag reason of the network video conference according to the network state of each user client.
[0033] Therefore, by matching the first video echo code and the first audio echo code with the second video echo code and the second audio echo code, the network state of each user client can be quickly determined, and the number of user clients in different network states can accurately determine the lag reason of the network video conference, reducing the waste of time in the network video conference and ensuring the normal progress of the conference process. BRIEF DESCRIPTION OF DRAWINGS
[0034] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings needed to be used in the embodiments of the application will be briefly introduced. Those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0035] Figure 1 is a flowchart of a network video conference lag reason positioning method provided by the embodiment of the application;
[0036] Figure 2 is a transmission diagram of an echo code provided by the embodiment of the application;
[0037] Figure 3 is a diagram for judging the network state of an echo code provided by the embodiment of the application;
[0038] Figure 4 is a user abnormal state distribution diagram provided by the embodiment of the application;
[0039] Figure 5 is another user abnormal state distribution diagram provided by the embodiment of the application;
[0040] Figure 6 is another user abnormal state distribution diagram provided by the embodiment of the application;
[0041] Figure 7 is a flowchart of another network video conference lag reason positioning method provided by the embodiment of the application;
[0042] Figure 8 This is a structural diagram of a device for locating the cause of network video conferencing freezes, provided by an embodiment of the present application;
[0043] Figure 9 This is a structural diagram of another device for locating the cause of network video conferencing freezes provided by an embodiment of the present application;
[0044] Figure 10 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.
[0046] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0047] It will be apparent to those skilled in the art that various modifications and variations can be made in this application without departing from the spirit or scope of this application. Therefore, this application is intended to cover modifications and variations of this application that fall within the scope of the corresponding claims (technical solutions claimed for protection) and their equivalents. It should be noted that the embodiments provided in the examples of this application can be combined with each other without contradiction.
[0048] According to the background technology, the existing technology has the following technical problems:
[0049] 1. There is a delay in discovering the problem
[0050] When the audience realizes the problem and thinks of using the public electronic board to feedback, it needs a period of time, and it also needs a period of time to find the message on the electronic board during the speech of the host, and it needs a period of time to determine the nature of the problem according to the number of people who leave messages, which causes the problem to be found late and the nature of the problem to be found out late, which seriously affects the conference.
[0051] 2. Cannot help the host determine whether the network problem is group or individual
[0052] It is particularly important to understand the nature and impact of the problem, because it will determine what the host should do. If it is the host's network problem, the host needs to adjust the network or restart the application, and the audience's problem should be handled by the audience themselves.
[0053] If the host cannot accurately determine whether it is his own problem or the problem of others, the host will have to spend some time to investigate and confirm in order to find out who has a problem, and even the person who has a problem does not handle it, but the person who does not have a problem handles it, which can easily affect the conference effect.
[0054] 3. When the user end does not cooperate in feedback in time, the host end cannot know the network status
[0055] If the host sends an oral or written instruction to all users to require them to feedback, if everyone does not actively cooperate, even if the network environment is not a problem, the host cannot know the network status clearly.
[0056] 4. The host always confirms the network status, which not only wastes time, but also affects the conference atmosphere
[0057] If the host spends a lot of time confirming the network, it will inevitably affect the conference process. At the same time, it is easy to interrupt the theme in the speech and affect the conference effect.
[0058] In order to solve the problems in the prior art, the embodiment of the application provides a network video conference lag reason positioning method, device, equipment and medium. First, the network video conference lag reason positioning method provided by the embodiment of the application is introduced.
[0059] Figure 1 The flowchart of the network video conference lag reason positioning method provided by the embodiment of the application is shown. As shown in Figure 1 The method is applied to the host client, and the method can specifically include the following steps:
[0060] S100, adding echo code in the encoding process of the audio and video data every preset adding time, generating an audio and video file, wherein the audio and video file comprises first video echo code and first audio echo code;
[0061] S200, continuously sending the audio and video file to at least one user client, so that the at least one user client sends second video echo code and second audio echo code to the host client after playing the audio and video file, wherein the second video echo code is video echo code generated by the user client after decoding the audio and video file, and the second audio echo code is audio echo code generated by the user client after decoding the audio and video file;
[0062] Optionally, in the embodiment of the present application, as shown in Figure 2 The so-called echo code refers to adding a certain amount of echo code every time interval (i.e., preset adding time) during the encoding of video and audio data, while normal audio and video encoding is performed. The present application does not limit the specific encoding format of the echo code, which can only be a piece of non-repetitive and randomly changing code representing the sending time. For example:
[0063] 1) using numbers to represent time: #S#HSM#-20231019154800123
[0064] 2) using random characters: #SEND#HSM#-whj93k49jg2k1l0
[0065] The echo code is mixed into the audio and video file, and the audio and video file is continuously transmitted from the host client to the user client. After receiving the audio and video file, the user client extracts the decoded echo code and processes it separately during the decoding process in the user client. In addition to the echo code, the audio and video file continues to be played normally. Since the user client can correctly decompress the echo code, it has proved that it can normally receive a batch of audio and video files, and then send the echo code back to the host client in a specific format, such as:
[0066] 1) using numbers to represent time: #USER-1-RECEIVE#HSM#-20231019154800123
[0067] 2) using random characters: #USER-1-RECEIVE#HSM#-whj93k49jg2k1l0
[0068] Since the echo code is continuously sent back from the user client to the host client in batches, the host client can determine the user receiving status based on the echo code. It should be noted that the user client must send the echo code back to the host client after playing the audio and video in the client.
[0069] It should be noted that even if the user closes the microphone and camera, the echo code is still transmitted to the host end in the audio and video channels. The user does not transmit audio and video data to the host end, but as long as the audio and video data from the host can be normally decoded and played, the echo code from the host end can be obtained. As long as the echo code is obtained, it must be returned to the host.
[0070] Regarding the change of the host, in some conference processes, the host may be changed according to the needs of the conference. When the host is changed, the echo code function will start again with the new host as the center, and the statistical information and judgment results will be displayed to the new host. The host change condition is whether the host voluntarily authorizes the host role to other users. Only when the host manually authorizes the conversion, the echo code data will start to be calculated again with the new host.
[0071] Optionally, in a specific implementation of the present application, first, an audio and video file containing the first video echo code and the first audio echo code is created. These echo codes can be used to verify whether the user client correctly decodes the audio and video file. Then, the sending time and target user are set: determine the preset sending time interval, for example, every 5 minutes, 10 minutes, etc. Determine the list of target user clients, these users will be the objects that receive the audio and video file. Then, a timing task is written, which uses a timer to trigger the sending operation periodically. In the timing task, write logic to execute the sending operation when the preset time interval arrives. When the timing task is triggered, iterate through the list of target user clients. For each user client, use the appropriate communication protocol to send the audio and video file to its corresponding receiving address.
[0072] And set up a listening mechanism on the host client to receive the second video echo code and the second audio echo code from the user client. When the user client decodes the audio and video file and generates the echo code, it will send these echo codes back to the host client. After the host client receives the echo code, it performs verification to ensure that they match the original echo code, thereby confirming that the user client correctly decodes the audio and video file. According to the verification result, the host client can perform corresponding operations, such as recording logs, sending notifications, etc.
[0073] Throughout the process, the embodiments of the present application should implement appropriate error handling mechanisms to handle possible network errors, decoding failures, etc. Detailed logs are recorded for troubleshooting and debugging when problems occur.
[0074] Select the appropriate communication protocol according to the application scenario and network environment. For large file transmission, consider using technologies such as chunked transfer, resuming transmission, etc. to improve transmission efficiency and stability. It is also necessary to ensure that the generation algorithm of the echo code is reliable and secure, and the verification algorithm should be able to accurately determine whether the user client has correctly decoded the audio and video files. During the process of sending and receiving data, the security of the data should be ensured, such as using encryption technology to protect the transmission and storage of data.
[0075] S300, according to the first preset time period, the first video echo code and the first audio echo code sent to each user client, and the second video echo code and the second audio echo code sent back by each user client, determine the network status of each user client in the first preset time period.
[0076] Optionally, in an optional embodiment of the present application, first, in the first preset time period, record the first video echo code and the first audio echo code sent to each user client. At the same time, collect the second video echo code and the second audio echo code sent back by each user client in the corresponding time. Then, match the sent first echo code with the received second echo code to ensure that they belong to the same echo code of the audio and video file. According to the time stamp of sending and receiving, align the data to analyze the network status of each user client at different time points.
[0077] Then, compare the first echo code and the second echo code to verify whether they are consistent. Consistency means that the user client has correctly decoded the audio and video file, and there is no serious data damage or loss in the network transmission process. According to the result of echo code verification, the proportion of user clients successfully decoding the audio and video file can be calculated. According to the network status indicators obtained by the above analysis, a network status label such as "good", "general", "poor" is determined for each user client.
[0078] Thresholds can be set to determine the network status, for example, if the packet loss rate exceeds a certain proportion, the network status is considered poor. Then, the network status of each user client obtained by analysis is displayed to the host or administrator in a visual way, so that they can understand the network situation. Network status data can be recorded in logs or databases for subsequent analysis and optimization.
[0079] If it is found in the analysis process that the network status of a certain user client is continuously poor, an abnormal processing mechanism can be triggered, such as sending a notification to the user or administrator, reminding them to check the network environment.
[0080] S400, according to the network status of each user client in the first preset time period, determine the reason for the lag of the network video conference.
[0081] Optionally, in one possible implementation of the present application, if it is found in the statistics that some users can normally play video and audio, i.e., their network status is displayed as normal, we can preliminarily infer that the network of the host end is also normal. This is because if there is a problem with the host end network, it will usually cause all or most users to have playback problems.
[0082] If multiple users (such as more than a certain percentage of participants, such as 30% or more) report that the audio, video, or all media channels are not normal, and these problems are consistent (i.e., multiple users have the same problem at the same time), it may mean that there is a common problem affecting multiple users.
[0083] In the analysis, those abnormal situations caused by individual user equipment failure, poor network environment, or active withdrawal should be excluded. These individual problems should not be considered as a factor in judging the overall network status of the meeting.
[0084] If the universal audio and video problems are not caused by individual user equipment or network, and there are enough users reporting the problem at the same time, it can be reasonably inferred that the host end is the problem. This may include insufficient network bandwidth of the host end, server overload, configuration error, or other network-related problems.
[0085] Once it is determined that the problem may be in the host end, the system should immediately notify the host and suggest that he or she make network adjustments. This may include checking network connections, optimizing server settings, increasing bandwidth, or taking other necessary measures to solve the problem.
[0086] After the host makes network adjustments, the system should continue to monitor the network status of the users to ensure that the problem is solved. If the problem still exists, further analysis may be needed to determine the cause and consider other possible solutions.
[0087] The entire analysis process and results should be recorded and formed into a report for subsequent analysis and improvement. This includes the collection of statistical data, the application of analysis methods, the logic of problem judgment, and the final solution, etc.
[0088] In this way, the cause of the network video conference lag can be effectively determined based on the network status data of each user client combined with the judgment logic. This helps to quickly identify problems and take appropriate measures to solve them, improving the stability of network video conferences and user experience.
[0089] The network video conference lag reason positioning method of the embodiment of the application, the host client adds echo code in the encoding process of the audio and video data every preset adding time, generates an audio and video file, and continuously sends the audio and video file to at least one user client. After the first video echo code and the first audio echo code in the audio and video file are decoded at the user client, the second video echo code and the second audio echo code are generated and returned to the host client. The host client can determine the network status of each user client by matching the first video echo code and the first audio echo code with the second video echo code and the second audio echo code, and finally quickly determine the lag reason of the network video conference according to the network status of each user client.
[0090] Therefore, by matching the first video echo code and the first audio echo code with the second video echo code and the second audio echo code, the network status of each user client can be quickly determined, and the number of user clients with different network statuses can accurately determine the lag reason of the network video conference, reduce the waste of time in the network video conference, and ensure the normal progress of the conference process.
[0091] In an embodiment, the step 300 can further perform the following steps:
[0092] S310, the first video echo code and the first audio echo code sent to each user client in the first preset time period, and the second video echo code and the second audio echo code returned by each user client are matched one by one to obtain a matching result, and the matching result is used to indicate whether the first video echo code and the first audio echo code correspond to the second video echo code and the second audio echo code.
[0093] S320, according to the matching result, determine the network status of each user client.
[0094] In these optional embodiments, by matching the sent and received echo codes one by one, the processing of the audio and video echo codes by the user client can be accurately determined, and the network status of each user client can be effectively determined. This matching mechanism not only improves the accuracy of network status detection, but also helps to discover network problems in time, providing a reliable data basis for subsequent lag reason analysis. At the same time, through automatic matching and state determination, the efficiency and stability of the entire network video conference system are also improved.
[0095] In an embodiment, the step 320 can further perform the following steps:
[0096] S321, in the case that the matching result indicates that the first video echo code and the second video echo code match, and the first audio echo code and the second audio echo code match, determining that the network state of the user client is a first state, the first state being used to indicate that the audio and video of the user client are both normal;
[0097] S322, in the case that the matching result indicates that the first video echo code and the second video echo code do not match, and the first audio echo code and the second audio echo code match, determining that the network state of the user client is a second state, the second state being used to indicate that the audio of the user client is normal, and the video is not normal;
[0098] S323, in the case that the matching result indicates that the first video echo code and the second video echo code match, and the first audio echo code and the second audio echo code do not match, determining that the network state of the user client is a third state, the third state being used to indicate that the video of the user client is normal, and the audio is not normal;
[0099] S324, in the case that the matching result indicates that the first video echo code and the second video echo code do not match, and the first audio echo code and the second audio echo code do not match, determining that the network state of the user client is a fourth state, the fourth state being used to indicate that the audio and video of the user client are both not normal.
[0100] In an embodiment, the step 320 can further perform the following step:
[0101] S325, in the case that the matching result indicates that the first video echo code and the second video echo code do not match, the first audio echo code and the second audio echo code do not match, and the second video echo code and the second audio echo code are both user offline echo codes, determining that the network state of the user client is a fifth state, the fifth state being used to indicate that the user client actively logs off.
[0102] Optionally, in a specific implementation manner of the present application, as shown in Figure 3 , after the moderator end receives the echo codes sent back by the user, the user echo archives are established according to the user information, for example, when user1 is online, the echo archive of user1 is established, when user2 is online, the echo archive of user2 is established, and so on until the conference ends. Thus, different user state scenarios are analyzed according to the echo archives of each user, and the current real state of the network is comprehensively analyzed according to the user state. If the audience point exits the button and actively exits the conference room in the middle of the way, a special exit code is fed back to the moderator, and then the user actively logs off (i.e. Figure 3 ).
[0103] Optionally, in the embodiments of the present application, according to the first audio echo code and the second audio echo code, it is judged whether the audio has missing. If there is intermittent missing, it can be determined that the audio has stuttering; if the audio is missing after a time and is missing in a continuous time period, it is considered that the audio interruption has not recovered; then according to the first video echo code and the second video echo code, it is judged whether the video has missing. If there is intermittent missing, it can be determined that the video has stuttering; if the video is missing after a time and is missing in a continuous time period, it is considered that the video interruption has not recovered, if the audio and the video are missing after a time and are missing in a continuous time period, it is considered that the user is offline abnormally.
[0104] Specifically, in the process of detecting the audio state, the system first compares the first audio echo code sent with the second audio echo code received. If it is found that the audio signal has intermittent missing phenomenon in the transmission process, it means that the audio stream has encountered stuttering in the transmission process. This stuttering may be caused by network jitter, insufficient bandwidth or device performance problems. In addition, if the system detects that the audio signal suddenly disappears at a certain time point and has not recovered in the subsequent continuous time period, it can be judged that the audio interruption has not recovered. This situation may indicate that the user's network connection has problems, or the user's device has been closed or entered the unresponsive state.
[0105] For the detection of the video state, the echo code comparison method is also adopted. If the system finds that the video frame has discontinuous or missing in the transmission process, it can be judged that the video stream has stuttering. Video stuttering may be caused by network delay, packet loss or video encoding and decoding problems. Similarly, if the video signal completely disappears after a certain time point and has not recovered for a period of time, it will be considered as video interruption and non-recovery. This may mean that the user's network connection is very unstable, or the user's device has been unable to normally receive and display the video stream.
[0106] Further, if the system detects that the audio and video signals have disappeared after a certain time and have not recovered in a continuous time period, it will be considered as a strong signal of the user's abnormal offline. This situation usually means that the user's network connection has been completely disconnected, or the user's device has been closed or cannot work normally. In this case, the system should timely notify other users or administrators for corresponding processing.
[0107] Through the above detailed detection and analysis, the transmission state of audio and video can be accurately judged, and problems can be found and located in time, so as to provide users with more stable and smooth network video conference experience. At the same time, the accurate judgment of the user's abnormal offline also helps to ensure the smooth progress of the meeting and the timely transmission of information.
[0108] In an embodiment, the step 400 can further comprise the following steps:
[0109] S410, in the case that the proportion of the user clients in the first state accounts for more than the first preset value of all user clients, determining that the cause of the network video conference lag is the first cause, the first cause being used to indicate that the user client network lags;
[0110] S420, in the case that the proportion of the user clients in the second state accounts for more than the second preset value of all user clients, or the proportion of the user clients in the third state accounts for more than the third preset value of all user clients, or the proportion of the user clients in the fourth state accounts for more than the fourth preset value of all user clients, determining that the cause of the network video conference lag is the second cause, the second cause being used to indicate that the host client network lags.
[0111] Optionally, in a specific implementation of the present application, when a certain audience joins the conference room, the statistical data of the user for the conference is established, and according to the current state of each user counted in the previous step, the current conference network problem is determined (active exit user does not need to be counted). The judgment logic is: when there is a client network state normal (both video and audio can be played normally), it can be basically determined that the host end network is normal. When there are many clients with audio, video or all media channels not normal, it can be determined that the host end has a problem, and the host needs to adjust the network.
[0112] Table 1 Correspondence between user client state and network state
[0113]
[0114] Specifically, in the embodiment of the present application, during the video conference, by monitoring the state of the user client in real time, the problems can be quickly found and determined. These problems mainly include two categories: inconsistency and consistency.
[0115] In the inconsistency problem, if only a small number of users have audio and video problems, the system determines that this is likely to be a network status or device problem of individual users, and is irrelevant to the host. Therefore, in this case, the host does not need to adjust his own device or network settings, and can continue to host the conference normally. However, if the inconsistency problem occurs in a large number of users, although it is not necessarily completely the host's problem, the number of problems is large, which has affected the normal progress of the conference. At this time, the system will prompt the host to deal with the network problem for a certain period of time to ensure that the conference can proceed smoothly.
[0116] In the case of consistency issues, if all users experience audio issues, the system will likely identify the host's audio device or settings as being malfunctioning. In this case, the host should promptly adjust their audio device to ensure proper audio transmission and maintain meeting quality. Similarly, if all users experience video issues, the system will likely identify the host's video device or settings as being malfunctioning, and the host should make appropriate adjustments.
[0117] The most serious issue is when all users experience audio and video issues. The system determines that this is likely due to poor network conditions or network settings for the host. In this case, the host should not only adjust their video equipment but also pay attention to the stability and speed of the network connection. If necessary, they can contact the network administrator or technical support staff for assistance.
[0118] This solution enables the system to accurately identify issues with different user client states and provide appropriate responses. This not only improves meeting efficiency and quality, but also reduces interruptions or delays caused by technical issues, enhancing the user experience. Furthermore, this solution demonstrates the system's ability to respond promptly and efficiently to user issues.
[0119] The following describes different abnormal distribution states:
[0120] ① Typical normal audience network status distribution diagram
[0121] like Figure 4 As shown, due to the Figure 4 The majority of users experience normal audio and video, so the host's network status can be assumed to be normal, and no adjustments are required. Individual viewers experiencing abnormalities should be considered to be due to their own network conditions and should resolve the issue on their own.
[0122] ② Typical single-tone (video) frequency abnormal state distribution
[0123] like Figure 5 As shown in the figure, if all users cannot receive audio, it should be considered that there is a problem with the host's network. Although the video can be transmitted successfully, it has already affected the meeting effect and the host needs to make adjustments. Similarly, if only the video is normal but the audio is not, the host also needs to make adjustments.
[0124] ③LAN normal status
[0125] like Figure 6As shown, the network state distribution feature of such a network is that abnormal users and normal users exist at the same time, but the proportion of abnormal users is high. This shows that the moderator may be in the same local area network region as the normal users, and not in the same local area network as other users. In this case, if you want the meeting to proceed stably, you need to be in the same local area network. Either the moderator and all normal users can be switched to the network of all abnormal users, or all abnormal users can be switched to the network of the moderator.
[0126] Figure 7 A flowchart of a positioning method for network video conference lag reasons provided by an embodiment of the application is shown. As shown in the figure, Figure 7 The method is applied to a user client, and the method can specifically include the following steps:
[0127] S500, receiving an audio and video file sent by a moderator client, the audio and video file including a first video echo code and a first audio echo code;
[0128] S600, decoding the audio and video file to obtain the first video echo code and the first audio echo code;
[0129] S700, generating a second video echo code and a second audio echo code according to the first video echo code and the first audio echo code;
[0130] S800, sending the second video echo code and the second audio echo code to the moderator client.
[0131] Optionally, in the embodiment of the application, the user client receives the audio and video file sent by the moderator client through a network connection (such as Wi-Fi or a mobile data network). This usually involves a network communication protocol to ensure reliable data transmission. The client will listen to a specific network port and start receiving data as soon as it detects a data packet from the moderator client. The received data can be audio and video files transmitted in the form of a stream, which may have been compressed and optimized for network transmission.
[0132] Subsequently, the audio and video decoder of the client will decode the received compressed audio and video file. The decoding process restores the compressed audio and video data into the original, playable audio and video signal. The decoder will decode according to the encoding format of the audio and video file. During the decoding process, the system will extract the first video echo code and the first audio echo code embedded in the audio and video stream. These echo codes are usually embedded as metadata in the audio and video stream for subsequent echo detection and state judgment.
[0133] Subsequently, after receiving and decoding the first video echo code and the first audio echo code, the client generates corresponding second video echo code and second audio echo code according to certain algorithms or rules. The process of generating the second echo code may involve processing and conversion of the first echo code, such as encryption, adding timestamp, calculating checksum, etc. These operations are designed to ensure that the second echo code can accurately reflect the state of the first echo code, and be used for subsequent state matching and judgment.
[0134] Finally, after generating the second video echo code and the second audio echo code, the client will package these echo codes into data packets and send them back to the host client through the network. The sending process also depends on the network communication protocol to ensure reliable transmission of data. The client may use the same network port as the receiving data or establish a new connection to send the echo code. In addition, in order to ensure the real-time and accuracy of the data, the sending process may need to consider network delay, packet loss, etc., and take corresponding measures to handle them.
[0135] Through the implementation of the above steps, the user client can receive, decode audio and video files, generate and send echo codes, so as to perform state matching and judgment with the host client. The implementation of these steps depends on the audio and video processing capability of the client, the network communication capability, and the corresponding algorithm and rule support.
[0136] In an embodiment, after the step 800, the method can further perform the following steps:
[0137] S810, determining the host volume data according to the decoded audio and video files, the host volume data being used to indicate the sound size of the network video conference host;
[0138] S820, determining the volume drop result according to the host volume data in the second time period, the volume drop result being used to indicate whether the host volume data is lower than the volume threshold;
[0139] S830, in the case that the volume drop result indicates that the host volume data is lower than the volume threshold, sending the volume drop reminder to the host client.
[0140] Optionally, in one specific implementation of the present application, a "human voice volume change detection module" is added to the user client. The echo codes from the user client are counted and analyzed at the host end, and a judgment opinion is formed. This module is divided into two steps: user data information counting and problem judgment. The specific implementation method includes the following steps:
[0141] (1) Counting the volume data from the host, and counting the sound size according to the volume waveform, etc.
[0142] (2) The first filtering, filtering out the period of time without volume or only with persistent background volume in the silence state.
[0143] (3) The second filtering, filtering out the audio files with human voice in the audio files after the first filtering. The technical method of identifying human voice is to use intelligent speech recognition technology to see if the voice can be converted into language text. If it can be converted into language text, it is human voice. At the same time, the content after the language content is converted into text is temporarily stored for later use.
[0144] (4) According to the human voice volume data, the average volume size change in a period of time is continuously tracked and counted. If the volume data suddenly drops and lasts for a period of time, and the program judges that the volume is too small and has fallen below the lower limit of the volume threshold that the human ear can clearly hear the semantic content, then the language text content before and after the volume drop is immediately marked and stored as "to-be-sent volume drop content". Then prompt the user of the user client and ask if you need to send a volume drop reminder to the host.
[0145] (5) The user of the user client confirms to send a volume drop reminder to the host and sends the "to-be-sent volume drop content" to the host.
[0146] (6) After the host gets the reminder, he quickly knows the event of volume drop and the content part of volume drop, and makes corresponding adjustment.
[0147] In other embodiments of the application, if the user can still express clearly even if the host's volume suddenly decreases, the minimum threshold of human voice volume can be appropriately reduced. In addition, if the volume is reduced by one level, the host's volume can be intelligently amplified to achieve the purpose of balancing the volume before and after.
[0148] In these optional embodiments, the application can effectively filter silence and background sound, identify and store human voice content, and when detecting a sudden volume drop below the threshold, prompt the user to send a reminder to the host. This not only improves the communication efficiency of the meeting, but also provides the host with a basis for timely adjusting the volume, ensuring clear communication of the meeting content. Enhances the interactivity and user experience of the meeting.
[0149] In a specific example, scenario one: the host end has a problem
[0150] Step 1. The host establishes a conference room and waits for other users to enter the conference room for a meeting.
[0151] Step 2. User 1 enters the conference room, and the host establishes statistical information for him and continuously sends echo code in the audio and video channels.
[0152] Step 3. User 1 receives the echo code of audio and video, processes the echo code, identifies the identity information, and sends it back to the host end.
[0153] Step 4. The host calculates the video and audio status of user 1 according to the received echo code information.
[0154] Step 5. User 2 and other users enter the conference room, and steps 2 to 4 are repeatedly performed.
[0155] Step 6. In the conference, all user audio abnormalities occur suddenly. The conference system reports to the host, prompting the host to solve the problem, and the host solves the network problem.
[0156] Scenario Two: User Client Problem
[0157] Step 1. The host establishes a conference room and waits for other users to enter the conference room for a meeting.
[0158] Step 2. User 1 enters the conference room, and the host establishes statistical information for him and continuously sends echo codes in the audio and video channels.
[0159] Step 3. User 1 receives the echo code of audio and video, processes the echo code, identifies the identity information, and sends it back to the host end.
[0160] Step 4. The host calculates the video and audio status of user 1 according to the received echo code information.
[0161] Step 5. User 2 and other users enter the conference room, and steps 2 to 4 are repeatedly performed.
[0162] Step 6. In the conference, individual user audio abnormalities occur suddenly. The conference system reports to the host, prompting the host to solve the problem, and the host solves the network problem.
[0163] Step 7. The conference system timely reminds the abnormal user outside the conference system through SMS and phone, etc.
[0164] Scenario Three: Sudden Decrease in Voice Volume
[0165] Step 1. While the host is continuously lecturing, the voice volume change detection module in the conference program of each client continuously detects the volume change.
[0166] Step 2. The voice volume change detection module detects a sudden decrease in the host's volume, pops up an interface to the client user, and asks whether the host's voice can be heard clearly and whether it needs to be reminded.
[0167] Step 3. The user selects "can hear clearly" and continues to listen to the host's speech, returning to step 1. If the user selects "cannot hear clearly and needs to send", step 4 is entered.
[0168] Step 4. The user client sends a volume abnormality prompt to the host end and sends the text content before and after the volume abnormality to the host.
[0169] Step 5. The host makes adjustments according to the content after receiving the prompt.
[0170] Optionally, in the embodiments of the present application, a scheme can be used to enable the host to discover network problems in a video conference in a timely manner and determine the nature of the network problems, so that the host can learn about the situation in a timely manner and make correct responses.
[0171] The present application does not rely on subjective feelings of people, but acquires network states in real time through technical means, without human judgment and thinking processes, so that problems can be discovered more objectively and timely, and a large amount of conference time is saved. Meanwhile, the host can learn about network state information before the user, and can make more scientific processing responses under various network conditions.
[0172] Since the present application provides a function of accurately determining problem points, problem ranges and problem natures, the determination of problems is no longer dependent on subjective feedback but is objectively and fairly determined, without the need for the user of the user client to sequentially cooperate with feedback.
[0173] Since the user of the user client does not need to make sounds and movements, this is particularly important in some special conferences that require a meeting atmosphere. The sequential feedback of the user of the client will not adversely affect the conference.
[0174] The present application adds a certain amount of echo code at intervals during the video and audio encoding of the host, produces an audio and video file, and transmits the audio and video file to the client of the participant. The client of the participant decodes the audio and video file, separately picks out the echo code, and plays the rest normally. The picked-out echo code is sent back to the host end to form statistical data. The condition of the echo code returned by the participant is analyzed statistically to form a judgment opinion, i.e., whether the network of the host is abnormal. If it is abnormal, the host is reminded to make network adjustments. If it is not abnormal, the corresponding participant is reminded to make network adjustments.
[0175] In addition, there are two advanced additional functions. One is that when it is determined that an individual network problem occurs in a certain participant client, a short message is sent to the participant to remind the participant to adjust the network in a timely manner to ensure the effect of the meeting. The other is that when the host end audio data reception is normal, but the speaking volume suddenly decreases, the host is reminded to adjust the sound in a timely manner.
[0176] Optionally, in the embodiment of the present application, whether it is a network on-off problem or a volume size problem, it is determined according to objective data, avoiding subjective dishonesty, subjective judgment, and other subjective factors. It can clearly define which party the conference network problem is most likely to come from, which is particularly important in some high-level conference scenarios or important conference scenarios.
[0177] Whether it is a network quality problem or a volume size problem, it can quickly analyze and locate the problem source, quickly report the problem, and make an objective overall judgment, which can greatly reduce the time spent on troubleshooting and ensure the smooth progress of the network conference. Moreover, the interference on the host during the reminding process is minimal. The host can observe the abnormality of some individual users and can choose to ignore them or wait until the network is restored before continuing to speak according to the actual situation. It can also remind the user with the problem to make adjustments in a timely manner. When the problem is found, the client user can be reminded in a timely manner to minimize the missed conference content.
[0178] Figure 8 The structure of the network video conference lag reason positioning device provided by another embodiment of the present application is shown. For the sake of convenience, only the part related to the embodiment of the present application is shown.
[0179] Reference Figure 8 The network video conference lag reason positioning device is applied to a host client. The device can include:
[0180] The generating module 801 is configured to add echo code every preset addition time in the encoding process of the audio and video data to generate an audio and video file, wherein the audio and video file includes first video echo code and first audio echo code.
[0181] The first sending module 802 is configured to send the audio and video file to at least one user client every preset sending time, wherein the audio and video file includes the first video echo code and the first audio echo code, so that the at least one user client sends second video echo code and second audio echo code to the host client after playing the audio and video file, the second video echo code is the video echo code generated by the user client after decoding the audio and video file, and the second audio echo code is the audio echo code generated by the user client after decoding the audio and video file.
[0182] The first determining module 803 is configured to determine the network status of each user client in the first preset time period according to the first video echo code and the first audio echo code sent to each user client in the first preset time period, and the second video echo code and the second audio echo code sent back by each user client.
[0183] The second determining module 804 is configured to determine the reason for the freezing of the network video conference according to the network status of each user client in the first preset time period.
[0184] In an embodiment, the first determining module 803 can include:
[0185] The first matching sub-module is configured to match the first video echo code and the first audio echo code sent to each user client in the first preset time period and the second video echo code and the second audio echo code sent back by each user client one by one to obtain a matching result, where the matching result is used to indicate whether the first video echo code and the first audio echo code correspond to the second video echo code and the second audio echo code.
[0186] The first determining sub-module is configured to determine the network status of each user client according to the matching result.
[0187] In an embodiment, the first determining sub-module can include:
[0188] The first determining unit is configured to determine that the network status of the user client is a first state in a case where the matching result indicates that the first video echo code and the second video echo code are matched and the first audio echo code and the second audio echo code are matched, where the first state is used to indicate that the audio and video of the user client are both normal.
[0189] The second determining unit is configured to determine that the network status of the user client is a second state in a case where the matching result indicates that the first video echo code and the second video echo code are not matched and the first audio echo code and the second audio echo code are matched, where the second state is used to indicate that the audio of the user client is normal and the video is not normal.
[0190] The third determining unit is configured to determine that the network status of the user client is a third state in a case where the matching result indicates that the first video echo code and the second video echo code are matched and the first audio echo code and the second audio echo code are not matched, where the third state is used to indicate that the video of the user client is normal and the audio is not normal.
[0191] The fourth determining unit is configured to determine that the network status of the user client is a fourth state in a case where the matching result indicates that the first video echo code and the second video echo code are not matched and the first audio echo code and the second audio echo code are not matched, where the fourth state is used to indicate that the audio and video of the user client are both not normal.
[0192] In an embodiment, the first determining sub-module can further include:
[0193] The fifth determining unit is configured to determine that the network state of the user client is a fifth state in a case where the matching result indicates that the first video echo code and the second video echo code are not matched, the first audio echo code and the second audio echo code are not matched, and the second video echo code and the second audio echo code are both offline echo codes of the user, and the fifth state is used to indicate that the user client is offline actively.
[0194] In an embodiment, the second determining module 803 can include:
[0195] The second determining sub-module is configured to determine that the cause of the lag of the network video conference is a first cause in a case where the proportion of the user clients in the first state to all user clients is greater than a first preset value, and the first cause is used to indicate that the user client network is lagged.
[0196] The third determining sub-module is configured to determine that the cause of the lag of the network video conference is a second cause in a case where the proportion of the user clients in the second state to all user clients is greater than a second preset value, or the proportion of the user clients in the third state to all user clients is greater than a third preset value, or the proportion of the user clients in the fourth state to all user clients is greater than a fourth preset value, and the second cause is used to indicate that the host client network is lagged.
[0197] Figure 9 FIG. 8 shows a structural schematic diagram of a device for locating a cause of lag of a network video conference according to another embodiment of the present application. For ease of illustration, only parts related to the embodiments of the present application are shown.
[0198] With reference to Figure 9 The device for locating a cause of lag of a network video conference is applied to a user client, and the device can include:
[0199] The receiving module 901 is configured to receive an audio and video file sent by a host client, and the audio and video file includes a first video echo code and a first audio echo code.
[0200] The decoding module 902 is configured to decode the audio and video file to obtain the first video echo code and the first audio echo code.
[0201] The generating module 903 is configured to generate a second video echo code and a second audio echo code according to the first video echo code and the first audio echo code.
[0202] The second sending module 904 is configured to send the second video echo code and the second audio echo code to the host client.
[0203] In an embodiment, the device for locating a cause of lag of a network video conference can further include:
[0204] The third determining module is configured to determine host volume data according to the decoded audio and video file, the host volume data being used to indicate the sound size of the host of the network video conference.
[0205] The fourth determining module is configured to determine a volume drop result according to the host volume data in the second time period, the volume drop result being used to indicate whether the host volume data is lower than the volume threshold.
[0206] The third sending module is configured to send a volume drop reminder to the host client in a case where the volume drop result indicates that the host volume data is lower than the volume threshold.
[0207] It should be noted that the information interaction and execution process between the above apparatuses / units are based on the same concept as the method embodiments of the application, and are corresponding apparatuses of the battery thermal runaway early warning method. All the implementation manners in the above method embodiments are applicable to the embodiments of the apparatus, and the specific functions and technical effects brought by the implementation manners can be referred to the method embodiments part, and will not be described here.
[0208] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the apparatus is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific names of the functional units and modules are only for convenient distinction, and do not limit the protection scope of the application. The specific working process of the units and modules in the system can refer to the corresponding process in the foregoing method embodiments, and will not be described here.
[0209] Figure 10 A hardware structure schematic diagram of an electronic device provided by an embodiment of the application is shown.
[0210] The device can include a processor 1001 and a memory 1002 storing program instructions.
[0211] The processor 1001 executes the program to implement the steps in any of the above method embodiments.
[0212] For example, the program can be divided into one or more modules / units, one or more modules / units are stored in the memory 1002, and are executed by the processor 1001 to complete the present application. One or more modules / units can be a series of program instruction segments capable of completing a specific function, which are used to describe the execution process of the program in the device.
[0213] Specifically, the processor 1001 described above can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0214] The memory 1002 can include a mass storage for data or instructions. By way of example and not limitation, the memory 1002 can include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. Where appropriate, the memory 1002 can include removable or non-removable (or fixed) media. Where appropriate, the memory 1002 can be internal or external to the integrated gateway disaster recovery device. In certain embodiments, the memory 1002 is non-volatile solid-state memory.
[0215] The memory can include read-only memory (ROM), random access memory (RAM), magnetic disk storage mediums, optical storage mediums, flash memory devices, electrical, optical, or other physical / tangible memory storage devices. Thus, in general, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software that, when executed, is operable to perform operations as described with reference to the methods according to the aspects of the present disclosure.
[0216] The processor 1001 implements any one of the above-described embodiments by reading and executing program instructions stored in the memory 1002.
[0217] In one example, the electronic device can further include a communication interface 1003 and a bus 1010. Wherein the processor 1001, the memory 1002, the communication interface 1003 are connected through the bus 1010 and complete the communication between each other.
[0218] The communication interface 1003 is mainly used to realize the communication between the modules, devices, units and / or equipment in the embodiments of the present application.
[0219] Bus 1010 includes hardware, software, or both, to couple components of the online data traffic metering device to each other and to couple components to other systems. For example, but not limited to, the bus can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand (IB) interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus or a combination of two or more of these. Where suitable, bus 1010 can include one or more buses. Although a particular bus arrangement is described and shown in the embodiments, the present application contemplates any suitable bus or interconnect.
[0220] In addition, in combination with the method in the above-mentioned embodiments, the embodiments of the present application can provide a storage medium for implementation. The storage medium has program instructions stored thereon; the program instructions are executed by a processor to implement any one of the methods in the above-mentioned embodiments.
[0221] The embodiments of the present application further provide a chip, which includes a processor and a communication interface, the communication interface is coupled to the processor, the processor is configured to execute programs or instructions, to implement various processes of the above-mentioned method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.
[0222] It should be understood that the chip mentioned in the embodiments of the present application can also be referred to as a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0223] The embodiments of the present application provide a computer program product, which is stored in a storage medium, and the program product is executed by at least one processor to implement various processes of the above-mentioned method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.
[0224] It should be understood that the present application is not limited to the specific configurations and processes described above and shown in the drawings. For the sake of brevity, detailed descriptions of well-known methods are omitted herein. In the above-mentioned embodiments, several specific steps are described and shown as examples. However, the method processes of the present application are not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order of the steps, after understanding the spirit of the present application.
[0225] The functional modules shown in the structural block diagram above can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and the like. When implemented in software, the elements of the present application are program or code segments that are used to perform the required tasks. The program or code segments can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave over a transmission medium or communication link. The "machine-readable medium" can include any medium that can store or transfer information. Examples of the machine-readable medium include an electronic circuit, a semiconductor memory device, a ROM, a flash memory, an erasable ROM (EROM), a floppy diskette, a CD-ROM, an optical disk, a hard disk, a fiber optic medium, a radio frequency (RF) link, and the like. The code segments can be downloaded via a computer network, such as the Internet, an intranet, and the like.
[0226] It should also be noted that the example embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiments, or in an order different from the embodiments, or several steps can be performed simultaneously.
[0227] The above describes aspects of the present disclosure with reference to flowcharts and / or block diagrams of methods, apparatus (systems) and program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams and combinations of blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus enable the implementation of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. The processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It should also be understood that each block of the block diagrams and / or flowcharts and combinations of blocks in the block diagrams and / or flowcharts can also be implemented by special-purpose hardware to perform the specified functions or acts, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0228] The above is only a specific implementation of the present application, and those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, module and unit can refer to the corresponding process in the foregoing method embodiments, which will not be described here. It should be understood that the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be covered within the protection scope of the present application.
Claims
1. A method for locating a network video conference stall cause, characterized in that, The method applied to the host client comprises: Adding echo codes in the encoding process of the audio and video data every preset adding time to generate an audio and video file, wherein the audio and video file comprises first video echo codes and first audio echo codes; Continuously sending the audio and video file to at least one user client to make the at least one user client send second video echo codes and second audio echo codes to the host client after playing the audio and video file, wherein the second video echo codes are video echo codes generated by the user client after decoding the audio and video file, and the second audio echo codes are audio echo codes generated by the user client after decoding the audio and video file; Determining network states of the user clients in a first preset time period according to the first video echo codes and the first audio echo codes sent to the user clients and the second video echo codes and the second audio echo codes sent back by the user clients in the first preset time period; Determining a reason for lag of the network video conference according to the network states of the user clients in the first preset time period.
2. The method of claim 1, wherein, The method for determining the network states of the user clients in the first preset time period comprises: Matching the first video echo codes and the first audio echo codes sent to the user clients and the second video echo codes and the second audio echo codes sent back by the user clients in the first preset time period one by one to obtain a matching result, wherein the matching result is used to indicate whether the first video echo codes and the first audio echo codes correspond to the second video echo codes and the second audio echo codes; Determining the network states of the user clients according to the matching result.
3. The method of claim 2, wherein, The method for determining the network states of the user clients according to the matching result comprises: In a case where the matching result indicates that the first video echo codes and the second video echo codes match and the first audio echo codes and the second audio echo codes match, determining that the network state of the user client is a first state, wherein the first state is used to indicate that the audio and video of the user client are normal; In a case where the matching result indicates that the first video echo codes and the second video echo codes do not match and the first audio echo codes and the second audio echo codes match, determining that the network state of the user client is a second state, wherein the second state is used to indicate that the audio of the user client is normal and the video is not normal; In a case where the matching result indicates that the first video echo codes and the second video echo codes match and the first audio echo codes and the second audio echo codes do not match, determining that the network state of the user client is a third state, wherein the third state is used to indicate that the video of the user client is normal and the audio is not normal. In a case where the matching result indicates that the first video echo code and the second video echo code are not matched, and the first audio echo code and the second audio echo code are not matched, the network state of the user client is determined as a fourth state, and the fourth state is used to indicate that the audio and video of the user client are both abnormal.
4. The method of claim 3, wherein, The determining the network state of each user client according to the matching result further includes: In a case where the matching result indicates that the first video echo code and the second video echo code are not matched, the first audio echo code and the second audio echo code are not matched, and the second video echo code and the second audio echo code are both offline echo codes of the user, the network state of the user client is determined as a fifth state, and the fifth state is used to indicate that the user client is actively offline.
5. The method of claim 4, wherein, The determining the freezing reason of the network video conference according to the network state of each user client in the first preset time period includes: In a case where the proportion of the user client in the first state to all the user clients is greater than a first preset value, the freezing reason of the network video conference is determined as a first reason, and the first reason is used to indicate that the user client network is frozen; In a case where the proportion of the user client in the second state to all the user clients is greater than a second preset value, or the proportion of the user client in the third state to all the user clients is greater than a third preset value, or the proportion of the user client in the fourth state to all the user clients is greater than a fourth preset value, the freezing reason of the network video conference is determined as a second reason, and the second reason is used to indicate that the host client network is frozen.
6. A method for locating a network video conference stall cause, characterized in that, The method applied to a user client includes, receiving an audio and video file sent by a host client, the audio and video file including a first video echo code and a first audio echo code; decoding the audio and video file to obtain the first video echo code and the first audio echo code; generating a second video echo code and a second audio echo code according to the first video echo code and the first audio echo code; sending the second video echo code and the second audio echo code to the host client, so that the host client determines the network state of each user client in a first preset time period according to the first video echo code and the first audio echo code, and the second video echo code and the second audio echo code in the first preset time period, and determines the freezing reason of the network video conference according to the network state of each user client in the first preset time period.
7. The method of claim 6, wherein, After the decoding of the audio and video file to obtain the first video echo code and the first audio echo code, the method further includes: determining host volume data according to the decoded audio and video file, the host volume data being used to indicate the sound size of the host of the network video conference; determining a volume drop result according to the host volume data in a second time period, the volume drop result being used to indicate whether the host volume data is lower than a volume threshold value; In a case where the volume reduction result indicates that the host volume data is lower than a volume threshold, a volume reduction reminder is sent to the host client.
8. A device for locating the cause of network video conference freezes, characterized in that: The device applied to the host client comprises: A generating module is configured to add echo codes in an encoding process of the audio and video data every preset adding time to generate an audio and video file, wherein the audio and video file comprises first video echo codes and first audio echo codes; A first sending module is configured to continuously send the audio and video file to at least one user client, so that the at least one user client sends second video echo codes and second audio echo codes to the host client after playing the audio and video file, wherein the second video echo codes are video echo codes generated by the user client after decoding the audio and video file, and the second audio echo codes are audio echo codes generated by the user client after decoding the audio and video file; A first determining module is configured to determine network states of the user clients in a first preset time period according to the first video echo codes and the first audio echo codes sent to the user clients and the second video echo codes and the second audio echo codes sent back by the user clients in the first preset time period; A second determining module is configured to determine a cause of a lag of the network video conference according to the network states of the user clients in the first preset time period.
9. A device for locating the cause of network video conferencing freezes, characterized in that: The device applied to the user client comprises: A receiving module is configured to receive an audio and video file sent by a host client, wherein the audio and video file comprises first video echo codes and first audio echo codes; A decoding module is configured to decode the audio and video file to obtain the first video echo codes and the first audio echo codes; A generating module is configured to generate second video echo codes and second audio echo codes according to the first video echo codes and the first audio echo codes; A second sending module is configured to send the second video echo codes and the second audio echo codes to the host client, so that the host client determines network states of the user clients in a first preset time period according to the first video echo codes and the first audio echo codes and the second video echo codes and the second audio echo codes in the first preset time period, and determines a cause of a lag of the network video conference according to the network states of the user clients in the first preset time period.
10. An electronic device, comprising: The device comprises a processor and a memory storing computer program instructions; The processor executes the computer program instructions to implement the method for locating a cause of a lag of a network video conference according to any one of claims 1-7.
11. A computer readable storage medium characterized by The computer program instructions are stored on the computer readable storage medium and are executed by the processor to implement the method for locating a cause of a lag of a network video conference according to any one of claims 1-7.
12. A computer program product, characterised in that, The instructions in the computer program product are executed by the processor of the electronic device to enable the electronic device to perform the method for locating a cause of a lag of a network video conference according to any one of claims 1-7.
Citation Information
Patent Citations
Live broadcast lagging prompting method and device, computer equipment and storage medium
CN114189700A
Verification method and device and terminal equipment
CN116361833A