An audio and video quality diagnosis method based on image recognition and network analysis

By constructing a data layer and shared nodes to collect and time-align audio and video data from multiple terminals, the problem of event asynchrony caused by information transmission delays from multiple terminals in video conferencing is solved, realizing synchronous transmission of audio and video data and improving the smoothness and user experience of video conferencing.

CN120812240BActive Publication Date: 2025-11-18GUANGZHOU YUNSHITONG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511308891.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-11-18
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

In video conferencing, information transmission delays caused by the different locations of multiple terminals can lead to asynchronous events during the video conference, resulting in inconsistencies between the beginning and end of the video conference and a lack of smoothness.

Method used

By using image recognition and network analysis methods, a data layer is constructed to collect and time-align audio and video data from multiple terminals. Shared nodes and data pools are used to share audio and video data and adjust bandwidth, ensuring synchronous transmission of audio and video data among multiple terminals.

Benefits of technology

It ensures smooth events during video conferences, guarantees the correspondence between audio and video data in real time, and improves the video conference experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120812240B_ABST
    Figure CN120812240B_ABST
Patent Text Reader

Abstract

The application discloses an audio and video quality diagnosis method based on image recognition and network analysis, relates to the technical field of multimedia communication, collects audio and video data of multiple terminals, analyzes the audio and video data of the multiple terminals based on image recognition, obtains diagnosis data, and prompts the terminals according to the diagnosis data; obtains transmission quality information of the audio and video data after sharing, the transmission quality information includes time delay information and resolution, and based on network analysis, transmission quality information that does not meet preset conditions is taken as abnormal information, and the bandwidth of a channel between a terminal corresponding to the abnormal information and a data layer is adjusted. The application can collect the audio and video data of the multiple terminals through the data layer and perform time alignment, synchronously provide the time-aligned audio and video data to the multiple terminals, thereby ensuring the smoothness of the conference video, enabling the audio and video data under the actual time to correspond, and improving the experience of the conference video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimedia communication technology, and more specifically to an audio and video quality diagnosis method based on image recognition and network analysis. Background Technology

[0002] As economic globalization deepens and corporate organizational structures become increasingly decentralized, video conferencing becomes necessary for internal management and departmental collaboration. However, during video conferences, the different locations of multiple devices can lead to delays in information transmission. This uneven latency can cause events to be out of sync during the video conference, resulting in inconsistencies and a lack of smoothness. Summary of the Invention

[0003] The purpose of this invention is to provide an audio and video quality diagnosis method based on image recognition and network analysis to address the shortcomings of the prior art.

[0004] To achieve the above objectives, the present invention provides the following technical solution: an audio and video quality diagnosis method based on image recognition and network analysis, comprising the following steps:

[0005] The participating terminals are identified, and audio and video data from these terminals are collected. The audio and video data includes both video and audio data. A data layer is constructed for each terminal, and the audio and video data from the terminals is shared based on this data layer.

[0006] Based on image recognition, audio and video data from multiple terminals are analyzed to obtain diagnostic data, and prompts are given to the corresponding terminals based on the diagnostic data.

[0007] The transmission quality information of the shared audio and video data is obtained. The transmission quality information includes latency information and resolution. Based on network analysis, the transmission quality information that does not meet the preset conditions is identified as abnormal information, and the bandwidth of the channel between the terminal corresponding to the abnormal information and the data layer is adjusted.

[0008] In a preferred embodiment, the step of constructing a data layer corresponding to multiple terminals and sharing audio and video data of multiple terminals based on the data layer includes:

[0009] A data pool is built based on the multiple terminals participating in the meeting, and multiple shared nodes are used as the data layer. The multiple shared nodes are connected to the data pool through an audio and video distribution chain. The number of shared nodes is the same as that of terminals, and they are connected one-to-one.

[0010] Establish a channel between multiple terminals and the data pool;

[0011] Based on the acquisition of corresponding audio and video data by the terminal, the audio and video data is provided to the data pool through the channel between multiple terminals and the data pool, and the audio and video data of multiple terminals are time-aligned in the data pool.

[0012] Based on the data pool, the aligned audio and video data, excluding the terminal corresponding to the shared node to be distributed, is distributed to the shared node.

[0013] The received audio and video data is provided to the corresponding terminal by the shared node.

[0014] In a preferred embodiment, the step of constructing a data pool based on the multiple terminals participating in the meeting and multiple shared nodes as a data layer, wherein the multiple shared nodes are connected to the data pool through an audio / video distribution chain, and the number of shared nodes and terminals is the same, and they are connected one-to-one, includes:

[0015] The network locations of multiple terminals are determined separately. Based on the network locations of the multiple terminals, shared nodes with the same transmission distance for the corresponding multiple terminals are selected. The network location of the data pool is arranged in the middle of the multiple shared nodes to obtain the data layer.

[0016] In the data pool, a data alignment space and multiple alignment chains corresponding to multiple terminals are set up. The alignment chain includes a loop channel, a loop chain, an access point, and a loop key. The loop key includes an exit loop point and an access loop point. The loop chain is composed of multiple virtual signal carriers connected in a chain to form a closed loop.

[0017] Multiple alignment chains in the data pool are paired with multiple shared nodes. Audio and video distribution chains are set between the corresponding alignment chains and shared nodes. The audio and video distribution chain includes a distribution loop channel, a distribution loop chain, a mobile access point, and a distribution loop key. The distribution loop key includes an exit distribution loop point and an access distribution loop point. The distribution loop chain is composed of multiple virtual signal carriers connected in a closed loop. The mobile access points in the multiple audio and video distribution chains are interconnected.

[0018] Connect the exit loop point in the alignment chain to the access loop point in the audio / video dispatch chain, and connect the access loop point to the exit loop point.

[0019] The shared node connects to the corresponding terminal via a mobile access point.

[0020] In a preferred embodiment, the step of collecting corresponding audio and video data based on terminals, providing the audio and video data to the data pool through the channel between multiple terminals and the data pool, and performing time alignment of the audio and video data of multiple terminals in the data pool includes:

[0021] The system collects audio and video data of the corresponding personnel through the terminal and marks them with timestamps, then transmits the timestamped audio and video data to the data alignment space in the data pool.

[0022] In the data alignment space, audio and video data collected from multiple terminals are mapped according to timestamps.

[0023] In a preferred embodiment, the step of distributing the aligned audio and video data (excluding the terminal corresponding to the shared node) to the shared node based on the data pool includes:

[0024] The audio and video data after time alignment in the data alignment space are copied as a whole according to the number of terminals. The audio and video data of multiple terminals are separated by removing the audio and video data of one different terminal, resulting in multiple audio and video data to be matched.

[0025] Match the multiple audio and video data to be matched with the alignment chains corresponding to the missing terminals;

[0026] Based on the same timestamp and corresponding interval, multiple audio and video data to be matched are segmented, transmitted to the loop channel through the access point of the corresponding alignment chain, and loaded into the virtual signal carrier of the loop chain in sequence.

[0027] The loop chain continuously transmits data in the loop channel. The virtual signal carrier is transmitted to the access distribution loop point in the audio and video distribution chain through the exit loop point on the loop chain, and then enters the distribution loop channel for continued transmission, and is provided to the mobile access point.

[0028] In a preferred embodiment, the step of providing the received audio and video data to the corresponding terminal according to the sharing node includes:

[0029] The mobile access point extracts the audio and video data from the received virtual signal carrier and transmits it to the connected sharing node, which then transmits the audio and video data to the corresponding terminal.

[0030] The virtual signal carriers through the mobile access points are transmitted in the distribution loop channel. When the transmission reaches the discharge distribution loop point, the virtual signal carriers are transmitted from the discharge distribution loop point to the access loop point and enter the loop channel.

[0031] The virtual signal carrier entering the loop channel continues to load audio and video data after time alignment in the data alignment space at the access point.

[0032] In a preferred embodiment, the step of analyzing audio and video data from multiple terminals based on image recognition to obtain diagnostic data, and then providing prompts to the corresponding terminals based on the diagnostic data, includes:

[0033] The video data corresponding to the terminal is subjected to image recognition to obtain image information, which includes the position of the person's face in the frame and the completeness of the person's face in the frame. Video data that does not meet the preset conditions is used as diagnostic data.

[0034] Anomalies are detected and displayed on the corresponding terminals based on diagnostic data.

[0035] In a preferred embodiment, the step of adjusting the bandwidth of the channel between the terminal corresponding to the abnormal information and the data layer includes:

[0036] Shared nodes in multiple audio and video distribution chains transmit the audio and video data received with the same timestamp and corresponding interval time from the virtual signal carrier to the corresponding terminal;

[0037] Obtain adjustment information for mobile access points in multiple audio and video dispatch chains when their network location exceeds the limit. The adjustment information includes the direction of movement of the mobile access points and the range of network movement.

[0038] The adjustment parameters are matched according to the adjustment information. The adjustment parameters include the adjustment direction and the corresponding channel bandwidth adjustment parameters. The bandwidth of the channel between the terminal and the data pool that exceeds the adjustment information is adjusted according to the adjustment parameters.

[0039] The technical effects and advantages provided by the present invention in the above technical solution are as follows:

[0040] This invention can collect and time-align audio and video data from multiple terminals through a data layer, and synchronously provide the time-aligned audio and video data to multiple terminals, ensuring the smooth sequence of events in video conferences and enabling correspondence between audio and video data at actual times, thereby improving the video conference experience. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0042] Figure 1 This is a flowchart of the method of the present invention.

[0043] Figure 2 This is a logical framework diagram of the data pool of the present invention. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0045] Example 1, please refer to Figure 1 and Figure 2 As shown in this embodiment, an audio / video quality diagnosis method based on image recognition and network analysis includes the following steps:

[0046] S1. Determine the multiple terminals participating in the meeting, collect audio and video data from the multiple terminals, including video data and audio data, construct a data layer for each terminal, and share the audio and video data of the multiple terminals based on the data layer;

[0047] S2. Analyze audio and video data from multiple terminals based on image recognition to obtain diagnostic data, and provide prompts to the corresponding terminals based on the diagnostic data;

[0048] S3. Obtain the transmission quality information of the shared audio and video data. The transmission quality information includes latency information and resolution. Based on network analysis, the transmission quality information that does not meet the preset conditions is regarded as abnormal information. The bandwidth of the channel between the terminal corresponding to the abnormal information and the data layer is adjusted.

[0049] As described in steps S1-S3 above, the data layer can collect and time-align audio and video data from multiple terminals, and synchronously provide the time-aligned audio and video data to multiple terminals. This ensures the smoothness of the conference video and enables the audio and video data to correspond in real time, thereby improving the conference video experience.

[0050] In one embodiment, step S1, which involves constructing a data layer corresponding to multiple terminals and sharing audio and video data among the multiple terminals based on the data layer, includes:

[0051] S11. Construct a data pool and multiple shared nodes as the data layer based on the multiple terminals participating in the meeting (multiple shared nodes are interconnected). Multiple shared nodes are connected to the data pool through an audio and video distribution chain. The number of shared nodes and terminals is the same, and they are connected one-to-one.

[0052] S12. Establish a channel between multiple terminals and the data pool;

[0053] S13. Collect corresponding audio and video data based on the terminal, provide the audio and video data to the data pool through the channel between the multiple terminals and the data pool, and perform time alignment on the audio and video data of the multiple terminals in the data pool.

[0054] S14. Based on the data pool, distribute the aligned audio and video data, excluding the terminal corresponding to the shared node to be distributed, to the shared node.

[0055] S15. Provide the received audio and video data to the corresponding terminal according to the shared node;

[0056] In one embodiment, step S11, which involves constructing a data pool based on multiple participating terminals and multiple shared nodes as a data layer, with the shared nodes connected to the data pool via an audio / video distribution chain, and the number of shared nodes being equal to the number of terminals and connected in a one-to-one correspondence, includes:

[0057] S111. Determine the network location of each of the multiple terminals, select a shared node with the same transmission distance for each of the multiple terminals based on their network locations, and arrange the network location of the data pool in the middle of the multiple shared nodes to obtain the data layer.

[0058] S112. Set up a data alignment space and multiple alignment chains corresponding to multiple terminals in the data pool. The alignment chain includes a loop channel, a loop chain, an access point, and a loop key. The loop key includes an exit loop point and an access loop point. The loop chain is composed of multiple virtual signal carriers connected in a closed loop (here, the loop channel is the transmission channel, the loop chain is the combination of virtual signal carriers in the transmission channel, for example, multiple virtual machines are connected in a chain, the access point is the network node for input data, and the exit loop point and access loop point of the loop key are two adjacent network nodes on the loop channel, used for data access and output).

[0059] S113. Assign each of the multiple alignment chains in the data pool to a pair of multiple shared nodes, and set up audio and video dispatch chains between the corresponding alignment chains and shared nodes. The audio and video dispatch chain includes a dispatch loop channel, a dispatch loop chain, a mobile access point, and a dispatch loop key. The dispatch loop key includes an exit dispatch loop point and an access dispatch loop point. The dispatch loop chain is composed of multiple virtual signal carriers connected in a closed loop. The mobile access points in the multiple audio and video dispatch chains are interconnected (the dispatch loop channel and the loop channel are the same. The dispatch loop channel is a ring channel that is closed by the mobile access point and the dispatch loop key. The dispatch loop chain and the loop chain are the same. They are both composed of multiple virtual machines connected in a chain. The mobile access point is a network node that can change network position. The dispatch loop key and the loop key are the same. They are both network nodes set on the channel).

[0060] S114. Connect the exit loop point in the alignment chain to the access loop point in the audio / video distribution chain, and connect the access loop point to the exit loop point.

[0061] S115. Connect the shared node to the corresponding terminal through the mobile access point.

[0062] In one embodiment, step S13, which involves collecting corresponding audio and video data based on terminals, providing the audio and video data to the data pool through a channel between multiple terminals and the data pool, and performing time alignment of the audio and video data from multiple terminals in the data pool, includes:

[0063] S131. Collect audio and video data of the corresponding personnel through the terminal and mark it with a timestamp, and transmit the audio and video data marked with the timestamp to the data alignment space in the data pool;

[0064] S132. In the data alignment space, the audio and video data collected by multiple terminals are matched according to the timestamp;

[0065] In one embodiment, step S14, which involves distributing the aligned audio and video data (excluding the terminal corresponding to the shared node) to the shared node based on the data pool, includes:

[0066] S141. Copy the time-aligned audio and video data in the data alignment space as a whole according to the number of terminals, and remove the audio and video data of one different terminal from the audio and video data of multiple multi-terminals to obtain multiple audio and video data to be matched.

[0067] S142. Match the multiple audio and video data to be matched with the alignment chains corresponding to the missing terminals respectively;

[0068] S143. Based on the same timestamp and the same corresponding time interval, the multiple audio and video data to be matched are segmented and transmitted to the loop channel through the access point of the corresponding alignment chain, and loaded into the virtual signal carrier of the loop chain in sequence.

[0069] S144. The loop chain continuously transmits in the loop channel. The virtual signal carrier is transmitted to the access distribution loop point in the audio and video distribution chain through the exit loop point on the loop chain, and then enters the distribution loop channel for continued transmission and is provided to the mobile access point.

[0070] In one embodiment, step S15, which provides the received audio and video data to the corresponding terminal according to the shared node, includes:

[0071] S151. The mobile access point extracts the audio and video data from the received virtual signal carrier and transmits it to the connected sharing node, which then transmits the audio and video data to the corresponding terminal.

[0072] S152. The virtual signal carrier through the mobile access point is transmitted in the distribution loop channel. When it is transmitted to the discharge distribution loop point, the virtual signal carrier is transmitted to the access loop point through the discharge distribution loop point and enters the loop channel.

[0073] S153. The virtual signal carrier entering the loop channel continues to load the audio and video data after time alignment in the data alignment space at the access point.

[0074] As described in steps S11-S15 above, to ensure the accuracy of the sequence of events between multiple terminals in a corporate video conference at the same time, the network location of the multiple terminals must first be determined. The network location represents the location where audio and video data will be shared. Based on the network location of the multiple terminals, a sharing node with the same transmission distance is selected for each terminal. This ensures that after the sharing node is in place, the audio and video data can be transmitted to the corresponding terminals according to the shared node with the same transmission distance. This guarantees the consistency of audio and video data reaching each terminal and avoids gaps in audio and video data between multiple terminals due to transmission discrepancies. For example, if someone on one terminal asks a question during the conference, the audio and video data may be transmitted, but due to transmission distance issues, some terminals may not receive the data. Some terminals receive data early, while others receive it late. For those who raise this issue during the meeting, the late-receiving terminals experience information mismatch. While others are moving on to the next topic, the late-receiving terminals may still be responding to the previous question in their audio / video recordings, resulting in serious time inaccuracies. Therefore, it's necessary to collect audio / video data generated simultaneously by all terminals and timestamp the collected data. Based on these timestamps, the collected audio / video data can be time-aligned. This ensures that, from the perspective of each terminal, all audio / video data from multiple shared terminals is synchronized and delayed, except for its own, which is ahead of schedule. This allows for managing the correspondence between audio / video data from all other terminals except itself. A data pool is placed at the center of multiple sharing nodes to ultimately obtain the data layer. A data alignment space and multiple alignment chains corresponding to multiple terminals are set up in the data pool. The data alignment space is used to collect audio and video data from multiple terminals and align the data according to timestamps to ensure the correspondence of subsequent data transmission. The alignment chains are used to arrange the audio and video data, dividing the time-aligned audio and video data from multiple terminals into time periods according to timestamps. A virtual signal carrier is used to store the audio and video data within a time period. Since audio and video data from multiple terminals are collected and transmitted to the data alignment space, it is not necessary to transmit the audio and video data of the local terminal when sharing audio and video data. For example, if there are three terminals... Terminals A, B, and C simultaneously collect audio and video data from all three terminals. After time alignment, the audio and video data from the three terminals are copied, resulting in three copies of multi-terminal audio and video data. Then, the audio and video data corresponding to terminal A is deleted from the first multi-terminal audio and video data, the audio and video data corresponding to terminal B is deleted from the second multi-terminal audio and video data, and the audio and video data corresponding to terminal C is deleted from the third multi-terminal audio and video data. In this way, the audio and video data of the other terminals, after deleting the audio and video data corresponding to terminal A, are transmitted together to the shared node corresponding to terminal A, and finally transmitted to terminal A.The specific transmission operation is as follows: Based on the same timestamp and corresponding interval time, multiple audio and video data to be matched are sequentially transmitted to the virtual signal carrier of the loop chain in the loop channel through the corresponding alignment link entry point. The loop chain continues to transmit in the loop channel. Through the exit loop point on the loop chain, the virtual signal carrier is transmitted to the access distribution loop point in the audio and video distribution chain, and enters the distribution loop channel for continued transmission, providing it to the mobile access point. At the mobile access point, the audio and video data in the virtual signal carrier that has passed through is extracted, and then the extracted audio and video data is provided to the sharing node. The sharing node transmits the received audio and video data to the corresponding terminal. The virtual signal carrier passing through the mobile access point continues to be transmitted in the distribution loop channel. When it reaches the exit distribution loop point, the virtual signal carrier is transmitted to the access loop point through the exit distribution loop point and enters the loop channel. The virtual signal carrier (virtual machine) entering the loop channel continues transmission. This virtual signal carrier is blank and does not store audio or video data. After reaching the access point, all virtual signal carriers in the alignment chain can store the same timestamp and corresponding interval from the data alignment space. In this way, the virtual signal carriers in the alignment chain and the virtual signal carriers in the audio and video distribution chain can perform an effective loop, which can ensure the synchronous transmission of audio and video data and provide it to the corresponding terminal, thus achieving smooth conferencing.

[0075] In one embodiment, step S2, which involves analyzing audio and video data from multiple terminals based on image recognition to obtain diagnostic data and then providing prompts to the corresponding terminals based on the diagnostic data, includes:

[0076] S21. Perform image recognition on the video data corresponding to the terminal to obtain image information, wherein the image information includes the position of the person's face in the frame and the completeness of the person's face in the frame, and video data that does not meet the preset conditions are used as diagnostic data.

[0077] S22. Provide an anomaly alert on the corresponding terminal based on the diagnostic data.

[0078] As described in steps S21-S22 above, the terminal performs image recognition on the collected video data. First, it determines the facial contour of the person in the video, obtains the position of the person's face in the video frame, and assesses the completeness of the face in the frame. The position of the center point of the facial contour is taken as the position of the person's face in the frame. Then, it analyzes whether the person's face is complete in the frame. The area inside the outer contour of the person's face is taken as the area of ​​the entire face. The completeness of the person's face in the frame is determined based on the area of ​​the facial contour displayed in the frame. Video data that is less than a preset condition is used as diagnostic data. The preset condition is a preset percentage of the area of ​​the facial contour displayed in the frame. For example, if half a face is displayed in the frame, it means that the percentage of the facial contour displayed in the frame is 50%. Percentages below the preset condition are used as diagnostic data, while percentages above the preset condition are left unchecked. Then, based on the diagnostic data, the terminal can be prompted to move towards the center of the video frame. Specific operations include moving the person's position and moving the camera position. The purpose is to display the facial video while ensuring that the proportion of the person's face is not less than the preset condition, thus facilitating better video conferencing. Additionally, additional video quality analysis functions can be added. For example, it supports video display output anomaly detection, which can detect abnormal video images such as black screens, green screens, and blue screens caused by camera malfunctions or incompatible decoding libraries. It supports image freeze detection, which can identify video stream freezes exceeding a set threshold, caused by network latency, encoding errors, or insufficient terminal performance, severely impacting the real-time performance of meetings. It supports image blur detection, which can detect abnormalities such as inaccurate camera focus or autofocus failure, resulting in blurry and unclear video surveillance images. It supports video image occlusion detection, which can detect abnormalities where the camera's view is obstructed by external objects, resulting in partial or complete obstruction of the camera's view. It supports video screen distortion detection, which can detect video image distortion caused by data packet loss due to video transmission network issues or incompatible video decoding libraries, rendering the video surveillance unusable and losing its monitoring value. It supports brightness anomaly detection, which can detect excessively high or low brightness in the overall or localized areas of the image, ensuring that faces and meeting content are clearly identifiable. Excessive brightness may lead to loss of detail, while excessive darkness will affect the recognition of facial expressions and environmental information, reducing communication efficiency.

[0079] In one embodiment, step S3, which adjusts the bandwidth of the channel between the terminal corresponding to the abnormal information and the data layer, includes:

[0080] S31. Shared nodes in multiple audio and video distribution chains transmit the audio and video data received from the same timestamp and the corresponding virtual signal carrier at the same time interval to the corresponding terminal.

[0081] S32. Obtain adjustment information for mobile access points in multiple audio and video dispatch chains when their network locations exceed the specified limits. The adjustment information includes the direction of movement of the mobile access points and the range of movement within the network.

[0082] S33. Match the corresponding adjustment parameters according to the adjustment information. The adjustment parameters include the adjustment direction and the corresponding channel bandwidth adjustment parameters. Adjust the bandwidth of the channel between the terminal and the data pool that exceeds the adjustment information according to the adjustment parameters.

[0083] As described in steps S31-S33 above, since multiple terminals are collecting and transmitting audio and video data, and the network quality of these terminals varies, delays in the transmission process can occur. When the data transmitted from multiple terminals arrives at the audio and video distribution chain and the corresponding data stored sequentially through the distribution loop is inconsistent, a mobile access point is used to move along the distribution loop. This allows the mobile access point to synchronously acquire audio and video data with the same timestamp and corresponding interval, ensuring that the audio and video data from multiple terminals are transmitted to the corresponding terminals in time alignment. Each terminal's display of audio and video data from other terminals is synchronized, improving the continuity of events in the video conference, the actual dialogue, the intervals between dialogues, and the dynamics within the video. The situation before and after the event is the same. For example, if there is a channel bandwidth resource difference between one terminal and the data pool, the data transmission delay will be severe. In this case, the audio and video data in the corresponding loop channel and the dispatch loop channel of other terminals will arrive in the pool early and fill up. In order to ensure that the timestamps of the audio and video data provided to the terminals by multiple terminals are corresponding, the mobile access point will move towards the access dispatch loop point on the dispatch loop channel. This way, it can obtain audio and video data with the same timestamp as other terminals. If the mobile access point is kept in one position on the dispatch loop channel, it will have to wait for the data to be transmitted to the mobile access point before it can be transmitted. In this case, if the channel bandwidth is unstable, it is impossible to guarantee that the audio and video data with the same timestamp is provided to the terminal. In this case, the audio and video data received by the terminal is not equal and has no time correspondence.Similarly, if the audio and video data transmission fills a large number of dispatch loop channels, the mobile access point will move closer to the exit dispatch loop point along the dispatch loop channel. This lengthens the dispatch loop channel between the mobile access point and the access dispatch loop point, resulting in a larger number of virtual signal carriers storing audio and video data. This buffers the fluctuations in channel bandwidth during movement. Multiple mobile access points communicate with each other and can provide audio and video data with the same timestamp to the shared node. Multiple fluctuation nodes are set on the dispatch loop channel. These fluctuation nodes are network-connected nodes that allow mobile nodes to move between them. This serves as a buffer range for mobile access points to obtain audio and video data with the same timestamp. Adjustment information is obtained for mobile access points in multiple audio and video dispatch chains that exceed network position limits. This adjustment information includes the mobile access point's position... The movement direction and network movement range of the mobile access point are defined. The network movement range is a range within the set area of ​​the fluctuating node. The movement direction of the mobile access point is either towards the discharge / dispatch loop point or towards the access / dispatch loop point. Based on the adjustment information, corresponding adjustment parameters are matched. These parameters include the adjustment direction and the corresponding channel bandwidth adjustment parameters. The bandwidth of the channel between the terminal and the data pool exceeding the adjustment information is adjusted according to these parameters. When the network movement range is exceeded, network analysis is performed to obtain the signal transmission quality of the channel. Then, based on the movement direction of the mobile access point, the bandwidth of the channel between the corresponding terminal and the data pool is adjusted. For example, if the mobile access point moves towards the access / dispatch loop point on the dispatch loop channel, it indicates poor bandwidth resources or the terminal being far from the data pool, resulting in transmission delay, requiring increased bandwidth resources. Conversely, reducing bandwidth resources aims to ensure that the audio and video data of all terminals are transmitted to the shared node at the same timestamp on the mobile access point, providing better multi-terminal alignment and synchronization capabilities for video conferencing and improving the user experience.

[0084] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for diagnosing audio and video quality based on image recognition and network analysis, characterized in that, Includes the following steps: The participating terminals are identified, and audio and video data from these terminals are collected. This audio and video data includes both video and audio data. A data layer is constructed for each terminal, and the audio and video data from the terminals is shared based on this data layer. A data pool is built based on the multiple terminals participating in the meeting, and multiple shared nodes are used as the data layer. The multiple shared nodes are connected to the data pool through an audio and video distribution chain. The number of shared nodes and terminals is the same, and they are connected one-to-one. The steps of constructing a data pool based on multiple participating terminals and multiple shared nodes as the data layer, with the multiple shared nodes connected to the data pool via an audio / video distribution chain, and the number of shared nodes and terminals being equal and connected one-to-one, include: The network locations of multiple terminals are determined separately. Based on the network locations of the multiple terminals, shared nodes with the same transmission distance for the corresponding multiple terminals are selected. The network location of the data pool is arranged in the middle of the multiple shared nodes to obtain the data layer. In the data pool, a data alignment space and multiple alignment chains corresponding to multiple terminals are set up. The alignment chain includes a loop channel, a loop chain, an access point, and a loop key. The loop key includes an exit loop point and an access loop point. The loop chain is composed of multiple virtual signal carriers connected in a chain to form a closed loop. Multiple alignment chains in the data pool are paired with multiple shared nodes. Audio and video distribution chains are set between the corresponding alignment chains and shared nodes. The audio and video distribution chain includes a distribution loop channel, a distribution loop chain, a mobile access point, and a distribution loop key. The distribution loop key includes an exit distribution loop point and an access distribution loop point. The distribution loop chain is composed of multiple virtual signal carriers connected in a closed loop. The mobile access points in the multiple audio and video distribution chains are interconnected. Connect the exit loop point in the alignment chain to the access loop point in the audio / video dispatch chain, and connect the access loop point to the exit loop point. The shared node is connected to the corresponding terminal through a mobile access point; Based on image recognition, audio and video data from multiple terminals are analyzed to obtain diagnostic data, and prompts are given to the corresponding terminals based on the diagnostic data. The transmission quality information of the shared audio and video data is obtained. The transmission quality information includes latency information and resolution. Based on network analysis, the transmission quality information that does not meet the preset conditions is identified as abnormal information, and the bandwidth of the channel between the terminal corresponding to the abnormal information and the data layer is adjusted.

2. The audio / video quality diagnosis method based on image recognition and network analysis according to claim 1, characterized in that, Establish a channel between multiple terminals and the data pool; Based on the acquisition of corresponding audio and video data by the terminal, the audio and video data is provided to the data pool through the channel between multiple terminals and the data pool, and the audio and video data of multiple terminals are time-aligned in the data pool. Based on the data pool, the aligned audio and video data, excluding the terminal corresponding to the shared node to be distributed, is distributed to the shared node. The received audio and video data is provided to the corresponding terminal by the shared node.

3. The audio / video quality diagnosis method based on image recognition and network analysis according to claim 2, characterized in that, The steps of collecting corresponding audio and video data based on terminals, providing the audio and video data to the data pool through the channel between multiple terminals and the data pool, and performing time alignment of the audio and video data of multiple terminals in the data pool include: The system collects audio and video data of the corresponding personnel through the terminal and marks them with timestamps, then transmits the timestamped audio and video data to the data alignment space in the data pool. In the data alignment space, audio and video data collected from multiple terminals are mapped according to timestamps.

4. The audio / video quality diagnosis method based on image recognition and network analysis according to claim 3, characterized in that, The step of distributing the aligned audio and video data (excluding the terminal corresponding to the shared node) to the shared node based on the data pool includes: The audio and video data after time alignment in the data alignment space are copied as a whole according to the number of terminals. The audio and video data of multiple terminals are separated by removing the audio and video data of one different terminal, resulting in multiple audio and video data to be matched. Match the multiple audio and video data to be matched with the alignment chains corresponding to the missing terminals; Based on the same timestamp and corresponding interval, multiple audio and video data to be matched are segmented, transmitted to the loop channel through the access point of the corresponding alignment chain, and loaded into the virtual signal carrier of the loop chain in sequence. The loop chain continuously transmits data in the loop channel. The virtual signal carrier is transmitted to the access distribution loop point in the audio and video distribution chain through the exit loop point on the loop chain, and then enters the distribution loop channel for continued transmission, and is provided to the mobile access point.

5. The audio / video quality diagnosis method based on image recognition and network analysis according to claim 4, characterized in that, The step of providing the received audio and video data to the corresponding terminal according to the shared node includes: The mobile access point extracts the audio and video data from the received virtual signal carrier and transmits it to the connected sharing node, which then transmits the audio and video data to the corresponding terminal. The virtual signal carriers through the mobile access points are transmitted in the distribution loop channel. When the transmission reaches the discharge distribution loop point, the virtual signal carriers are transmitted from the discharge distribution loop point to the access loop point and enter the loop channel. The virtual signal carrier entering the loop channel continues to load audio and video data after time alignment in the data alignment space at the access point.

6. The audio / video quality diagnosis method based on image recognition and network analysis according to claim 1, characterized in that, The step of analyzing audio and video data from multiple terminals based on image recognition to obtain diagnostic data, and then providing prompts to the corresponding terminals based on the diagnostic data, includes: The video data corresponding to the terminal is subjected to image recognition to obtain image information, which includes the position of the person's face in the frame and the completeness of the person's face in the frame. Video data that does not meet the preset conditions is used as diagnostic data. Anomalies are detected and displayed on the corresponding terminals based on diagnostic data.

7. The audio / video quality diagnosis method based on image recognition and network analysis according to claim 5, characterized in that, The step of adjusting the bandwidth of the channel between the terminal corresponding to the abnormal information and the data layer includes: Shared nodes in multiple audio and video distribution chains transmit the audio and video data received with the same timestamp and corresponding interval time from the virtual signal carrier to the corresponding terminal; Obtain adjustment information for mobile access points in multiple audio and video dispatch chains when their network location exceeds the limit. The adjustment information includes the direction of movement of the mobile access points and the range of network movement. The adjustment parameters are matched according to the adjustment information. The adjustment parameters include the adjustment direction and the corresponding channel bandwidth adjustment parameters. The bandwidth of the channel between the terminal and the data pool that exceeds the adjustment information is adjusted according to the adjustment parameters.

Citation Information

Patent Citations

  • Audio and video equipment resource scheduling method and system

    CN115914539A

  • Audio sharing method and device, computer equipment and storage medium

    CN117896469A