Audio-visual synchronization detection method and apparatus
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-29
- Publication Date
- 2026-08-11
AI Technical Summary
然而,前者仅能提供统计层面的间接映射,易受瞬时网络波动干扰,无法区分“网络指标差但内容同步”与“网络指标正常但内容不同步”等复杂情形;后者则过度依赖时间戳的完整性与准确性,在路测场景中常见的多设备时钟偏差、时间戳损坏或编码异常情况下容易失效
[0010] The audio-visual synchronization detection method and apparatus provided in this application, by matching the expected image features corresponding to the audio amplitude features with the actual image features corresponding to the video frames, can eliminate the reliance on easily interfered network indicators or timestamps in the prior art, and can directly judge audio-visual synchronization based on audio and video content, which can significantly improve the accuracy of audio-visual synchronization detection in road test scenarios.
Smart Images

Figure CN122554619A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a method and apparatus for detecting audio-visual synchronization. Background Technology
[0002] In the field of wireless communication, with the large-scale deployment of 5G / 4G networks, video services (such as road test videos, in-vehicle 5G video calls, and mobile terminal video playback) have become core scenarios for network quality assessment and user experience optimization. Road testing, as a crucial means for operators to obtain real-world network performance data, requires the simultaneous acquisition of audio and video signals in dynamically changing road environments and accurate determination of audio-visual synchronization status to ensure end-to-end quality of video services. However, road test scenarios face complex interference such as signal fluctuations, multi-device clock deviations, environmental noise, and lighting changes, posing significant challenges to the robustness and accuracy of audio-visual synchronization detection methods.
[0003] Currently, mainstream audio-visual synchronization detection technologies in the industry are mainly divided into two categories: one is an indirect evaluation method based on network indicators, such as RSRP (Reference Signal Received Power), SINR (Signal to Interference plus Noise Ratio), latency, and packet loss rate. This method establishes a correlation between network parameters and video quality through statistical modeling, thereby inferring the synchronization status. The other is a comparison method based on audio and video timestamps, such as PTS (Presentation Time Stamp), which determines synchronization by parsing the timestamp differences embedded during the encoding stage. However, the former can only provide an indirect mapping at the statistical level and is susceptible to interference from instantaneous network fluctuations, failing to distinguish between complex situations such as "poor network indicators but synchronized content" and "normal network indicators but out-of-sync content." The latter relies excessively on the integrity and accuracy of timestamps, and is prone to failure in common scenarios such as multi-device clock deviations, timestamp corruption, or encoding anomalies during drive testing. Therefore, existing technologies cannot directly determine audio-visual synchronization based on audio and video content, resulting in low accuracy in audio-visual synchronization detection. Summary of the Invention
[0004] This application provides a method and apparatus for detecting audio-visual synchronization.
[0005] According to a first aspect of the embodiments of this application, a method for detecting audio-visual synchronization is provided, the method comprising: Acquire the target video file in the road test scenario, extract the audio signal from the target video file, and calculate the amplitude characteristics of the audio frames in the audio signal; Extract the actual image features of video frames from the target video file; Time-align audio frames with video frames to establish a correspondence between them; Obtain the feature mapping model, which contains the correspondence between multiple continuous amplitude intervals and expected image features. The feature mapping model is trained based on normally synchronized audio and video samples. For target audio frames and target video frames that have a corresponding relationship, the expected image features are determined by the feature mapping model based on the amplitude range in which the amplitude features of the target audio frame are located. Calculate the matching degree between the actual image features and the expected image features of the target video frame, and determine whether the target audio frame and the target video frame are synchronized based on the matching degree.
[0006] According to a second aspect of the embodiments of this application, an audio-visual synchronization detection device is provided, the device comprising: The file acquisition module is used to acquire target video files in road test scenarios; The audio feature acquisition module is used to extract audio signals from the target video file and calculate the amplitude features of audio frames in the audio signal; The actual image feature acquisition module is used to extract the actual image features of video frames from the target video file; The correspondence establishment module is used to time-align audio frames with video frames and establish the correspondence between audio frames and video frames; The model acquisition module is used to acquire the feature mapping model. The feature mapping model contains the correspondence between multiple continuous amplitude intervals and expected image features. The feature mapping model is trained based on normally synchronized audio and video samples. The expected image feature acquisition module is used to determine the corresponding expected image features based on the amplitude range of the amplitude features of the target audio frame and the target video frame that have a corresponding relationship, through a feature mapping model. The audio-visual synchronization judgment module is used to calculate the matching degree between the actual image features and the expected image features of the target video frame, and to determine whether the target audio frame and the target video frame are synchronized based on the matching degree.
[0007] According to a third aspect of the embodiments of this application, an electronic device is provided. The electronic device includes a memory and a processor, wherein a computer program is stored in the memory, and the processor executes the program to implement the method described above.
[0008] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the methods described above in this application.
[0009] According to a fifth aspect of the embodiments of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods described above in this application.
[0010] The audio-visual synchronization detection method and apparatus provided in this application, by matching the expected image features corresponding to the audio amplitude features with the actual image features corresponding to the video frames, can eliminate the reliance on easily interfered network indicators or timestamps in the prior art, and can directly judge audio-visual synchronization based on audio and video content, which can significantly improve the accuracy of audio-visual synchronization detection in road test scenarios. Attached Figure Description
[0011] Further details, features, and advantages of this application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 A schematic diagram of the system architecture provided for an exemplary embodiment of this application; Figure 2 A flowchart of an audio-visual synchronization detection method provided in an exemplary embodiment of this application; Figure 3 A schematic block diagram of the functional modules of an audio-visual synchronization detection device provided in an exemplary embodiment of this application; Figure 4 A structural block diagram of an electronic device provided in an exemplary embodiment of this application; Figure 5 A structural block diagram of a computer system provided for an exemplary embodiment of this application. Detailed Implementation
[0012] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of this application.
[0013] It should be understood that the steps described in the method embodiments of this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.
[0014] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this application are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0015] It should be noted that the terms "a" and "a plurality of" used in this application are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more". The names of the messages or information exchanged between multiple devices in the embodiments of this application are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0016] It is understood that before using the technical solutions disclosed in the various embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0017] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this application's technical solution, based on the prompt message.
[0018] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device. It is understood that the above notification and user authorization process is merely illustrative and does not limit the implementation of this application; other methods that comply with relevant laws and regulations may also be applied to the implementation of this application.
[0019] The following will combine Figure 1 The technical solutions of the embodiments of this application will be described in detail. Figure 1This diagram illustrates the device architecture of the audio-visual synchronization detection system 10 provided in an embodiment of this application. The audio-visual synchronization detection system 10 is used for audio-visual synchronization detection in drive testing scenarios and can further achieve geolocation and network root cause tracing. Specifically, the audio-visual synchronization detection system 10 can be deployed in an electronic device, which can be a server or a terminal; the embodiments are not limited thereto.
[0020] like Figure 1 As shown, the audio-visual synchronization detection system 10 may include: an audio-visual split output module 11, a video content recognition module 12, an audio content extraction module 13, an audio-visual content mapping module 14, a matching degree calculation and synchronization determination module 16, a geographic association and root cause tracing module 15, and a root cause location output module 17.
[0021] The connections and functions of each module are as follows: The audio / video splitting output module 11 is connected to the video content recognition module 12 and the audio content extraction module 13, respectively. The audio / video splitting output module 11 receives the input video stream, separates the audio signal from the video signal in the video stream, outputs the separated audio signal to the audio content recognition module 13, and outputs the separated video signal to the video content extraction module 12.
[0022] The video content extraction module 12 is connected to the audio / video split output module 11 and the matching degree calculation and synchronization determination module 16, respectively. The video content extraction module 12 is used to receive the video signal from the audio / video split output module 11, extract the actual image features of each video frame (for example, extract the RGB components of each pixel in the central region of the video frame and calculate the mean), and output the extracted actual image features to the matching degree calculation and synchronization determination module 16.
[0023] The audio content recognition module 13 is connected to both the audio / video split output module 11 and the audio / video content mapping module 14. The audio content recognition module 13 receives audio signals from the audio / video split output module 11, performs frame-by-frame processing on the audio signals, calculates the amplitude characteristics of each audio frame (e.g., using the RMS (Root Mean Square) algorithm), and outputs the calculated amplitude characteristics to the audio / video content mapping module 14.
[0024] The audio-video content mapping module 14 is connected to the audio content recognition module 13 and the matching degree calculation and synchronization determination module 16, respectively. The audio-video content mapping module 14 internally stores a feature mapping model, which contains the correspondence between multiple continuous amplitude intervals and expected image features. This feature mapping model is trained through supervised learning based on a large number of normally synchronized audio and video samples. The audio-video content mapping module 14 receives amplitude features from the audio content recognition module 13, queries the corresponding expected image features based on the amplitude interval of the amplitude features, and outputs the expected image features to the matching degree calculation and synchronization determination module 16.
[0025] The matching degree calculation and synchronization determination module 16 is connected to the video content extraction module 12, the audio-video content mapping module 14, and the geographic association and root cause tracing module 15, respectively. The matching degree calculation and synchronization determination module 16 receives actual image features from the video content extraction module 12 and expected image features from the audio-video content mapping module 14. It calculates the matching degree between the two (e.g., Euclidean distance) and determines whether the corresponding audio and video frames are synchronized based on the relationship between the matching degree and a preset threshold. Simultaneously, the matching degree calculation and synchronization determination module 16 outputs the occurrence time and related information of the audio-video desynchronization event to the geographic association and root cause tracing module 15.
[0026] The geographic association and root cause tracing module 15 is connected to the matching degree calculation and synchronization determination module 16 and the root cause location output module 17, respectively. The geographic association and root cause tracing module 15 receives test data (i.e., drive test data) input (e.g., CSV / Excel files containing fields such as timestamp, longitude, latitude, RSRP, and SINR) and receives audio-visual desynchronization event information from the matching degree calculation and synchronization determination module 16. Based on the event's occurrence time, the geographic association and root cause tracing module 15 searches for matching timestamps in the drive test data, extracts the corresponding geographic location (longitude, latitude) and network indicators (RSRP, SINR), determines the event's root cause based on the comparison results between the network indicators and preset root cause thresholds, and outputs the root cause analysis results to the root cause location output module 17.
[0027] The root cause localization output module 17 is connected to the geographic association and root cause tracing module 15. It is used to output or display the received root cause analysis results (including the geographic location of audio-visual desynchronization events, associated network indicators, and root cause conclusions) for use by network optimization personnel.
[0028] The above modules work together to achieve a complete functional chain from audio and video input to synchronization determination, and then to geolocation and network root cause tracing.
[0029] This application embodiment achieves accurate detection of audio-visual synchronization in road test scenarios by directly utilizing the root mean square amplitude features of audio frames and the RGB image features of the central region of video frames, and calculating the Euclidean distance matching degree based on a pre-trained amplitude-RGB quantitative mapping model. It does not rely on easily interfered network indicators or timestamp metadata, significantly enhancing the robustness of the detection method; the synchronization judgment error can reach within 20ms, which is better than the prior art. At the same time, this application embodiment does not rely on audio peaks or video motion features, and is suitable for edge scenarios that are difficult to handle by traditional methods, such as static images and stable audio. In addition, by associating with road test data files, it can accurately locate the geographical location of audio-visual desynchronization events, and associate with network indicators such as RSRP and SINR to trace the root causes of weak coverage or high quality and poor performance. It provides a complete closed-loop support for network optimization from content detection to problem location and optimization decision-making, and has good engineering implementation capabilities.
[0030] Therefore, as a specific implementation of the above embodiments, this application also provides an audio-visual synchronization detection method, which can be applied to the above-described audio-visual synchronization detection system 10, such as... Figure 2 As shown, the method may include the following steps: In step S210, the target video file in the road test scenario is obtained, the audio signal is extracted from the target video file, and the amplitude characteristics of the audio frames in the audio signal are calculated.
[0031] In this embodiment, the target video file can be a pre-recorded road test video in a road test scenario. This road test video can be video content collected by recording devices mounted on mobile vehicles (such as cars, trains, drones, etc.) during wireless network (such as 5G communication) road tests, including but not limited to videos of in-vehicle conversations, road environment videos, or test videos played on mobile terminal screens. These videos share the common characteristic that, during recording, the system simultaneously collects wireless network metrics (such as RSRP and SINR) and geographic location information, thereby forming a correlation data between audio / video content and network quality and spatial location.
[0032] By acquiring the target video file in the road test scenario, the audio signal can be extracted from the target video file, and the amplitude characteristics of the audio frames in the audio signal can be calculated. For example, FFmpeg tools can be used to extract the audio signal. The audio encoding format can be PCM_S16LE, the sampling rate can be 44100Hz, and it can be mono. The extracted audio signal is processed into frames, and the frame length can be N sampling points (N is a positive integer, for example, N=1024), with no overlapping frames. For each frame of audio signal, its root mean square amplitude (RMS) can be calculated using formula (1): (1) in, For the first i The first frame of audio k The amplitude of each sampling point can be... RMS i As the first i Amplitude characteristics of frame audio.
[0033] In step S220, the actual image features of the video frames are extracted from the target video file.
[0034] In this embodiment, to avoid edge interference and improve efficiency, only the RGB features of the central region of the video frame can be extracted; however, this embodiment is not limited to this. Let the video frame width be... W Height is H The central area is R × R A rectangle of pixels (R=50) is defined. The R, G, and B components of all pixels within this region are extracted. The mean of each component is calculated to obtain the RGB triplet (R...). avg G avg B avg () serves as the actual image feature of the video frame.
[0035] For example, by extracting pixels from the central region: ( frame (Video frames in BGR format).
[0036] Calculate the average RGB color: After converting from BGR format to RGB format, calculate the average RGB components of all pixels. (2) Where M is the total number of pixels in the central region. , , The first p The RGB component values of each pixel.
[0037] In step S230, the audio frames and video frames are time-aligned to establish a correspondence between them.
[0038] Assuming the video frame rate is F The audio sampling rate is S The audio frame length is N The duration of a single audio frame is T audio = N / S The duration of a single video frame is T video =1 / F For each audio frame, obtain its start timestamp. t audio Calculate the corresponding video frame indexf = t audio × F This establishes a one-to-one correspondence between "audio frame amplitude characteristics" and "video frame index".
[0039] In step S240, a feature mapping model is obtained; wherein, the feature mapping model contains the correspondence between multiple continuous amplitude intervals and expected image features, and the feature mapping model is trained based on normally synchronized audio and video samples.
[0040] The mapping relationship between audio amplitude and video RGB in this embodiment can be pre-established by manually controlling audio and video parameters during the shooting stage. Alternatively, it can be based on a large number of samples confirmed to be synchronized between audio and video in real road test scenarios, and semantic-level correspondence patterns can be automatically obtained through supervised learning. In this way, the feature mapping model is essentially a quantitative expression of the joint distribution of audio and video features under synchronized conditions, without relying on any active encoding or human intervention. Therefore, it is applicable to any passively recorded road test video and has good generalization ability and engineering practicality.
[0041] For example, this feature mapping model can be trained through supervised learning based on a large number of normally synchronized audio and video samples (e.g., 1000 sets of samples covering different road scenes). The feature mapping model can divide the continuous audio amplitude value range into multiple continuous and non-overlapping intervals, each interval corresponding to a unique expected RGB color (target RGB). For example, the amplitude interval (0.052, 0.056] corresponds to RGB(174, 0, 1), (0.10, 0.12] corresponds to RGB(232, 0, 1), ..., (0.496, 0.499] corresponds to RGB(2, 30, 92). The mapping relationship reflects the semantic design that "audio energy is positively correlated with visual salience".
[0042] For example, the mapping relationship is shown in Table 1.
[0043] Table 1:
[0044] As shown in Table 1, the feature mapping model divides the continuous range of audio root mean square amplitude values into 9 consecutive and non-overlapping intervals (it can also be 8 or 10 groups, etc., the embodiments are not limited to this), and each interval uniquely corresponds to a target RGB color value. Among them, the serial numbers 1 to 9 are used to identify the order of each mapping entry, and the combination ID (such as S1, S2, ..., S9) is the internal label of each mapping entry, which is convenient for quick reference in system implementation or data transmission; the correspondence between amplitude intervals and target RGB is determined through training with a large number of normal synchronization samples, aiming to maximize the visual distinguishability of the expected colors corresponding to different amplitude intervals, thereby improving the accuracy of audio-visual synchronization detection.
[0045] As shown in Table 1, the feature mapping model can be trained using AI-supervised learning to automatically learn the correlation features between audio amplitude and video RGB in road test scenarios, thereby improving the scene adaptability and detection accuracy of the mapping. The specific construction process is as follows: First, a large number of normally synchronized road test audio and video samples were collected, covering different road types (urban roads, highways, tunnels, etc.), different lighting conditions, and different audio types (speech, environmental noise, etc.). For each sample, the root mean square amplitude (RMS) of the audio frame and the average RGB value of the central region of the corresponding video frame were extracted to form the training dataset.
[0046] Secondly, the training data is analyzed using supervised learning algorithms (such as clustering or regression-based models) to divide the continuous range of audio RMS values into nine consecutive and non-overlapping amplitude intervals. Each interval corresponds to a unique expected RGB color value, which is determined by the statistical center (such as the mean or mode) of the RGB values of the corresponding video frames in all synchronized samples falling within that interval. During the division and assignment process, the goal is to maximize the distinguishability between the corresponding RGB colors in each interval, ensuring that the expected colors corresponding to different amplitude intervals have significant visual differences, thereby improving the sensitivity of subsequent matching detection.
[0047] The core design basis of this feature mapping model is as follows: audio amplitude is positively correlated with human perception of "sound intensity," that is, the larger the amplitude, the stronger the sound perceived by the human ear; at the same time, the brightness and saturation of RGB colors are positively correlated with human visual perception of "visual prominence," that is, the brighter and more saturated the color, the more visually prominent it is. Through the above mapping, this embodiment establishes a semantic association between "audio energy" and "video visual features" that conforms to human perception habits: when the audio energy is high (louder sound), the corresponding expected RGB color tends to be brighter or more saturated; when the audio energy is low (lower sound), the expected RGB color tends to be darker or paler. This design not only ensures the rationality of the feature mapping model in terms of physical semantics, but also enhances the robustness and accuracy of detection by maximizing color distinguishability.
[0048] In step S250, for target audio frames and target video frames that have a corresponding relationship, the corresponding expected image features are determined by a feature mapping model based on the amplitude range in which the amplitude features of the target audio frame are located.
[0049] The expected image features can be obtained based on the amplitude features and feature mapping model of the target audio frame. As shown in Table 1, for example, if the RMS value of an audio frame is 0.276, then it falls into the interval (0.275, 0.277], and the corresponding expected RGB is (149, 221, 84).
[0050] In step S260, the matching degree between the actual image features and the expected image features of the target video frame is calculated, and the matching degree is used to determine whether the target audio frame and the target video frame are synchronized.
[0051] In this embodiment, the matching degree can be calculated by Euclidean distance, Manhattan distance or cosine similarity between the actual image features and the expected image, but the embodiment is not limited to this.
[0052] Specifically, taking Euclidean distance calculation as an example, let's assume the actual RGB values are ( R a , G a , B a ), expected RGB is ( R e , G e , B e ), calculate Euclidean distance D : (3) The smaller the value of D, the higher the matching degree between the actual RGB and the target RGB, and the better the audio-visual synchronization.
[0053] The synchronization threshold can be determined through training with a large number of samples. T For example, select 1000 sets of normally synchronized audio and video samples (covering different road scenes and different audio types), and calculate the matching distance of each set of samples. D The threshold is determined using the 3σ criterion: .in σ is the mean of the matching distances for all samples, and σ is the standard deviation. Experiments have shown that a threshold of T = 50 can be selected, meaning that when D < 50, the audio and video are considered synchronized; when D ≥ 50, the audio and video are considered out of sync.
[0054] The audio-visual synchronization detection method provided in this application, by matching the expected image features corresponding to the audio amplitude features with the actual image features corresponding to the video frames, can eliminate the reliance on easily interfered network indicators or timestamps in the prior art, and can directly judge audio-visual synchronization based on audio and video content, which can significantly improve the accuracy of audio-visual synchronization detection in road test scenarios.
[0055] Based on the above embodiments, in another embodiment provided in this application, the method may further include the following steps: In step S271, a road test data file is obtained. The road test data file contains a timestamp and the corresponding longitude, latitude, reference signal received power (RSRP), and signal-to-interference-plus-noise ratio (SINR).
[0056] In this embodiment, the road test data file can be a data file recorded during the road test (e.g., CSV or Excel format), which contains a timestamp and the corresponding longitude, latitude, RSRP, and SINR. This road test data file can be acquired through a road test terminal and shares the same time reference as the target video file.
[0057] In step S272, when it is determined that there is an event of audio-visual desynchronization, the occurrence time of the event is obtained.
[0058] Specifically, when an audio frame and its corresponding video frame are determined to be out of sync based on the matching degree, the timestamp of the audio frame (or video frame) corresponding to the out-of-sync event is recorded as the time of the event.
[0059] In step S273, the timestamp that matches the occurrence time is found in the road test data file, and the corresponding longitude and latitude are extracted as the geographical location of the event, as well as the corresponding RSRP and SINR are extracted.
[0060] In this embodiment, if a timestamp that precisely matches the time of the event exists in the road test data file, the corresponding geographic coordinates and network metrics are directly obtained; if no precise match exists, the difference in seconds between the timestamps and all timestamps is calculated, and the data corresponding to the record with the smallest difference (e.g., difference ≤ 10 seconds) and the closest time is selected.
[0061] In step S274, root cause determination is performed based on the extracted RSRP and SINR, and the determination result is output in association with the geographic location.
[0062] Specifically, the extracted RSRP is compared with a preset weak coverage threshold (e.g., -105dBm): if RSRP ≤ -105dBm, the root cause is determined to be weak coverage, indicating that there is a poor signal.
[0063] The extracted SINR is compared with a preset high quality threshold (e.g., -3dB): if SINR ≤ -3dB, the root cause is determined to be high quality (i.e., strong interference), such as the possibility of mutual interference between signals in the same frequency band.
[0064] In this embodiment, the root cause determination result is output along with the geographical location (longitude and latitude) for use by network optimization personnel. For example, the output format may be: "Longitude 116.397, Latitude 39.908, audio-visual desynchronization event occurred, root cause: weak coverage (RSRP=-110dBm)".
[0065] This embodiment, by correlating audio-visual synchronization detection results with time-geographic-network indicators in drive test data files, can automatically locate the specific geographical location where audio-visual desynchronization events occur. Based on RSRP and SINR thresholds, it identifies network root causes such as weak coverage or poor network quality, thus achieving a complete closed loop from "content-level synchronization detection" to "problem spatial location" and then to "network root cause analysis." This solution effectively addresses the shortcomings of existing technologies that "only detect without tracing the source," providing a direct and actionable decision-making basis for wireless network optimization and significantly improving the efficiency and accuracy of drive test data analysis.
[0066] Based on the above embodiments, in another embodiment provided in this application, in order to describe in detail how to calculate the amplitude characteristics of audio frames in an audio signal, the above step S210 may specifically include the following steps: In step S211, the extracted audio signal is processed into frames to obtain multiple audio frames; wherein each audio frame contains multiple sampling points.
[0067] Since audio signals are continuous time-domain signals, they need to be divided into multiple short frames for ease of analysis. Specifically, a fixed frame length can be used to divide the audio signal into frames, with each audio frame containing multiple sampling points. For example, a frame length of N = 1024 sampling points can be set. Frames can be non-overlapping, or a certain percentage of overlap (such as 50%) can be set to enhance continuity. This embodiment uses non-overlapping frames as an example. After framing, a series of audio frames arranged in chronological order are obtained.
[0068] In step S212, the root mean square amplitude of each audio frame is calculated, and the root mean square amplitude is used as the amplitude feature of the corresponding audio frame; wherein, the root mean square amplitude is calculated based on the amplitude of each sampling point within the audio frame.
[0069] The root mean square amplitude is an effective indicator of the energy intensity of an audio frame. For the first... i 1 audio frame, assuming it contains 1 audio frame N There are 3 sampling points, and the amplitude of each sampling point is denoted as x. i,1 ,x i,2 ,...,xi,N The root mean square amplitude is calculated as follows: first, calculate the square of the amplitude at each sampling point; then, calculate the average of these squared values; finally, take the square root of this average value, which is the root mean square amplitude (see Formula 1 above). The calculated RMS... i As the first i The amplitude characteristics of each audio frame are used for subsequent audio-visual synchronization matching.
[0070] This embodiment, through frame segmentation and root mean square amplitude calculation, can transform continuous audio signals into discrete, easily processed amplitude feature sequences. Root mean square amplitude, as a measure of audio energy, has good physical meaning and is computationally simple. It can effectively reflect the intensity changes of audio frames and has a certain smoothing effect on short-term fluctuations, avoiding interference from noise at individual sampling points. This provides a stable and reliable audio feature input for subsequent quantitative mapping with video RGB features, thereby improving the accuracy and robustness of overall audio-visual synchronization detection.
[0071] Based on the above embodiments, in another embodiment provided in this application, step S220 may specifically include the following steps: In step S221, the center region of the video frame in the target video file is determined.
[0072] In this embodiment, considering that the edge regions of the video frame may be affected by distortion from the shooting device, encoding compression artifacts, or edge noise interference from the road test environment, and that the central region usually contains the most important visual content, this embodiment only extracts the central region of the video frame as the feature extraction object. Specifically, assuming the width of the video frame is W pixels and the height is H pixels, a square region with a side length of R pixels centered on the frame center is taken as the central region, where R is a preset value (e.g., R=50 pixels). This size has been experimentally verified to achieve a good balance between anti-interference and computational efficiency.
[0073] In step S222, the R component, G component and B component of all pixels in the central region are extracted.
[0074] For each pixel within the central region, obtain the component values of its three color channels: red (R), green (G), and blue (B). If the original video frame format is BGR, convert it to RGB format before extraction.
[0075] In step S223, the mean values of the R component, G component, and B component of all pixels within the central region are calculated respectively.
[0076] Suppose there are M pixels in the central region, and the R, G, and B component values of the p-th pixel are respectively Rp , Gp , Bp The mean of the R components is The mean values of the G and B components are calculated in the same way.
[0077] In step S224, the combination of the mean values of the R component, the G component, and the B component is used as the actual image features of the video frame.
[0078] The actual image features of this video frame can be represented as a triplet ( R avg , G avg , B avg ),in R avg , G avg , B avg These are the average values of the three components calculated in step S223. This combined feature can effectively characterize the overall color tendency and brightness level of the central region of the video frame.
[0079] This embodiment extracts only the mean RGB components of the central region of the video frame as the actual image feature, avoiding interference from edge noise and coding artifacts, reducing computational complexity, and preserving the most crucial visual information for audio-visual synchronization detection. The size of the central region is optimized (e.g., 50×50 pixels) to achieve a good balance between anti-interference and computational efficiency. This feature extraction method is simple and stable, and has good semantic correspondence potential with audio amplitude features, providing a reliable data foundation for subsequent matching degree calculation with expected image features, thereby improving the overall performance of audio-visual synchronization detection.
[0080] Based on the above embodiments, in another embodiment provided in this application, in order to detail how to time-align audio frames and video frames and establish a correspondence between them, the above step S230 may specifically include the following steps: In step S231, the video frame rate, audio sampling rate, and audio frame length of the target video file are obtained.
[0081] In this embodiment, the video frame rate F (unit: frames / second) represents the number of video frames per second; the audio sampling rate S (unit: Hz) represents the number of audio sample points collected per second; and the frame length N (unit: sample points) of the audio frame is the number of sample points contained in each audio frame set during frame processing in step S211. These parameters can be read from the metadata of the video file or obtained directly according to the settings during preprocessing.
[0082] In step S232, the duration of a single audio frame is determined based on the audio sampling rate and the frame length of the audio frame, and the duration of a single video frame is determined based on the video frame rate.
[0083] Duration of a single audio frame T audio = N / S (Unit: seconds) Duration of a single video frame T video =1 / F (Unit: seconds). For example, when N=1024 and S=44100Hz, T audio ≈ 23.2 Taudio ≈23.2 milliseconds; when F=30 frames / second, T video ≈33.3 milliseconds.
[0084] In step S233, the start timestamp of each audio frame is obtained, and the corresponding video frame index is calculated based on the start timestamp and the video frame rate.
[0085] Each audio frame has a corresponding start timestamp on the audio timeline. t audio (Unit: seconds), this timestamp can be obtained by multiplying the audio frame number by the duration of a single audio frame: t audio =( i 1)× T audio Then, the corresponding video frame index is calculated. f = t audio × F ,in This indicates rounding down. The formula means finding the video frame on the video timeline that is closest to and no later than the start time of the audio frame.
[0086] In step S234, a temporal correspondence is established between the amplitude features of the audio frame and the image features of the corresponding video frame based on the video frame index.
[0087] Specifically, the first i Amplitude characteristics of each audio frame RMS i With index f Image features of video frames ( R avg , G avg , B avg)They are associated to form a one-to-one corresponding feature pair for subsequent matching degree calculation.
[0088] In this embodiment, through an accurate time alignment algorithm, a corresponding relationship is established between discrete audio frames and video frames on the time axis to ensure the spatio-temporal consistency of subsequent feature matching. This method is based on the mathematical relationship between the start timestamp of the audio frame and the video frame rate. By adopting a floor strategy, it ensures that each audio frame can uniquely correspond to a video frame, avoiding misalignment problems caused by frame rate mismatch or timing deviation. This alignment method does not depend on the integrity of the timestamp (even if there are minor errors in the timestamp, frame-level alignment can still work stably), and the computational complexity is extremely small, making it suitable for real-time or offline processing in road test scenarios, providing an accurate time reference for subsequent amplitude-RGB mapping matching.
[0089] Based on the above embodiment, in another embodiment provided by this application, the above step S260 may specifically include the following steps: In step S261, calculate the Euclidean distance D between the actual image feature and the expected image feature of the target video frame, and use the Euclidean distance D as the matching degree.
[0090] Specifically, the Euclidean distance D between the actual image feature and the expected image feature of the target video frame can be calculated through the above formula (3).
[0091] The smaller the distance value D, the closer the actual RGB color is to the expected RGB color, that is, the higher the audio-visual synchronization degree; conversely, the larger the distance, the lower the synchronization degree. Use the calculated D as the matching degree between the current audio frame and the corresponding video frame.
[0092] In step S262, when the distance is less than the preset threshold, it is determined that the target audio frame and the target video frame are audio-visual synchronized.
[0093] The preset threshold T can be determined through the statistical distribution of a large number of normal synchronization samples. For example, the 3σ criterion is adopted: collect 1000 groups of audio-visual samples with normal synchronization, calculate the matching distance of each group of samples, obtain the mean μ and the standard deviation σ, then the threshold T = μ + 3σ. Through experimental verification, the preferred threshold T = 50. When D < T, it is determined as audio-visual synchronization.
[0094] In step S263, when the distance is not less than the preset threshold, it is determined that the target audio frame and the target video frame are not audio-visual synchronized.
[0095] That is, when D ≥ T, it is determined as not audio-visual synchronized. At this time, this asynchronous event and its corresponding time information can be recorded for subsequent geographical positioning and root cause tracing.
[0096] This embodiment uses the Euclidean distance between the actual RGB and the expected RGB values as the matching degree to achieve a quantitative assessment of audio-visual synchronization. The Euclidean distance comprehensively reflects the overall difference among the three color components and has the advantages of simple calculation and clear physical meaning. A threshold determined based on the 3σ criterion is used for judgment, fully utilizing the statistical characteristics of normal synchronization samples. This adaptively balances missed detections and false alarms, avoiding the subjectivity of manually setting thresholds. This judgment method does not rely on any network metrics or timestamps, but directly relies on the characteristics of audio and video content. It has high detection accuracy (synchronization error ≤20ms) and is applicable to various complex drive test scenarios such as static images and stable audio, providing reliable synchronization status input for subsequent geolocation and root cause analysis.
[0097] By dividing each functional module according to its corresponding function, this application provides an audio-visual synchronization detection device, which can be a server, a terminal, or a chip applied to a server. Figure 3 This is a schematic block diagram of the functional modules of an audio-visual synchronization detection device provided for an exemplary embodiment of this application. Figure 3 As shown, the audio-visual synchronization detection device includes: File acquisition module 31 is used to acquire target video files in road test scenarios; The audio feature acquisition module 32 is used to extract audio signals from the target video file and calculate the amplitude features of audio frames in the audio signal; The actual image feature acquisition module 33 is used to extract the actual image features of video frames from the target video file; The correspondence establishment module 34 is used to time-align audio frames and video frames and establish the correspondence between audio frames and video frames. The model acquisition module 35 is used to acquire the feature mapping model. The feature mapping model contains the correspondence between multiple continuous amplitude intervals and expected image features. The feature mapping model is trained based on normally synchronized audio and video samples. The expected image feature acquisition module 36 is used to determine the corresponding expected image features based on the amplitude range of the amplitude features of the target audio frame and the target video frame that have a corresponding relationship, through a feature mapping model. The audio-visual synchronization judgment module 37 is used to calculate the matching degree between the actual image features and the expected image features of the target video frame, and to judge whether the target audio frame and the target video frame are synchronized based on the matching degree.
[0098] In another embodiment provided in this application, the device further includes a root cause analysis module, specifically used for: Obtain the drive test data file, which contains timestamps and the corresponding longitude, latitude, reference signal received power (RSRP), and signal-to-interference-plus-noise ratio (SINR). When an event of audio-visual desynchronization is confirmed, the time of occurrence of the event is obtained; Find the timestamp that matches the occurrence time in the road test data file, and extract the corresponding longitude and latitude as the geographical location of the event, as well as the corresponding RSRP and SINR. Root cause analysis is performed based on the extracted RSRP and SINR, and the results are output in association with geographic location.
[0099] In another embodiment provided in this application, the audio feature acquisition module 32 is specifically used for: The extracted audio signal is segmented into frames to obtain multiple audio frames; each audio frame contains multiple sampling points. Calculate the root mean square amplitude of each audio frame and use the root mean square amplitude as the amplitude feature of the corresponding audio frame; where the root mean square amplitude is calculated based on the amplitude of each sampling point within the audio frame.
[0100] In another embodiment provided in this application, the video frame is an RGB image, and the actual image feature acquisition module 33 is specifically used for: Determine the center region of the video frames in the target video file; Extract the R, G, and B components of all pixels within the central region; Calculate the mean R component, mean G component, and mean B component of all pixels within the central region respectively; The combination of the mean values of the R component, G component, and B component is used as the actual image features of the video frame.
[0101] In another embodiment provided in this application, the correspondence establishment module 34 is specifically used for: Obtain the video frame rate, audio sampling rate, and audio frame length of the target video file; The duration of a single audio frame is determined based on the audio sampling rate and the frame length of the audio frame, and the duration of a single video frame is determined based on the video frame rate. Obtain the start timestamp of each audio frame, and calculate the corresponding video frame index based on the start timestamp and video frame rate; Establish a temporal correspondence between the amplitude features of audio frames and the image features of the corresponding video frames based on video frame indexing.
[0102] In another embodiment provided in this application, the audio-visual synchronization judgment module 37 is specifically used for: Calculate the Euclidean distance between the actual image features and the expected image features of the target video frame, and use the Euclidean distance as the matching degree; When the distance is less than a preset threshold, the target audio frame and the target video frame are determined to be in audio-visual synchronization. Alternatively, if the distance is not less than a preset threshold, it can be determined that the target audio frame and the target video frame are out of sync.
[0103] The audio-visual synchronization detection device provided in this application, by matching the expected image features corresponding to the audio amplitude features with the actual image features corresponding to the video frames, can eliminate the reliance on easily interfered network indicators or timestamps in the prior art, and can directly judge audio-visual synchronization based on audio and video content, which can significantly improve the accuracy of audio-visual synchronization detection in road test scenarios.
[0104] This application also provides an electronic device, including: at least one processor; a memory for storing executable instructions of the at least one processor; wherein the at least one processor is configured to execute the instructions to implement the method disclosed in the embodiments of this application.
[0105] Figure 4 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application. For example... Figure 4 As shown, the electronic device 1800 includes at least one processor 1801 and a memory 1802 coupled to the processor 1801. The processor 1801 can perform the corresponding steps in the methods disclosed in the embodiments of this application.
[0106] The processor 1801 described above can also be called a central processing unit (CPU), which can be an integrated circuit chip with signal processing capabilities. Each step in the method disclosed in this application can be implemented by the integrated logic circuitry in the hardware of the processor 1801 or by instructions in software form. The processor 1801 can be a general-purpose processor, a digital signal processor (DSP), an ASIC (Application Specific Integrated Circuit), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can be located in the memory 1802, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The processor 1801 reads information from the memory 1802 and, in conjunction with its hardware, completes the steps of the above method.
[0107] Furthermore, the various operations / processes according to this application, when implemented via software and / or firmware, can be transmitted from a storage medium or network to a computer system with a dedicated hardware architecture, such as... Figure 5 The computer system 1900 shown is equipped with the programs that constitute the software. When various programs are installed, the computer system is able to perform various functions, including those described above. Figure 5 A structural block diagram of a computer system provided for an exemplary embodiment of this application.
[0108] Computer System 1900 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices can also represent various forms of mobile devices, such as cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present application described and / or claimed herein.
[0109] like Figure 5As shown, the computer system 1900 includes a computing unit 1901, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 1902 or a computer program loaded from a storage unit 1908 into a random access memory (RAM) 1903. The RAM 1903 may also store various programs and data required for the operation of the computer system 1900. The computing unit 1901, ROM 1902, and RAM 1903 are interconnected via a bus 1904. An input / output (I / O) interface 1905 is also connected to the bus 1904.
[0110] Multiple components in computer system 1900 are connected to I / O interface 1905, including: input unit 1906, output unit 1907, storage unit 1908, and communication unit 1909. Input unit 1906 can be any type of device capable of inputting information into computer system 1900. Input unit 1906 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of the electronic device. Output unit 1907 can be any type of device capable of presenting information and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 1908 may include, but is not limited to, hard disks and optical disks. Communication unit 1909 allows computer system 1900 to exchange information / data with other devices via a network such as the Internet, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0111] The computing unit 1901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1901 performs the various methods and processes described above. For example, in some embodiments, the methods disclosed in the embodiments of this application can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 1908. In some embodiments, part or all of the computer program can be loaded and / or installed on an electronic device via ROM 1902 and / or communication unit 1909. In some embodiments, the computing unit 1901 can be configured to perform the methods disclosed in the embodiments of this application by any other suitable means (e.g., by means of firmware).
[0112] This application also provides a computer-readable storage medium, wherein when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is able to perform the methods disclosed in this application.
[0113] The computer-readable storage medium in this application embodiment may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. The aforementioned computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specifically, the aforementioned computer-readable storage medium may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing.
[0114] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0115] This application also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the methods disclosed in the embodiments of this application.
[0116] In embodiments of this application, computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer.
[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0118] The modules, components, or units described in the embodiments of this application can be implemented in software or hardware. The names of the modules, components, or units do not necessarily constitute a limitation on the module, component, or unit itself.
[0119] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0120] The above description is merely an embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.
[0121] While specific embodiments of this application have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and are not intended to limit the scope of this application. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this application. The scope of this application is defined by the appended claims.
Claims
1. A method for detecting audio-visual synchronization, characterized in that, The method includes: Obtain the target video file in the road test scenario, extract the audio signal from the target video file, and calculate the amplitude characteristics of the audio frames in the audio signal; Extract the actual image features of the video frames from the target video file; The audio frames and video frames are time-aligned to establish a correspondence between them; A feature mapping model is obtained, which contains the correspondence between multiple continuous amplitude intervals and expected image features. The feature mapping model is trained based on normally synchronized audio and video samples. For target audio frames and target video frames that have the aforementioned correspondence, the corresponding expected image features are determined by the feature mapping model based on the amplitude range in which the amplitude features of the target audio frame are located. Calculate the matching degree between the actual image features of the target video frame and the expected image features, and determine whether the target audio frame and the target video frame are synchronized based on the matching degree.
2. The method according to claim 1, characterized in that, The method further includes: Obtain a road test data file, which includes a timestamp and the corresponding longitude, latitude, reference signal received power (RSRP), and signal-to-interference-plus-noise ratio (SINR). When an event of audio-visual asynchrony is determined, the occurrence time of the event is obtained; The timestamp matching the occurrence time is found in the road test data file, and the corresponding longitude and latitude are extracted as the geographical location of the event, as well as the corresponding RSRP and SINR are extracted. Root cause analysis is performed based on the extracted RSRP and SINR, and the analysis results are output in association with the geographic location.
3. The method according to claim 1, characterized in that, The calculation of the amplitude characteristics of audio frames in the audio signal includes: The extracted audio signal is segmented into frames to obtain multiple audio frames; each audio frame contains multiple sampling points. Calculate the root mean square amplitude of each audio frame and use the root mean square amplitude as the amplitude feature of the corresponding audio frame; wherein the root mean square amplitude is calculated based on the amplitude of each sampling point within the audio frame.
4. The method according to claim 1, characterized in that, The video frame is an RGB image, and the extraction of the actual image features of the video frame from the target video file includes: Determine the central region of the video frames in the target video file; Extract the R, G, and B components of all pixels within the central region; Calculate the mean R component, mean G component, and mean B component of all pixels within the central region respectively; The combination of the mean values of the R component, G component, and B component is used as the actual image feature of the video frame.
5. The method according to claim 1, characterized in that, The step of aligning the audio frames with the video frames in time and establishing a correspondence between the audio frames and the video frames includes: Obtain the video frame rate, audio sampling rate, and audio frame length of the target video file; The duration of a single audio frame is determined based on the audio sampling rate and the frame length of the audio frame, and the duration of a single video frame is determined based on the video frame rate. Obtain the start timestamp of each audio frame, and calculate the corresponding video frame index based on the start timestamp and the video frame rate; Based on the video frame index, a temporal correspondence is established between the amplitude features of the audio frame and the image features of the corresponding video frame.
6. The method according to claim 1, characterized in that, The step of calculating the matching degree between the actual image features and the expected image features of the target video frame, and determining whether the target audio frame and the target video frame are synchronized based on the matching degree, includes: Calculate the Euclidean distance between the actual image features and the expected image features of the target video frame, and use the Euclidean distance as the matching degree; When the distance is less than a preset threshold, it is determined that the target audio frame and the target video frame are in audio-visual synchronization. Alternatively, if the distance is not less than a preset threshold, it is determined that the target audio frame and the target video frame are out of sync.
7. A device for detecting audio-visual synchronization, characterized in that, The device includes: The file acquisition module is used to acquire target video files in road test scenarios; An audio feature acquisition module is used to extract audio signals from the target video file and calculate the amplitude features of audio frames in the audio signals; The actual image feature acquisition module is used to extract the actual image features of video frames from the target video file; The correspondence establishment module is used to time-align the audio frames with the video frames and establish a correspondence between the audio frames and the video frames; The model acquisition module is used to acquire a feature mapping model, which contains the correspondence between multiple continuous amplitude intervals and expected image features. The feature mapping model is trained based on normally synchronized audio and video samples. The expected image feature acquisition module is used to determine the corresponding expected image features for target audio frames and target video frames that have the aforementioned correspondence, based on the amplitude range in which the amplitude features of the target audio frame are located, through the feature mapping model. The audio-visual synchronization judgment module is used to calculate the matching degree between the actual image features of the target video frame and the expected image features, and to determine whether the target audio frame and the target video frame are in audio-visual synchronization based on the matching degree.
8. An electronic device, characterized in that, include: At least one processor; Memory for storing the at least one processor-executable instruction; The at least one processor is configured to execute the instructions to implement the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method as described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-6.