Methods, equipment and media for synchronous synthesis and adaptive display of audio and video in rail flaw detection

By establishing a global time reference and time synchronization processing, the problem of time synchronization deviation in the merging of audio and video in rail flaw detection was solved, achieving millisecond-level synchronization of audio and video and ensuring the accuracy of flaw detection judgment.

CN122093590APending Publication Date: 2026-05-26SICHUAN CRUNGOO INFORMATION ENG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN CRUNGOO INFORMATION ENG CO LTD
Filing Date
2026-04-07
Publication Date
2026-05-26

Smart Images

  • Figure CN122093590A_ABST
    Figure CN122093590A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of rail inspection and maintenance technology, and discloses a method, equipment, and medium for synchronized audio-visual synthesis and adaptive display of rail flaw detection. The method includes: parsing audio and video files to determine a global time reference; performing time synchronization processing on various types of files; determining the size of the composite canvas, video position, and text label position based on layout configuration rules; superimposing the time-synchronized video data onto the background canvas according to the determined layout position to generate a composite video; and merging the composite video with the time-synchronized unified audio data to generate a merged flaw detection video. This invention, by establishing a globally unified time reference, calculating the time difference and implementing synchronization schemes for audio and video separately, and combining standardized preprocessing with adaptive image synthesis, solves the time synchronization deviation problem existing in the merging of audio and video for rail flaw detection in the prior art, significantly improving the accuracy of rail flaw detection judgment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of track inspection and maintenance technology, and specifically discloses a method, equipment and medium for synchronous synthesis and adaptive display of audio and video for rail flaw detection. Background Technology

[0002] Railway rail flaw detection operations require recording multiple media files, including flaw detection waveform video, on-site operation video, and audio. The merged audio and video are then played back for damage assessment. The time synchronization accuracy of the audio and video directly affects the expert's judgment on the presence of rail damage. Currently, when merging rail flaw detection audio and video, inconsistencies in start times and discontinuous recording periods from multiple devices lead to time synchronization discrepancies in the merged audio and video. This fails to meet the flaw detection experts' requirements for comparing waveforms, probe locations, and weld positions, easily resulting in missed or misjudged rail damage, threatening train operation safety. Summary of the Invention

[0003] In view of this, the present invention provides a method, device and medium for synchronous synthesis and adaptive display of audio and video for rail flaw detection, in order to solve the above problems.

[0004] The specific implementation of this invention is as follows: A method for synchronous synthesis and adaptive display of multi-channel audio and video in rail flaw detection includes: The time and duration information of audio and video files are analyzed, and a global time base is determined based on the time and duration information. Classify files into video and audio types, and perform time synchronization processing on each type of file. Based on the number of media types of the videos to be merged, preset layout configuration rules are established; based on the layout configuration rules, the size of the compositing canvas, the layout position of each video, and the display position of the corresponding video type text labels are determined. Based on the determined canvas size and the total duration corresponding to the global time base, a background canvas is created; the time-synchronized video data from each channel is superimposed onto the background canvas according to the determined layout position to generate a composite video. The synthesized video is merged with the time-synchronized unified audio data to generate a flaw detection merged video.

[0005] As an optional method, the global time base includes the earliest start time, the latest end time, and the total duration, where the total duration is the difference between the latest end time and the earliest start time; the time information is the start time of each audio and video file.

[0006] As an optional method, time synchronization processing for video files includes: Get the time offset of the first video relative to the earliest start time of the global time base, the time offset of the subsequent videos relative to the end time of the previous video, and the end time offset of the last video relative to the latest end time of the global time base. If a time offset exists, generate a compensation video that matches the offset duration and stitch the two corresponding consecutive videos together; If the last video has an end time offset, add a compensation video of the corresponding duration at the end of it to make the total duration of the processed videos of the same type consistent with the total duration of the global time base.

[0007] As an optional method, time synchronization processing for audio files includes: Obtain the time offset of each audio file relative to the earliest start time of the global time base. If a time offset exists, add a silent segment of the corresponding duration at the beginning of the audio file to form time-synchronized audio data. Mix all time-synchronized audio data to generate unified audio data.

[0008] As an optional method, the compensation video is a black screen video containing preset identification information; after the black screen video is generated, all original video files and the black screen compensation video are processed with preset uniform resolution and uniform color space.

[0009] As an optional approach, video aspect ratio adaptation processing is also included, including: Obtain the original aspect ratio of each video and compare it with the preset standard aspect ratio; The video image is scaled proportionally to the aspect ratio of the preset standard resolution, and black borders of equal width are added to fill the difference in size between the scaled and preset standard resolution areas.

[0010] As an optional approach, the layout configuration rules include layout specifications, spacing specifications, text label area specifications, and canvas boundary specifications. The layout specifications include row number configuration and column number configuration, and the spacing specifications include horizontal and vertical spacing configuration between videos.

[0011] As an optional method, the size of the composite canvas, the layout of each video stream, and the display position of the corresponding video type text identifier are determined, including: Extract the number of rows and columns, spacing specifications, text area specifications, and canvas boundary specifications from the preset layout configuration rules; Based on the above specifications and the unified video specifications, calculate the total width and total height of the composite canvas; Calculate the layout position of each video stream sequentially by row and column, and then obtain the display coordinates of the corresponding text label based on the layout position of each video stream.

[0012] On the other hand, the present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of a method for synchronous synthesis and adaptive display of multi-channel audio and video in rail flaw detection.

[0013] On the other hand, the present invention also provides a storage medium storing a computer program, wherein when the computer program is executed by a processor, the steps of a method for synchronous synthesis and adaptive display of multi-channel audio and video in rail flaw detection are implemented.

[0014] The beneficial effects of this invention are as follows: This invention solves the time synchronization deviation problem in existing rail flaw detection audio-video merging technologies by establishing a globally unified time reference, calculating the time difference and implementing synchronization schemes for audio and video separately, and combining standardized preprocessing with adaptive image synthesis. The processed audio and video achieve millisecond-level time synchronization, ensuring that flaw detection experts can compare waveforms, probe positions, and weld conditions across multiple video streams during playback. This avoids missed or incorrect damage detection due to time deviations, significantly improving the accuracy of rail flaw detection judgment. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the process for the rail flaw detection multi-channel audio and video synchronous synthesis and adaptive display method of the present invention; Figure 2 This is a schematic diagram of a screen layout effect according to the present invention. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been presented in the various embodiments of the present invention to enable the reader to better understand the present invention. However, the technical solutions claimed in the present invention can be implemented even without these technical details and various changes and modifications based on the following embodiments.

[0017] Example 1 During rail flaw detection operations, multiple audio and video files are generated from different recording devices, each with varying start times and durations. To overcome the drawbacks of manual or network synthesis, it is necessary to define the temporal attributes of each file. In this embodiment, the specific duration of each audio and video file is extracted, and the end time of each file is calculated based on its start time. Based on this extracted and calculated time data, this embodiment establishes a global time reference.

[0018] Specifically, in this embodiment, the processing medium (computer or flaw detection equipment) aggregates the start times of all files and selects the earliest time as the global earliest start time. This time serves as the actual start node (time) of this flaw detection operation. Simultaneously, it aggregates the end times of all files and selects the latest time as the global latest end time. This time corresponds to the actual end node (time) of this flaw detection operation. The duration between the global earliest start time and the global latest end time is the global total duration. This total duration will serve as the standard duration for the merged audio and video, providing a unified time reference for the subsequent synchronous processing of all audio and video.

[0019] After establishing a global time baseline, this embodiment categorizes video files by media type. Video files of the same media type are arranged chronologically, and time differences are calculated using a chained approach. When processing the first video, its start time is compared with the earliest global start time. If a difference exists, a black screen compensation video is generated. The compensation video maintains the same specifications as other videos, displaying a text message indicating no video in the center of the screen, and its duration precisely fills the gap. When processing subsequent videos, the start time of the current video is compared with the end time of the previous video. If an interval exists, a corresponding black screen compensation video is inserted to ensure temporal continuity and consistency between video segments of the same media type. When processing the last video of the media type, its end time is compared with the latest global end time. If there is remaining time, a black screen compensation video is added to the end of that video, ensuring that the total duration of the entire media type's video sequence perfectly matches the global total duration.

[0020] Through the above steps, the originally scattered and inconsistent video clips are integrated into continuous video data. All videos are aligned with the global time reference, avoiding the problems of time discontinuity or overlap caused by disrupting time continuity, and providing basic resource preparation for subsequent adaptive layout and audio-visual synthesis.

[0021] As an optional method, the audio and video files to be processed are all named according to standardized naming conventions, which may include device serial numbers, job location information, and start recording time. During time parsing, the filenames are split according to this naming rule, and the start recording time is extracted as the basic time information. At the same time, the precise duration of each file is read using corresponding tools or software, and the end time of each file is further calculated by combining the start time, ensuring that the time range of each file is clear. The extracted start recording time is standardized and converted to the format of year, month, day, hour, minute, second, and millisecond. According to the general implementation scenario of this embodiment, only valid files with complete millisecond precision are retained.

[0022] Based on the time ranges of all files obtained through parsing, preparations begin to establish a global time benchmark. By summarizing the start times of all files, the earliest time is selected as the global earliest start time, corresponding to the actual start node of this flaw detection operation. Similarly, by summarizing the end times of all files, the latest time is selected as the global latest end time, corresponding to the actual end node of the operation. The difference between the two is the global total duration, which becomes the standard duration of the merged audio and video. The design logic of the global time benchmark is to establish a unique time scale, providing a unified reference for the synchronous processing of all audio and video files, thereby fundamentally avoiding synchronization deviation problems caused by inconsistent recording start and end times.

[0023] After establishing the timeline, video files are categorized by media type. When processing video files of a specific media type, it's first determined whether the current video is the first file of that type. If it is, the earliest global start time is used as a reference to calculate the time offset between its start time and the baseline's beginning, ensuring the first file is anchored to the global timeline. If it's not the first file, the end time of the previous video is used as a reference to calculate the time offset. This chained reference method ensures that video files of the same type are sequentially linked. When the calculated time offset is greater than 0, it indicates a time interval between the current video and the previous time point. In this case, a black screen compensation video that perfectly matches the time offset is generated to fill the gap. The size of the compensation video is consistent with the subsequently standardized video, and a preset text prompt is added to the center of the screen. This avoids black screen periods without prompts during playback and clearly informs experts that there is no valid video data for that period.

[0024] After stitching the current video and any compensation video (if applicable), the stitched result is used as the basis for processing the next video. This process of judgment and stitching is repeated until the last video of this type is processed. For the last video, the end time offset between its end time and the global latest end time needs to be calculated. If this offset is greater than 0, it means there is still remaining time after the last video ends. A black screen compensation video of corresponding duration needs to be added to the end of this video. Ultimately, all video files of this media type are stitched together to form a continuous video data sequence, and the total duration of the sequence is completely consistent with the global total duration. Throughout this process, all video files also undergo specification standardization processing simultaneously. The image is scaled proportionally to a preset standard resolution, black borders are added to areas with size differences, and the color space standard is unified to ensure consistent display specifications for all videos during subsequent adaptive layout.

[0025] The audio synchronization process follows the same logic as video synchronization. After file parsing and grouping, audio files are categorized according to media type identifiers to form an ordered audio sequence. At this point, each audio stream knows its own start time, end time, and duration. However, since multiple audio streams come from different recording devices, their start recording times may differ. Directly merging them would lead to sound misalignment. Therefore, targeted synchronization processing based on a global time reference is required.

[0026] When processing a single audio stream, the time offset between the audio stream's start time and the earliest global start time is calculated, using the earliest global start time as a reference. If an offset exists, it means the audio was recorded later than the start time of the task. In this case, a silent segment is generated that perfectly matches the offset duration. The silent segment uses a zero-value audio signal, maintaining the same sampling rate and channel configuration as the original audio to avoid audio quality conflicts during mixing. The silent segment is then spliced ​​to the beginning of the original audio, aligning the effective sound portion of the audio with the global timeline, ensuring that the corresponding audio content can be heard synchronously from the earliest global start time.

[0027] After completing the time synchronization of a single audio stream, the multi-channel audio mixing stage begins. This embodiment considers the need to simultaneously preserve multiple audio streams, including ambient sounds and equipment operation sounds, during flaw detection operations. The mixing process must maintain the sound details of each audio stream and avoid mutual interference. Audio mixing technology integrates all time-synchronized audio streams into a single unified audio stream. During mixing, parameters such as the audio sampling rate and number of channels remain consistent to ensure clear sound quality of the output unified audio stream, fully reproducing the sound scene of the flaw detection operation. At this point, the total duration of the unified audio stream strictly matches the overall global duration, preventing premature sound termination and the generation of unnecessary silent audio segments.

[0028] Having completed the aforementioned audio and video pre-processing, this embodiment will now describe the adaptive screen playback process. The purpose of adaptive screen playback is to address the layout chaos caused by the variable number and specifications of video streams during flaw detection operations. Through dynamic calculation, the screen automatically adapts to different numbers of videos, ensuring a neat layout and clear viewing, thus providing a good visual experience for flaw detection experts.

[0029] Creating a background canvas is fundamental to the steps described above. Its purpose is to provide a unified and stable platform for the orderly overlay of multiple video streams, ensuring a neat layout and complete time coverage. The canvas size should strictly adhere to the calculations for adaptive screen layout. This means the total width is calculated by adding the left and right margins, the width of all videos, and the horizontal gaps; the total height is calculated by adding the bottom margin, the height of all text areas, and the video heights. This size design perfectly accommodates all videos, text labels, and spacing, without any redundant space or elements exceeding the canvas's limits. The canvas background color is set to white by default. This choice clearly presents the video content and text labels, avoiding color interference that could affect the viewing experience of inspection experts. Furthermore, the canvas's duration is completely consistent with the total duration corresponding to the global time base, ensuring a complete canvas coverage from the earliest start time to the latest end time, without any time gaps.

[0030] The above process requires first clarifying the various basic configuration parameters. These parameters can be preset to default values ​​for common flaw detection scenarios, or flexibly adjusted according to actual viewing needs. Please refer to... Figure 2 To achieve the display effect shown in the figure, the basic configuration parameters used in this embodiment are explained as follows: Define basic configuration parameters: The number of rows refers to the number of horizontal rows in the layout, and the number of columns refers to the number of vertical rows. Together, they constitute the layout style of row and column combinations. The width and height of a single video are in pixels, indicating the standard display size of each video. The horizontal gap is the horizontal spacing between adjacent videos, and the vertical gap is the vertical spacing between adjacent videos. The text area height is the space above each video used to display the type identifier, ensuring that experts can quickly identify the video source. The boundary spacing is the distance between the outermost video and the edge of the canvas, avoiding videos that are too close to the edge of the canvas and affect the viewing experience. All of the above spacing parameters are in pixels.

[0031] Based on typical viewing requirements, the canvas must be able to fully accommodate all video, text areas, and spacing elements. Elements cannot exceed the canvas's limits, nor can excessive redundant space negatively impact the viewing experience. Therefore, after determining the basic parameters, the next step is to calculate the total size of the composite canvas.

[0032] The calculation of the total canvas width needs to comprehensively consider the left and right boundary spacing, the width of all videos, and the horizontal gap between videos. A boundary spacing is reserved on each side, and the middle portion is the sum of the width of the number of videos in each column and the number of columns minus one horizontal gap. This calculation method ensures that all horizontal elements are tightly and reasonably arranged, without omissions or overlaps. The calculation of the total canvas height must include the bottom boundary spacing, the height of the text area in all rows, and the height of all videos. A corresponding text area needs to be reserved above each row of videos, so the text area height needs to be accumulated according to the number of rows, plus the height of all videos and the bottom boundary spacing, ensuring that the vertical elements are also neatly arranged and coordinated with the horizontal layout, resulting in a smooth viewing experience when multiple videos are played.

[0033] Next, the coordinates of each video are calculated one by one. As an optional method, in this embodiment, the coordinates are based on the top left corner of the canvas, with the horizontal axis as the X-axis and the vertical axis as the Y-axis. For the video located in the i-th row and j-th column, the X-coordinate of its top left corner is calculated starting from the left boundary interval, plus the sum of the width of the previous j-1 columns of videos and the horizontal gap, ensuring that the videos in the same row are arranged to the right with a fixed interval and are horizontally aligned. The Y-coordinate is calculated based on the height of the text area, plus the sum of the height of the text area of ​​the previous i-1 rows, the height of the video, and the vertical gap, reserving space for the text area above each row of videos while ensuring uniform vertical spacing between different rows of videos and orderly vertical arrangement.

[0034] After determining the video coordinates, the coordinates of the text area need to be linked to the corresponding video to ensure that the text label intuitively corresponds to the video content (equivalent to a title), facilitating quick identification by experts. The width of the text area in each video is exactly the same as the width of the corresponding video, and its horizontal position is aligned with the video. Therefore, the X-coordinate of the text area is the same as the X-coordinate of the corresponding video. The text area is located directly above the video, so its Y-coordinate is the Y-coordinate of the corresponding video minus the height of the text area. This ensures that the text title does not overlap with the video content and is linked to the content, improving the viewing experience.

[0035] After creating the background canvas, video overlay begins. The synchronized video data is presented on the canvas according to a preset layout. The overlay process, based on the previously calculated coordinates of each video stream, places the consecutive video data sequences one by one into their corresponding rows and columns, using the top-left corner of the canvas as the origin. This ensures uniform horizontal spacing within the same row and consistent vertical spacing within the same column, with no overlap or offset in the video layout. For cases where a video file of a certain media type is missing, a pre-generated black screen compensation video is overlaid according to its corresponding coordinates. This compensation video conforms to the standard video specifications, and a text prompt in the center of the screen indicates that there is no valid video data at that location to avoid misunderstanding. This embodiment uses layer overlay technology, with each video stream placed as an independent layer on the background canvas. When a video stream ends or a time interval occurs, the playback of other videos remains unaffected, maintaining the integrity and continuity of the image, ultimately forming a well-structured and clearly presented composite video.

[0036] After the synthesized video is generated, it needs to be merged with the time-synchronized unified audio data to achieve millisecond-level audio-visual synchronization, which is generally required in the scenario described in this embodiment. This is crucial for ensuring the accuracy of flaw detection judgment. The unified audio data is the result of time synchronization and mixing of multiple audio streams, containing all sound information from the flaw detection site, and its total duration perfectly matches the global total duration. During the merging process, a video coding standard with strong compatibility and high compression efficiency (such as Advanced Video Coding, High Efficiency Video Coding, VP8, VP9, ​​AV1, etc.) and pixel format (such as YUV4:2:0, YUV4:2:2, etc.) can be selected; this embodiment does not impose any restrictions on this. The timestamp alignment technology used during merging strictly matches each frame of the synthesized video with the corresponding time point of the unified audio data, completely eliminating audio-visual misalignment problems caused by network latency and differences in device performance. The final generated flaw detection merged video not only achieves an orderly layout of multiple video streams and clear text labeling, but also achieves millisecond-level audio-visual synchronization. Furthermore, the video file name follows a standardized format and includes key information such as equipment serial number, operation location, and timestamp, which facilitates subsequent storage, archiving, and association of flaw detection records. This ensures that experts can accurately compare waveforms, probes, and weld locations during playback and make accurate damage judgments.

[0037] Example 2 This embodiment further explains the file merging and video compensation in Embodiment 1. After collecting the required video files, the files are initially standardized in a corresponding processing medium (which, depending on the scenario, could be a computer, flaw detection equipment, etc.).

[0038] This embodiment parses media files according to a standardized naming format. The parsed filename format is set to "devicerecord-device serial number-weld number-line number-rail number-mileage-track-side-timestamp-file type", extracting information such as device serial number, weld number, line number, rail number, mileage, track, side, timestamp, and file type to establish a complete file metadata structure including location, time, and type. This structured design allows the script to parse in batches according to fixed rules, avoiding recognition errors caused by chaotic identifier formats.

[0039] In one optional method, the filenames for a specific work scenario uniformly follow a fixed format: "devicerecord-device serial number-weld number-track number-rail number-mileage marker-track number-side marker-start recording time-media type". Each field is separated by a hyphen, comprising 10 components, as detailed below: The prefix "devicerecord" is used to distinguish audio and video files specifically for flaw detection operations, preventing them from being confused with other unrelated files. Device serial number: uniquely identifies the recording device and prevents files from different flaw detection devices from being mixed up; Weld number, track number, rail number, mileage marker, track number, side marker: rail information used in flaw detection operations to distinguish different rail markings; Start recording time: This is the time when video recording begins, used to identify the uniform start time of the video; Media type: Differentiate file attributes and use fixed numerical encoding. 00 is on-site audio, 01 is left camera video, 02 is right camera video, 03 is flaw detector screen video, 04 is endoscope video, and so on to identify different video types.

[0040] Next, video merging is performed. First, a list of files to be processed is obtained through a preset file path (local save directory for non-flaw detection equipment, and the device's preset storage path for flaw detection equipment). Then, the parsing operation is performed according to the following steps: 1. Perform initial filtering on the obtained file list, automatically excluding files whose filenames contain the "Converted" prefix. In this embodiment, the "Converted" prefix is ​​used to identify already processed original source files to avoid duplicate reading. After filtering, a list of files to be parsed is generated, which will be the processing objects for subsequent parsing operations.

[0041] 2. By splitting the string, the filename of each file to be parsed is split according to the hyphen, generating a parameter array `parts_array` containing 10 elements. The splitting logic is based on a fixed structure of standardized naming rules, with each element corresponding to a preset information field.

[0042] For example, if a file is named "devicerecord-SN001-W001-L001-R001-M100-T1-L-20250201110313010-01.mp4", the automatic parsing process in this embodiment is as follows: After performing string splitting, the parameter array is: parts_array=['devicerecord','SN001','W001','L001','R001','M100','T1','L','20250201110313010','01']; The information corresponding to each element is as follows: prefix devicerecord, device serial number SN001, weld number W001, line number L001, rail number R001, mileage marker M100, track number T1, side marker L, start recording time 20250201110313010, media type 01 (left camera / video); Based on the file extension ".mp4", the file can be fully identified as: February 1, 2025, 11:03:13:010 milliseconds, recorded by device with serial number SN001, corresponding to the left camera video of line L001, weld W001, mileage M100, track T1, and side L.

[0043] Furthermore, the start recording time field is a fundamental standard for time synchronization. This embodiment first checks the length of the timestamp string, retaining only strings with a length of 17 characters. The design logic is that a 17-character string corresponds to the complete format of year (first 4 characters), month (5-6 characters), day (7-8 characters), hour (9-10 characters), minute (11-12 characters), second (13-14 characters), and millisecond (15-17 characters). If the length is less than 17 characters, it indicates that the time information of the file is missing or incomplete during recording, and such files will be automatically excluded to avoid affecting subsequent synchronization and merging.

[0044] The specific steps are as follows: For example, in the above 20250201110313010, the first 4 digits are determined to be the year, the 5th and 6th digits to be the month, the 7th and 8th digits to be the day, the 9th and 10th digits to be the hour, the 11th and 12th digits to be the minute, the 13th and 14th digits to be the second, and the 15th and 17th digits to be the millisecond. After multiplying the millisecond part by 1000 to convert it to microsecond precision, a standardized time object is constructed based on the split time fields to form a unified time base data.

[0045] Considering that different files from the same work site may have overlapping or consecutive recording times, after obtaining the standardized time object, this embodiment also needs to further obtain the accurate duration, start time, and complete time interval of the file. Optionally, a video length extraction tool can be called to read the precise duration of each audio and video file, the extracted file duration can be added to the standardized start time, the end time of the file can be calculated, and finally a complete time interval object containing the start time, duration, and end time can be generated.

[0046] For example, let's take a video file with a start time of 2025-02-01 11:03:13.100000 and an extraction duration of 30.5 seconds as an example: The end time is calculated as 11:03:13.100 + 30.5 seconds = 11:03:43.600; The time interval object is calculated as: Time interval object = [2025-02-01 11:03:13.100, 30.5, 2025-02-01 11:03:43.600].

[0047] Based on the technical field of this embodiment, in actual scenarios, flaw detection operations are usually carried out in segments according to specific lines / welds (spatial dispersion), with multiple devices recording simultaneously within the same work segment (multiple data sources), and the recording process may generate scattered files due to device switching or pausing (discontinuous time). Therefore, the script in this embodiment will also perform three-level grouping to organize the above-mentioned disordered audio and video files into an ordered set of the same work area, the same flaw detection time period, and the same type of data.

[0048] Specifically, for location grouping, this embodiment aims to isolate audio and video files from different work areas to ensure that all subsequently merged files correspond to the same rail work point. This avoids mismatches between merged files and flaw detection records caused by mixed processing, which directly affects the accuracy of subsequent damage verification. This embodiment constructs equipment information key values ​​by extracting seven key parameters from the parameter array: First, the standardized identifier of each file to be processed is parsed. The seven types of parameters mentioned above are extracted through string segmentation to construct a composite identification key, `device_key` (device serial number, weld number, line number, rail number, mileage, track number, and side identifier). Then, files with completely identical identification keys are grouped into the same location group. For example, all files with the parsed identifier parameter "SN001-W001-L001-R001-M100-T1-L" will be aggregated into one location group; while files with the parameter "SN001-W001-L002-R001-M100-T1-L" will be grouped into another location group due to their different line numbers. The reason for choosing these seven types of parameters is that a single parameter cannot uniquely pinpoint the work point (the same weld may correspond to different lines, and the same mileage may involve different tracks), while the combination of the seven types of parameters can achieve the location of the work point from three dimensions: equipment, line, and specific location.

[0049] The purpose of time grouping is to aggregate all files belonging to the same flaw detection period within the same location group. Since the start and end times of each video file are not the same, there will inevitably be one file recorded at the beginning of flaw detection and another file recorded at the end of flaw detection. Furthermore, flaw detection at the same work point is a continuous process, but multi-device recording may have differences in start / stop times, or multiple scattered files may be generated due to equipment pauses. Although these files are not completely continuous in time, they belong to the same flaw detection task and need to be included in the same time period for processing.

[0050] Therefore, this embodiment uses a disjoint-set data structure (DFS) algorithm to achieve time grouping. Firstly, the DFS algorithm is suitable for the transitive time overlap problem in this embodiment's scenario (e.g., files A and B overlap, B and C overlap, even if A and C do not directly overlap, they still need to be grouped into the same time period), which aligns with the actual scenario of continuous flaw detection. Secondly, the DFS algorithm can achieve query and merging operations with approximately constant time, enabling rapid processing of massive numbers of files, which is superior to traditional traversal comparison methods. Specifically, the time overlap determination rule in this embodiment is: for any two files, if the start time of file A is less than or equal to the end time of file B, and the start time of file B is less than or equal to the end time of file A, then the two files are determined to have time overlap. This rule covers cases including partial overlap, complete inclusion, and endpoint connection, ensuring that no files within the same continuous flaw detection time period are missed.

[0051] As an optional approach, during implementation, based on the start timestamp parsed from the file identifier and combined with the file duration obtained from the video length extraction tool, the precise time interval [start time, end time] (accurate to milliseconds) for each file is calculated. For example, a file with a start time of 11:03:13.010 and a duration of 30.5 seconds has a time interval of [11:03:13.010, 11:03:43.510].

[0052] Subsequently, a disjoint-set data structure is initialized within the same location group, all file pairs are traversed, and the grouping is determined based on overlap rules to determine whether to merge the groups. A specific example is as follows: File A (11:00:00.000-11:10:00.000) and File B (11:05:00.000-11:15:00.000): If the start time of A is less than or equal to the end time of B, and the start time of B is less than or equal to the end time of A, they are considered to overlap and are merged into Group 1. File B and file C (11:12:00.000-11:20:00.000): The start time of B is less than or equal to the end time of C, and the start time of C is less than or equal to the end time of B. They are considered to be overlapping. Since B is already in group 1, C is assigned to group 1. The final group 1 contains files A, B, and C, covering the entire continuous flaw detection period (11:00:00.000-11:20:00.000).

[0053] Media type grouping categorizes files within the same time group according to their data attributes. In this embodiment, the grouping is based on the "media type" field in the file standardization identifier. This field uses a fixed numerical code for automatic identification: 00 represents on-site audio, 01 represents left camera video, 02 represents right camera video, 03 represents flaw detector screen video, 04 represents endoscope video, and so on.

[0054] In one implementation, all files are traversed within the same time group, the media type code of each file is extracted, and the files are divided into different subgroups according to the code: files with code 00 are assigned to the audio subgroup, and files with codes 01-04, etc., are assigned to the corresponding video subgroups (left camera subgroup, right camera subgroup, etc.).

[0055] Subsequently, the files within each subgroup are sorted by their start timestamp: the audio subgroup generates a time-sorted sequence of audio files, with each file assigned the identifier "A-{serial number}" (A-{1}, A-{2}); each video subgroup generates a time-sorted sequence of video files, with each file assigned the identifier V-{camera ID}-{serial number} (e.g., V-{01}-{1}, V-{02}-{2}). The reason for this sorting is that files of the same type may generate multiple fragments due to recording pauses; sorting by time allows for chained splicing, ensuring the continuity of audio and video streams of the same type on the timeline.

[0056] For example, a time group contains the following files: two audio files (coded 00), three left camera files (coded 01), and one flaw detector screen display file (coded 03). After grouping and sorting, these files form three sequences: Audio sequence: A-{1} (starting at 11:00:00.000), A-{2} (starting at 11:08:00.000); Left camera video sequence: V-{01}-{1} (starting at 11:00:00.000), V-{01}-{2} (starting at 11:05:00.000), V-{01}-{3} (starting at 11:12:00.000); The flaw detector screen displays the video sequence: V-{03}-{1} (starting at 11:00:00.000).

[0057] Therefore, the grouping design in this embodiment fully considers the actual operation scenario of rail flaw detection, which not only avoids the problem of document confusion at different operation points and at different times, but also the grouping process is fully automated without human intervention, thus solving the problem of low efficiency and reliance on manual document organization in the prior art.

[0058] After grouping the files, a problem arises: although audio and video files within the same time group belong to the same flaw detection period, their recording start / stop times differ, and fragmented segments may occur due to equipment pauses. Therefore, it is necessary to align all audio and video streams on the same timeline through unified benchmark calibration and interval compensation. To solve this problem, this embodiment sets a global time benchmark to provide a unique time reference for all audio and video files within the same time group, eliminating time start deviations between different files. This undoubtedly demonstrates that this embodiment can overcome practical problems such as potential start delays during multi-device recording or discontinuous file times caused by segmented recording on a single device.

[0059] The implementation method of this embodiment is as follows: Using milliseconds as the precision unit, the standardized identifiers of audio and video files within the same group are parsed, the start times of all files within the time group are traversed, and the minimum value is extracted as the earliest start time within the group (so that the timeline of all files can cover the entire flaw detection period and avoid the loss of early content caused by individual files starting late). Iterate through the end times of all files in the time group (calculated by adding the file duration to the start time), extract the maximum value as the latest end time in the group (to ensure that the timeline of all files can continue until the end of the flaw detection task, and to avoid the loss of later content due to the early termination of individual files); The total global duration is calculated as follows: Latest end time within the group - Earliest start time within the group. This total global duration is the standard duration for the final merged files.

[0060] After determining the global time base, this embodiment needs to perform time compensation on video and audio files separately. For video files, considering that videos from the same perspective need to maintain temporal continuity in order to present a smooth picture after merging, this embodiment will use chained time difference calculation to identify and fill in the time intervals between segments to avoid scene jumps. Finally, the scattered video segments of the same media type will be completed into a continuous video stream with the same total global duration.

[0061] Specifically, for the first video in the sequence, its time offset is equal to the video's start time minus the earliest start time within the group. This calculation is used to fill in the blank period from the global baseline to the start of the video. If the video's start time is the earliest start time within the group, then its start time offset is 0.

[0062] For subsequent videos in the sequence: Time offset = Start time of the current video - End time of the previous video. This calculation is used to fill in the blank period between the end of the previous video and the start of the current video.

[0063] For the last video in the sequence: End offset = latest end time within the group - end time of this video. This calculation is used to fill in the gap between the end of the current video and the global baseline endpoint.

[0064] The chained computation described above does not directly align with the global reference because videos of the same media type may be recorded in multiple segments (due to a brief device malfunction or accidental recording followed by a restart after a few seconds). The method described in this embodiment can fill in each segment interval. Directly aligning with the global reference in this case might disrupt the inherent temporal flow between segments.

[0065] Based on the time deviation calculation described above, this embodiment also requires automatic compensation operations to be performed on video sequences of the same media type according to the above rules. The steps include: Initialize the video stream buffer to store the completed continuous video data; Iterate through the video sequence and determine if the currently processed video is the first video in the sequence: If it is the first video: Calculate the time offset. If the time offset > 0, generate a black screen compensation video (optional parameters: resolution 640×480 pixels, duration = time offset, pure black background, add "No Video" white text prompt in the center of the screen, font size 25 pixels), and concatenate the compensation video with the current video in chronological order and store it in the cache; if the time offset ≤ 0, directly store the current video in the cache, no compensation is needed.

[0066] If it is a subsequent video: Calculate the time offset. If the time offset > 0, generate a black screen compensation video with the same parameters as above, first splice the compensation video to the end of the existing video stream in the cache, and then splice the current video; if the time offset ≤ 0, directly splice the current video to the end of the cached video stream to ensure time continuity.

[0067] When processing up to the last video in the sequence, calculate the time offset of the last video: If the time offset of the last video is greater than 0, generate a black screen compensation video with the same parameters, and stitch it to the end of the cached video stream to fill the gap between the end of the last video and the global endpoint. If the time offset of the last video is less than or equal to 0, no additional compensation is needed, as the video stream in the cache is already consistent with the total global duration.

[0068] Assumption: The global baseline parameters for a certain time group are: earliest start time within the group = 0s, latest end time within the group = 90s, and total global duration = 90s; the left camera video sequence is one continuous segment (0s-90s), the right camera video sequence is one segment (5s-80s), and the endoscope video sequence is two segments (10s-30s and 70s-90s). The compensation processing results are as follows: Left camera: Time offset = 0s, time offset of the last video = 0s, no compensation, output 90s continuous video stream; Right camera: The time offset of the first video is 5s (5s-0s), generating a 5s black screen compensation video; the time offset of the last video is 10s (90s-80s), generating a 10s black screen compensation video; the final output is a 90s continuous stream consisting of 5s black screen + 75s video + 10s black screen. Endoscopy: The first video time offset is 10s (10s-0s), generating a 10s black screen compensation video, which is then stitched together with the first 30s video; the second video time offset is 40s (70s-30s), generating a 40s black screen compensation video, which is then stitched together with the second 20s video; the last video time offset is 0s; the final output is a 90s continuous stream consisting of 10s black screen + 30s video + 40s black screen + 20s video.

[0069] The audio file time compensation uses a silence segment generation method to ensure that each audio stream is aligned with the global time reference, and then a unified audio track is formed through mixing. Furthermore, this method is superior to the final mixing of multiple audio streams, as it eliminates the need to maintain the continuity of individual audio segments; each stream only needs to be aligned with the global time reference, thus avoiding sound misalignment after mixing. Therefore: For each audio file (media type encoded as "00") within the same time group, perform time compensation independently. The specific steps include: The time offset for each audio stream is calculated as: start time of that audio stream - earliest start time within the group. Unlike video, audio uses a global baseline directly without chained calculations because multiple audio streams will eventually be mixed.

[0070] If the time offset is greater than 0, a silent segment with the same duration is generated. The silent segment uses a zero-value audio signal (to ensure no noise interference). To avoid audio quality distortion or format incompatibility during mixing, its sampling rate and number of channels must be consistent with the original audio file. At this point, the silent segment is spliced ​​to the beginning of the original audio, thus aligning the audio start point with the earliest global start time.

[0071] Calculate the end offset for each audio track = latest end time within the group - end time of that audio track. If the end offset > 0, generate a silent segment of the corresponding duration and append it to the end of the original audio track to ensure that the total duration of each audio track is consistent with the global total duration.

[0072] Specific steps: This embodiment uses the above time group reference parameters as an example: (earliest start time in the group = 0s, latest start time in the group = 90s). The audio sequence contains 3 files: Audio 1 (0s-90s), Audio 2 (30s-90s), and Audio 3 (20s-50s). The compensation processing results are as follows: Audio 1: Time offset = 0s, end offset = 0s, no compensation, output 90s of original audio; Audio 2: Time offset = 30s (30s-0s), generate a 30s silent segment and splice it to the front end; End offset = 0s, finally output a 90s audio stream consisting of 30s of silence + 60s of original audio; Audio 3: Time offset = 20s (20s-0s), generate a 20s silent segment and splice it to the beginning; End offset = 40s (90s-50s), generate a 40s silent segment and splice it to the end; The final output is a 90s audio stream consisting of 20s silence + 30s original audio + 40s silence.

[0073] After all audio compensation is completed, the multi-channel aligned audio streams are merged into a single mixed audio stream using a mixing filter, thus automatically synchronizing the sound and picture during playback.

[0074] Example 3 This embodiment further explains the canvas creation process in Embodiment 1. The basic configuration parameters used in this calculation are consistent with those in a typical scenario; please refer to the documentation again. Figure 2 In this embodiment, the number of rows is set to 2, the number of columns is set to 2, the standard display size of a single video is 640×480, the horizontal and vertical spacing between videos is 20 pixels, the height of the text area reserved above each video is 60 pixels, and the spacing between the video boundary and the canvas edge is 20 pixels.

[0075] The total canvas size must fully accommodate the four video feeds, corresponding text areas, and all spacing elements. The total canvas width is calculated using the logic from Example 1, with a 20-pixel margin on each side. The middle section includes the width of two video columns and the number of columns minus one horizontal margin. Specifically, the calculation is twice the margin, plus two video widths, plus one horizontal margin: 2 × 20 + 2 × 640 + 20 = 1340 pixels. This width ensures the two video columns are horizontally evenly arranged with symmetrical left and right margin spacing. The total canvas height must cover the 20-pixel margin at the bottom, the height of the two lines of text areas, and the height of the two lines of video. Specifically, the margin plus twice the text area height, plus twice the video height: 20 + 2 × 60 + 2 × 480 = 1100 pixels. This height allows for a reasonable vertical distribution of the two lines of video and text areas, with consistent vertical spacing.

[0076] The video coordinates are calculated by establishing a coordinate system with the top left corner of the canvas as the origin, the horizontal axis as the X-axis, and the vertical axis as the Y-axis. The position of each video is derived sequentially according to the row and column order.

[0077] The video in the first row and first column corresponds to the video displayed on the flaw detector screen. Its upper left X coordinate is calculated from the left boundary interval, which is 20 pixels; the Y coordinate is the height of the text area, which is 60 pixels. The starting coordinates of the flaw detector screen display video are (20, 60).

[0078] The video in the first row and second column corresponds to the endoscope video. Its X coordinate is the sum of the left boundary interval plus one video width and one horizontal interval, which is 20 + (2-1) × (640 + 20) = 680 pixels. Its Y coordinate is consistent with the video in the first row and first column, which is 60 pixels. The starting coordinates of the endoscope video are (680, 60).

[0079] The video in the second row and first column corresponds to the video from the left camera. Its X coordinate is the same as that of the video in the first row and first column, which is 20 pixels. The Y coordinate is the sum of the height of the text area in the first row, the height of the video, and the vertical spacing, plus the height of the text area in the current row, which is 60 + (2-1) × 480 + 60 = 600 pixels. The starting coordinates of the left camera video are (20, 600). The video in the second row and second column corresponds to the video from the right camera. Its X coordinate is the same as that of the video in the first row and second column, which is 680 pixels; its Y coordinate is the same as that of the video in the second row and first column, which is 600 pixels. The starting coordinates of the right camera are (680, 600). At this point, the entire grid layout is neatly aligned.

[0080] The coordinates of the text area are bound to the corresponding video, ensuring that the labels are intuitive and do not obscure the video content. The width of the text area for each video is exactly the same as the width of the corresponding video, and it is horizontally aligned with the video. Therefore, the X coordinate of the text area is the same as the X coordinate of the corresponding video. The text area is located directly above the video, and its Y coordinate is the Y coordinate of the corresponding video minus the height of the text area. That is, the Y coordinate of the text area for the two videos in the first row is 60-60=0 pixels, and the Y coordinate of the text area for the two videos in the second row is 620-60=560 pixels. The text area will display labels for "Flaw Detector Screen Display Video", "Endoscope Video", "Left Camera Video", and "Right Camera Video", respectively. If a video is missing, the text label will be adjusted accordingly to prompts such as "Flaw Detector Screen Display Not Received, No Video Available" or "Endoscope Not Activated, No Video Available".

[0081] After the coordinate calculations are completed, a background canvas will be created based on the above parameters. The canvas size will be fixed at 1340×1100 pixels, the background color will be set to white, and the duration will be exactly the same as the total duration corresponding to the global time base, ensuring that it can fully cover the entire time range of the merged video. When overlaying videos subsequently, each video will be placed at the calculated coordinate positions, and synchronous display will be achieved through layer overlay technology. Even if there is a time interval or missing video, the corresponding black screen compensation video will be overlaid at the same coordinates, ensuring that the layout remains neat and orderly.

[0082] Example 4 The solution described in Example 1 already illustrates that the basic configuration parameters when creating the canvas can be preset to default values ​​for common flaw detection scenarios, or flexibly adjusted according to actual viewing needs. Therefore, this example illustrates the process of dynamically adjusting the above solution based on the actual number of video streams during flaw detection operations: First, it is clear that all layout calculations are based on a unified basic parameter system. In this embodiment, the default size of a single video is 640 pixels wide and 480 pixels high. The horizontal and vertical spacing between videos is 20 pixels. The height of the text area above each video is 60 pixels. The boundary spacing between the video and the canvas edge is 20 pixels. The above parameters can be adjusted according to actual viewing needs, and the adapted layout can still be derived through the same logic after adjustment.

[0083] When the video to be merged is a single stream, the layout automatically adopts a 1x1 column centered display format. This layout maximizes the display area of ​​the single video stream, making it easier for experts to focus on the details of a single frame. The calculation of the total canvas width follows the established formula in Example 1, with a 20-pixel margin on each side, and only one video width in the middle, with no horizontal gap between videos. Therefore, the total width is 2×20+1×640+20×(1-1)=680 pixels. The total canvas height includes the 20-pixel margin at the bottom, the height of one line of text area, and the height of one video, calculated as 20+1×60+1×480=560 pixels. The coordinates of the top left corner of the single video stream are taken from the top left corner of the canvas as the origin, with the X coordinate being 20 pixels from the left margin and the Y coordinate being 60 pixels from the height of the text area, ensuring that the video is centered and that there is reserved space above for text. The corresponding text area has the same X coordinate as the video, which is 20 pixels, and the Y coordinate is 60-60=0 pixels, aligning the text with the video and not obscuring the content.

[0084] When there are two videos to be merged, the layout automatically switches to a 1x2 left-right arrangement. This layout allows the two videos to evenly divide the canvas space horizontally, facilitating comparison. The total canvas width is calculated with a 40-pixel gap between the left and right edges, including two video widths and one horizontal gap, totaling 2×20 + 2×640 + 20×(2-1) = 1340 pixels. The total canvas height remains the same as the single-video layout, at 560 pixels. The X-coordinate of the video in the first row and first column is 20 pixels, and the Y-coordinate is 60 pixels; the X-coordinate of the video in the first row and second column is 20 + (2-1)×(640+20) = 680 pixels, and the Y-coordinate is also 60 pixels. The horizontal spacing between the two videos is uniform, with no overlap or uneven spacing. The corresponding text area coordinates are (20, 0) and (680, 0), respectively.

[0085] When there are three videos to be merged, the layout automatically adjusts to a triangular arrangement, i.e., a basic framework of 2 rows and 2 columns, with the third video placed in the bottom center (or the top center). This layout balances the display area of ​​the three videos, preventing any video from being inconvenient to view due to improper placement. The total canvas width is still calculated based on 2 columns, which is 1340 pixels, ensuring sufficient horizontal space to accommodate the centered third video; the total height is calculated based on 2 rows, which is 20 + 2 × 60 + 2 × 480 = 1100 pixels, providing ample space for the two vertical rows of elements. The video coordinates in the first row and first column are (20, 60), and the video coordinates in the first row and second column are (680, 60), consistent with the positions of the top two videos in a four-channel layout. The third video in the second row needs to be centered, with an X coordinate of (1340-640)÷2=350 pixels and a Y coordinate of 60+(2-1)×(60+480+20)=620 pixels, ensuring a uniform vertical spacing with the top two videos. The corresponding text area coordinates are (20, 0), (680, 0), and (350, 560), respectively. The text labels are bound to the corresponding videos and do not affect the integrity of the image.

[0086] When there are four videos to be merged, they will be displayed in the style of Example 3 above.

[0087] Through this layout calculation logic based on the number of channels, regardless of whether the number of video channels to be merged is 1, 2, 3 or 4, the appropriate screen layout can be automatically generated by adjusting the number of rows and columns and combining a unified size and coordinate calculation formula, without the need for manual adjustment of layout parameters.

[0088] On the other hand, this embodiment also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of a method for synchronous synthesis and adaptive display of multi-channel audio and video in rail flaw detection.

[0089] On the other hand, this embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of a method for synchronous synthesis and adaptive display of multi-channel audio and video in rail flaw detection.

[0090] The embodiments of the present invention have been described in detail above. The various embodiments in the specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section. It should be noted that those skilled in the art can make several improvements and modifications to the present invention without departing from the principles of the invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A method for synchronous synthesis and adaptive display of multi-channel audio and video in rail flaw detection, characterized in that, include: The time and duration information of the audio and video files are analyzed, and a global time reference is determined based on the time and duration information. Classify files into video and audio types, and perform time synchronization processing on each type of file. Based on the number of media types of the videos to be merged, preset layout configuration rules are established; based on the layout configuration rules, the size of the compositing canvas, the layout position of each video, and the display position of the corresponding video type text identifier are determined. Create a background canvas based on the determined canvas size and the total duration corresponding to the global time base; The synchronized video data from each stream is overlaid onto the background canvas according to the determined layout to generate a composite video. The synthesized video is merged with the time-synchronized unified audio data to generate a flaw detection merged video.

2. The method for synchronous synthesis and adaptive display of multi-channel audio and video in rail flaw detection according to claim 1, characterized in that, The global time base includes the earliest start time, the latest end time, and the total duration, where the total duration is the difference between the latest end time and the earliest start time; the time information is the start time of each audio and video file.

3. The method for synchronous synthesis and adaptive display of multi-channel audio and video in rail flaw detection according to claim 1, characterized in that, Time synchronization processing for video files includes: Get the time offset of the first video relative to the earliest start time of the global time base, the time offset of the subsequent videos relative to the end time of the previous video, and the end time offset of the last video relative to the latest end time of the global time base. If the time offset exists, generate a compensation video that matches the offset duration and stitch the two corresponding consecutive videos together; If the last video has an end time offset, add a compensation video of the corresponding duration at the end of it to make the total duration of the processed videos of the same type consistent with the total duration of the global time base.

4. The method for synchronous synthesis and adaptive display of multi-channel audio and video in rail flaw detection according to claim 1, characterized in that, Time synchronization processing for audio files includes: Obtain the time offset of each audio file relative to the earliest start time of the global time base. If the time offset exists, add a silent segment of the corresponding duration at the beginning of the audio file to form time-synchronized audio data. Mix all time-synchronized audio data to generate unified audio data.

5. The method for synchronous synthesis and adaptive display of multi-channel audio and video in rail flaw detection according to claim 3, characterized in that, The compensation video is a black screen video, which contains preset identification information. After the black screen video is generated, all original video files and the black screen compensation video are processed with preset unified resolution and unified color space.

6. The method for synchronous synthesis and adaptive display of multi-channel audio and video in rail flaw detection according to claim 1, characterized in that, It also includes video aspect ratio adaptation processing, including: Obtain the original aspect ratio of each video and compare it with the preset standard aspect ratio; The video image is scaled proportionally to the aspect ratio of the preset standard resolution, and black borders of equal width are added to fill the difference in size between the scaled and preset standard resolution areas.

7. The method for synchronous synthesis and adaptive display of multi-channel audio and video in rail flaw detection according to claim 1, characterized in that, The layout configuration rules include layout specifications, spacing specifications, text label area specifications, and canvas boundary specifications. The layout specifications include row number configuration and column number configuration, and the spacing specifications include horizontal and vertical spacing configuration between videos.

8. The method for synchronous synthesis and adaptive display of multi-channel audio and video in rail flaw detection according to claim 7, characterized in that, The process of determining the size of the composite canvas, the layout of each video stream, and the display position of the corresponding video type text identifier includes: Extract the number of rows and columns, spacing specifications, text area specifications, and canvas boundary specifications from the preset layout configuration rules; Based on the above specifications and the unified video specifications, calculate the total width and total height of the composite canvas; Calculate the layout position of each video stream sequentially by row and column, and then obtain the display coordinates of the corresponding text label based on the layout position of each video stream.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor is configured to execute the computer program to implement the steps of the rail flaw detection multi-channel audio and video synchronous synthesis and adaptive display method as described in any one of claims 1-8.

10. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the steps of the rail flaw detection multi-channel audio and video synchronous synthesis and adaptive display method as described in any one of claims 1-8.