Audio and video semantic enhancement processing method based on artificial intelligence

By synchronously sampling the spectrum and edge displacement changes of audio and video streams, identifying the dominant mode of interference peaks and setting weight ratios, the problem of semantic ambiguity in multi-source audio and video signals is solved, and semantic expression is enhanced and consistent.

CN121438817BActive Publication Date: 2026-04-28BEIJING LIUJINSUIYUE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING LIUJINSUIYUE TECH CO LTD
Filing Date
2025-11-10
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately determine the source of semantic conflicts when processing multi-source audio and video signals, leading to ambiguous semantic understanding and misidentification. This is especially true in scenarios with high background noise or blurred motion, where semantic fusion imbalance and frame-level content mismatch affect the integrity and consistency of the output.

Method used

By synchronously sampling the spectral variation amplitude of audio and video streams and the edge displacement variation amplitude of video action frames, modal variation comparison data is constructed, the dominant modality of interference peaks is identified, the semantic fusion weight ratio is set, and cross-modal repair processing is performed to generate semantic enhancement results.

Benefits of technology

It enhances the consistency and integrity of semantic expression in the context of multimodal signal interference, and improves the temporal stability and cross-modal consistency of frame-level semantic annotation by aligning modal features with temporal positions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121438817B_ABST
    Figure CN121438817B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of speech signal processing, in particular to an audio and video semantic enhancement processing method based on artificial intelligence, a sampling audio and video change amplitude is constructed to build a comparison sequence, a trend is analyzed to identify a dominant mode, a fusion weight is set and an enhancement interval is marked, a mode semantic repair is executed, a time frame is aligned and a semantic enhancement result is output. In the present application, through synchronous sampling of audio spectrum and video frame edge displacement and construction of change amplitude sequence, high-precision correlation of multi-modal content in time domain can be realized, after dynamic identification and determination of the dominant mode in the abnormal change section, the weight proportion of semantic fusion is set according to the type of the dominant mode, the information interference section can be targetedly strengthened, on the basis of clear mode dominance, structure repair or clarity compensation is carried out on the fuzzy area or spectrum weakening area, and the consistency and integrity of multi-modal semantic expression are effectively enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech signal processing technology, and in particular to an artificial intelligence-based method for enhancing audio and video semantics. Background Technology

[0002] The field of speech signal processing technology mainly involves the acquisition, analysis, recognition and conversion of human speech signals, and is the foundation for realizing functions such as human-computer interaction, language understanding and voice control.

[0003] Among them, audio and video semantic enhancement processing methods refer to semantic supplementation, reasoning and expression optimization of multimodal content containing speech and video information, in order to solve the problem of incomplete semantic content expression or ambiguous semantic understanding in multi-source audio and video signals.

[0004] Because it relies solely on the raw semantic layer content of audio and video streams for processing, existing technologies struggle to accurately determine the source of semantic conflicts when frame-level asynchrony or significant differences in intermodal expression occur. They also cannot dynamically identify the dominant modality affected by interference, resulting in unclear semantic interpretation of abrupt segments. Furthermore, due to the lack of a joint analysis mechanism for intermodal change trends, it is prone to misidentification in scenarios with cross-interference from multiple sources. For example, in high background noise or motion blur scenarios, the system may mistakenly identify non-dominant modal changes as the primary information source, leading to semantic fusion imbalance and frame-level content mismatch, thus affecting the integrity and consistency of the final semantic output. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing an artificial intelligence-based audio and video semantic enhancement processing method.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: an artificial intelligence-based audio and video semantic enhancement processing method, comprising the following steps:

[0007] S1: Synchronously sample the sequence of speech spectrum change amplitude and video action frame edge displacement change amplitude in the audio and video stream at a specified time, and form a comparison sequence according to the corresponding frames to output modal change comparison data;

[0008] S2: Input the modal change comparison data into the sliding frame window, determine the continuous change trend of the speech and video modalities respectively, identify the dominant source of the interference peak, and output the dominant modal of the interference peak;

[0009] S3: Set the semantic fusion weight ratio of speech and video according to the type of dominant mode of the interference peak segment, mark the interference peak segment frame as the enhancement target, and output the modality weight adjustment result;

[0010] S4: Based on the modality weight adjustment results, perform semantic repair processing of the corresponding modality according to the dominant modality type to obtain cross-modality repair results;

[0011] S5: Align the temporal frame positions in the speech modality and the image modality based on the cross-modal repair results, and output the semantic enhancement results.

[0012] As a further aspect of the present invention, the modal change comparison data includes the trajectory of spectral change amplitude, the trajectory of edge displacement change amplitude, and the inter-frame modal correspondence. The dominant modality of the interference peak segment specifically includes the dominant modality category, the corresponding time frame segment, and the interference abrupt change trend characteristics. The modal weight adjustment results include the modal fusion ratio configuration, the dominant modality marking interval, and the auxiliary modality ratio compensation. The cross-modal repair results specifically include the enhanced modal structure, time frame mapping information, and the inter-modal feature alignment relationship. The semantic enhancement results include joint semantic labels, master-slave modality indications, and frame-level semantic annotation sequences.

[0013] As a further aspect of the present invention, the step of obtaining the modal change comparison data specifically includes:

[0014] S111: Synchronously sample audio and video streams within a specified time period, acquire speech signal data and divide it into continuous frame segments, extract the difference in the main frequency amplitude value between speech frames based on the short-time Fourier transform amplitude spectrum within each frame, and generate a speech spectrum change amplitude sequence.

[0015] S112: Based on the speech spectrum change amplitude sequence, collect video action frame image data at the corresponding frame position, call the optical flow method to extract the displacement rate of edge pixels in adjacent frame images and extract the degree of direction change, and obtain the video action frame edge displacement change amplitude sequence.

[0016] S113: Based on the speech spectrum change amplitude sequence and the video action frame edge displacement change amplitude sequence, perform corresponding matching and combination according to the frame index order to obtain modal change comparison data.

[0017] As a further aspect of the present invention, the step of obtaining the dominant mode of the interference peak segment specifically includes:

[0018] S211: Input the modal change comparison data into a sliding frame window, extract the speech spectrum change amplitude value within a continuous frame interval, analyze the trend and direction of change amplitude based on the frame sequence order, and obtain the speech modality upward trend interval;

[0019] S212: Input the modal change comparison data into the sliding frame window, extract the amplitude value of the edge displacement change of the video action frame in the corresponding frame interval, identify the range and trend direction of the continuous change of direction between adjacent frames, and obtain the video modal upward trend interval.

[0020] S213: Based on the number of consecutive frames and the magnitude of change of the rising trend interval of the speech modality and the rising trend interval of the video modality, determine the dominant change performance of the two modalities in the current time period. If the number of consecutive rising frames of the target modality exceeds the set frame threshold and the magnitude of change is greater than that of the other modality, then extract the target modality type as the dominant modality of the interference peak segment.

[0021] As a further aspect of the present invention, the step of obtaining the modal weight adjustment result specifically includes:

[0022] S311: Obtain the sequence number of the time frame corresponding to the type of the dominant mode of the interference peak segment, and extract the mode marker information and time distribution data within the frame segment range under the current mode type to generate the dominant mode frame segment marker;

[0023] S312: Call the dominant modality frame segment marker, select the current fusion parameters of the speech modality or video modality within the frame segment, set the fusion weight of speech and video according to the dominant modality type, and generate the semantic fusion weight ratio configuration result.

[0024] S313: Based on the semantic fusion weight ratio configuration result, the frame segment covered by the dominant modality is taken as the target interval for enhancement processing, and the modality type, time position and fusion parameter configuration content of the current frame segment are recorded simultaneously to generate the modality weight adjustment result.

[0025] As a further aspect of the present invention, the step of obtaining the cross-modal repair result specifically includes:

[0026] S411: Based on the dominant modal type recorded in the modal weight adjustment result, identify the modal category of the current enhancement target interval, extract the frame range and modal feature content corresponding to the current modal type, and generate dominant modal enhancement input frame data;

[0027] S412: Call the dominant modality enhancement input frame data. If the dominant modality is video, extract the blurred area of ​​the image edge as the structural reconstruction target. If the dominant modality is speech, extract the weakened area of ​​the main frequency band of the spectrum as the sharpness restoration target. Perform corresponding image structure reconstruction or spectrum sharpness restoration through the neural network to generate the dominant modality restoration result.

[0028] S413: Based on the dominant modality repair result, obtain the feature data of the auxiliary modality at the same frame position and perform time frame alignment. Construct a cross-modal semantic fusion structure through the temporal mapping relationship between the dominant modality and the auxiliary modality to generate a cross-modal repair result.

[0029] As a further aspect of the present invention, the step of obtaining the semantic enhancement result specifically includes:

[0030] S511: Based on the cross-modal repair results, extract the time position sequence of the dominant modality repair frame segment, obtain the synchronization reference mark of each frame in the sequence, and construct a cross-modal time frame alignment index by combining the original time frame index of the auxiliary modality.

[0031] S512: Based on the cross-modal time frame alignment index, synchronize and associate the contents of the speech modality and the image modality that are at the same time frame position to generate a speech-image frame synchronization relationship set;

[0032] S513: Establish a semantic tag fusion structure based on the speech-image frame synchronization relationship set, label the master-slave modality attributes for each group of time frame synchronization items, and generate semantic enhancement results.

[0033] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0034] In this invention, by synchronously sampling the audio spectrum and video frame edge displacement and constructing a change amplitude sequence, high-precision association of multimodal content in the time domain can be achieved. After dynamically identifying and determining the dominant modality in abnormal change segments, the weight ratio of semantic fusion is set according to the dominant modality type, which can achieve targeted enhancement of information interference segments. Based on the clear modality dominance, structural repair or clarity compensation is performed on ambiguous or spectrally weakened regions, effectively enhancing the consistency and integrity of multimodal semantic expression. Furthermore, by aligning modal features with temporal positions, a semantic fusion tag structure is constructed, making frame-level semantic annotation more temporally stable and cross-modal consistent, enhancing the semantic intelligibility and expressive integrity of speech and video in the context of multi-source signal interference. Attached Figure Description

[0035] Figure 1 This is a schematic diagram of the main steps of the present invention;

[0036] Figure 2 This is a flowchart of step S1 of the present invention;

[0037] Figure 3 This is a flowchart of step S2 of the present invention;

[0038] Figure 4 This is a flowchart of step S3 of the present invention;

[0039] Figure 5 This is a flowchart of step S4 of the present invention;

[0040] Figure 6 This is a flowchart of step S5 of the present invention. Detailed Implementation

[0041] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0042] Please see Figure 1 This invention provides a technical solution: an audio and video semantic enhancement processing method based on artificial intelligence, comprising the following steps:

[0043] S1: Synchronously sample the sequence of speech spectrum change amplitude and video action frame edge displacement change amplitude in the audio and video stream at a specified time, and form a comparison sequence according to the corresponding frames to output modal change comparison data;

[0044] S2: Input the modal change comparison data into the sliding frame window, determine the continuous change trend of speech and video modalities respectively, identify the dominant source of interference peaks, and output the dominant modality of interference peaks;

[0045] S3: Set the semantic fusion weight ratio of speech and video according to the type of dominant mode of interference peak segment, mark the interference peak segment frame as the enhancement target, and output the modality weight adjustment result;

[0046] S4: Based on the modality weight adjustment results, perform semantic repair processing of the corresponding modality according to the dominant modality type to obtain cross-modality repair results;

[0047] S5: Align the temporal frame positions in the speech modality and the image modality based on the cross-modal restoration results, and output the semantic enhancement results.

[0048] The modal change comparison data includes the trajectory of spectral change amplitude, the trajectory of edge displacement change amplitude, and the inter-frame modal correspondence. The dominant modality of the interference peak segment is specifically the dominant modality category, the corresponding time frame segment, and the interference abrupt change trend characteristics. The modal weight adjustment results include the modal fusion ratio configuration, the dominant modality marking interval, and the auxiliary modality ratio compensation. The cross-modal repair results are specifically the enhanced modal structure, time frame mapping information, and the inter-modal feature alignment relationship. The semantic enhancement results include joint semantic labels, master-slave modality indicators, and frame-level semantic annotation sequences.

[0049] Please see Figure 2 The specific steps for obtaining modal change comparison data are as follows:

[0050] S111: Synchronously sample audio and video streams within a specified time period, acquire speech signal data and divide it into continuous frame segments, extract the difference in the main frequency amplitude value between speech frames based on the short-time Fourier transform amplitude spectrum within each frame, and generate a speech spectrum change amplitude sequence.

[0051] After synchronously sampling audio and video streams within a specified time period, the speech signal is acquired and processed into frames based on a fixed time window of 20 milliseconds. In each frame segment, a short-time Fourier transform is applied to extract the frequency domain amplitude spectrum. The amplitude of the frequency component corresponding to the energy peak in the spectrum of the current frame is extracted as the dominant frequency amplitude value of the frame. Subsequently, the dominant frequency amplitude values ​​of adjacent frames are subtracted in chronological order to obtain the inter-frame energy fluctuation and construct a speech dominant frequency difference sequence. To eliminate the influence of background noise or non-linguistic fluctuations, a benchmark range for the variation difference needs to be set. Referring to the statistical average difference of stable segments without pronunciation transitions in a large-scale call speech sample, the variation amplitude is set to no more than 0.3 as the benchmark value. If the difference of a frame exceeds 0.3, it is identified as a semantic change frame. This benchmark value needs to be obtained by the average amplitude difference of multiple stable pronunciation segments within the sample.

[0052] S112: Based on the speech spectrum change amplitude sequence, collect video action frame image data at the corresponding frame position, call the optical flow method to extract the displacement rate of edge pixels in adjacent frame images and extract the degree of direction change, and obtain the video action frame edge displacement change amplitude sequence.

[0053] Based on the frame position index in the speech spectrum variation amplitude sequence, video action frame image data at the corresponding time point is extracted. Optical flow is used to process the image frame sequence. First, Canny edge detection is applied to identify the image structure edges. Then, the position changes of edge pixels in adjacent frames are tracked, and the displacement velocity of pixels between frames is calculated in pixels per second. At the same time, the angle change of the movement direction of the edge points is obtained as the direction change parameter. This angle represents the continuity of the action direction between frames. A threshold for judging abrupt changes in direction needs to be set. Based on the statistical direction change pattern in the human action video samples, the angle change is set to less than 15 degrees as consistent direction and greater than 45 degrees as abrupt change in direction. The boundary range comes from the angle distribution characteristics between static and violent action transitions in the real dataset. After processing each frame of action image in the above way, a numerical sequence of edge displacement velocity and its direction change between consecutive frames is obtained, which finally forms the video action frame edge displacement variation amplitude sequence.

[0054] S113: Based on the speech spectrum change amplitude sequence and the video action frame edge displacement change amplitude sequence, perform corresponding matching and combination according to the frame index order to obtain modal change comparison data;

[0055] Based on the amplitude change sequences of the two modalities, speech and video, a one-to-one matching relationship is established at the index position of each frame. The difference in the main frequency amplitude in the speech frame and the edge displacement amplitude of the corresponding video frame are combined into a pair of data items to construct a joint modal sequence structure. In the actual matching process, it is necessary to verify the consistency of the number and time alignment of the two modal frames. If a frame is missing, a linear interpolation frame filling operation is performed on the original sampling time axis. The interpolation is based on the amplitude trend line of adjacent frames to determine the filling value. After the construction is completed, a continuous frame pair structure is formed. Each frame contains the amplitude change values ​​of speech and image. A modal consistency judgment threshold is set for subsequent judgment of the degree of modal difference. Based on the statistical distribution of the ratio of speech and image amplitude change in the speech-action consistent data segment, the ratio is set to be between 0.5 and 2.0 as the consistency interval. Frame pairs below 0.5 or above 2.0 are regarded as modal abrupt frame points. The entire pairing structure forms modal change comparison data.

[0056] Please see Figure 3 The specific steps for obtaining the dominant mode of the interference peak segment are as follows:

[0057] S211: Input the modal change comparison data into the sliding frame window, extract the speech spectrum change amplitude value within the continuous frame interval, analyze the trend and direction of change amplitude based on the frame sequence order, and obtain the speech modality upward trend interval;

[0058] After inputting the modal change comparison data into the sliding frame window, the amplitude value of the corresponding speech spectrum change within each sliding interval is first extracted. This value comes from the speech spectrum difference sequence completed in the previous step. Then, the amplitude values ​​within the interval are arranged in chronological order to construct the judgment path for inter-frame change trends. The directionality of continuous rises or falls in the frame sequence is identified. During processing, the sign change of the difference between adjacent frames needs to be judged. For example, if the speech spectrum difference is continuously positive in consecutive frames without a sudden drop, it constitutes an upward trend. The minimum number of frames for the trend to continue can be set to 4 frames. (Refer to the source for settings.) Based on the distribution of time frame lengths from the average pronunciation start to the change segment, the pronunciation changes most frequently in the 4-6 frame interval of the speech corpus. Therefore, 4 frames are set as the starting point for trend recognition. In the actual extraction process, when a frame segment meets the condition that the number of consecutively rising frames is not less than the threshold and the cumulative change value of the amplitude is greater than the preset amplitude benchmark (for example, the average rise amplitude is set to 0.3 based on the statistics of stable speech segments), the frame segment is identified as the rising trend interval of the speech modality. The trend boundary is determined by the position of the first rise and the position of the trend termination point. Finally, the change interval with continuous enhancement features in the speech modality is obtained.

[0059] S212: Input the modal change comparison data into the sliding frame window, extract the amplitude value of the edge displacement change of the video action frame within the corresponding frame interval, identify the range and trend direction of the continuous change of direction between adjacent frames, and obtain the video modal upward trend interval.

[0060] Modal change comparison data is used to locate video frame segments corresponding to the speech modality time index using a sliding frame window method, and the amplitude of image edge displacement changes is extracted. Edge displacement is obtained based on edge pixel coordinate change information extracted using optical flow in the image sequence. First, edge detection is performed to extract the edge contours in the current frame image, forming an edge pixel set. For each edge pixel, its position coordinates in the current frame are recorded. In the next frame, the motion matching point corresponding to the pixel is found using the optical flow vector field. Connecting these two position coordinates constructs a direction vector. The angle of the direction vector is defined as the direction of motion of the pixel from the current frame to the next frame. After completing this operation for two consecutive frame segments, the two direction vectors of the pixel in two adjacent frame pairs can be obtained. Furthermore, the angle between these two vectors is used to calculate the amplitude of the direction change. The angle calculation is accomplished using the cosine theorem. Essentially, it constructs a two-dimensional vector set based on the movement direction of pixels in consecutive frames, and extracts angular relationships on this set. By statistically analyzing the average angle change value of all edge pixels, a directional change feature index for each frame is constructed. When setting the angle threshold standard, it is necessary to refer to the statistical range of continuous edge movement direction changes in typical human action segments. In the experimental samples, below 15 degrees can be defined as directional stability, and above 45 degrees can be defined as directional abrupt change. The specific values ​​are calculated based on the statistical boundaries of continuous human actions and directional change actions. Angle fluctuation is judged for each frame interval and combined with the trend of edge displacement amplitude changes. If the directional change remains within 15 degrees and the displacement amplitude continues to increase within a sliding frame window, and the number of consecutive frames reaches more than 4 frames, then the segment is identified as the video modality upward trend interval.

[0061] S213: Based on the number of consecutive frames and the magnitude of change of the rising trend interval of the speech modality and the rising trend interval of the video modality, determine the dominant change performance of the two modalities in the current time period. If the number of consecutive rising frames of the target modality exceeds the set frame threshold and the magnitude of change is greater than that of the other modality, then extract the target modality type as the dominant modality of the interference peak segment.

[0062] Based on the frame sequence information of the rising trend intervals of the speech modality and the video modality, two key indicators, the number of consecutive frames and the cumulative amplitude value, are extracted as the basis for modal change performance. First, it is determined whether the number of consecutive frames of the rising trend of each modality reaches the set threshold standard. The threshold standard is set to 4 consecutive frames based on sample statistical analysis. In the call video dataset, this threshold can cover the minimum change cycle of most speech semantic changes or the start of actions. At the same time, the amplitude increment within the rising segment of each modality is compared. If the number of consecutive frames of a certain modality reaches 4 and its cumulative amplitude increase value is higher than that of the other modality, then the modality can be determined as the dominant changing modality in the current time period. The extraction of the dominant type requires the two indicators to be met simultaneously, that is, the number of consecutive frames meets the requirement and the amplitude change is dominant. If either one is not met, it is not considered dominant. This judgment structure is used to avoid misjudging the dominant modality due to a sudden increase in a single frame or short-term fluctuation. Finally, the target modality type that meets the requirements is extracted and its frame segment is marked.

[0063] Please see Figure 4 The specific steps for obtaining the modal weight adjustment results are as follows:

[0064] S311: Obtain the sequence number of the time frame corresponding to the type of the dominant mode of the interference peak segment, and extract the mode marker information and time distribution data within the frame segment range under the current mode type to generate the dominant mode frame segment marker;

[0065] After obtaining the time frame segment corresponding to the dominant mode of the interference peak, it is necessary to first extract the continuous frame index information within the frame segment range under that mode type as the data basis for subsequent processing. In the operation, a complete time series number set is constructed with the start and end frame numbers of the interference peak as the boundary. At the same time, the modal label information corresponding to this frame segment is extracted. This information comes from the dominant type identified in the preceding modal change comparison analysis results. For example, if it is a speech modality, it corresponds to the speech spectrum change characteristics; if it is a video modality, it corresponds to the image edge displacement change characteristics. In terms of temporal distribution, it is based on each frame. The sampling timestamps of the data are extracted to ensure the continuity of the timeline and the correspondence between the modality determination. The timestamp extraction is calculated using the sampling start point and the inter-frame interval parameter. The inter-frame interval parameter must be consistent with the fixed sampling rate of the acquisition device. For example, if the frame rate is 50 frames per second, the time interval is 0.02 seconds. The sampling duration multiplied by the frame rate can be used to inversely deduce the total number of frames. In this way, a labeling structure with a dominant modality type identifier, frame sequence number, and time distribution record can be formed. This structure is used to explicitly describe the dominant modality frame segment and locate the data position, ultimately forming the dominant modality frame segment label.

[0066] S312: Call the dominant modality frame segment marker, select the current fusion parameters of the speech modality or video modality within the frame segment, set the fusion weight of speech and video according to the dominant modality type, and generate the semantic fusion weight ratio configuration result.

[0067] After invoking the dominant modality frame segment marker, the dominant modality type corresponding to the frame segment is first determined. Based on the dominant modality determination result, the average variation amplitude of the speech modality or video modality within the frame segment is extracted as the fusion parameter input. The average variation amplitude of the speech modality is derived from the average of the dominant frequency difference values ​​of all frames in the speech spectrum variation amplitude sequence for that segment. The average variation amplitude of the video modality is derived from the average edge displacement amplitude of all frames in the video motion frame edge displacement variation amplitude sequence for that segment. After extraction, a normalized ratio allocation is performed between the two modalities. Let the average variation amplitude of the speech modality be... The mean amplitude of video modal variation is The current dominant mode is The fusion weights are speech modal weights. With video modal weights According to the dominant mode type The weighting formula is as follows, depending on the differences:

[0068] If the dominant modal type is speech, that is =voice, then: , ;

[0069] If the dominant modality is video, i.e. =Video, then: , ;

[0070] in, It represents the average amplitude of the change in the speech modality within the dominant frame segment, reflecting the fluctuation of the dominant frequency intensity of the speech signal within that segment; It represents the average amplitude of edge displacement change of video modes within the same segment, and is used to measure the intensity of dynamic changes in image structure; and These are the final weights assigned to the speech modality and the video modality in the semantic fusion calculation; It is a compensation coefficient designed to further enhance the weight assignment of a dominant mode in the weight calculation when the dominant mode has already achieved dominance.

[0071] Compensation coefficient The setting is based on the following: the higher the "modal dominance strength" of the dominant mode, that is, the higher the proportion of the change amplitude of the dominant mode relative to the other mode, then... The larger the value, the better. To ensure that the compensation coefficient is not arbitrarily amplified or excessively affects the fusion balance, this value needs to be determined by collecting a large amount of speech and video alignment sample data in a typical training corpus, and statistically analyzing the ratio of subjective consistency scores to modality change amplitude for different dominant modalities. Statistical analysis revealed that in frames where the ratio of dominant modality change amplitude to non-dominant modality amplitude reaches 1.5 or higher, users have the highest accuracy in judging modality dominance labels, and the consistency scores are concentrated in the range of 0.6 to 0.75. Therefore, the median value of this distribution range is taken and shifted upwards by one standard deviation as the empirical compensation amount, and the final value is set as follows. As a constant input.

[0072] The compensation coefficient actually reflects the elasticity of adjusting the semantic fusion weight when the dominant modality changes in magnitude. Through the above function structure, the fusion weight value pair of speech and video can be dynamically calculated to guide the determination of the subsequent modality enhancement target region and the fusion execution of the semantic reconstruction process, and finally generate the semantic fusion weight ratio configuration result.

[0073] Suppose that in a detected interference peak segment, the system identifies the dominant modality as "speech," and within this frame segment: the mean amplitude variation of the dominant frequency of the speech modality is... The mean amplitude of edge displacement change in the video modality is Dominant Enhancement Compensation Coefficient (Based on training samples).

[0074] Since the dominant modality is speech, the fusion weights need to be calculated according to the following formula:

[0075] , ;

[0076] Calculate the ratio of the amplitude of speech variation to the total amplitude: ;

[0077] Add compensation coefficient ;

[0078] Calculate the fusion weights for video modalities: ;

[0079] In this frame segment, the fusion weights calculated by the system are: speech modality weights. Video modal weights This means that in the subsequent semantic fusion process, the speech modality will dominate the expression in this frame segment, and the speech and image semantic features will be weighted at a ratio of 3:1 during the fusion process.

[0080] S313: Based on the semantic fusion weight ratio configuration result, the frame segment covered by the dominant modality is taken as the target interval for enhancement processing, and the modality type, time position and fusion parameter configuration of the current frame segment are recorded simultaneously to generate the modality weight adjustment result;

[0081] Based on the semantic fusion weight ratio configuration results, the frame range covered by the current dominant modality is first located. The start and end frame numbers in the previous dominant modality frame segment markers are used as the target interval for enhancement processing. Within this interval, the fusion weight parameters are read synchronously. These parameters are numerical pairs of speech modality weights and video modality weights, derived from the output of the fusion weight calculation function in the previous step. During operation, this interval needs to be bound to the current modality type and recorded as the dominant modality type field. At the same time, the start and end time positions are calculated based on the time sampling index of each frame. The time position information is obtained by converting the sampling start point and the frame rate. For example, if the sampling start point is 0 seconds and the frame rate is 50 frames per second, then the time corresponding to the 10th frame is 0.2 seconds. This time information is used to provide a time reference for data synchronization and alignment in subsequent modality repair operations. Subsequently, the fusion weight configuration, modality type field, and time start and end positions are recorded as configuration items in a unified structure to generate an adjustment configuration structure that integrates dominant modality attributes, time interval boundaries, and fusion ratio data. The final output is the modality weight adjustment result.

[0082] Please see Figure 5 The specific steps for obtaining cross-modal repair results are as follows:

[0083] S411: Based on the dominant modal type recorded in the modal weight adjustment results, identify the modal category of the current enhancement target interval, extract the frame range and modal feature content corresponding to the current modal type, and generate dominant modal enhancement input frame data;

[0084] Based on the dominant modality type information recorded in the modality weight adjustment results, the modality category corresponding to the current enhancement target interval is identified, and data content matching this modality type is extracted accordingly. If the dominant modality is speech, continuous spectrum frame data corresponding to the marked frame segment needs to be obtained from the speech spectrum change amplitude sequence. The extracted data should cover the entire time range of the marked frame segment, including feature information such as the energy distribution of the main frequency band and the frequency change trend between frames. If the dominant modality is video, video image frames corresponding to the frame segment range need to be selected from the edge displacement change amplitude sequence, and edge detection operations are performed on each frame to extract image features such as edge intensity value and edge structure continuity. The extraction of the time frame segment range needs to be based on the start frame number and end frame number recorded in the weight adjustment results, combined with the sampling start point and frame rate information to calculate the actual time range corresponding to the frame segment, ensuring that the extracted frame segment data is complete and continuous on the time axis. The feature content extraction method adopts different strategies according to different modality types: speech modality uses the amplitude change of the main frequency band of the spectrum as the core feature, while video modality uses the spatial contour and displacement direction of the edge region as the main content. All extracted feature data will serve as the input data set for the dominant modality, forming the model input for subsequent enhancement stages, and ultimately generating dominant modality enhancement input frame data.

[0085] S412: Call the dominant modality to enhance the input frame data. If the dominant modality is video, extract the blurred area at the edge of the image as the target for structural reconstruction. If the dominant modality is speech, extract the weakened area of ​​the main frequency band of the spectrum as the target for sharpness restoration. Perform corresponding image structure reconstruction or spectrum sharpness restoration through the neural network to generate the dominant modality restoration result.

[0086] After invoking the dominant modality to enhance the input frame data, the corresponding repair target region is determined according to the dominant modality type. If the dominant modality is video, the image frames within the selected frame segment are processed sequentially. Edge detection is performed on each frame, and regions with gradient values ​​significantly lower than adjacent edge segments are identified as blurred structure regions. Simultaneously, the pixel structure of the two frames before and after the current frame at the same image position is collected to construct neighboring frame guiding content to assist in repair. If the dominant modality is speech, regions with a dominant frequency band energy level lower than 20% of the average dominant frequency of the preceding and following frames are extracted from the spectrogram as sharpness-weakened segments. Shape mapping is performed using the frequency trajectory of the dominant frequency region in adjacent frames. To unify the processing of image modality and speech modality repair tasks, this stage adopts the U-Net neural network structure to complete the semantic reconstruction process by encoding and decoding blurred segments.

[0087] In the image structure reconstruction process, blurred regions are mapped as input to feature maps, which are then decoded by the network and reconstructed into texture-enhanced images. The output image contains coordinates of these regions. Repair pixel values The calculation is as follows:

[0088] ;

[0089] in, The reconstructed image is in the first... line, number The pixel value at the column position; The number of channels in the feature map is typically set to 3 or more to include multi-scale contextual semantics. Indicates channel Image in offset , The neighboring pixel feature values ​​are used to collect the context around the blurred region; For the convolution kernel in the channel Below, relative position The weight values; the summation symbol arrive arrive The description is based on the current pixel as the center, and extends to its top, bottom, left, and right. Perform sliding convolution processing over a range of pixels. The kernel radius is typically set to 1 or 2, determining the receptive field size of the network. The weighted sum of the convolution results across all channels is the repair output at the current position.

[0090] For speech spectrum clarity restoration, the network input is a stitched image of blurred spectrum frames. After enhancement by a convolutional neural network, the output is a clear spectrum image, with the time position on the spectrum... Frequency points Repair value The expression is as follows:

[0091] ;

[0092] in, To enhance the spectrogram in time frames Frequency points Energy value; This indicates the number of feature map channels in the convolutional network, usually set to 16 or 32, to accommodate multiple frequency domain feature sub-maps. It is a passage In time ,frequency The feature map pixel values ​​reflect the frequency change trend and the degree of blur. It is the corresponding channel The kernel parameters at this position are used to learn frequency-enhanced features; As a bias term, the result is shifted globally by incorporating nonlinear transformations; As the activation function, ReLU or LeakyReLU is used to perform nonlinear enhancement processing on the spectral energy result, making the output more closely resemble the spectral shape of real speech.

[0093] The two restoration processes described above target blurred image edges and weakened speech spectrum regions, respectively. Through contextual information fusion and cross-frame feature mapping, and leveraging the U-Net network structure, they achieve structural reconstruction and sharpness restoration, ultimately generating the dominant modality restoration result. The parameters in the model are as follows: The settings are adjusted based on the repair error performance of the experimental sample set, and the configuration with the minimum error is usually selected through cross-validation.

[0094] This paper describes the restoration of an image frame and a speech frame within an interference peak segment, with the dominant mode being "video." This requires structural reconstruction of the blurred region in the image and also demonstrates the enhancement process of the speech mode's spectrum. It assumes the system has already identified the blurred or weakened target region and input it into a U-Net neural network, using standard configuration for forward propagation. The actual calculation processes for the image structure restoration formula and the speech spectrum enhancement formula are shown below.

[0095] Image structure restoration; ;

[0096] in, The image has three channels: RGB. The convolution kernel radius is 1, corresponding to The feeling of the wild; : Current coordinates of the pixel to be repaired; : For the input image, the first Location in the passage Pixel values ​​(normalized); : No. The convolution kernel weights of the channels;

[0097] For channels (Red Channel):

[0098] Neighbor pixel value : respectively ;

[0099] Convolution kernel weights : respectively ;

[0100] Calculate the convolution result ;

[0101] For the channel and Repeat the same calculation process, assuming the results are 0.49 and 0.46 respectively;

[0102] Sum the results from the three channels: ;

[0103] These are the standardized pixel reconstruction values, which can be used to recover the target image at its location. The blurred area texture structure.

[0104] Speech spectrum enhancement: ;

[0105] in, This indicates that spectral enhancement is performed using 4 feature map channels; , : Indicates the frequency point 80 in the 25th frame; :aisle The input feature values; :aisle The convolution weights; : Bias term, set to a fixed small value; : ReLU activation function.

[0106] Channel input feature value ;

[0107] convolution kernel parameters ;

[0108] plus bias :

[0109] ;

[0110] The enhancement value of the spectrum at the 25th frame and the 80th frequency point is 0.82, which corresponds to the output energy value after the clarity restoration of the main frequency band.

[0111] The two calculation formulas mentioned above are used for video image modality restoration and speech spectrum modality restoration, respectively, with each calculation result undertaking a specific semantic reconstruction function. In image structure reconstruction, the final calculated result represents the reconstruction intensity value of each pixel in the blurred region of the image frame. This value is obtained by fusing the context information of its neighboring pixels with convolution weights, and is used to restore the image edge structure, contour shape, and texture details, thereby completing the restoration of local structural defects caused by abrupt action changes or edge weakening in the image. In speech spectrum enhancement, the calculation result represents the enhanced energy value of each time frame and frequency point in the spectrum map. It is the result of nonlinear reconstruction of blurred frequency bands, which is reflected in the clearer boundaries of the main frequency band and more coherent frequency transitions. Its restoration process relies on frequency domain pattern learning of the input feature map and network response optimization to low-energy segments.

[0112] S413: Based on the dominant modality repair result, obtain the feature data of the auxiliary modality at the same frame position and perform temporal frame alignment. Construct a cross-modal semantic fusion structure through the temporal mapping relationship between the dominant modality and the auxiliary modality to generate cross-modal repair results.

[0113] Based on the dominant modality restoration result, the time index of the restored frame segment is first used to locate the precise position of each time frame on the time axis. At this position, the corresponding feature data of the auxiliary modality is extracted. If the dominant modality is video, the auxiliary modality is speech, and features such as the dominant frequency energy distribution and frequency trajectory changes of the speech spectrogram at that time point need to be extracted. If the dominant modality is speech, the auxiliary modality is video, and the edge structure direction and image texture density of the video frame at that time point need to be extracted. The time frame alignment operation requires consistency in the sampling start point and frame rate between the two modalities. The sampling index is converted, and each frame of the dominant modality segment is converted into an actual timestamp. The closest time frame in the auxiliary modality is located. Subsequently, a one-to-one frame location mapping relationship is established between the dominant and auxiliary modalities, and combined with the feature values ​​of each modality at that moment, a joint speech-image feature representation is formed at the same time point. Based on this, semantic fusion is performed on all frame pairs. A bimodal joint representation structure is generated through methods such as concatenation, feature weighting, and alignment label construction, and summarized into a continuous time series representation. The output is the cross-modal restoration result.

[0114] Please see Figure 6 The specific steps for obtaining the semantic enhancement results are as follows:

[0115] S511: Based on the cross-modal repair results, extract the time position sequence of the dominant modality repair frame segment, obtain the synchronization reference mark of each frame in the sequence, and construct the cross-modal time frame alignment index by combining the original time frame index of the auxiliary modality.

[0116] Based on the cross-modal repair results, the system first identifies the continuous frame segments that have been repaired in the dominant modality, extracts their corresponding time position sequences on the timeline, and extracts the previously recorded synchronization reference markers for each time frame. These markers contain alignment anchor information generated during the enhancement process of the dominant modality repair frame, such as the time start point, frame number range, and repair strategy number. Simultaneously, the system calls the original frame index sequence of the auxiliary modality and selects the auxiliary modality frame with the smallest absolute time difference as the alignment frame by comparing the relative difference between the dominant modality timestamp and the auxiliary modality sampling time. The system generates a cross-modal time frame alignment index based on the time frame mapping relationship between all dominant and auxiliary modalities. This index data structure contains the auxiliary modality frame position corresponding to each frame of the dominant modality, along with time difference and similarity judgment labels. The time difference is set to not exceed the average sampling interval between the two modal frames. If the difference is within this threshold, the two frames are considered suitable for cross-modal alignment, thus completing the binding of the frame positions of the dominant and auxiliary modalities on the timeline.

[0117] S512: Based on the cross-modal time frame alignment index, synchronize and associate the contents of the speech modality and the image modality that are at the same time frame position to generate a speech-image frame synchronization relationship set.

[0118] Based on the cross-modal time frame alignment index, the system extracts the current modal content of the speech and image modal at each matching frame position. It reads feature information such as texture density, edge direction, and grayscale distribution of the frame in the image modality, and feature items such as the dominant frequency, frequency concentration, and average frame energy of the spectrogram in the speech modality. The system binds each pair of time frame content to a speech-image synchronization frame pair according to its time label. A semantic alignment framework is then constructed based on this, ensuring all frame pairs remain consistent on the time axis. Alignment is based on the sampling time difference between modalities not exceeding the standard frame interval setting, typically determined by the sampling frequency. For example, if the video frame rate is 25 frames per second and the speech frame rate is 50 frames per second, the maximum allowable alignment error is 0.02 seconds. During frame pair construction, if the difference between the auxiliary modality frame position and the dominant modality timestamp exceeds the set range, the frame pair is discarded and marked as mismatched. All frame pairs that meet the time alignment conditions are written into the speech-image synchronization relationship set.

[0119] S513: Establish a semantic tag fusion structure based on the speech-image frame synchronization relationship set, label the master and slave modal attributes for each group of time frame synchronization items, and generate semantic enhancement results;

[0120] Based on the audio-image frame synchronization relationship set, the system performs semantic tag fusion processing on synchronized frame pairs in chronological order. Within each frame pair, the dominant modality type is identified, and master-slave labeling is applied to the feature items of the two modalities according to the fusion strategy. Specifically, the label structure marks the dominant semantic source, the auxiliary modality's role type, and the fusion method number at that time point. The assignment logic for master-slave modality attributes relies on the preceding modality judgment result. If the dominant modality is determined to be video in the interference peak segment identification at that time point, then the label structure of that frame pair sets the image modality as the master modality and the audio modality as the slave modality, and vice versa. Each frame pair in the semantic fusion structure corresponds to a semantic enhancement label item, which includes a modality identifier, fusion strategy type, time position number, and correlation score. All label structures are sorted and summarized into a semantic enhancement result, which serves as the temporal semantic output after audio-video modality fusion.

[0121] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. An audio-visual semantic enhancement processing method based on artificial intelligence, characterized in that, Includes the following steps: S1: Synchronously sample the sequence of speech spectrum change amplitude and video action frame edge displacement change amplitude in the audio and video stream at a specified time, and form a comparison sequence according to the corresponding frames to output modal change comparison data; S2: Input the modal change comparison data into the sliding frame window, determine the continuous change trend of the speech and video modalities respectively, identify the dominant source of the interference peak, and output the dominant modal of the interference peak; S3: Set the semantic fusion weight ratio of speech and video according to the type of dominant mode of the interference peak segment, mark the interference peak segment frame as the enhancement target, and output the modality weight adjustment result; S4: Based on the modality weight adjustment results, perform semantic repair processing of the corresponding modality according to the dominant modality type to obtain cross-modality repair results; S5: Align the temporal frame positions in the speech modality and the image modality based on the cross-modal repair results, and output the semantic enhancement results.

2. The audio and video semantic enhancement processing method based on artificial intelligence according to claim 1, characterized in that, The modal change comparison data includes the trajectory of spectral change amplitude, the trajectory of edge displacement change amplitude, and the inter-frame modal correspondence. The dominant modality of the interference peak segment specifically includes the dominant modality category, the corresponding time frame segment, and the interference abrupt change trend characteristics. The modal weight adjustment results include the modal fusion ratio configuration, the dominant modality marking interval, and the auxiliary modality ratio compensation. The cross-modal repair results specifically include the enhanced modal structure, time frame mapping information, and the inter-modal feature alignment relationship. The semantic enhancement results include joint semantic labels, master-slave modality indicators, and frame-level semantic annotation sequences.

3. The audio and video semantic enhancement processing method based on artificial intelligence according to claim 1, characterized in that, The specific steps for obtaining the modal change comparison data are as follows: S111: Synchronously sample audio and video streams within a specified time period, acquire speech signal data and divide it into continuous frame segments, extract the difference in the main frequency amplitude value between speech frames based on the short-time Fourier transform amplitude spectrum within each frame, and generate a speech spectrum change amplitude sequence. S112: Based on the speech spectrum change amplitude sequence, collect video action frame image data at the corresponding frame position, call the optical flow method to extract the displacement rate of edge pixels in adjacent frame images and extract the degree of direction change, and obtain the video action frame edge displacement change amplitude sequence. S113: Based on the speech spectrum change amplitude sequence and the video action frame edge displacement change amplitude sequence, perform corresponding matching and combination according to the frame index order to obtain modal change comparison data.

4. The audio and video semantic enhancement processing method based on artificial intelligence according to claim 3, characterized in that, The specific steps for obtaining the dominant mode of the interference peak segment are as follows: S211: Input the modal change comparison data into a sliding frame window, extract the speech spectrum change amplitude value within a continuous frame interval, analyze the trend and direction of change amplitude based on the frame sequence order, and obtain the speech modality upward trend interval; S212: Input the modal change comparison data into the sliding frame window, extract the amplitude value of the edge displacement change of the video action frame in the corresponding frame interval, identify the range and trend direction of the continuous change of direction between adjacent frames, and obtain the video modal upward trend interval. S213: Based on the number of consecutive frames and the magnitude of change of the rising trend interval of the speech modality and the rising trend interval of the video modality, determine the dominant change performance of the two modalities in the current time period. If the number of consecutive rising frames of the target modality exceeds the set frame threshold and the magnitude of change is greater than that of the other modality, then extract the target modality type as the dominant modality of the interference peak segment.

5. The audio and video semantic enhancement processing method based on artificial intelligence according to claim 4, characterized in that, The specific steps for obtaining the modal weight adjustment results are as follows: S311: Obtain the sequence number of the time frame corresponding to the type of the dominant mode of the interference peak segment, and extract the mode marker information and time distribution data within the frame segment range under the current mode type to generate the dominant mode frame segment marker; S312: Call the dominant modality frame segment marker, select the current fusion parameters of the speech modality or video modality within the frame segment, set the fusion weight of speech and video according to the dominant modality type, and generate the semantic fusion weight ratio configuration result. S313: Based on the semantic fusion weight ratio configuration result, the frame segment covered by the dominant modality is taken as the target interval for enhancement processing, and the modality type, time position and fusion parameter configuration content of the current frame segment are recorded simultaneously to generate the modality weight adjustment result.

6. The audio and video semantic enhancement processing method based on artificial intelligence according to claim 5, characterized in that, The specific steps for obtaining the cross-modal repair results are as follows: S411: Based on the dominant modal type recorded in the modal weight adjustment result, identify the modal category of the current enhancement target interval, extract the frame range and modal feature content corresponding to the current modal type, and generate dominant modal enhancement input frame data; S412: Call the dominant modality enhancement input frame data. If the dominant modality is video, extract the blurred area of ​​the image edge as the structural reconstruction target. If the dominant modality is speech, extract the weakened area of ​​the main frequency band of the spectrum as the sharpness restoration target. Perform corresponding image structure reconstruction or spectrum sharpness restoration through the neural network to generate the dominant modality restoration result. S413: Based on the dominant modality repair result, obtain the feature data of the auxiliary modality at the same frame position and perform time frame alignment. Construct a cross-modal semantic fusion structure through the temporal mapping relationship between the dominant modality and the auxiliary modality to generate a cross-modal repair result.

7. The audio and video semantic enhancement processing method based on artificial intelligence according to claim 6, characterized in that, The steps for obtaining the semantic enhancement result are as follows: S511: Based on the cross-modal repair results, extract the time position sequence of the dominant modality repair frame segment, obtain the synchronization reference mark of each frame in the sequence, and construct a cross-modal time frame alignment index by combining the original time frame index of the auxiliary modality. S512: Based on the cross-modal time frame alignment index, synchronize and associate the contents of the speech modality and the image modality that are at the same time frame position to generate a speech-image frame synchronization relationship set; S513: Establish a semantic tag fusion structure based on the speech-image frame synchronization relationship set, label the master-slave modality attributes for each group of time frame synchronization items, and generate semantic enhancement results.

Citation Information

Patent Citations

  • Audio and video auxiliary tactile signal reconstruction method based on cloud edge collaboration

    CN113642604A

  • Semantic comprehension driven cross-modal information fusion and retrieval method and system

    CN120448563A