Audio-visual content synchronous sound effect synthesis method

By constructing a visual event recognition model and sound effect synthesis method, the problem of inaccurate sound effects recognition in the prior art is solved, and accurate matching of sound effects and immersive experience of audio-visual content is achieved.

CN120358379AActive Publication Date: 2025-07-22COMMUNICATION UNIVERSITY OF CHINA
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510868599.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-22
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify local events in images and synchronize matching sound effects, resulting in inaccurate sound effects and degraded user immersion experience in audio-visual content.

Method used

By extracting image feature sets and using pre-trained visual event recognition models, the spatial audio-visual positioning index and emotional sound effect adjustment index are analyzed, and the sound effect segments are matched and generated. The dynamic sound effect synthesis path is constructed based on the image brightness mean, color mutation rate and inter-frame optical flow vector mode value.

Benefits of technology

It realizes accurate identification of local events and synchronous generation of sound effects, improves the synchronous expression of audio-visual content and user immersion, avoids delays, misalignments or omissions in sound effects, and enhances the sound response accuracy of local events in multi-frame images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358379A_ABST
    Figure CN120358379A_ABST
Patent Text Reader

Abstract

The invention discloses an audio-visual content synchronous sound effect synthesis method, and relates to the technical field of audio-visual content processing. According to the audio-visual content synchronous sound effect synthesis method, feature extraction is performed on each frame of image data to be synthesized, a key frame needing to generate a sound effect is recognized in combination with a pre-trained visual event recognition model, trigger event information of the key frame is extracted, and a spatial sound image positioning index and an emotional sound effect adjustment index of a corresponding frame of image are further analyzed. On this basis, the sound effect matching index of each frame of image is calculated, interval matching is carried out on the sound effect matching index and a preset sound effect fragment, and accurate selection and sound effect synthesis are achieved. The image feature set is introduced, and by means of the pre-trained visual event recognition model, a complete recognition path facing image event dynamics is constructed; event frames with sound effect requirements can be accurately identified under the condition that visual information does not have significant scene switching, and structured information such as event area image data, event types and triggering time can be output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio-visual content processing, and in particular to a method for synchronously synthesizing audio effects for audio-visual content. Background Art

[0002] In the context of the rapid development of multimedia content generation and processing technologies, how to improve the perceptual consistency between images and audio has become an important research direction in the field of audio-visual content enhancement. Currently, many video editing or intelligent synthesis systems rely on manual annotation of events and manual matching of audio effect segments, which have problems such as low efficiency, poor adaptability, and strong subjectivity. Especially in scenes involving sudden, emotional, or spatially dynamic images, existing audio effect generation methods often struggle to accurately reflect the spatial location and emotional intensity of image events, resulting in distorted expression of audio-visual content and affecting the immersive experience of users. With the integrated application of technologies such as image recognition, audio synthesis, and multi-modal learning, constructing a technical path that can automatically identify key events in images and synchronously synthesize matching audio effects has become the key breakthrough point for enhancing the expressive power and emotional rendering effect of video content.

[0003] The limitations of existing technologies at least include the following problems. Most audio effect synthesis methods mainly rely on fixed timestamps, preset rules, or global image change trends to trigger audio effect generation, but ignore the actual occurrence time and regional differences of local events in the image. This rough time-driven or full-frame processing method is difficult to capture the key frames in the image sequence that truly require audio effects. In audio-visual content, the events with audio effect significance are often local events, such as instantaneous triggers of object collisions, character turns, and action starts. The spatial positions and emotional changes of these events are often highly local and non-persistent. If only low-resolution indicators such as average frame brightness and global action amplitude are used to judge whether to generate an audio effect, a large number of non-key frames will be misjudged as audio effect frames, or key frames will be missed, directly resulting in mis-timing or missing of audio effects. Summary of the Invention

[0004] Aiming at the deficiencies of the existing technology, the present invention provides a method for synchronously synthesizing audio effects for audio-visual content, which solves the problem that the existing technology lacks the ability to accurately identify the frames where image events occur, resulting in inaccurate audio effect matching and fragmented experience in audio-visual content.

[0005] To achieve the above object, the present invention is implemented by the following technical solutions: An audio-visual content synchronous sound effect synthesis method, comprising the following steps: obtaining a plurality of frames of image data to be synthesized, and respectively performing feature extraction to obtain an image feature set of each frame of image to be synthesized; inputting the image feature set of each frame of image to be synthesized into a pre-trained visual event recognition model to identify a plurality of frames of images that need to generate sound effects, and outputting corresponding trigger event frame information; based on the image feature set and the trigger event frame information, analyzing the spatial sound image localization index and the emotional sound effect adjustment index of each frame of image that needs to generate sound effects; based on the spatial sound image localization index and the emotional sound effect adjustment index, matching the sound effect segments of each frame of image that needs to generate sound effects, and performing sound effect synthesis processing.

[0006] Further, each frame of image data includes pixel values and two-dimensional coordinates of a plurality of pixel points, and the image feature set includes an image brightness mean value, a color mutation rate, and an inter-frame optical flow vector modulus value.

[0007] Further, the specific steps of obtaining the image feature set of each frame of image to be synthesized are as follows: respectively performing average gray processing on each frame of image to be synthesized to obtain the image brightness mean value of each frame of image to be synthesized; performing color space histogram analysis processing on adjacent frames of images to be synthesized to obtain the color mutation rate of each frame of image to be synthesized; performing optical flow estimation average processing on adjacent frames of images to be synthesized to obtain the inter-frame optical flow vector modulus value of each frame of image to be synthesized.

[0008] Further, the trigger event frame information includes a trigger time, an event type, and event area image data, the event area image data includes pixel values and two-dimensional coordinates of a plurality of event pixel points, and the visual event recognition model includes a frame-level feature encoding layer, a time series modeling layer, an attention focusing layer, and an event classification and localization output layer.

[0009] Further, the specific steps of identifying a plurality of frames of images that need to generate sound effects and outputting corresponding trigger event frame information are as follows: in the frame-level feature encoding layer of the visual event recognition model, receiving the image feature set of each frame of image to be synthesized, and performing per-field encoding to output an image encoding feature vector of the corresponding frame; in the time series modeling layer of the visual event recognition model, arranging the image encoding feature vectors of each frame of image in time order to construct a time series feature tensor, and extracting a time correlation feature vector of each frame of image; in the attention focusing layer of the visual event recognition model, reading the time correlation feature vector of each frame of image, and performing a difference score on the feature change rate between it and adjacent frames, and extracting the frames with a difference score higher than a set threshold as candidate trigger event frames; in the event classification and localization output layer of the visual event recognition model, based on the image encoding feature vector of the candidate trigger event frame, analyzing and outputting the corresponding trigger event frame information of a plurality of frames of images that need to generate sound effects.

[0010] Further, the specific steps for analyzing the spatial audio image localization index of each frame of the image for which a sound effect needs to be generated are as follows: Based on the event area image data, analyze the edge perturbation intensity and the centroid point coordinates of the event area of each frame of the image for which a sound effect needs to be generated; Based on the two-dimensional coordinates of several pixel points of each frame of the image for which a sound effect needs to be generated, analyze the image center coordinates of the corresponding frame of the image, and combine with the centroid point coordinates of the event area of the corresponding frame of the image to analyze the center coordinate offset value; Based on the event area image data, analyze the ratio of the event area of each frame of the image for which a sound effect needs to be generated; Read the modulus value of the inter-frame optical flow vector of each frame of the image for which a sound effect needs to be generated, and perform a comprehensive analysis in combination with the edge perturbation intensity, the center coordinate offset value, and the ratio of the event area of the corresponding frame of the image to obtain the spatial audio image localization index of each frame of the image for which a sound effect needs to be generated.

[0011] Further, the specific formula for calculating the spatial audio image localization index of a certain frame of the image for which a sound effect needs to be generated is as follows: ; Wherein, is the spatial audio image localization index of a certain frame of the image for which a sound effect needs to be generated, is the edge perturbation intensity of a certain frame of the image for which a sound effect needs to be generated, is the edge perturbation response coefficient stored in the database, is the modulus value of the inter-frame optical flow vector of a certain frame of the image for which a sound effect needs to be generated, is the optical flow response coefficient stored in the database, is the center coordinate offset value of a certain frame of the image for which a sound effect needs to be generated, is the center offset response coefficient stored in the database, is the ratio of the event area of a certain frame of the image for which a sound effect needs to be generated, is the area ratio response coefficient stored in the database.

[0012] Further, the specific steps for analyzing the emotional sound effect adjustment index of each frame of the image for which a sound effect needs to be generated are as follows: Based on the event area image data, analyze the local texture perturbation ratio and the maximum color gradient value of each frame of the image for which a sound effect needs to be generated; Read the average image brightness and the color mutation rate of each frame of the image for which a sound effect needs to be generated, and perform a comprehensive analysis in combination with the texture perturbation ratio and the maximum color gradient value of the corresponding frame of the image to obtain the emotional sound effect adjustment index of each frame of the image for which a sound effect needs to be generated.

[0013] Further, the specific formula for calculating the emotional sound effect adjustment index of a certain frame of the image for which a sound effect needs to be generated is as follows: ; Wherein, is the emotional sound effect adjustment index of a certain frame of the image for which a sound effect needs to be generated, is the color mutation rate of a certain frame of image for which a sound effect needs to be generated, is the color mutation response coefficient stored in the database, is the texture perturbation ratio of a certain frame of image for which a sound effect needs to be generated, is the texture perturbation response coefficient stored in the database, is the interaction response coefficient stored in the database, is the maximum color gradient value of a certain frame of image for which a sound effect needs to be generated, is the color gradient response coefficient stored in the database, is the average image brightness of a certain frame of image for which a sound effect needs to be generated, is the image brightness response coefficient stored in the database.

[0014] Furthermore, based on the spatial sound image localization index and the emotional sound effect adjustment index, the specific steps for matching the sound effect segment of each frame of image for which a sound effect needs to be generated are as follows: Read the spatial sound image localization index and the emotional sound effect adjustment index of each frame of image for which a sound effect needs to be generated, and perform weighted analysis to obtain the sound effect matching index of each frame of image for which a sound effect needs to be generated; Judge and analyze the sound effect matching index of each frame of image for which a sound effect needs to be generated with a preset number of sound effect matching intervals respectively, and each sound effect matching interval corresponds to a sound effect segment respectively; Take the sound effect segment corresponding to when the sound effect matching index is within a preset sound effect matching interval as the generated sound effect of the image.

[0015] The present invention has the following beneficial effects: (1). This audio-visual content synchronous sound effect synthesis method constructs a complete recognition path for image event dynamics from frame-level feature encoding, time series modeling, attention focusing to event classification and localization output by introducing an image feature set composed of the average image brightness, color mutation rate and the modulus value of the inter-frame optical flow vector, and with the help of a pre-trained visual event recognition model, which is significantly better than the existing passive triggering methods based only on frame difference or threshold judgment. This method can not only accurately identify the event frames with sound effect requirements in the case of no significant scene switching in visual information, but also support the refined modeling of subsequent sound image mapping and emotional adjustment calculation by outputting structured information such as event area image data, event type and trigger time, effectively avoiding problems such as sound effect triggering delay, misalignment or omission, improving the sound response accuracy of local events in multiple frames of images, and enhancing the synchronous expressiveness and user immersion of the final audio-visual content.

[0016] (2) The method for synchronizing audio effects synthesis of audiovisual content constructs a spatial sound image localization index and an emotional audio effect adjustment index. The former is modeled by comprehensively considering the edge perturbation intensity, centroid offset, regional area ratio, and optical flow vector modulus value, and the latter is modeled by integrating features such as local texture perturbation ratio, maximum color gradient, image brightness, and color mutation, significantly improving the spatial matching and emotional perception ability of audio effect segments. By introducing response coefficients, such as edge perturbation response coefficient, texture perturbation response coefficient, etc., this method avoids the limitations of traditional static mapping-based audio effect synthesis in regional adaptation and emotional differentiation, realizes the fine-grained expression of local dynamic features and their actual quantification in sound effect design, enables the audio effect content to not only match the visual position but also adapt to the emotional tension brought by the change of the picture, and effectively supports the multi-dimensional sound-picture linkage under complex content.

[0017] (3) The method for synchronizing audio effects synthesis of audiovisual content, after extracting the spatial sound image localization index and the emotional audio effect adjustment index, sets weighted fusion to form an audio effect matching index, and performs mapping judgment with multiple preset audio effect matching intervals. Each matching interval is bound to a specific audio effect segment. This mechanism realizes the dynamic allocation of audio effect segments and the optimal scheduling of resources while ensuring the rationality of judgment, avoiding problems such as segment repetition and synthesis imbalance caused by fixed rule matching. During the audio effect synthesis process, the selected segments will be inserted into the target audio track according to the time of their corresponding trigger events, and maintain time sequence synchronization and audio envelope continuity, ultimately realizing the full-process linkage from event recognition to audio generation, improving the expression richness and rhythm control ability of audiovisual content in complex scenarios, and significantly enhancing the scalability and practicality of the overall system.

[0018] Of course, it is not necessary for any product implementing the present invention to simultaneously achieve all the above-mentioned advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flowchart of a method for synchronizing audio effects synthesis of audiovisual content according to the present invention.

[0020] Figure 2 It is a specific step flowchart for obtaining the image feature set of each frame of image to be synthesized in a method for synchronizing audio effects synthesis of audiovisual content according to the present invention.

[0021] Figure 3 It is a specific step flowchart for outputting the corresponding trigger event frame information in a method for synchronizing audio effects synthesis of audiovisual content according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] Please refer to Figure 1 , an embodiment of the present invention provides a method for synchronizing audio effects synthesis of audiovisual content, including the following steps: Obtain a number of frame image data to be synthesized, and perform feature extraction on each of them respectively to obtain the image feature set of each frame of the image to be synthesized; input the image feature set of each frame of the image to be synthesized into a pre-trained visual event recognition model to identify several frames of images that need to generate sound effects, and output the corresponding trigger event frame information; based on the image feature set and the trigger event frame information, analyze the spatial sound image localization index and the emotional sound effect adjustment index of each frame of the image that needs to generate sound effects; based on the spatial sound image localization index and the emotional sound effect adjustment index, match the sound effect segments of each frame of the image that needs to generate sound effects, and perform sound effect synthesis processing.

[0023] Among them, the specific steps of performing sound effect synthesis processing are as follows: according to the trigger event frame information of each frame of the image, extract its trigger time field, and according to the matched sound effect segment, perform time-axis alignment processing on the start playback position of the sound effect segment and the trigger time of the corresponding image frame; subsequently, call the audio resource path of the sound effect segment in the synthesized audio track, and according to the spatial sound image localization index of the current frame of the image, set the spatial sound image offset parameter of the sound effect segment in the audio track to adjust the spatial display position of the sound effect in the left and right channels; at the same time, according to the emotional sound effect adjustment index of this frame of the image, set the loudness control value and the rhythm fluctuation curve parameter of the sound effect segment to adjust the playback intensity and rhythm change of the segment in the synthesized output; after completing the above parameter configuration, insert the sound effect segment into the specified position of the target audio track, and perform mixing synthesis according to the set sound effect hierarchical structure to generate an audio track segment containing the sound effect content corresponding to the target frame of the image.

[0024] Specifically, each frame of image data includes the pixel values and two-dimensional coordinates of a number of pixels, and the image feature set includes the average image brightness, the color mutation rate, and the modulus value of the inter-frame optical flow vector.

[0025] As Figure 2 shown, the specific steps to obtain the image feature set of each frame of the image to be synthesized are as follows: perform average gray processing on each frame of the image to be synthesized respectively to obtain the average image brightness of each frame of the image to be synthesized, specifically: first, read the pixel values of all pixels of each frame of the image to be synthesized and convert them into gray values. The gray value of each pixel is calculated by the standard weighting method. The commonly used weighting factors are: the red channel is multiplied by 0.2989, the green channel is multiplied by 0.5870, and the blue channel is multiplied by 0.1140. Then, calculate the average value of the gray values of all pixels in the current frame of the image to obtain the average image brightness of this frame of the image. This brightness average value reflects the overall brightness of the image.

[0026] Perform color space histogram analysis on adjacent frames of images to be synthesized, and obtain the color mutation rate of each frame of image to be synthesized. Specifically: First, obtain the current frame of image to be synthesized and its adjacent frame of image. By analyzing the color space histograms of these two frames of images, calculate the degree of difference in color distribution of each frame of image. The degree of difference can be measured by methods such as chi-square distance or histogram intersection. In this way, the color mutation rate of each frame of image is obtained, and this mutation rate reflects the change amplitude of the image in terms of color.

[0027] Perform optical flow estimation averaging on adjacent frames of images to be synthesized, and obtain the magnitude of the inter-frame optical flow vector of each frame of image to be synthesized. Specifically: First, obtain the current frame of image and its previous frame of image. Use the optical flow method to estimate the displacement (motion vector) of each pixel point between these two frames of images. The optical flow algorithm calculates the motion vector (u, v) of each pixel through the motion of the pixel point in space. Then, calculate the magnitude of the optical flow vector of each pixel point, which represents the motion intensity of this pixel. Finally, average the optical flow magnitudes of all pixels to obtain the magnitude of the inter-frame optical flow vector of this frame of image, which reflects the motion intensity and change speed of the image.

[0028] In this implementation plan, by separately extracting the image brightness mean value, color mutation rate, and magnitude of the inter-frame optical flow vector to form an image feature set, the comprehensive perception ability of the dynamic change features of the image is significantly improved. The image brightness mean value reflects the overall visual brightness level and helps to identify the lighting features of scene changes; the color mutation rate evaluates the change in color distribution through histogram differences and accurately captures sudden or significant visual events in the picture; the magnitude of the inter-frame optical flow vector quantifies the motion intensity at the pixel level in the image and effectively perceives the dynamic behavior of objects in the scene. The three work together to construct a stable, sensitive, and directional image feature basis, enhancing the accuracy and response timeliness of subsequent event recognition and sound effect generation.

[0029] Specifically, the trigger event frame information includes the trigger time, event type, and event area image data. The event area image data includes the pixel values and two-dimensional coordinates of several event pixel points. The visual event recognition model includes a frame-level feature encoding layer, a time series modeling layer, an attention focusing layer, and an event classification and localization output layer.

[0030] Among them, the pre-training steps of the visual event recognition model are as follows: First, construct an image sequence event dataset, which consists of multiple short-term image sequence samples. Each sample is composed of continuously captured image frames. Each frame of the image contains complete pixel values and two-dimensional coordinate fields, and is accompanied by a trigger event field annotated frame by frame. Specifically, each frame of the image is manually or semi-automatically annotated whether it contains a trigger event for which a sound effect needs to be generated. If it does, the event type label (such as "fall", "break", "explosion", etc.) and the image data of the event area, that is, the pixel values and their two-dimensional coordinates of all pixel points in the corresponding area, are additionally annotated. Each image sequence sample in the training dataset is also equipped with a real-time axis for accurately calibrating the event trigger time to ensure that the time field in the subsequent model output is traceable and verifiable.

[0031] Before training, the parameter initialization of each structural layer of the visual event recognition model is required. The input of the frame-level feature encoding layer is three fields: the mean image brightness, the color mutation rate, and the modulus of the inter-frame optical flow vector. A field-level encoding network is used to construct a linear mapping layer and a non-linear activation unit for each field respectively, and a vector embedding of a unified length is output; the initialization parameters adopt the Xavier initialization method, and a mini-batch of image samples is used to pre-train this layer separately. The training objective is to minimize the reconstruction error of each field in the same frame to enhance the expression independence and discriminative ability between fields. Only the frame-level feature encoding layer is trained in this stage, and its output serves as the basic input for the subsequent modeling layer.

[0032] After completing the pre-training of the feature encoding layer, the training process is extended to the time series modeling layer, the attention focusing layer, and the event classification and localization output layer. An end-to-end joint training method is adopted. First, a bidirectional gated recurrent unit (Bi-GRU) is used to construct the time modeling layer to perform time correlation modeling on the encoded features of the image frame sequence; then, the attention focusing layer performs local difference analysis on the time features to extract potential trigger nodes; finally, the event classification and localization output layer performs event type classification and region localization processing on the trigger frame. The training objectives include two items: one is to minimize the event classification loss function (using cross-entropy loss), and the other is to minimize the event region localization error (using a weighted combination of IoU loss and bounding box offset loss). The entire model is based on the Adam optimizer during the training process and gradually converges to a state where it can accurately output the information of the trigger event frame.

[0033] As Figure 3As shown below, the specific steps for identifying several frames of images that require sound effects generation and outputting the corresponding trigger event frame information are as follows: In the frame-level feature encoding layer of the visual event recognition model, receive the image feature set of each frame to be synthesized, and perform per-field encoding to output the image encoding feature vector for the corresponding frame. Specifically, it is to receive the image feature set of each frame to be synthesized, and perform independent encoding processing on the three fields of the average image brightness, color mutation rate, and the modulus of the inter-frame optical flow vector in sequence. The encoding method is a field-level transformation network constructed based on a fully connected mapping structure, which maps the original numerical value of each field to an embedded vector of a unified dimension and then splices and fuses them to form the image encoding feature vector of the current frame image, which is used to represent the multi-dimensional expression features of this frame at the static attribute level.

[0034] In the time series modeling layer of the visual event recognition model, arrange the image encoding feature vectors of each frame in chronological order to construct a time series feature tensor, and extract the time correlation feature vector of each frame. Specifically, arrange the image encoding feature vectors of each frame output by the frame-level feature encoding layer in frame order to construct a time series feature tensor; then use a sequence modeling network based on a bidirectional gated recurrent unit (Bi-GRU) to perform bidirectional propagation processing on this feature tensor, extract the dynamic correlation features between each frame image and its time neighborhood frames, and output the time correlation feature vector of each frame, which represents the dynamic state of this frame in the event evolution process.

[0035] In the attention focusing layer of the visual event recognition model, read the time correlation feature vector of each frame, and perform a difference score on the feature change rate between it and the adjacent frames, and extract the frames with a difference score higher than the set threshold as candidate trigger event frames. Specifically, input the time correlation feature vector of each frame, and use a scoring model based on a weighted attention mechanism to perform a difference score on the feature change rate between it and the front and rear frames; the scoring model calculates the modulus of the time correlation feature difference vector between the current frame and the adjacent frames, and combines the attention weight mapping mechanism to give a higher score to the frames with significant difference changes; finally, extract the frame images with a difference score higher than the set threshold as candidate trigger event frames, which indicates that this frame may correspond to the key node of the event mutation in the image.

[0036] In the event classification and localization output layer of the visual event recognition model, according to the image encoding feature vector of the candidate trigger event frame, the corresponding trigger event frame information of several frames of images that need to generate sound effects is analyzed and output. Specifically, it is as follows: Receive the image encoding feature vector of the candidate trigger event frame and input it into at most a classification event recognition network and a region localization network for processing respectively; among them, the event recognition network performs feature discrimination on the image encoding feature vector based on the multi-layer perceptron structure and outputs the event type category label to which the frame of image belongs; the region localization network, based on the pixel-level attention mapping of the candidate frame image and the image back-projection mechanism, extracts the event region image data in the frame, including the pixel values of the event pixel points and their two-dimensional coordinates; finally, the trigger event frame information including the trigger time, event type and event region image data is output for subsequent sound effect matching and synthesis processing.

[0037] In this implementation plan, by constructing a multi-level visual event recognition model, the accurate recognition of key image frames that need to generate sound effects and the extraction of event regions are realized. The model pre-training process is carefully stratified. First, static and dynamic features such as image brightness, color mutation rate, and inter-frame optical flow are encoded, and then time modeling is performed through Bi-GRU to capture the temporal features in event evolution. Subsequently, the attention mechanism is used to accurately focus on the frames with significant feature changes, enhancing the response sensitivity to sudden events. Finally, the classification and localization output layer can not only identify the event type, but also extract the pixel data of the event region, providing reliable semantic support and spatial anchor points for subsequent sound effect generation. This processing flow significantly improves the accuracy of event recognition and the precision of time localization, avoiding the sound effect mismatch problem caused by fuzzy event boundaries and inaccurate frame-level determination in traditional methods, and providing a high-efficiency and highly robust technical foundation for audiovisual content synthesis.

[0038] Specifically, the specific steps for analyzing the spatial sound image localization index of each frame of image that needs to generate sound effects are as follows: Based on the event region image data, analyze the edge perturbation intensity and the centroid point coordinates of the event region of each frame of image that needs to generate sound effects; based on the two-dimensional coordinates of several pixel points of each frame of image that needs to generate sound effects, analyze the image center coordinates of the corresponding frame of image (that is, calculate the arithmetic mean of all x-direction and y-direction coordinate values), and combine the centroid point coordinates of the event region of the corresponding frame of image to analyze the center coordinate offset value; based on the event region image data, analyze the ratio of the event region area of each frame of image that needs to generate sound effects (that is, the ratio of the number of event region pixel points to the number of image pixel points); read the modulus value of the inter-frame optical flow vector of each frame of image that needs to generate sound effects, and perform comprehensive analysis in combination with the edge perturbation intensity, center coordinate offset value, and ratio of the event region area of the corresponding frame of image to obtain the spatial sound image localization index of each frame of image that needs to generate sound effects.

[0039] Among them, the specific steps for analyzing the edge perturbation intensity of each frame image for which sound effects need to be generated are as follows: Read the image data of the recognized event area in the current frame image, and extract the pixel values and corresponding two-dimensional coordinate values of all pixel points in this area; Subsequently, use an edge detection algorithm (such as the Sobel operator) to perform edge extraction on this area, obtain the gradient direction and intensity distribution of edge pixels, and calculate the standard deviation of the change rate of all edge pixels in the gradient direction as the edge perturbation intensity of this frame image.

[0040] The specific steps for analyzing the centroid point coordinates of the event area of each frame image for which sound effects need to be generated are as follows: Based on the two-dimensional coordinates of all pixel points in this event area, calculate the arithmetic mean of all x-direction and y-direction coordinate values respectively to obtain the centroid point coordinates of this event area, which are used as the structural feature representing the spatial center position of this area.

[0041] The specific formula for calculating the spatial sound image localization index of a certain frame image for which sound effects need to be generated is as follows: ; Among them, is the spatial sound image localization index of a certain frame image for which sound effects need to be generated, is the edge perturbation intensity of a certain frame image for which sound effects need to be generated, is the edge perturbation response coefficient stored in the database, is the modulus value of the inter-frame optical flow vector of a certain frame image for which sound effects need to be generated, is the optical flow response coefficient stored in the database, is the central coordinate offset value of a certain frame image for which sound effects need to be generated, is the central offset response coefficient stored in the database, is the ratio of the area of the event area of a certain frame image for which sound effects need to be generated, is the area ratio response coefficient stored in the database.

[0042] It should be noted that the specific steps for obtaining the edge perturbation response coefficient , optical flow response coefficient , central offset response coefficient , area ratio response coefficient stored in the database are as follows: For the edge perturbation response coefficient: Select the image frames marked as those requiring sound effect generation in the sample set. Sequentially extract the image data of the event area in each sample image, call an edge detection operator (such as Sobel) to calculate the edge gradient values of all pixels in the event area, and calculate the mean and standard deviation of the rate of change of the gradient direction of all edge pixels in this area. The calculation result is the edge perturbation intensity value of the current sample. Subsequently, based on the fitting relationship between the edge perturbation intensities of all samples and their actual sound image offset amounts, perform a regression analysis using the least squares method, and extract the linear weight term of the edge perturbation intensity in the regression equation as the edge perturbation response coefficient. The data type is a floating point number with three decimal places of precision.

[0043] The method for obtaining the optical flow response coefficient is as follows: Perform frame-by-frame optical flow estimation processing on each group of training images containing consecutive frames, obtain the motion vector values of all pixel points in the entire frame image, and calculate the full-frame average value of the optical flow modulus as the inter-frame motion intensity index of this frame. Then, establish a one-to-one correspondence between this motion intensity and the amplitude of the sound image offset direction of the finally generated sound effect, fit its variation relationship, and extract the fitting coefficient of the optical flow intensity term as the optical flow response coefficient. The data type is a fixed-point real number with four significant figures retained.

[0044] The method for obtaining the center offset response coefficient is as follows: Read all pixel coordinate values of the event area in each frame of the image, calculate its regional center point coordinates, and calculate the Euclidean distance from them to the geometric center coordinates of the entire frame image. After normalization processing, it is used as the center offset value. Then, perform a linear fit on the center offset values of all training samples and their sound image positioning error amplitudes. The offset slope in the fitting result is the center offset response coefficient, with the unit being the percentage of the sound image offset corresponding to each pixel offset, and the precision is retained to the 0.001 level.

[0045] The method for obtaining the area ratio response coefficient is as follows: Count the number of valid pixels in the event area of the training sample, calculate the ratio with the total number of pixels in this frame of the image to obtain the area ratio of the event area. Then, establish an associated model between this area ratio value and the loudness gain parameter of the finally synthesized sound effect segment, fit its trend using a logarithmic function, and extract the logarithmic curve slope of the area ratio factor as the area ratio response coefficient.

[0046] Among them, a specific implementation example of calculating the spatial sound image positioning index of a certain frame of the image requiring sound effect generation is as follows, including the following parameters: The edge perturbation intensity of a certain frame of the image requiring sound effect generation is approximately: 1.215.

[0047] The edge perturbation response coefficient stored in the database is approximately: 0.395.

[0048] The modulus value of the inter-frame optical flow vector of a certain frame of the image requiring sound effect generation is approximately: 0.982.

[0049] The optical flow response coefficient stored in the database is approximately: 0.336.

[0050] The central coordinate offset value of a certain frame of image for which a sound effect needs to be generated is approximately: 0.654.

[0051] The central offset response coefficient stored in the database is approximately: 0.431.

[0052] The ratio of the area of the event region of a certain frame of image for which a sound effect needs to be generated is approximately: 0.148.

[0053] The area ratio response coefficient stored in the database is approximately: 0.572.

[0054] Substitute the above data into the specific formula for calculating the spatial sound image localization index of a certain frame of image for which a sound effect needs to be generated, and we get: The spatial sound image localization index of a certain frame of image for which a sound effect needs to be generated = (((1.215)^0.395) / (1 + exp(-0.336×0.982)))×ln(1 + 0.431×((0.654)^2))×(0.572×((0.148)^(1 / 2))) ≈ 0.023391.

[0055] In this implementation scheme, by constructing the spatial sound image localization index, the system integrates multiple physically measurable image structure parameters such as edge perturbation intensity, inter-frame optical flow modulus, central coordinate offset value, and the ratio of the area of the event region, effectively enhancing the mapping logic between image events and spatial sound effects. This index not only considers the edge complexity and dynamic change intensity of the event region in the image, but also combines the spatial distribution offset and coverage ratio of the event region in the image frame, comprehensively reflecting the spatial salience of the event and the sound image triggering characteristics. The weight coefficients of each parameter are obtained through mathematical fitting based on a large number of labeled samples, with clear data sources and response mechanisms, ensuring the stability and traceability of index calculation. Through this structured modeling method, the sound image direction and sound field focusing position that should be matched for each frame of image can be accurately predicted, significantly improving the spatial fidelity and immersion of sound effect synthesis, and avoiding problems such as sound image drift or positioning errors caused by the lack of event space information in traditional schemes.

[0056] Specifically, the specific steps for analyzing the emotional sound effect adjustment index of each frame of image for which a sound effect needs to be generated are as follows: Based on the image data of the event region, analyze the local texture perturbation ratio and the maximum color gradient value of each frame of image for which a sound effect needs to be generated; read the average image brightness and color mutation rate of each frame of image for which a sound effect needs to be generated, and conduct a comprehensive analysis in combination with the texture perturbation ratio and the maximum color gradient value of the corresponding frame of image to obtain the emotional sound effect adjustment index of each frame of image for which a sound effect needs to be generated.

[0057] Among them, the specific steps for analyzing the local texture perturbation ratio of each frame of the image for which the sound effect needs to be generated are as follows: Read the image data of the event area in the current frame of the image for which the sound effect needs to be generated, and extract the pixel values and corresponding two-dimensional coordinate values of all pixel points in this area; for obtaining the local texture perturbation ratio, use a sliding window method to perform the gray-level co-occurrence matrix (GLCM) calculation on the event area, extract three texture metrics of contrast, entropy value, and energy in each local window, and count the proportion of windows in the entire event area where each metric fluctuates beyond the set threshold as the local texture perturbation ratio of this area.

[0058] The specific steps for analyzing the maximum color gradient value of each frame of the image for which the sound effect needs to be generated are as follows: Convert the event area image to the HSV color space, then calculate the gradient values of the hue (H) channel in the horizontal and vertical directions respectively, and select the maximum value of all hue gradient values in this area as the maximum color gradient value of the current frame of the image.

[0059] The specific formula for calculating the emotional sound effect adjustment index of a certain frame of the image for which the sound effect needs to be generated is as follows: ; Among them, is the emotional sound effect adjustment index of a certain frame of the image for which the sound effect needs to be generated, is the color mutation rate of a certain frame of the image for which the sound effect needs to be generated, is the color mutation response coefficient stored in the database, is the texture perturbation ratio of a certain frame of the image for which the sound effect needs to be generated, is the texture perturbation response coefficient stored in the database, is the interaction response coefficient stored in the database, is the maximum color gradient value of a certain frame of the image for which the sound effect needs to be generated, is the color gradient response coefficient stored in the database, is the average image brightness of a certain frame of the image for which the sound effect needs to be generated, is the image brightness response coefficient stored in the database.

[0060] It should be noted that the specific steps for obtaining the color mutation response coefficient , texture perturbation response coefficient , interaction response coefficient , color gradient response coefficient , and image brightness response coefficient stored in the database are as follows: For the color mutation response coefficient: Select the image frames marked as needing to generate sound effects in the sample set, calculate the difference in color histograms between each frame of the image and its adjacent previous frame image, and calculate the average value of the change amplitude of the color values of each pixel point by channel to obtain the color mutation rate of the current sample image; Subsequently, statistically analyze the corresponding relationship between the color mutation rates of all samples and the emotional intensity levels in the sound effect synthesis, fit its changing trend through a non-linear regression model, and extract the sensitive slope coefficient corresponding to the color mutation variable in the fitting equation as the color mutation response coefficient. The numerical type is floating-point, and the precision is reserved to three decimal places.

[0061] For the texture perturbation response coefficient: Perform gray-scale texture analysis on the event area image data in the sample image, extract the gray-scale gradients around each pixel, and calculate the texture variance value of this area as the perturbation index. Further, statistically analyze the numerical mapping relationship between the texture perturbation values of all samples and the emotional fluctuation rate in the generated sound effects, and extract the non-linear response coefficient corresponding to the texture variable after fitting with a polynomial regression model as the texture perturbation response coefficient.

[0062] For the interaction response coefficient: On the premise of knowing the color mutation rate and texture perturbation ratio of the sample image, calculate their product as the collaborative perturbation index, and then perform paired analysis with the sound effect rhythm adjustment intensity value corresponding to this frame of the image in the historical annotation. Fit its collaborative effect trend with a logarithmic function and extract the non-linear weight of the collaborative product term as the interaction response coefficient, which is used to control the comprehensive impact of the interaction between color and texture on the sound effect output rhythm.

[0063] For the color gradient response coefficient: Read the color channel gradients of the pixel points in the event area of the sample image, statistically analyze its maximum gradient value, and fit it with the emotional high-frequency change intensity in the corresponding sound effect to extract the influence amplitude of the maximum gradient variable on the emotional intensity as the color gradient response coefficient.

[0064] For the image brightness response coefficient: Statistically analyze the overall average brightness of the sample image, and fit it with the emotional rise amplitude in the sound effect segment corresponding to this image to extract the inhibition coefficient of the image brightness on the upper limit value of the sound effect emotion as the image brightness response coefficient.

[0065] In this implementation, by constructing an emotional sound effect adjustment index, the system integrates multi-dimensional image-emotion correlation features such as the average image brightness, color mutation rate, local texture perturbation ratio, and maximum color gradient value, effectively realizing the coupling adjustment between the image context and the sound effect emotion parameters. This index not only considers the impact of static image attributes (such as brightness and color changes) on the emotional atmosphere, but also comprehensively considers the stimulating effects of local texture perturbation and color gradient changes on emotional fluctuations, and introduces a color-texture interaction factor to precisely control the rhythm, intensity, and frequency change trend of the sound effect through response coefficients. Each response coefficient is obtained by fitting a large number of labeled samples, with a clear physical source and data statistical basis, ensuring the interpretability of the calculation results and the control accuracy. By introducing this index, dynamic adjustment of the sound effect in terms of rhythm, emotional intensity, and climax control can be achieved, improving the synchronization and immersion of the sound effect and visual content in emotional expression, which is significantly better than the traditional static scheme that only generates sound effects based on event categories.

[0066] Specifically, based on the spatial audio image positioning index and the emotional sound effect adjustment index, the specific steps for matching the sound effect segments of each frame of the image for which the sound effect needs to be generated are as follows: Read the spatial audio image positioning index and the emotional sound effect adjustment index of each frame of the image for which the sound effect needs to be generated, and perform weighted analysis to obtain the sound effect matching index of each frame of the image for which the sound effect needs to be generated; Judge and analyze the sound effect matching index of each frame of the image for which the sound effect needs to be generated with a number of preset sound effect matching intervals, and each sound effect matching interval corresponds to a sound effect segment respectively; The sound effect segment corresponding to when the sound effect matching index is in a preset sound effect matching interval is used as the generated sound effect of the image.

[0067] Among them, the sound effect segments corresponding to the sound effect matching intervals include but are not limited to the following examples: Match the frame images with the sound effect matching index value in the interval [0.0, 0.4) to the "basic background type" sound effect segments (such as environmental wind sounds, mild vibration sounds); Match the frame images with the sound effect matching index value in the interval [0.4, 0.7) to the "medium dynamic type" sound effect segments (such as object sliding sounds, mild collision sounds); Match the frame images with the sound effect matching index value in the interval [0.7, 1.0] to the "high-intensity event type" sound effect segments (such as bursting sounds, rapid whistling sounds).

[0068] The sound effect segment resources are stored in the sound effect resource index table in a one-to-one binding manner with their corresponding matching intervals, and quick retrieval and insertion processing are performed according to the matching results during operation.

[0069] In this implementation, by constructing a joint sound effect matching mechanism based on the spatial audio image localization index and the emotional sound effect adjustment index, the comprehensive perception of the spatial characteristics and emotional attributes of the event area in the image is realized, and based on this, the precise matching and calling of sound effect segments are driven. Compared with the traditional method of simply mapping sound effects based on event types, this method introduces the sound effect matching index as an intermediate variable, taking into account both the position, offset degree, and perturbation intensity of the event area in the image, and also integrating information such as color mutation and texture fluctuation of the emotional component, thereby improving the situational adaptability and expression accuracy of sound effect synthesis. By setting multiple sound effect matching intervals and binding them to sound effect segments one by one, progressive sound effect adaptation from "background type" to "high-intensity event type" can be achieved, with good continuity and difference, avoiding abrupt switching or distortion and mismatch. At the same time, sound effect resources are bound and dynamically retrieved according to the index table within the matching interval, ensuring that the system has high-efficiency response and fast insertion capabilities during real-time generation, greatly improving the fluency and realism of the generated sound effects, and providing precise support for multi-scene image enhancement and audio-visual synthesis.

[0070] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0071] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.

Claims

1. An audiovisual content synchronous sound effect synthesis method, characterized in that, Including the following steps: Obtain several frames of image data to be synthesized, and perform feature extraction on each of them respectively to obtain the image feature set of each frame to be synthesized; Input the image feature set of each frame to be synthesized into a pre-trained visual event recognition model to identify several frames of images for which sound effects need to be generated, and output the corresponding trigger event frame information; Based on the image feature set and the trigger event frame information, analyze the spatial sound image localization index and the emotional sound effect adjustment index of each frame of image for which sound effects need to be generated; Based on the spatial sound image localization index and the emotional sound effect adjustment index, match the sound effect segments of each frame of image for which sound effects need to be generated, and perform sound effect synthesis processing.

2. The method for synthesizing synchronized sound effects of audiovisual content according to claim 1, wherein Each frame of image data includes the pixel values of several pixel points and two-dimensional coordinates, and the image feature set includes the average image brightness, the color mutation rate, and the modulus value of the inter-frame optical flow vector.

3. The method for synthesizing synchronized sound effects of audiovisual content according to claim 2, wherein, The specific steps to obtain the image feature set of each frame to be synthesized are as follows: Perform average gray processing on each frame to be synthesized respectively to obtain the average image brightness of each frame to be synthesized; Perform color space histogram analysis processing on adjacent frames to be synthesized to obtain the color mutation rate of each frame to be synthesized; Perform average processing of optical flow estimation on adjacent frames to be synthesized to obtain the modulus value of the inter-frame optical flow vector of each frame to be synthesized.

4. The method for synthesizing synchronous sound effects of audiovisual content according to claim 1, characterized in that The trigger event frame information includes the trigger time, the event type, and the event area image data. The event area image data includes the pixel values of several event pixel points and two-dimensional coordinates. The visual event recognition model includes a frame-level feature encoding layer, a time series modeling layer, an attention focusing layer, and an event classification and localization output layer.

5. The audio-visual content synchronous sound effect synthesis method according to claim 4, characterized in that, The specific steps to identify several frames of images for which sound effects need to be generated and output the corresponding trigger event frame information are as follows: In the frame-level feature encoding layer of the visual event recognition model, receive the image feature set of each frame to be synthesized, and perform field-by-field encoding to output the image encoding feature vector of the corresponding frame; In the time series modeling layer of the visual event recognition model, arrange the image encoding feature vectors of each frame in chronological order to construct a time series feature tensor, and extract the time correlation feature vector of each frame; In the attention focusing layer of the visual event recognition model, read the time correlation feature vector of each frame, and perform a difference score on the feature change rate between it and the adjacent frames, and extract the frames with a difference score higher than the set threshold as candidate trigger event frames; In the event classification and localization output layer of the visual event recognition model, based on the image encoding feature vector of the candidate trigger event frame, analyze and output the corresponding trigger event frame information of several frames of images for which sound effects need to be generated.

6. The audio-visual content synchronous sound effect synthesis method according to claim 2, wherein The specific steps to analyze the spatial sound image localization index of each frame of image for which sound effects need to be generated are as follows: Based on the event area image data, analyze the edge perturbation intensity and the centroid point coordinates of the event area of each frame of image for which sound effects need to be generated; Based on the two-dimensional coordinates of several pixel points of each frame of image for which sound effects need to be generated, analyze the image center coordinates of the corresponding frame, and combine the centroid point coordinates of the event area of the corresponding frame to analyze the center coordinate offset value; Based on the event area image data, analyze the ratio of the event area occupancy of each frame of image for which sound effects need to be generated; Read the modulus of the inter-frame optical flow vector of each frame of the image for which the sound effect needs to be generated, and conduct a comprehensive analysis in combination with the edge perturbation intensity, the central coordinate offset value, and the ratio of the event area to the total area of the corresponding frame of the image, so as to obtain the spatial sound image positioning index of each frame of the image for which the sound effect needs to be generated.

7. The method for synthesizing synchronous sound effects of audiovisual content according to claim 6, characterized in that, The specific formula for calculating the spatial sound image positioning index of a certain frame of the image for which the sound effect needs to be generated is as follows: ; Among them, , , , , are, in sequence, the spatial sound image localization index, edge perturbation intensity, frame - to - frame optical flow vector modulus value, central coordinate offset value, and the ratio of the area of the event region of a certain frame image for which the sound effect needs to be generated, , , , are, in sequence, the edge perturbation response coefficient, optical flow response coefficient, central offset response coefficient, and area ratio response coefficient stored in the database.

8. The method for synthesizing synchronized sound effects of audiovisual content according to claim 1, characterized in that, The specific steps for analyzing the emotional sound effect adjustment index of each frame of the image for which the sound effect needs to be generated are as follows: Based on the event area image data, analyze the local texture perturbation ratio and the maximum color gradient value of each frame of the image for which the sound effect needs to be generated; Read the average image brightness and the color mutation rate of each frame of the image for which the sound effect needs to be generated, and conduct a comprehensive analysis in combination with the texture perturbation ratio and the maximum color gradient value of the corresponding frame of the image, so as to obtain the emotional sound effect adjustment index of each frame of the image for which the sound effect needs to be generated.

9. The method for synthesizing synchronous sound effects of audiovisual content according to claim 1, wherein The specific formula for calculating the emotional sound effect adjustment index of a certain frame of the image for which the sound effect needs to be generated is as follows: ; Among them, , , , , are, in sequence, the emotional sound effect adjustment index, color mutation rate, texture perturbation ratio, maximum color gradient value, and average image brightness of a certain frame of image for which the sound effect needs to be generated, , , , , are, in sequence, the color mutation response coefficient, texture perturbation response coefficient, interaction response coefficient, color gradient response coefficient, and image brightness response coefficient stored in the database.

10. The method for synthesizing synchronized sound effects of audiovisual content according to claim 1, wherein Based on the spatial sound image positioning index and the emotional sound effect adjustment index, the specific steps for matching the sound effect segment of each frame of the image for which the sound effect needs to be generated are as follows: Read the spatial sound image positioning index and the emotional sound effect adjustment index of each frame of the image for which the sound effect needs to be generated, and conduct a weighted analysis to obtain the sound effect matching index of each frame of the image for which the sound effect needs to be generated; Respectively judge and analyze the sound effect matching index of each frame of the image for which the sound effect needs to be generated with several preset sound effect matching intervals, and each sound effect matching interval corresponds to a sound effect segment respectively; Take the sound effect segment corresponding to the case where the sound effect matching index is within a preset sound effect matching interval as the generated sound effect of the image.

Citation Information

Patent Citations

  • Video sound effect generation method and device, sound effect generation model training method and device and medium

    CN117793400A

  • Audio-visual data set making method and system

    CN118245226A

  • Method for generating sound effect video and electronic equipment

    CN119496954A

  • Method for improving video semantic understanding based on space-time convolution

    CN120014516A

  • System and method for audio-visual content synthesis

    CN1860504A