Real-time Fusion Method and System for Audio and Video Based on Historical Data
Through the audio and video fusion method based on historical data, multi-source audio and video data is automatically identified and processed, the problem of inefficiency in the existing technology is solved, real-time audio and video fusion and efficient image and audio switching are realized, and processing efficiency and the quality of generating audio and video are improved.
Patent Information
- Application Number
- CN202510380308.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-28
AI Technical Summary
In multi-source audio and video fusion processing, the existing technology is inefficient and cannot realize real-time processing, especially when there are many audio sources or image sources, manual orchestration efficiency is low, and real-time fusion of audio and video cannot be realized.
By obtaining historical audio and video data, extracting and associating feature groups of stored image frames and audio frames, establishing scene models and processing algorithms, automatically identifying and processing audio and video data of target objects, and realizing automatic switching and fusion of images and audio.
It improves the efficiency of audio and video fusion processing, realizes real-time processing and automatic switching of multi-source audio and video data, reduces manual intervention, and ensures the quality of generated audio and video.
Smart Images

Figure CN119906842B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of multi-source data processing, and in particular to a method and system for real-time fusion of audio and video based on historical data. Background Art
[0002] In practice, audio data is usually collected by a microphone, and image data is collected by a camera. After obtaining audio frames and image frames, through audio processing and image processing, an audio file and a video file are formed, and the two are arranged and fused according to certain rules to form a common video image with sound.
[0003] Generally, when there is only one sound source and one image source, it is only necessary to synchronously collect audio and video data, and after encoding, the audio packet queue and the video packet queue are directly integrated in sequence to obtain the target audio and video. However, when there are multiple sound sources and image sources, it is necessary to manually process the collected audio data and image data one by one, such as associating each video image with the sound, and finally through a specific arrangement order, so that the sound can correspond to the video image, and finally generate a video image in which the sound can follow the video camera switching.
[0004] Obviously, when there are many sound sources or image sources, the efficiency of manually arranging audio and video is very low, and the effect of real-time processing and fusion of audio and video cannot be achieved. Summary of the Invention
[0005] Aiming at the problems of low efficiency in multi-source audio and video fusion processing and inability to achieve real-time processing in practical applications, the first object of this application is to provide a method for real-time fusion of audio and video based on historical data, which uses historical big data to accurately determine the association between different audio data and image data, automatically fuses the associated audio and video, and can automatically switch images and sounds following the actual scene switching, without manual intervention, with high efficiency and can perform real-time processing on audio and video data. To achieve the above method, the second object of this application is to propose a system for real-time fusion of audio and video based on historical data, and the third object is to propose to protect a computer-readable storage medium. The specific solutions are as follows:
[0006] A method for real-time fusion of audio and video based on historical data, comprising:
[0007] S100, obtaining the historical audio and video data of each target object, respectively extracting a first feature group and a second feature group for identifying the image and audio of each target object from the image frame and the audio sampling frame, and storing the two in association;
[0008] S200, establishing and storing the evolution features of the audio and video data of the target object itself and between each target object in each scene model, and storing the target object name in association with the above evolution features and the scene model;
[0009] S300. Establish and associate the storage of images and audio processing algorithms of each target object in each scene model;
[0010] S400. Obtain multiple audio-visual data at the current moment, identify and select at least one target object from the audio-visual data as the current target object;
[0011] S500. Determine and, based on the evolution characteristics of the audio-visual data of the current target object, determine the scene model, and determine the image and audio processing algorithms according to the scene model;
[0012] S600. Generate image frames and audio sampling frames from the audio-visual data of the current target object, process the above images and audio based on the processing algorithm, and then encode and fuse them to form a video segment;
[0013] S700. Use the video segment as the starting segment or splice it with the video segment formed at the previous moment to generate the target audio-visual, and repeat steps S400 - S700;
[0014] Among them, the first feature group is a combination of specific image points in the image frame;
[0015] The second feature group is a combination of specific sampling values in the audio sampling frame;
[0016] The evolution characteristics include the change trend of the audio-visual data of the current target object itself and the association trend between the audio-visual data of the current target object and the audio-visual data of other target objects before transitioning from the current target object to the next target object.
[0017] By adopting the above technical solution, according to the characteristics of the audio-visual data of each target object, the audio-visual data of one target object is automatically selected from multiple audio-visual data for processing, and the evolution characteristics of the current audio-visual data are identified and confirmed during the processing. Thus, the scene model and related processing algorithms can be determined, realizing the automatic processing of images and audio, and continuous processing of the received audio-visual data, with high efficiency.
[0018] Furthermore, establishing and associating the storage of images and audio processing algorithms of each target object in each scene model further includes:
[0019] Store the rendering rules and background audio of the image frames of each target object in each scene model;
[0020] Store the importance ranking of each target object in each scene model;
[0021] The image and audio processing algorithms include:
[0022] Render the image frames of the target object according to the set rendering rules;
[0023] Perform filtering and enhancement processing on the audio of the target object, perform filtering and attenuation processing on the audio of the remaining target objects based on the importance ranking of each target object, or add background audio to the audio of the target object according to the scene model.
[0024] Furthermore, identifying and selecting at least one target object as the current target object from the audio-visual data includes:
[0025] Configure a processing priority for each target object and store it in association with the name of each target object;
[0026] At the same time, search for matches of the first feature group and the second feature group from the image frames and audio sampling frames of multiple audio-visual data to confirm the target objects included in the current audio-visual data;
[0027] Analyze the audio-visual data corresponding to each target object to generate the duration of the audio-visual;
[0028] Select the target objects whose audio-visual duration exceeds the set value, sort the target objects based on the processing priority of each target object to form a queue to be processed;
[0029] Select at least one target object with a higher ranking from the queue to be processed as the current target object.
[0030] By adopting the above technical solution, a target object can be selected for processing according to the audio-visual duration of each target object and the preset processing priority, avoiding the situation of frequent switching of later images and audio caused by too short audio-visual duration, and ensuring the playback quality of the later target video.
[0031] Furthermore, identifying and selecting at least one target object as the current target object from the audio-visual data further includes:
[0032] Analyze the audio-visual data corresponding to each target object, identify the position coordinates of the first feature group from each image frame of the target object, generate the movement trajectory of the first feature group, determine the speed and direction of the above movement trajectory, if the above movement trajectory exceeds the set interval, delete the corresponding target object from the queue to be processed; and / or
[0033] Analyze the audio-visual data corresponding to each target object, determine the intensity and amplitude of the audio of the target object from each audio sampling frame of the target object, if the above intensity and amplitude are lower than the set value, delete the corresponding target object from the queue to be processed.
[0034] By adopting the above technical solution, it is possible to avoid the situation of blurred picture quality and unclear sound in the generated target audio-visual content later, and ensure the quality of the generated target audio-visual content.
[0035] Further, determining the evolution characteristics of the current target object's audio-visual data includes:
[0036] Extracting multiple image frames and audio sampling frames of the target object within a set time period before the current moment;
[0037] Identifying and storing the position coordinates of the first feature group from multiple image frames, generating the action trajectory of the target object based on the above position coordinates, and determining the evolution characteristics of the target object's audio-visual data according to the action trajectory; and / or
[0038] Identifying and storing the second feature group from multiple audio sampling frames, generating the audio change trend of the target object based on the second feature group, and determining the evolution characteristics of the target object's audio-visual data according to the audio change trend.
[0039] Further, the target object includes a person or object that can make a sound and has a visible appearance in the audio-visual recording scene;
[0040] The scene model includes an audio-visual recording scene with multiple image sources and sound sources.
[0041] In a second aspect, the present application provides a real-time audio and video fusion system based on historical data, including:
[0042] A data acquisition unit configured to acquire the historical audio-visual data of each target object and the audio-visual data to be processed;
[0043] A target object marking unit configured to, based on the historical audio-visual data, respectively extract the first feature group and the second feature group for identifying the image and audio of each target object from the image frames and audio sampling frames, and perform identity marking on the target objects in the current audio-visual data based on the first feature group and the second feature group;
[0044] A data storage unit configured to associatively store the evolution characteristics of the target objects themselves and the audio-visual data between the target objects in each scene model, associatively store the image and audio processing algorithms of each target object in each scene model, and store the already generated video segments;
[0045] A data processing unit configured to identify and select at least one target object from the current audio-visual data as the current target object, determine and confirm the scene model according to the evolution characteristics of the current target object's audio-visual data, determine the image and audio processing algorithms according to the scene model, process the image and audio of the current target object based on the processing algorithms, and then encode and fuse them to form a video segment;
[0046] A video segment splicing unit, configured to obtain the above-mentioned video segments, use the video segments as starting segments or splice them with the already generated video segments to generate a target audio-visual and output it;
[0047] Wherein, the first feature group is a combination of specific image points in an image frame;
[0048] The second feature group is a combination of specific sampling values in an audio sampling frame.
[0049] Further, the data processing unit is configured with:
[0050] A first target object screening module, configured to:
[0051] Analyze the audio-visual data of each target object associated with the current audio-visual data, generate the duration of the audio-visual corresponding to each target object, select the target objects whose audio-visual duration exceeds a set value, and sort the target objects based on the processing priority of each target object to form a to-be-processed queue;
[0052] Select at least one top-ranked target object from the to-be-processed queue as the current target object; and
[0053] A second target object screening module, configured to:
[0054] Analyze the audio-visual data corresponding to each target object, identify the position coordinates of the first feature group from each image frame of the target object, generate the movement trajectory of the first feature group, determine the speed and direction of the movement trajectory, and if the movement trajectory exceeds a set interval, delete the corresponding target object from the to-be-processed queue; and / or
[0055] Analyze the audio-visual data corresponding to each target object, determine the intensity and amplitude of the audio of the target object from each audio sampling frame of the target object, and if the intensity and amplitude are lower than a set value, delete the corresponding target object from the to-be-processed queue.
[0056] Finally, there is also provided a computer-readable storage medium, in which a program module for implementing the above-mentioned real-time audio and video fusion method based on historical data is loaded.
[0057] In summary, the present application includes at least one of the following beneficial technical effects: This audio-visual fusion system obtains the characteristics of the audio-visual data of each target object according to historical data, automatically selects the audio-visual data of one target object from multiple audio-visual data for processing, and identifies and confirms the evolution characteristics of the current audio-visual data during the processing, thereby determining the scene model and related processing algorithms, realizing the automatic processing of images and audio, and significantly improving the processing efficiency compared with manual processing. Brief Description of the Drawings
[0058] Figure 1 This is the overall schematic diagram of the audio and video fusion method of this application;
[0059] Figure 2 This is the schematic diagram of the method for selecting target objects from audio and video data;
[0060] Figure 3 This is the functional block diagram of the system of this application.
[0061] Reference Signs: 1. Data acquisition unit; 2. Target object marking unit; 3. Data storage unit; 4. Data processing unit; 5. Video segment splicing unit. Detailed Embodiments
[0062] The following details the embodiments of this application, and the examples of the embodiments are shown in the drawings.
[0063] In the description of this specification, the description with reference to the terms "certain embodiments", "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiments or examples. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0064] The embodiments of this application disclose a real-time audio and video fusion method based on historical data. Referring to Figure 1 , it mainly includes the following steps:
[0065] S100. Obtain the historical audio and video data of each target object, respectively extract the first feature group and the second feature group for identifying each target object from the image frames and audio sampling frames, and associate and store the first feature group, the second feature group with the target object names;
[0066] S200. Establish and store the evolution features of the target objects themselves and the audio and video data between the target objects in each scene model, and associate and store the target object names with the above evolution features and scene models;
[0067] S300. Establish and associate and store the image and audio processing algorithms of each target object in each scene model;
[0068] S400. Obtain multiple audio and video data at the current moment, and identify and select at least one target object from the audio and video data as the current target object;
[0069] S500, determine and based on the evolution characteristics of the current target object's audio-visual data, confirm the scene model and determine the processing algorithms for images and audio according to the scene model;
[0070] S600, generate image frames and audio sampling frames from the audio-visual data of the current target object, process the above-mentioned images and audio based on the processing algorithm, and then encode and fuse them to form a video segment;
[0071] S700, use the video segment as the starting segment or splice it with the video segment formed at the previous moment to generate the target audio-visual, and repeat steps S400 - S700.
[0072] In step S100, the target object includes, but is not limited to, people or objects that can make sounds and have visible shapes in the audio-visual recording scene. The first feature group is a combination of specific image points in the image frame, such as a combination of specific pixel points. The above combination can accurately reflect the target object to which the current image frame belongs. The second feature group is a combination of specific sampling values in the audio sampling frame, such as voice frequency characteristics, voiceprint characteristics, etc. Based on the above second feature group, the target object to which the sound belongs can be judged from the sound. By storing the above first feature group and the second feature group, the target objects contained in the current audio-visual data can be obtained.
[0073] In step S200, the evolution characteristics include the change trend of the current target object's own audio-visual data when the image and audio of the current target object transition from the current moment to the next moment, or the correlation trend between the audio-visual data of the current target object and other target objects before transitioning from the current target object to the next target object. For example, the change trend of the sound amplitude of the target object or the change trend of the target object's actions. The above correlation trend includes the coordinate position relationship between different target objects in the image, the ratio relationship of frequencies or amplitudes between different target objects in the audio, etc. In specific practice, the change trend of the target object's image or actions can be obtained through data analysis of the position coordinates of the first feature group in a set number of consecutive frames before and after. The change trend of the target object's audio can be obtained through the analysis of the second feature group.
[0074] In step S200, the scene model refers to the audio-visual recording environment where each target object is located. In the embodiments of the present application, it specifically refers to a scene with multiple image sources and sound sources, such as a video conference scene, a variety show recording scene, etc. Since each target object in different scene models has different image and audio characteristics, including the aforementioned evolution characteristics, the above evolution characteristics and scene model can be associated and stored, and the scene model where the current target object is located can be inferred from the evolution characteristics.
[0075] In step S300, establish and associate and store the image and audio processing algorithms of each target object in each scene model. Specifically, it includes:
[0076] S310 stores the rendering rules and background audio of each target object image frame in each scene model. By storing in advance the image rendering rules of different target objects in different scene models and the background audio to be fused, the algorithm matching time required for subsequent audio-visual data processing can be greatly shortened.
[0077] S320 stores the importance ranking of each target object in each scene model. The above ranking refers to the processing priority of each target object in different scene models. For example, when there are 3 target objects in the same scene model, the above 3 target objects can be ranked, so as to determine the priority of audio-visual data processing, first process the audio-visual data of a certain important target object, and the audio-visual data of the remaining target objects is processed later or directly deleted.
[0078] In the embodiment of the present application, the image and audio processing algorithms include:
[0079] S330 performs rendering processing on the image frames of the target object according to the set rendering rules, such as adjusting and rendering the image color, etc.;
[0080] S340 performs filtering and enhancement processing on the audio of the target object, and performs filtering and attenuation processing on the audio of the remaining target objects based on the importance ranking of each target object, or adds background audio to the audio of the target object according to the scene model. The audio filtering and attenuation processing in the above steps includes filtering out the high-frequency and low-frequency components in the audio, reducing the size of the audio data volume, and at the same time reducing the amplitude of the sound, etc.
[0081] Through the above step S300, different adjustments can be made to the images and audios of different target objects, so that the levels of the target audio-visual synthesized later are clearer, highlighting the images and sounds of the target objects.
[0082] In step S400, at least one target object is identified and selected from the audio-visual data as the current target object, as Figure 2 shown, and further includes:
[0083] S410 configures a processing priority for each target object and stores it in association with the name of each target object;
[0084] S420 searches for and matches the first feature group and the second feature group from the image frames and audio sampling frames of multiple audio-visual data to confirm the target objects included in the current audio-visual data;
[0085] S430 analyzes the audio-visual data corresponding to each target object to generate the duration of the audio-visual;
[0086] S440. Select target objects whose audio - video duration exceeds a set value, sort the target objects based on the processing priority of each target object, and form a queue to be processed.
[0087] S450. Select at least one target object with a higher ranking from the queue to be processed as the current target object.
[0088] The above - mentioned technical solution can select a target object for processing according to the audio - video duration of each target object and the preset processing priority, avoiding the frequent switching of later - stage images and audio caused by too short audio - video duration, and ensuring the playback quality of the later - stage target video.
[0089] In order to ensure the quality of the later - generated target audio - video, in the embodiment of the present application, when identifying and selecting at least one target object as the current target object from the audio - video data, it further includes:
[0090] S460. Analyze the audio - video data corresponding to each target object, identify the position coordinates of the first feature group from each image frame of the target object, generate the movement trajectory of the first feature group, determine the speed and direction of the above - mentioned movement trajectory. If the above - mentioned movement trajectory exceeds the set interval, delete the corresponding target object from the queue to be processed; and / or
[0091] S470. Analyze the audio - video data corresponding to each target object, determine the intensity and amplitude of the target object's audio from each audio sampling frame of the target object. If the above - mentioned intensity and amplitude are lower than the set value, delete the corresponding target object from the queue to be processed.
[0092] In a specific embodiment, the position coordinates of the first feature group refer to the pixel coordinates of each image site in the image frame.
[0093] In step S500, determining the evolution characteristics of the audio - video data of the current target object further includes:
[0094] S510. Extract multiple image frames and audio sampling frames of the target object within a set duration before the current moment.
[0095] S521. Identify and store the position coordinates of the first feature group from multiple image frames, generate the action trajectory of the target object based on the above - mentioned position coordinates, and determine the evolution characteristics of the audio - video data of the target object according to the action trajectory; and / or
[0096] S522. Identify and store the second feature group from multiple audio sampling frames, generate the audio change trend of the target object based on the above - mentioned second feature group, and determine the evolution characteristics of the audio - video data of the target object according to the above - mentioned audio change trend.
[0097] The present application provides a real-time audio and video fusion system based on historical data, as Figure 3 shown, mainly including: a data acquisition unit 1, a target object marking unit 2, a data storage unit 3, a data processing unit 4, and a video segment splicing unit 5.
[0098] The data acquisition unit 1 is configured to acquire historical data and audio-visual data to be processed.
[0099] The target object marking unit 2 is configured to, based on historical audio-visual data, extract a first feature group and a second feature group for identifying each target object from an image frame and an audio sampling frame respectively, and perform identity marking on the target objects in the current audio-visual data based on the first feature group and the second feature group, that is, configure an identity recognition reference for each target object. In practice, the above target object marking unit 2 includes an image recognition module and an audio recognition module, which are respectively used to identify specific graphics or pixel sites in the image frame, and voice features in the audio, such as voiceprint features, voice frequency features, etc.
[0100] The data storage unit 3 is data-connected to the data acquisition unit 1 and the target object marking unit 2, and is configured to store the historical audio-visual data of each target object, associate and store the evolution features of the target object itself and the audio-visual data between each target object in each scenario model, associate and store the image and audio processing algorithms of each target object in each scenario model, associate and store the data of each target object and its corresponding first feature group and second feature group, and store the generated video segments.
[0101] The data processing unit 4 is configured to identify and select at least one target object from the current audio-visual data as the current target object, determine and confirm the scenario model according to the evolution features of the audio-visual data of the current target object, determine the image and audio processing algorithms according to the scenario model, process the image and audio of the current target object based on the processing algorithms, and then encode and fuse them to form a video segment to be spliced.
[0102] The video segment splicing unit 5 is data-connected to the data processing unit 4, and is configured to obtain the video segment generated in the data processing unit 4, and then use the video segment as the starting segment or splice it with the already stored video segments to generate the target audio-visual and output it.
[0103] In order to delete the audio-visual data of secondary target objects from the audio-visual data to be processed, while improving the processing efficiency and ensuring the quality of the target audio-visual generated later, specifically, the above data processing unit is configured with a first target object screening module and a second target object screening module.
[0104] The first target object screening module is configured to: analyze the audio and video data of each target object associated with the current audio and video data, generate the duration of the audio and video corresponding to each target object, select the target objects whose audio and video duration exceeds a set value, and sort the target objects based on the processing priority of each target object to form a queue to be processed;
[0105] Then, at least one target object with a high ranking is selected from the queue to be processed as the current target object.
[0106] The second target object screening module is configured to analyze the audio and video data corresponding to each target object, identify the position coordinates of the first feature group from each image frame of the target object, generate a movement trajectory of the first feature group, determine the speed and direction of the movement trajectory, and if the movement trajectory exceeds a set interval, delete the corresponding target object from the queue for processing; and / or
[0107] The audio and video data corresponding to each target object is analyzed, and the intensity and amplitude of the target object audio are determined from each audio sampling frame of the target object. If the intensity and amplitude are lower than the set value, the corresponding target object is deleted from the waiting queue.
[0108] The present application also provides a computer-readable storage medium storing a computer program for executing the method for real-time audio and video fusion based on historical data described in any of the above embodiments. The computer-readable storage medium includes, but is not limited to, a disk memory, a CD-ROM, an optical memory, and the like.
[0109] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A real-time fusion method for audio and video based on historical data, characterized in that, include: S100, acquiring historical audio and video data of each target object, extracting a first feature group and a second feature group for identifying the image and audio of each target object from the image frame and the audio sampling frame, respectively, and storing the first feature group and the second feature group in association with each other; S200, establishing and storing the evolution characteristics of the target object itself and the audio and video data between the target objects in each scene model, and associating the target object name with the above evolution characteristics and the scene model for storage; S300, establishing and associating storage of image and audio processing algorithms for each target object in each scene model; S400, acquiring a plurality of audio and video data at a current moment, identifying and selecting at least one target object from the audio and video data as a current target object; S500, determining and confirming a scene model based on the evolution characteristics of the audio and video data of the current target object and determining an image and audio processing algorithm based on the scene model; S600, generating image frames and audio sample frames from the audio and video data of the current target object, processing the image and audio based on the processing algorithm, and then encoding and fusing them to form a video segment; S700, taking the video segment as the starting segment or splicing it with the video segment formed at the previous moment to generate the target audio and video, and repeating steps S400-S700; Wherein, the first feature group is a combination of specific image points in the image frame; The second feature group is a combination of specific sampling values in the audio sampling frame; The evolution characteristics include a change trend of the audio and video data of the current target object itself before the transition from the current target object to the next target object, and a correlation trend between the audio and video data of the current target object and other target objects.
2. The real-time audio and video fusion method based on historical data according to claim 1, wherein, Establish and associate storage of image and audio processing algorithms for each target object in each scene model, including: Storing rendering rules and background audio for each target object image frame in each scene model; Store the importance ranking of each target object in each scene model; The image and audio processing algorithms include: Render the image frame of the target object according to the set rendering rules; Perform filtering and enhancement processing on the audio of the target object, and perform filtering and attenuation processing on the audio of the remaining target objects based on the importance ranking of each target object, or add background audio to the audio of the target object according to the scene model.
3. The real-time audio and video fusion method based on historical data according to claim 1, wherein Identifying and selecting at least one target object from the audio and video data as a current target object includes: Configure a processing priority for each target object and store it in association with the name of each target object; Simultaneously searching for matching first feature groups and second feature groups from image frames and audio sample frames of a plurality of audio and video data to identify a target object contained in the current audio and video data; Analyze the audio and video data corresponding to each target object to generate the duration of the audio and video; Select target objects whose audio and video duration exceeds a set value, sort the target objects based on their processing priority, and form a queue to be processed; At least one target object with a high ranking is selected from the queue to be processed as the current target object.
4. The real-time audio and video fusion method based on historical data according to claim 2, characterized in that, Identifying and selecting at least one target object from the audio and video data as the current target object further includes: Analyze the audio-visual data corresponding to each target object, identify the position coordinates of the first feature group from each image frame of the target object, generate the movement trajectory of the first feature group, determine the speed and direction of the above movement trajectory, and if the above movement trajectory exceeds the set interval, delete the corresponding target object from the queue to be processed; and / or Analyze the audio-visual data corresponding to each target object, and determine the intensity and amplitude of the target object's audio from each audio sampling frame of the target object. If the above intensity and amplitude are lower than the set value, delete the corresponding target object from the queue to be processed.
5. The real-time audio and video fusion method based on historical data according to claim 1, characterized in that Determine the evolution characteristics of the audio-visual data of the current target object, including: Extract multiple image frames and audio sampling frames of the target object within a set time period before the current moment; Identify and store the position coordinates of the first feature group from multiple image frames, generate the action trajectory of the target object based on the above position coordinates, and determine the evolution characteristics of the target object's audio-visual data according to the action trajectory; and / or Identify and store the second feature group from multiple audio sampling frames, generate the audio change trend of the target object based on the above second feature group, and determine the evolution characteristics of the target object's audio-visual data according to the above audio change trend.
6. The real-time audio and video fusion method based on historical data according to claim 1, characterized in that The target object includes a person or object that can make a sound and has a visible appearance in an audio-visual recording scene.
7. The real-time audio and video fusion method based on historical data according to claim 1, characterized in that The scene model includes an audio-visual recording scene with multiple image sources and sound sources.
8. A real-time audio and video fusion system based on historical data, characterized in that, Including: A data acquisition unit (1) configured to obtain the historical audio-visual data of each target object and the audio-visual data to be processed; A target object marking unit (2) configured to, based on the historical audio-visual data, extract the first feature group and the second feature group for identifying the images and audio of each target object from the image frames and audio sampling frames respectively, and perform identity marking on the target objects in the current audio-visual data based on the above first feature group and second feature group; A data storage unit (3) configured to store the historical audio-visual data of each target object, associatively store the evolution characteristics of the target object itself and the audio-visual data between target objects in each scene model, associatively store the image and audio processing algorithms of each target object in each scene model, and store the generated video segments; A data processing unit (4) configured to identify and select at least one target object from the current audio-visual data as the current target object, determine and confirm the scene model according to the evolution characteristics of the current target object's audio-visual data, determine the image and audio processing algorithms according to the scene model, process the image and audio of the current target object based on the processing algorithms, and then encode and fuse them to form a video segment; A video segment splicing unit (5) configured to obtain the above video segment, use the video segment as the starting segment or splice it with the generated video segments to generate and output the target audio-visual; Among them, the first feature group is a combination of specific image points in the image frame; The second feature group is a combination of specific sampling values in the audio sampling frame; The evolution characteristics include the change trend of the current target object's own audio-visual data and the correlation trend between the current target object's audio-visual data and the audio-visual data of other target objects before transitioning from the current target object to the next target object.
9. The real-time audio and video fusion system based on historical data according to claim 8, wherein The data processing unit is configured with: A first target object screening module, configured to: Analyze the audio-visual data of each target object associated with the current audio-visual data, generate the duration of the audio-visual data corresponding to each target object, select the target objects whose audio-visual duration exceeds the set value, and sort the target objects based on the processing priorities of each target object to form a to-be-processed queue; Select at least one target object with a higher ranking from the to-be-processed queue as the current target object; And A second target object screening module, configured to: Analyze the audio-visual data corresponding to each target object, identify the position coordinates of the first feature group from each image frame of the target object, generate the movement trajectory of the first feature group, determine the speed and direction of the movement trajectory, and if the movement trajectory exceeds the set interval, delete the corresponding target object from the to-be-processed queue; and / or Analyze the audio-visual data corresponding to each target object, determine the intensity and amplitude of the audio of the target object from each audio sampling frame of the target object, and if the intensity and amplitude are lower than the set value, delete the corresponding target object from the to-be-processed queue.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for real-time fusion of audio and video based on historical data according to any one of claims 1-7.
Citation Information
Patent Citations
Image processing method and device and hardware device
CN111489769A
Music recommendation method and device and readable storage medium
CN113569088A