An integrated intelligent audio and video playback system
By building an audio-visual matching recognition model and using deep learning network to identify audio-visual and environmental features, the picture stability problem of audio-visual playback system in dynamic environments in the existing technology is solved, intelligent recommendation and optimization are achieved, and the stability and effect of audio-visual playback are improved.
Patent Information
- Application Number
- CN202510371259.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-27
AI Technical Summary
The existing intelligent audio and video playback system cannot effectively predict the dynamic changes in audio and video content and environmental status in a dynamic environment, resulting in lagging picture stability adjustment and unable to adapt to the shaking trend during ship navigation.
By building an audio-visual fit recognition model, using deep learning network to identify audio-visual and environmental features, calculate brightness, sound and shaking fit scores, optimize the audio-visual playback time and environmental data, and achieve intelligent recommendation and optimization.
It improves the picture stability and playback effect of audio and video playback in complex water environments, ensuring that the audio and video data and environmental data tend to stabilize during each optimization cycle.
Smart Images

Figure CN119893187B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent audio and video optimization, and in particular to an integrated intelligent audio and video playback system. Background Art
[0002] With the rapid development of multimedia technology, the application demand for intelligent audio and video playback systems in mobile carriers is growing. However, in dynamic environments, especially in scenes affected by complex external conditions such as surface ships, the audio and video playback effect is affected by many factors such as ship movement and environmental changes.
[0003] Existing systems mostly rely on static environmental parameters to adjust audio and video, and lack the ability to predict dynamic environmental changes; they ignore the deep correlation between environmental parameters and audio and video content characteristics, and are unable to adapt to the dynamic changes of audio and video content and environmental conditions; in addition, during the navigation of the ship, the roll, pitch and heave motions caused by the navigation of the ship and the waves have significant time-varying characteristics. Traditional methods cannot predict the shaking trend in advance, resulting in a lag in the adjustment of picture stability.
[0004] To this end, an integrated intelligent audio and video playback system is proposed. Summary of the invention
[0005] The object of the present invention is to provide an integrated intelligent audio and video playback system, including collecting environmental data to obtain historical audio and video environment data; making predictions based on the historical audio and video environment data to obtain predicted audio and video environment data; acquiring stored playable audio and video data and identifying audio and video features of the playable audio and video data; identifying playback environment data based on the playback duration of the playable audio and video data and the predicted audio and video environment data; identifying the playback environment data to obtain playback environment features; obtaining an audio and video recommendation score based on the audio and video features and the playback environment features; making recommendations based on the audio and video recommendation score and determining the audio and video data to be played; acquiring the audio and video features and the playback environment data of the audio and video data to be played, and optimizing the audio and video playback based on the playback environment data and the audio and video features; and realizing intelligent recommendation and optimization of audio and video through audio and video data and ship motion data.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] An integrated intelligent audio and video playback system, comprising:
[0008] The audio-visual environment recognition module collects environment data to obtain historical audio-visual environment data; performs prediction based on the historical audio-visual environment data to obtain predicted audio-visual environment data; the predicted audio-visual environment data includes predicted brightness data, predicted sound data, and predicted shaking data;
[0009] The audio-visual playback recommendation module obtains the stored playable audio-visual data and identifies the audio-visual characteristics of the playable audio-visual data; the audio-visual characteristics include audio-visual brightness characteristics, audio-visual sound characteristics, and audio-visual shaking characteristics; according to the playback duration of the playable audio-visual data and the predicted audio-visual environment data, the playback environment data is identified; the playback environment data is identified to obtain the playback environment characteristics; according to the audio-visual characteristics and the playback environment characteristics, the audio-visual recommendation score is identified.
[0010] The audio-visual playback optimization module makes recommendations according to the audio-visual recommendation score and determines the audio-visual data to be played; obtains the audio-visual characteristics and the playback environment data of the audio-visual data to be played, and optimizes the audio-visual playback according to the playback environment data and the audio-visual characteristics.
[0011] The process of identifying the predicted audio-visual environment data according to the historical audio-visual environment data is as follows:
[0012] Obtain the navigation trajectory data of the ship on the water surface, and the navigation trajectory data includes navigation coordinate points, navigation speed, and navigation time points; collect the historical environment data of the area passed by the navigation trajectory data as the historical audio-visual environment data.
[0013] Make a prediction according to the historical audio-visual environment data, the navigation speed, and the navigation time point to obtain the predicted environment data when the ship passes through the navigation coordinate point, which is used as the predicted audio-visual environment data.
[0014] Identify the playable audio-visual data and the playback environment data through the audio-visual fit recognition model; the audio-visual fit recognition model includes: a data feature recognition layer, a first fit recognition layer, a second fit recognition layer, a third fit recognition layer, and a recommendation score calculation layer.
[0015] The data feature recognition layer identifies the audio-visual characteristics according to the playable audio-visual data, including audio-visual brightness characteristics, audio-visual sound characteristics, and audio-visual shaking characteristics; identifies the playback environment characteristics according to the playback environment data, including environmental brightness characteristics, environmental sound characteristics, and environmental shaking characteristics.
[0016] The first fit recognition layer identifies according to the audio-visual brightness characteristics and the environmental brightness characteristics to obtain the brightness fit score.
[0017] The second fit recognition layer identifies according to the audio-visual sound characteristics and the environmental sound characteristics to obtain the sound fit score.
[0018] The third fit recognition layer identifies according to the audio-visual shaking characteristics and the environmental shaking characteristics to obtain the stability fit score.
[0019] The recommendation score calculation layer calculates the audio-visual recommendation score according to the brightness fit score, the sound fit score, and the stability fit score.
[0020] The formula for calculating the audio - video recommendation score based on the brightness matching score, sound matching score, and stability matching score is as follows:
[0021] ;
[0022] ;
[0023] Among them, represents the audio - video recommendation score; represents the brightness matching score; represents the sound matching score; represents the stability matching score; represents the first matching weight; represents the second matching weight; represents the third matching weight.
[0024] The audio - video playback optimization process includes dividing the playback duration of the to - be - played audio - video data to obtain an optimization period; the process of obtaining the optimization period includes:
[0025] Obtain the audio - video characteristics and playback environment data of the to - be - played audio - video data;
[0026] Identify the audio - video characteristics through a clustering algorithm, and perform clustering analysis based on the audio - video brightness characteristics, audio - video sound characteristics, and audio - video shaking characteristics to obtain the first - period division data;
[0027] Identify the playback environment data through a clustering algorithm, and perform clustering analysis based on the predicted brightness data, predicted sound data, and predicted shaking data to obtain the second - period division data;
[0028] Divide the playback duration according to the first - period division data and the second - period division data to obtain the optimization period.
[0029] Within the optimization period, perform playback optimization according to the playback environment data and audio - video characteristics, including brightness optimization, sound optimization, and stability optimization;
[0030] The brightness optimization includes identifying the brightness optimization data according to the predicted brightness data in the playback environment data and the audio - video brightness characteristics, and optimizing the brightness of the audio - video according to the brightness optimization data;
[0031] The sound optimization includes identifying the sound optimization data according to the predicted sound data in the playback environment data and the audio - video sound characteristics, and optimizing the sound of the audio - video according to the sound optimization data;
[0032] The stability optimization includes identifying the stability optimization data according to the predicted shaking data in the playback environment data and the audio - video shaking characteristics, and optimizing the picture of the audio - video according to the stability optimization data.
[0033] The acquisition of the stable and optimized data includes:
[0034] Identify the predicted sway data to obtain the characteristics of the predicted sway data; the predicted sway data includes the roll parameter, pitch parameter, and heave parameter of the vessel;
[0035] Detect the motion vectors within the picture of the audio-visual data to be played, identify the picture sway and change frequency, and obtain the audio-visual sway characteristics;
[0036] Obtain the characteristics of the predicted sway data and the audio-visual sway characteristics corresponding to the same optimization period, superimpose the sway motion of the vessel and the sway motion of the audio-visual picture, calculate the motion direction and amplitude to be compensated, and obtain the stable and optimized data.
[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0038] 1. Based on the deep learning network, the present invention constructs an audio-visual matching recognition model; identifies the audio-visual data that can be played and the playback environment data to obtain the audio-visual characteristics and the playback environment characteristics; identifies the brightness matching score according to the audio-visual brightness characteristics and the environment brightness characteristics, identifies the sound matching score according to the audio-visual sound characteristics and the environment sound characteristics, and identifies the stable matching score according to the audio-visual sway characteristics and the environment sway characteristics; calculates the audio-visual recommendation score through the brightness matching score, the sound matching score, and the stable matching score; thereby obtaining the matching degree between the audio-visual data and the playback environment data.
[0039] 2. The present invention obtains the audio-visual characteristics of the audio-visual data to be played and the playback environment data; identifies the first cycle division data through the clustering algorithm for the audio-visual characteristics; identifies the second cycle division data through the clustering algorithm for the playback environment data; divides the playback duration according to the first cycle division data and the second cycle division data to obtain the optimization period; thereby ensuring that within each optimization period, the audio-visual playback data and the playback environment data tend to be stable.
[0040] 3. The present invention identifies the predicted sway data to obtain the characteristics of the predicted sway data; detects the motion vectors within the picture of the audio-visual data to be played, identifies the picture sway and change frequency, and obtains the audio-visual sway characteristics; obtains the stable and optimized data according to the characteristics of the predicted sway data and the audio-visual sway characteristics within the same optimization period; thereby accurately compensating and optimizing the audio-visual picture according to the vessel sway data and the audio-visual picture sway, and improving the picture playback effect of the audio-visual in a complex water surface environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a schematic structural diagram of an integrated intelligent audio-visual playback system of the present invention;
[0042] Figure 2 Schematic diagram of the audio - video recommendation process of the present invention;
[0043] Figure 3 Schematic diagram of the structure of the audio - video matching recognition model of the present invention;
[0044] Figure 4 Schematic diagram of the audio - video optimization process of the present invention. Detailed implementation manners
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] Embodiment 1
[0047] An integrated intelligent audio - video playing system, the structure of which is as Figure 1 shown, including:
[0048] An audio - video environment recognition module, which collects environmental data to obtain historical audio - video environmental data; predicts according to the historical audio - video environmental data to obtain predicted audio - video environmental data; the predicted audio - video environmental data includes predicted brightness data, predicted sound data and predicted shaking data.
[0049] The process of identifying the predicted audio - video environmental data according to the historical audio - video environmental data is as follows:
[0050] Obtain the navigation track data of the ship on the water surface, the navigation track data includes navigation coordinate points, navigation speed and navigation time points; collect the historical environmental data of the area passed by the navigation track data as the historical audio - video environmental data;
[0051] Predict according to the historical audio - video environmental data, navigation speed and navigation time points to obtain the predicted environmental data when the ship passes through the navigation coordinate point as the predicted audio - video environmental data.
[0052] The present invention obtains the navigation track data of the ship, including navigation coordinate points, navigation speed and navigation time points; collects the historical audio - video environmental data of the area passed by the navigation track, and predicts according to the historical audio - video environmental data, navigation speed and navigation time points to obtain the predicted environmental data when the ship passes through the navigation coordinate point; thus accurately collecting the environmental data during the ship's navigation.
[0053] The audio - video playback recommendation module obtains the stored playable audio - video data and identifies the audio - video characteristics of the playable audio - video data; the audio - video characteristics include audio - video brightness characteristics, audio - video sound characteristics, and audio - video shaking characteristics; according to the playback duration of the playable audio - video data and the predicted audio - video environment data, the playback environment data is identified; the playback environment data is identified to obtain the playback environment characteristics; according to the audio - video characteristics and the playback environment characteristics, the audio - video recommendation score is identified. The process of audio - video recommendation is as Figure 2 shown.
[0054] Identify the playable audio - video data and the playback environment data through the audio - video fitting recognition model; the audio - video fitting recognition model is constructed based on a deep learning network, and its structure is as Figure 3 shown, including: a data feature recognition layer, a first fitting recognition layer, a second fitting recognition layer, a third fitting recognition layer, and a recommendation score calculation layer;
[0055] The data feature recognition layer identifies the audio - video characteristics according to the playable audio - video data, including audio - video brightness characteristics, audio - video sound characteristics, and audio - video shaking characteristics; identifies the playback environment characteristics according to the playback environment data, including environment brightness characteristics, environment sound characteristics, and environment shaking characteristics;
[0056] The first fitting recognition layer identifies based on the audio - video brightness characteristics and the environment brightness characteristics to obtain the brightness fitting score;
[0057] The second fitting recognition layer identifies based on the audio - video sound characteristics and the environment sound characteristics to obtain the sound fitting score;
[0058] The third fitting recognition layer identifies based on the audio - video shaking characteristics and the environment shaking characteristics to obtain the stability fitting score;
[0059] The recommendation score calculation layer calculates the audio - video recommendation score according to the brightness fitting score, the sound fitting score, and the stability fitting score.
[0060] The formula for calculating the audio - video recommendation score according to the brightness fitting score, the sound fitting score, and the stability fitting score is:
[0061] ;
[0062] ;
[0063] Among them, represents the audio - video recommendation score; represents the brightness fitting score; represents the sound fitting score; represents the stability fitting score; represents the first fitting weight; Represents the second matching weight; Represents the third matching weight.
[0064] Based on a deep learning network, the present invention constructs a video and audio matching recognition model; recognizes playable video and audio data and playback environment data to obtain video and audio features and playback environment features; recognizes a brightness matching score based on video and audio brightness features and environmental brightness features, recognizes a sound matching score based on video and audio sound features and environmental sound features, and recognizes a stability matching score based on video and audio shaking features and environmental shaking features; calculates a video and audio recommendation score through the brightness matching score, the sound matching score, and the stability matching score; thereby obtaining the matching degree between the video and audio data and the water surface playback environment data.
[0065] The video and audio playback optimization module makes recommendations according to the video and audio recommendation score and determines the video and audio data to be played; obtains the video and audio features and playback environment data of the video and audio data to be played, and optimizes the video and audio playback according to the playback environment data and the video and audio features. The optimization process of the video and audio is as Figure 4 shown.
[0066] The video and audio playback optimization process includes dividing the playback duration of the video and audio data to be played to obtain an optimization period; the obtaining process of the optimization period includes:
[0067] Obtain the video and audio features and playback environment data of the video and audio data to be played;
[0068] Identify the video and audio features through a clustering algorithm, and perform clustering analysis based on the video and audio brightness features, the video and audio sound features, and the video and audio shaking features to obtain the first period division data;
[0069] Identify the playback environment data through a clustering algorithm, and perform clustering analysis based on the predicted brightness data, the predicted sound data, and the predicted shaking data to obtain the second period division data;
[0070] Divide the playback duration according to the first period division data and the second period division data to obtain the optimization period.
[0071] The present invention obtains the video and audio features and playback environment data of the video and audio data to be played; identifies the video and audio features through a clustering algorithm to obtain the first period division data; identifies the playback environment data through a clustering algorithm to obtain the second period division data; divides the playback duration according to the first period division data and the second period division data to obtain the optimization period; thereby ensuring that within each optimization period, the video and audio playback data and the playback environment data tend to be stable.
[0072] During the optimization period, perform playback optimization according to the playback environment data and the video and audio features, including brightness optimization, sound optimization, and stability optimization;
[0073] The brightness optimization includes: identifying brightness optimization data according to the predicted brightness data and video and audio brightness characteristics in the playback environment data, and optimizing the brightness of the video and audio according to the brightness optimization data;
[0074] The sound optimization includes: identifying sound optimization data according to the predicted sound data and video and audio sound characteristics in the playback environment data, and optimizing the sound of the video and audio according to the sound optimization data;
[0075] The stability optimization includes: identifying stability optimization data according to the predicted shake data and video and audio shake characteristics in the playback environment data, and optimizing the picture of the video and audio according to the stability optimization data.
[0076] The present invention identifies brightness optimization data according to the predicted brightness data of the playback environment data and the video and audio brightness characteristics; identifies sound optimization data according to the predicted sound data of the playback environment data and the video and audio sound characteristics; identifies stability optimization data according to the predicted shake data and video and audio shake characteristics in the playback environment data; and optimizes the video and audio according to the brightness optimization data, sound optimization data and stability optimization data, effectively improving the playback effect of the video and audio in the water surface environment.
[0077] The acquisition of the stability optimization data includes:
[0078] Identifying the predicted shake data to obtain the characteristics of the predicted shake data; the predicted shake data includes the roll parameter, pitch parameter and heave parameter of the ship;
[0079] Detecting the motion vectors in the picture of the video and audio data to be played, identifying the picture shake and change frequency, and obtaining the video and audio shake characteristics;
[0080] Obtaining the characteristics of the predicted shake data and the video and audio shake characteristics corresponding to the same optimization period, superimposing the shake motion of the ship and the shake motion of the video and audio picture, and calculating the motion direction and amplitude to be compensated to obtain the stability optimization data.
[0081] The present invention identifies the predicted shake data to obtain the characteristics of the predicted shake data; detects the motion vectors in the picture of the video and audio data to be played, identifies the picture shake and change frequency, and obtains the video and audio shake characteristics; obtains the stability optimization data according to the characteristics of the predicted shake data and the video and audio shake characteristics within the same optimization period; thereby accurately compensating and optimizing the video and audio picture according to the ship shake data and the video and audio picture shake, and improving the picture playback effect of the video and audio in the complex water surface environment.
[0082] The present invention collects environmental data to obtain historical audio-visual environmental data; makes predictions based on the historical audio-visual environmental data to obtain predicted audio-visual environmental data; acquires the storable playable audio-visual data and identifies the audio-visual characteristics of the playable audio-visual data; identifies the playback environmental data according to the playback duration of the playable audio-visual data and the predicted audio-visual environmental data; identifies the playback environmental characteristics from the playback environmental data; identifies the audio-visual recommendation score according to the audio-visual characteristics and the playback environmental characteristics; makes recommendations based on the audio-visual recommendation score and determines the audio-visual data to be played; acquires the audio-visual characteristics and the playback environmental data of the audio-visual data to be played, and optimizes the audio-visual playback according to the playback environmental data and the audio-visual characteristics, realizing intelligent recommendation and optimization of audio-visual in a complex water surface environment.
[0083] Embodiment 2
[0084] An integrated intelligent audio-visual playback system, the structure of which is as Figure 1 shown, including:
[0085] An audio-visual environment recognition module that collects environmental data to obtain historical audio-visual environmental data; makes predictions based on the historical audio-visual environmental data to obtain predicted audio-visual environmental data; the predicted audio-visual environmental data includes predicted brightness data, predicted sound data, and predicted shaking data.
[0086] The historical audio-visual environmental data includes environmental brightness data, environmental sound data, and environmental shaking data;
[0087] The environmental brightness data is the brightness parameter in the audio-visual playback environment;
[0088] The environmental sound data is the sound parameter in the audio-visual playback environment, including the noise of the ship itself, the water surface noise, and the air noise; specifically, the noise of the ship itself includes engine noise, gear noise, and ship whistle, etc., the water surface noise includes water wave noise, the collision noise between the ship and the water wave, and the air noise includes wind noise, seabird noise, and the noise of other ships;
[0089] The environmental shaking data reflects the motion state of the ship itself, including roll parameters, pitch parameters, and heave parameters; among them, the roll parameters include roll angle, the pitch parameters include pitch angle, and the heave parameters include heave displacement.
[0090] The roll angle, pitch angle, and heave displacement of the ship are important parameters describing the motion state of the ship during navigation; the roll angle refers to the angle of the ship's left and right inclination, mainly caused by side waves or wind, and the ship speed and structure affect its amplitude; the pitch angle refers to the up and down movement of the bow and stern of the ship, which is related to the wave frequency and ship speed and affects the navigation performance; the heave displacement refers to the displacement of the ship in the vertical direction, caused by changes in waves and ship speed.
[0091] The process of identifying predicted audio-visual environment data based on the historical audio-visual environment data is as follows:
[0092] Obtain the navigation trajectory data of the vessels on the water surface. The navigation trajectory data includes navigation coordinate points, navigation speeds, and navigation time points. Collect the historical environment data of the area through which the navigation trajectory data passes as the historical audio-visual environment data.
[0093] Make a prediction based on the historical audio-visual environment data, navigation speed, and navigation time point to obtain the predicted environment data when the vessel passes through the navigation coordinate point, which is used as the predicted audio-visual environment data.
[0094] In the prediction of the predicted audio-visual environment data, a dual prediction mechanism of LSTM and trajectory correction is adopted.
[0095] The trajectory curvature describes the degree of bending of the navigation path and is used to quantify the severity of the vessel's turning or path change in the dynamic environment prediction, which directly affects the prediction accuracy of the environment data.
[0096] The present invention obtains the navigation trajectory data of the vessel, including navigation coordinate points, navigation speeds, and navigation time points; collects the historical audio-visual environment data of the area through which the navigation trajectory passes, and makes a prediction based on the historical audio-visual environment data, navigation speed, and navigation time point to obtain the predicted environment data when the vessel passes through the navigation coordinate point, thereby accurately collecting the environment data during the vessel's navigation.
[0097] The audio-visual playback recommendation module obtains the storable playable audio-visual data and identifies the audio-visual features of the playable audio-visual data. The audio-visual features include audio-visual brightness features, audio-visual sound features, and audio-visual shaking features. Identify the playback environment data based on the playback duration of the playable audio-visual data and the predicted audio-visual environment data. Identify the playback environment features from the playback environment data. Identify the audio-visual recommendation score based on the audio-visual features and the playback environment features.
[0098] Identify the playable audio-visual data and the playback environment data through the audio-visual fitness recognition model. The audio-visual fitness recognition model is constructed based on a deep learning network and includes: a data feature recognition layer, a first fitness recognition layer, a second fitness recognition layer, a third fitness recognition layer, and a recommendation score calculation layer.
[0099] The data feature recognition layer identifies the audio-visual features from the playable audio-visual data, including audio-visual brightness features, audio-visual sound features, and audio-visual shaking features. Identify the playback environment features from the playback environment data, including environment brightness features, environment sound features, and environment shaking features.
[0100] The first fit recognition layer performs recognition based on the audio-visual brightness feature and the environmental brightness feature, and obtains the expected environmental brightness by recognizing the audio-visual brightness feature; the brightness fit score is obtained according to the similarity between the expected environmental brightness and the environmental brightness feature.
[0101] The second fit recognition layer performs recognition based on the audio-visual sound feature and the environmental sound feature, and obtains the expected environmental sound by recognizing the audio-visual sound feature; the sound fit score is obtained according to the similarity between the expected environmental sound and the environmental sound feature.
[0102] The third fit recognition layer performs recognition based on the audio-visual shaking feature and the environmental shaking feature, and obtains the expected shaking data by recognizing the audio-visual shaking feature; the stability fit score is obtained according to the similarity between the expected shaking data and the environmental shaking feature.
[0103] The recommended score calculation layer calculates the audio-visual recommended score according to the brightness fit score, the sound fit score and the stability fit score.
[0104] The formula for calculating the audio-visual recommended score according to the brightness fit score, the sound fit score and the stability fit score is:
[0105] ;
[0106] ;
[0107] Among them, represents the audio-visual recommended score; represents the brightness fit score; represents the sound fit score; represents the stability fit score; represents the first fit weight; represents the second fit weight; represents the third fit weight.
[0108] Identify the playable audio-visual data to obtain Data Table 1.
[0109] Table 1 Data Table for Calculating Audio-Visual Fit Score
[0110] Playable audio-visual data Brightness fit score Sound fit score Stability fit score 1 88 76 81 2 69 72 90 3 83 92 74 4 92 87 86 5 79 83 81
[0111] Based on a deep learning network, the present invention constructs a video-audio matching recognition model; recognizes playable video-audio data and playback environment data to obtain video-audio features and playback environment features; recognizes a brightness matching score based on the video-audio brightness feature and the environment brightness feature, recognizes a sound matching score based on the video-audio sound feature and the environment sound feature, and recognizes a stability matching score based on the video-audio shaking feature and the environment shaking feature; calculates a video-audio recommendation score through the brightness matching score, the sound matching score, and the stability matching score; thereby obtaining the matching degree between the video-audio data and the water surface playback environment data.
[0112] A video-audio playback optimization module recommends according to the video-audio recommendation score and determines the video-audio data to be played; obtains the video-audio features and playback environment data of the video-audio data to be played, and optimizes the video-audio playback according to the playback environment data and the video-audio features.
[0113] The video-audio playback optimization process includes dividing the playback duration of the video-audio data to be played to obtain an optimization period; the process of obtaining the optimization period includes: obtaining the video-audio features and playback environment data of the video-audio data to be played; recognizing the video-audio features through a clustering algorithm, and performing clustering analysis according to the video-audio brightness feature, the video-audio sound feature, and the video-audio shaking feature to obtain first-period division data; recognizing the playback environment data through a clustering algorithm, and performing clustering analysis according to the predicted brightness data, the predicted sound data, and the predicted shaking data to obtain second-period division data; dividing the playback duration according to the first-period division data and the second-period division data to obtain an optimization period.
[0114] The present invention obtains the video-audio features and playback environment data of the video-audio data to be played; recognizes the video-audio features through a clustering algorithm to obtain first-period division data; recognizes the playback environment data through a clustering algorithm to obtain second-period division data; divides the playback duration according to the first-period division data and the second-period division data to obtain an optimization period; thereby ensuring that within each optimization period, the video-audio playback data and the playback environment data tend to be stable.
[0115] During the optimization period, playback optimization is performed according to the playback environment data and the video-audio features, including brightness optimization, sound optimization, and stability optimization;
[0116] The brightness optimization includes recognizing brightness optimization data according to the predicted brightness data in the playback environment data and the video-audio brightness feature, and optimizing the brightness of the video-audio according to the brightness optimization data;
[0117] The sound optimization includes identifying sound optimization data based on the predicted sound data and video and audio sound characteristics in the playback environment data, and optimizing the sound of the video and audio according to the sound optimization data. The sound optimization also includes building an air film audio system driven by AI. The air film audio system can sense environmental noise in real time through intelligent algorithms and adaptively adjust the noise reduction level, and is applicable to large-space application scenarios. Its working principle is as follows: based on the spatial learning of the air film, dynamically adjust to achieve the combination of active noise reduction and passive noise reduction. At the same time, based on the AI algorithm to learn the user's hearing curve and preferences, provide personalized noise reduction solutions, and can switch between noise reduction and ambient sound to provide extreme noise reduction effects.
[0118] The stability optimization includes identifying stability optimization data based on the predicted shaking data and video and audio shaking characteristics in the playback environment data, and optimizing the picture of the video and audio according to the stability optimization data.
[0119] The present invention identifies brightness optimization data according to the predicted brightness data and video and audio brightness characteristics of the playback environment data, identifies sound optimization data according to the predicted sound data and video and audio sound characteristics of the playback environment data, and identifies stability optimization data according to the predicted shaking data and video and audio shaking characteristics in the playback environment data. The video and audio are optimized according to the brightness optimization data, sound optimization data and stability optimization data, effectively improving the playback effect of the video and audio in the water surface environment.
[0120] The acquisition of the stability optimization data includes: identifying the predicted shaking data to obtain the predicted shaking data characteristics. The predicted shaking data includes the roll parameter, pitch parameter and heave parameter of the ship. Using the optical flow method to detect the motion vector in the picture of the video and audio data to be played, identify the picture shaking and change frequency to obtain the video and audio shaking characteristics. Obtain the predicted shaking data characteristics and video and audio shaking characteristics corresponding to the same optimization period, superimpose the shaking motion of the ship and the shaking motion of the video and audio picture, and calculate the motion direction and amplitude to be compensated to obtain the stability optimization data.
[0121] The present invention identifies the predicted shaking data to obtain the predicted shaking data characteristics, detects the motion vector in the picture of the video and audio data to be played, identifies the picture shaking and change frequency to obtain the video and audio shaking characteristics, and obtains the stability optimization data according to the predicted shaking data characteristics and video and audio shaking characteristics within the same optimization period. Therefore, the video and audio picture can be accurately compensated and optimized according to the ship shaking data and video and audio picture shaking, improving the picture playback effect of the video and audio in the complex water surface environment.
[0122] In order to verify the effect of the audio - video optimization of the present invention, a comparative optimization scheme is adopted for testing; the comparative optimization scheme is adjusted based on static environmental parameters, and the similarity coefficient between the actually played audio - video data and the standard audio - video data collected at a certain distance is used as a measure of the optimization effect, and the data obtained is shown in Table 2.
[0123] Table 2 Verification Data Table of Audio - Video Optimization Scheme
[0124] Comparison optimization plan Optimization plan of this application Improvement effect Audio-visual brightness 0.78 0.85 8.97% Audio-visual sound 0.82 0.90 9.75% Audio-visual picture 0.77 0.89 15.58%
[0125] It can be seen from Table 2 that the brightness optimization, sound optimization and stability optimization of the optimization scheme of the present application are all superior to the comparative optimization scheme.
[0126] The present invention collects environmental data to obtain historical audio - video environmental data; makes predictions based on the historical audio - video environmental data to obtain predicted audio - video environmental data; obtains the storable playable audio - video data and identifies the audio - video characteristics of the playable audio - video data; identifies the playback environmental data according to the playback duration of the playable audio - video data and the predicted audio - video environmental data; identifies the playback environmental characteristics from the playback environmental data; identifies the audio - video recommendation score according to the audio - video characteristics and the playback environmental characteristics; makes recommendations according to the audio - video recommendation score and determines the audio - video data to be played; obtains the audio - video characteristics and the playback environmental data of the audio - video data to be played, and optimizes the audio - video playback according to the playback environmental data and the audio - video characteristics, so as to realize the intelligent recommendation and optimization of audio - video in a complex water surface environment.
[0127] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An integrated intelligent audio and video playback system, characterized in that, Including: A video and audio environment recognition module, which collects environmental data according to the navigation track data to obtain historical video and audio environment data; predicts according to the historical video and audio environment data to obtain predicted video and audio environment data; the predicted video and audio environment data includes predicted brightness data, predicted sound data, and predicted shaking data; A video and audio playback recommendation module, which obtains the storable video and audio data and identifies the video and audio characteristics of the storable video and audio data; the video and audio characteristics include video and audio brightness characteristics, video and audio sound characteristics, and video and audio shaking characteristics; identifies the playback environment data according to the playback duration of the storable video and audio data and the predicted video and audio environment data; identifies the playback environment characteristics from the playback environment data; identifies the video and audio recommendation score according to the video and audio characteristics and the playback environment characteristics; A video and audio playback optimization module, which makes recommendations according to the video and audio recommendation score and determines the video and audio data to be played; obtains the video and audio characteristics and the playback environment data of the video and audio data to be played, and optimizes the video and audio playback according to the playback environment data and the video and audio characteristics; During the optimization period, optimize the playback according to the playback environment data and the video and audio characteristics, including brightness optimization, sound optimization, and stability optimization; the stability optimization includes identifying the stability optimization data according to the predicted shaking data and the video and audio shaking characteristics in the playback environment data, and optimizing the video and audio picture according to the stability optimization data; The acquisition of the stability optimization data includes: Identifying the predicted shaking data to obtain the predicted shaking data characteristics; the predicted shaking data includes the roll parameter, pitch parameter, and heave parameter of the ship; Detecting the motion vectors in the picture of the video and audio data to be played, identifying the picture shaking and change frequency, and obtaining the video and audio shaking characteristics; Obtaining the predicted shaking data characteristics and the video and audio shaking characteristics corresponding to the same optimization period, superimposing the shaking motion of the ship and the shaking motion of the video and audio picture, and calculating the motion direction and amplitude to be compensated to obtain the stability optimization data.
2. An integrated intelligent video and audio playback system according to claim 1, characterized in that: The process of identifying the predicted video and audio environment data from the historical video and audio environment data is: Obtaining the navigation track data of the ship on the water surface, the navigation track data including navigation coordinate points, navigation speed, and navigation time points; collecting the historical environment data of the area passed by the navigation track data as the historical video and audio environment data; Predicting according to the historical video and audio environment data, navigation speed, and navigation time points to obtain the predicted environment data when the ship passes through the navigation coordinate point as the predicted video and audio environment data.
3. An integrated intelligent video and audio playback system according to claim 1, characterized in that: Identifying the storable video and audio data and the playback environment data through a video and audio matching recognition model; the video and audio matching recognition model includes: a data feature recognition layer, a first matching recognition layer, a second matching recognition layer, a third matching recognition layer, and a recommendation score calculation layer; The data feature recognition layer recognizes video and audio features based on playable video and audio data, including video and audio brightness features, video and audio sound features, and video and audio shaking features; it recognizes playback environment features based on playback environment data, including environment brightness features, environment sound features, and environment shaking features. The first fit recognition layer recognizes based on the video and audio brightness feature and the environment brightness feature to obtain a brightness fit score. The second fit recognition layer recognizes based on the video and audio sound feature and the environment sound feature to obtain a sound fit score. The third fit recognition layer recognizes based on the video and audio shaking feature and the environment shaking feature to obtain a stability fit score. The recommended score calculation layer calculates a video and audio recommendation score based on the brightness fit score, the sound fit score, and the stability fit score.
4. The integrated intelligent video and audio playback system according to claim 3, wherein: The formula for calculating the video and audio recommendation score based on the brightness fit score, the sound fit score, and the stability fit score is: ; ; Among them, represents the audio-visual recommendation score; represents the brightness matching score; represents the sound matching score; represents the stability matching score; represents the first matching weight; represents the second matching weight; represents the third matching weight.
5. The integrated intelligent video and audio playback system according to claim 1, wherein: The video and audio playback optimization process includes dividing the playback duration of the video and audio data to be played to obtain an optimization period; the process of obtaining the optimization period includes: Obtaining the video and audio features and playback environment data of the video and audio data to be played. Identifying the video and audio features through a clustering algorithm, and performing cluster analysis based on the video and audio brightness feature, the video and audio sound feature, and the video and audio shaking feature to obtain first-period division data. Identifying the playback environment data through a clustering algorithm, and performing cluster analysis based on the predicted brightness data, the predicted sound data, and the predicted shaking data to obtain second-period division data. Dividing the playback duration according to the first-period division data and the second-period division data to obtain an optimization period.
Citation Information
Patent Citations
Method and device for reducing video frame jitter
CN106101527A
Systems and methods for providing ar / VR content based on vehicle conditions
US20200120371A1