Cloud big data-based audio and video matching method and system, and storage medium

By using cloud-based big data audio and video matching methods and leveraging deep neural networks and incremental update strategies, we have achieved real-time playback and personalized content recommendation of holographic naked-eye 3D videos on smart speakers. This solves the problems of insufficient matching and latency in existing technologies and improves the user experience.

CN120835166APending Publication Date: 2025-10-24深圳心沃科技有限公司
View PDF -1 Cites 0 Cited by

Patent Information

Application Number
CN202510649679.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing smart speakers have shortcomings in intelligent matching and real-time playback of holographic naked-eye 3D videos. They struggle to match and play corresponding video content in real time while playing music, and personalized video content recommendations are limited.

Method used

It adopts a cloud-based big data-driven audio and video matching method, constructs an audio and video semantic matching model through a deep neural network, generates a video playback order library by combining user terminal music data and cloud models, and achieves millisecond-level node localization and light field re-rendering through Bloom filters and incremental update strategies, supporting real-time switching of multimodal interaction via voice commands.

Benefits of technology

It achieves real-time intelligent matching between music playback and holographic naked-eye 3D video, eliminates the problem of mismatch between traditional SD card pre-stored videos and music rhythm, provides personalized video content recommendations, reduces video stream transmission latency, and supports naked-eye 3D display on different terminal devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120835166A_ABST
    Figure CN120835166A_ABST
Patent Text Reader

Abstract

The invention provides an audio and video matching method and system based on cloud big data and a storage medium, which are applied to the technical field of audio and video interaction, realize real-time intelligent matching of music playing and holographic naked-eye 3D videos, eliminate the problem of dislocation of videos pre-stored in a traditional SD card and music rhythms, and improve the user experience. Music components are dynamically analyzed through a cloud big data model, adaptive 3D light field contents are synchronously generated, and a personalized visual theme library is automatically recommended in combination with user historical behavior data; a voice instruction is supported to switch a multi-mode interaction scene in real time, and a dynamic parallax compensation technology is adopted to be compatible with naked eye 3D display parameters of different terminal devices; video stream transmission delay is reduced to a millisecond level through an edge computing node, and a static album cover is converted into a dynamic three-dimensional scene in real time by means of an AI generation engine.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio-video interaction, in particular to a cloud big data-based audio-video matching method, system and storage medium. BACKGROUND

[0002] As a multifunctional device, the smart speaker can be connected with smart devices such as mobile phones and tablet computers through Bluetooth, play songs and support voice interaction function. However, in the prior art, the smart speaker has deficiencies in multimedia interaction, especially in real-time playing and intelligent matching of holographic naked-eye 3D video, and the user experience needs to be improved.

[0003] The existing smart speaker mainly realizes audio playing and voice interaction function through a Bluetooth module, which can realize basic music playing and voice command recognition, but there is room for improvement in the following aspects:

[0004] Multimedia interaction: the existing technology has deficiencies in video playing, especially in intelligent matching and real-time playing of holographic naked-eye 3D video, which usually places song information or corresponding video of the song in an SD card to be played by naked-eye 3D video, which can be synchronized with the music rhythm;

[0005] Real-time: the existing technology is difficult to match and play corresponding video content in real time while playing music;

[0006] Personalized experience: the existing technology has limitations in providing personalized video content according to user preferences and scene requirements. SUMMARY

[0007] The present application aims to solve the technical problems of the existing technology in video playing, especially in intelligent matching and real-time playing of holographic naked-eye 3D video, which usually places song information or corresponding video of the song in an SD card to be played by naked-eye 3D video, which can be synchronized with the music rhythm; the existing technology is difficult to match and play corresponding video content in real time while playing music; the existing technology has limitations in providing personalized video content according to user preferences and scene requirements, and provides a cloud big data-based audio-video matching method, system and storage medium.

[0008] To solve the technical problems, the present application adopts the following technical means:

[0009] A cloud big data-based audio-video matching method, system and storage medium, the method comprising:

[0010] Obtaining music data of a user terminal to form a first music playing sequence library;

[0011] According to the music data in the first playback order library, the components of each piece of music data are analyzed;

[0012] According to the music data components combined with the cloud model interaction, corresponding video is generated;

[0013] According to the first music playback order library, a first video playback order library is formed;

[0014] If the user changes the music playback order library, the first video playback order library is changed synchronously.

[0015] Further, in the step of obtaining music data of the user terminal and forming the first music playback order library,

[0016] According to the user terminal input music arrangement combination of the player, the first music playback order library is formed, and the first music playback order library changes with the user terminal.

[0017] Further, in the step of analyzing the components of each piece of music data according to the music data in the first music playback order library,

[0018] According to the music in the first music playback order library, the music is analyzed according to different elements, including but not limited to song name, singer, lyrics and style.

[0019] Further, in the step of generating video according to the music data components combined with the cloud model interaction,

[0020] Through the deep neural network, an audio-video semantic matching model is constructed, and the music label, lyrics semantic vector and mel spectrum feature are cross-modally aligned.

[0021] Further, in the step of generating a plurality of videos according to the first music playback order library to form a first video playback order library,

[0022] According to the first music playback order library, the first video playback order library is formed, and the corresponding video of the unrecorded song is supplemented into the video storage library in the cloud.

[0023] Further, in the step of changing the first video playback order library synchronously if the user changes the music playback order library,

[0024] Based on the Bloom filter, the music-video hash index is quickly matched to realize millisecond-level node positioning, through the incremental update strategy, only the 3D light field local re-rendering of the changed segment is triggered, and layered encoding transmission is adopted to optimize bandwidth occupation.

[0025] Further, in the step of triggering the 3D light field local re-rendering of the changed segment only through the incremental update strategy and adopting layered encoding transmission to optimize bandwidth occupation,

[0026] The deployment of the light field parameter snapshot storage mechanism supports rolling back to the historical best viewing state through the gyroscope attitude, and synchronously updating the dynamic parallax compensation matrix to adapt to the adjusted music rhythm intensity.

[0027] A cloud big data audio-video matching system based on,

[0028] A music acquisition module acquires music data of a user terminal to form a first music play order library.

[0029] A music analysis module analyzes the components of each piece of music data according to the music data in the first play order library.

[0030] A video generation module generates videos according to the music data components and cloud model interaction.

[0031] A video play order module generates a first video play order library according to the first music play order library.

[0032] An audio-video change module synchronously changes the first video play order library if the user changes the music play order library.

[0033] A computer readable storage medium stores programs or instructions, which make a computer execute the steps of the cloud big data audio-video matching method according to any one of claims 1 to 8.

[0034] The application provides a cloud big data audio-video matching method, system and storage medium, which has the following beneficial effects:

[0035] The application realizes real-time intelligent matching of music playing and holographic naked-eye 3D video, eliminates the misalignment problem of traditional SD card pre-stored video and music rhythm, dynamically analyzes music components through a cloud big data model, and synchronously generates adaptive 3D light field content, and automatically recommends a personalized visual theme library combined with historical behavior data of a user.

[0036] The application supports real-time switching of multi-modal interactive scenes through voice instructions, and uses dynamic parallax compensation technology to be compatible with naked-eye 3D display parameters of different terminal devices.

[0037] The application reduces video stream transmission delay to the millisecond level through an edge computing node, and converts a static album cover into a dynamic three-dimensional scene in real time through an AI generation engine. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1 The application provides a flowchart of an embodiment of a cloud big data audio-video matching method, system and storage medium.

[0039] The objectives, functional features and advantages of the present application will be further described with reference to the embodiments, with reference to the accompanying drawings. DETAILED DESCRIPTION

[0040] It should be understood that the specific embodiments described herein merely exemplify the present application and do not limit the present application.

[0041] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0042] It should be noted that the terms "comprise", "contain", and "have" and any variations thereof in the specification and claims of the present application and the above-mentioned drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but optionally further comprises steps or units not listed, or optionally further comprises other steps or units inherent to the process, method, product or device. In the claims, specification and drawings of the present application, the terms such as "first" and "second" and the like relationship terms are only used to distinguish one entity / operation / object from another entity / operation / object, and do not necessarily require or imply any actual relationship or order between the entities / operations / objects.

[0043] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment can be included in at least one embodiment of the present application. The appearance of the phrase in various places in the specification does not necessarily all refer to the same embodiment, nor does it necessarily refer to a particular embodiment in an exclusive sense. It is explicitly and implicitly understood that the embodiments described herein can be combined.

[0044] Reference is made to the accompanying drawings that form a part of this disclosure Figure 1 The flowchart of the cloud-based big data audio-video matching method, system and storage medium in an embodiment of the present application is shown in the figure;

[0045] Embodiment one

[0046] A cloud-based big data audio-video matching method, system and storage medium, the method comprising:

[0047] Obtaining music data of a user terminal to form a first music playing order library;

[0048] According to the music data in the first playback order library, the components of each music data are parsed;

[0049] According to the music data component combined with the cloud model interaction corresponding video generation;

[0050] Generate several videos according to the first music playback order library to form a first video playback order library;

[0051] If the user changes the music playback order library, the first video playback order library is changed synchronously.

[0052] In this embodiment, in the step of obtaining music data of the user terminal and forming a first music playback order library,

[0053] According to the user terminal input music permutation and combination of the player, the first music playback order library is formed, and the first music playback order library changes with the user terminal.

[0054] Specifically,

[0055] First, the user uses a mobile phone, tablet or computer to connect the Bluetooth in the terminal to the sound box speaker, and then the user plays the music APP in the terminal, and the AI sound box transmits the audio electrical signal to the sound box speaker through Bluetooth decoding. In the process of Bluetooth music transmission, or automatically obtain various information in the music, among them, the next X songs of the current playing music in the user's music APP are obtained to form a first music playback order library.

[0056] In this embodiment, in the step of parsing the components of each music data according to the music data in the first playback order library,

[0057] According to the music in the first music playback order library, different elements are analyzed, including but not limited to song name, singer, lyrics and style.

[0058] Specifically,

[0059] Based on the first music playback order library formed above, the AI sound box transmits the first music playback order library to the cloud, and the cloud analyzes the X music songs in the first music playback order library.

[0060] In the step of generating video according to the music data component combined with the cloud model interaction corresponding,

[0061] Through the deep neural network, the audio-video semantic matching model is constructed, and the music label, lyrics semantic vector and mel spectrum feature are aligned across modalities.

[0062] The cloud analyzes the X music songs in the first music playback order library to generate video, wherein,

[0063] The first music playing order library received through the cloud interacts with AI intelligence models including but not limited to GPT, deepseek, etc.

[0064] For example,

[0065] The user plays music in the terminal APP for singer A-song A;

[0066] Input data:

[0067] Song metadata: BPM = 72 (Song A), singer style label (popular Chinese style)

[0068] Lyrics text: "The little yellow flower has been floating since the year of birth";

[0069] Spectrum features: high frequency band (8-12kHz) energy ratio 35% (prelude wind chime sound)

[0070] Processing flow:

[0071] BERT model extracts lyrics sentiment vector (dimension 512): [nostalgia: 0.78, fresh: 0.65]

[0072] CNN network analyzes mel spectrum to generate 128-dimensional rhythm feature vector;

[0073] Cross-modal alignment layer fuses the above features to generate joint semantic vector;

[0074] AI intelligence generates scene description:

[0075] Summer beach scene, golden particles form little yellow flowers floating in the wind;

[0076] Sea wave fluctuation frequency synchronizes with BPM, raindrop particles burst when sinking chorus;

[0077] Stable Diffusion 3D generation:

[0078] Input: joint semantic vector + GPT scene description

[0079] Output:

[0080] Basic scene: low-polygon beach model (12 million faces)

[0081] Dynamic elements:

[0082] Little yellow flower particle system (100,000 particles, drift speed 0.8 m / s)

[0083] Sea wave vertex animation (amplitude 0.3 m, frequency 1.2 Hz)

[0084] Light field parameter dynamic adjustment:

[0085] Style matching: Chinese style -> grating density adjusted to 120 ppi (cylindrical lens spacing 0.21 mm)

[0086] Rhythm mapping: BPM 72 -> particle refresh rate 86 fps

[0087] Emotional feedback: nostalgia index > 0.7 -> depth of field range locked at 58 mm (anti-dizziness mode)

[0088] In this embodiment, in the step of generating a plurality of videos to form a first video playback order library according to a first music playback order library,

[0089] The first video playback order library formed according to the first music playback order library supplements the corresponding video of the unrecorded song into the video storage library in the cloud.

[0090] Specifically,

[0091] When it is detected that a certain song is missing in the video library, automatically match the existing template with a style similarity of > 85%; three-level cache architecture

[0092] Storage level response time capacity

[0093] L1 (edge node) <50ms 200TB high-frequency use template (weekly play count > 10,000 times)

[0094] L2 (regional center) <150ms 5PB medium-frequency content (weekly play count 1,000-10,000 times)

[0095] L3 (core cloud) <500ms 50PB long-tail content and original material

[0096] Real-time generation acceleration engine

[0097] Mixed precision calculation (FP16+INT8) is adopted:

[0098] Particle system generation: NVIDIA OptiX 8.0 real-time ray tracing, single-frame rendering speed increased by 3.2 times

[0099] Light field compression: self-developed SparseVQ encoder, 8K light field data compression rate increased to 1:24

[0100] Distributed task scheduling: based on user GPS location to dynamically allocate the nearest rendering node (time delay <80ms)

[0101] Among them, the weekly play count of L1 is greater than 10,000 times, and the weekly play count of L2 is 1,000-10,000 times.

[0102] In the step of synchronously changing the first video playing order library if the user changes the music playing order library,

[0103] Based on the Bloom filter fast matching music-video hash index to realize millisecond-level node positioning, through the incremental update strategy, only trigger the 3D light field local re-rendering of the changed segment, and adopt layered encoding transmission to optimize bandwidth occupation.

[0104] In the step of synchronously changing the first video playing order library if the user changes the music playing order library,

[0105] Deploying a light field parameter snapshot storage mechanism supports rolling back to the historical best viewing state through a gyroscope attitude, and synchronously updating a dynamic parallax compensation matrix to adapt to the adjusted music rhythm intensity.

[0106] A cloud big data audio-video matching system based on,

[0107] A music acquisition module acquires music data of a user terminal to form a first music playing order library;

[0108] A music analysis module analyzes the components of each piece of music data according to a plurality of pieces of music data in the first playing order library;

[0109] A video generation module generates videos according to the music data components and cloud model interaction;

[0110] A video playing order module generates a plurality of videos to form a first video playing order library according to the first music playing order library;

[0111] An audio-video change module synchronously changes the first video playing order library if the user changes the music playing order library.

[0112] A computer-readable storage medium stores programs or instructions, which make a computer execute the steps of the cloud big data audio-video matching method according to any one of claims 1 to 8.

[0113] Embodiment two

[0114] A cloud big data audio-video matching method, system and storage medium, the method comprising:

[0115] Acquiring music data of a user terminal to form a first music playing order library;

[0116] Analyzing the components of each piece of music data according to a plurality of pieces of music data in the first playing order library;

[0117] Generating videos according to the music data components and cloud model interaction;

[0118] generate a first video play sequence library according to the first music play sequence library;

[0119] If the user changes the music play sequence library, the first video play sequence library is changed synchronously.

[0120] In the step of obtaining the music data of the user terminal and forming the first music play sequence library,

[0121] According to the user terminal input music arrangement combination of the player, the first music play sequence library is formed, and the first music play sequence library changes with the user terminal.

[0122] In the step of analyzing the components of each music data according to the music data in the first music play sequence library,

[0123] According to the music in the first music play sequence library, the music is analyzed according to different elements, including but not limited to song name, singer, lyrics and style.

[0124] In the step of generating a video corresponding to the music data components combined with the cloud model interaction,

[0125] The audio-video semantic matching model is constructed by a deep neural network, and the music label, lyrics semantic vector and mel spectrum feature are cross-modal aligned.

[0126] In the step of generating a first video play sequence library according to the first music play sequence library,

[0127] According to the first music play sequence library, the first video play sequence library is formed, and the corresponding video of the song not recorded is supplemented into the video storage library in the cloud.

[0128] In the step of changing the first video play sequence library synchronously if the user changes the music play sequence library,

[0129] Based on the bloom filter, the music-video hash index is quickly matched to realize millisecond-level node positioning, through the incremental update strategy, only the 3D light field local re-rendering of the changed segment is triggered, and layered encoding transmission is adopted to optimize bandwidth occupation.

[0130] In the step of triggering the 3D light field local re-rendering of the changed segment only through the incremental update strategy and adopting layered encoding transmission to optimize bandwidth occupation,

[0131] The light field parameter snapshot storage mechanism is deployed to support rolling back to the historical best viewing state through the gyroscope attitude, and the dynamic parallax compensation matrix is updated synchronously to adapt to the adjusted music rhythm intensity.

[0132] Specifically,

[0133] The user transmits a sound wave signal to the AI sound box through the microphone, the AI sound box converts it into an audio electrical signal and transmits it to the AI intelligence, and follows the user's music switching, fast forwarding, and changes in the play mode, and the first video play sequence library will also be changed accordingly, specifically:

[0134] When the user switches the music playlist through the mobile phone APP;

[0135] Such as switching the original playlist from singer A's "Track A" and "Track B" to singer B's "Track C" and "Track D";

[0136] The generation and synchronization process of the first video sequence library is specifically as follows:

[0137] The system immediately captures the metadata of the new playlist (such as the 128BPM rhythm of "Track C", the Chinese instrument sampling characteristics);

[0138] Through the cloud cross-modal model, a corresponding 3D video template is generated within 80ms (such as the particle rain thread special effect of the water village scene of Track C synchronized with the guzheng audio track waveform)

[0139] Trigger the incremental update engine to only re-render the light field data related to the changed track (such as the raindrop particle system of "Track C" in the first 10 seconds of the prelude), while retaining the pre-rendered cache of the unchanged song (the beach scene of "Track A" reuses the original template);

[0140] Through hierarchical coding technology, the new video stream is compressed to 35% of the original bandwidth, specifically, the basic layer transmits the static white wall and tile model, the enhanced layer dynamically updates the rain thread density parameter, and the copyright verification module is activated to scan the authorization status of the new song;

[0141] Among them, if "Track D" has no usage rights, it will be automatically replaced by the system default starry sky animation;

[0142] Optimize the recommendation weight combined with user historical behavior data;

[0143] For example: the user has adjusted the depth of field range of the chorus part of "Track C" many times, this time the preset depth of field fluctuation amplitude is reduced to ±3mm, and the edge node is preloaded with the 3D song lyrics floating special effect elements of the next "Track D", finally the millisecond-level switching of the entire video sequence library is completed without user awareness, and the version change log is generated and stored in the blockchain node to ensure that the operation is traceable;

[0144] In order to achieve the purpose of real-time changes in the first video sequence library and gradual updating in naked-eye 3D when the user switches songs.

[0145] A cloud-based big data audio and video matching system,

[0146] a music acquisition module, which acquires music data of a user terminal to form a first music play order library;

[0147] a music analysis module, which analyzes components of each piece of music data according to the music data in the first play order library;

[0148] a video generation module, which generates videos according to the music data components and cloud model interaction;

[0149] a video play order module, which generates a first video play order library according to the first music play order library;

[0150] a video-audio change module, which synchronously changes the first video play order library if the user changes the music play order library.

[0151] A computer-readable storage medium storing programs or instructions, which cause a computer to execute the steps of the cloud big data audio-video matching method according to any one of claims 1 to 8.

[0152] Embodiment Three

[0153] A cloud big data audio-video matching method, system and storage medium, the method comprising:

[0154] acquiring music data of a user terminal to form a first music play order library;

[0155] analyzing components of each piece of music data according to the music data in the first play order library;

[0156] generating videos according to the music data components and cloud model interaction;

[0157] generating a first video play order library according to the first music play order library;

[0158] synchronously changing the first video play order library if the user changes the music play order library.

[0159] In the step of acquiring music data of a user terminal to form a first music play order library,

[0160] forming the first music play order library according to the user terminal input music arrangement combination of the player, and the first music play order library changes with the user terminal.

[0161] In the step of analyzing components of each piece of music data according to the music data in the first play order library,

[0162] According to the music in the first music playing order library, the music is analyzed according to different elements, including but not limited to song name, singer, lyrics and style.

[0163] In the step of generating a video according to the interaction of the music data components and the cloud model,

[0164] Through the deep neural network, an audio-video semantic matching model is constructed to align the music label, lyrics semantic vector and mel-frequency spectrum feature in a cross-modal manner.

[0165] In the step of generating a video according to the interaction of the music data components and the cloud model,

[0166] According to the first music playing order library, the first video playing order library is formed, and the corresponding video of the unrecorded song is supplemented into the cloud video storage library.

[0167] In the step of changing the first video playing order library if the user changes the music playing order library,

[0168] Based on the Bloom filter, the music-video hash index is quickly matched to realize millisecond-level node positioning, and through the incremental update strategy, only the 3D light field local re-rendering of the changed segment is triggered, and layered encoding transmission is adopted to optimize bandwidth occupation.

[0169] In the step of changing the first video playing order library if the user changes the music playing order library,

[0170] The light field parameter snapshot storage mechanism is deployed to support rolling back to the historical best viewing state through the gyroscope attitude, and the dynamic parallax compensation matrix is updated synchronously to adapt to the adjusted music rhythm intensity.

[0171] A cloud-based big data audio-video matching system,

[0172] A music acquisition module acquires music data of a user terminal to form a first music playing order library;

[0173] A music analysis module analyzes the components of each piece of music data according to the music data in the first playing order library;

[0174] A video generation module generates a video according to the interaction of the music data components and the cloud model;

[0175] A video playing order module generates a first video playing order library according to the first music playing order library;

[0176] An audio-video change module synchronously changes the first video playing order library if the user changes the music playing order library.

[0177] A computer readable storage medium stores a program or instructions, which cause a computer to perform the steps of the cloud big data audio-video matching method according to any one of claims 1 to 8.

[0178] Especially important is,

[0179] Real-time intelligent matching of music playing and holographic naked-eye 3D video is realized, the problem of misalignment of traditional SD card pre-stored video and music rhythm is eliminated, and through a cloud big data model, music components are dynamically analyzed and 3D light field content adapted thereto is synchronously generated, and personalized visual theme libraries are automatically recommended in combination with historical behavior data of a user;

[0180] Real-time switching of multi-modal interactive scenes is supported through voice instructions, and dynamic parallax compensation technology is adopted to be compatible with naked-eye 3D display parameters of different terminal devices;

[0181] Video stream transmission delay is reduced to the order of milliseconds through an edge computing node, and a static album cover is converted into a dynamic three-dimensional scene in real time by relying on an AI generation engine.

[0182] Those skilled in the art will understand that embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0183] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in a flow or multiple flows and / or blocks Figure 1 The functions specified in a flow or multiple flows and / or blocks

[0184] These computer program instructions can also be stored in a computer-readable memory that can direct the computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including instruction devices that implement the functions specified in the flowcharts and / or block diagrams. Figure 1one or more processes and / or blocks Figure 1 the function specified in the one or more blocks.

[0185] These computer program instructions can also be loaded into computer or other programmable data processing devices, so that a series of operation steps are performed on the computer or other programmable data processing devices to generate computer-implemented processes, so that the instructions executed on the computer or other programmable data processing devices provide processes for implementing the flow Figure 1 one or more processes and / or blocks Figure 1 the function specified in the one or more blocks.

[0186] Although embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, changes and alterations can be made by those skilled in the art without departing from the principles and spirit of the present application, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A cloud-based big data audio-video matching method, system and storage medium, characterized in that, The method comprises: Obtain music data of the user terminal to form a first music play sequence library; According to the music data in the first play sequence library, analyze the composition of each music data; According to the music data composition combined with the cloud model interaction corresponding video generation; Generate several videos according to the first music play sequence library to form a first video play sequence library; If the user changes the music play sequence library, the first video play sequence library is changed synchronously. 2.The cloud-based big data audio-video matching method, system and storage medium of claim 1, wherein, In the step of obtaining music data of the user terminal to form a first music play sequence library, According to the user terminal input music arrangement combination of the player, the first music play sequence library is formed, and the first music play sequence library changes with the user terminal. 3.The cloud big data based audio-video matching method, system and storage medium of claim 1, wherein, In the step of analyzing the composition of each music data according to the music data in the first play sequence library, According to the music in the first music play sequence library, different elements are analyzed, including but not limited to song name, singer, lyrics and style. 4.The cloud big data based audio-video matching method, system and storage medium of claim 1, wherein, In the step of generating video according to music data composition combined with cloud model interaction, Through deep neural network, audio and video semantic matching model is constructed, music label, lyrics semantic vector and mel spectrum feature are cross-modal aligned. 5.The cloud big data based audio-video matching method, system and storage medium of claim 1, wherein, In the step of generating several videos according to the first music play sequence library to form a first video play sequence library, According to the first video play sequence library formed by the first music play sequence library, the corresponding video of the song not recorded is supplemented into the video storage library in the cloud. 6.The cloud big data based audio-video matching method, system and storage medium of claim 1, wherein, In the step of changing the first video play sequence library synchronously if the user changes the music play sequence library, Based on the bloom filter, the music-video hash index is matched quickly to realize millisecond level node positioning, through the incremental update strategy, only the 3D light field local re rendering of the changed segment is triggered, and layered coding transmission is adopted to optimize bandwidth occupation. 7.The cloud-based big data audio-video matching method, system and storage medium of claim 6, wherein, In the step of triggering the 3D light field local re rendering of the changed segment only through the incremental update strategy and adopting layered coding transmission to optimize bandwidth occupation, Deploy light field parameter snapshot storage mechanism to support rolling back to the historical best viewing state through gyroscope attitude and synchronously updating dynamic parallax compensation matrix to adapt to the adjusted music rhythm intensity. 8.A cloud big data audio and video matching system, characterized in that A music acquisition module obtains music data of the user terminal to form a first music play sequence library; A music analysis module analyzes the composition of each music data according to the music data in the first play sequence library; A video generation module generates video according to music data composition combined with cloud model interaction; A video play sequence module generates several videos according to the first music play sequence library to form a first video play sequence library; An audio and video change module changes the first video play sequence library synchronously if the user changes the music play sequence library.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores programs or instructions, which make the computer execute the steps of the cloud big data audio and video matching method according to any one of claims 1 to 8.