Artificial intelligence-based audio and video content automatic generation and analysis platform

By building an AI-based platform for the automatic generation and analysis of audio and video content, the problems of functional fragmentation and insufficient intelligence in existing technologies have been solved. It has achieved integrated processing of the entire process from understanding user intent to content generation, multimodal analysis and efficient retrieval, thereby improving the automation level of audio and video content processing and user experience.

CN122120571APending Publication Date: 2026-05-29BEIJING HUARUI JIAYIN TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING HUARUI JIAYIN TECH CO LTD
Filing Date
2026-02-26
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing audio and video processing systems suffer from fragmented functions, insufficient intelligence, and difficulty in achieving integrated solutions for content generation, multimodal analysis, and efficient retrieval, resulting in a poor user experience.

Method used

Build an AI-based platform for automatic generation and analysis of audio and video content. The platform understands user intent through a request analysis module, calls up a material library or generation engine to generate content, performs multi-dimensional feature extraction and recognition, executes automated editing, and establishes a hybrid search index to support semantic retrieval.

Benefits of technology

It achieves integrated processing across the entire process, from understanding user intent to content generation, multimodal analysis, and efficient retrieval, significantly improving the automation level of audio and video content processing and the accuracy of retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122120571A_ABST
    Figure CN122120571A_ABST
Patent Text Reader

Abstract

The application discloses an audio and video content automatic generation and analysis platform based on artificial intelligence and belongs to the technical field of audio and video processing. The content generation request of a user is received and semantic analysis is carried out, the user intention type and target content description information are identified, new materials are synthesized according to the intention and a preset material library generation engine, multi-dimensional feature extraction is carried out on the obtained audio and video materials, a content recognition module is inputted to carry out scene classification, object detection and emotion analysis, and a recognition result is generated, automatic editing is carried out on the materials based on the recognition result and the user intention, target audio and video content is generated, a vector index is constructed based on multi-dimensional features, an inverted index is constructed based on structured metadata, and a hybrid retrieval index is formed to support subsequent semantic retrieval. The application realizes the whole-process integrated processing from user intention understanding, content generation, multi-modal analysis, intelligent editing to semantic retrieval, and significantly improves the automation level and retrieval accuracy of audio and video content processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention discloses an artificial intelligence-based platform for automatic generation and analysis of audio and video content, belonging to the field of audio and video processing technology. Background Technology

[0002] With the rapid development of the internet and multimedia technologies, the production and consumption of audio and video content have experienced explosive growth. Traditional manual editing and processing methods can no longer meet the demands for rapid generation and efficient management of massive amounts of content. In recent years, artificial intelligence technology has been increasingly applied in the audio and video field. Generative adversarial networks based on deep learning can automatically synthesize images, audio, and even video clips. Convolutional neural networks have made breakthroughs in tasks such as image classification and object detection. Recurrent neural networks and the Transformer architecture have performed well in speech recognition and natural language processing. These technologies have laid a solid foundation for the automated processing of audio and video content. However, most existing technologies focus on the development of single functional modules, such as independent video generation tools, image recognition systems, or speech analysis engines, lacking comprehensive solutions that integrate content generation, multimodal analysis, intelligent editing, and efficient retrieval.

[0003] In practical applications, existing audio and video processing systems generally suffer from functional fragmentation and insufficient intelligence. On the one hand, content generation tools often require users to provide a large number of parameters and materials, making it difficult to automatically complete creation based on complex intentions described in natural language. On the other hand, content analysis modules typically process audio, video, and text separately, ignoring the correlation between multimodal information, which limits the accuracy of scene recognition, object detection, and sentiment analysis. Furthermore, the storage and retrieval of audio and video content still rely mainly on manually labeled keywords, failing to support deep retrieval based on semantic similarity or example content, making it difficult for users to quickly locate the required segments from massive amounts of material. When users need to complete a series of tasks from content creation to analysis and retrieval, they often have to switch between multiple independent systems, resulting in a cumbersome and inefficient process.

[0004] To address the aforementioned issues, this invention proposes an AI-based platform for the automatic generation and analysis of audio and video content. It aims to construct an integrated solution encompassing the entire process from user intent understanding, content generation or acquisition, multimodal feature extraction and recognition, automated editing, to content storage and semantic retrieval. The platform automatically retrieves audio and video content by parsing the intent and descriptive information in user requests, and then calls upon a media library or generation engine. Through the comprehensive extraction and fusion of audio, visual, and textual features, it achieves scene classification, object detection, and sentiment analysis of the content. Based on the recognition results and user intent, it automatically performs editing operations such as cropping, splicing, and adding special effects. Finally, it constructs a hybrid retrieval index using the generated content and its multimodal features, supporting semantic-level retrieval based on text and examples. This invention effectively solves the problems of fragmented functions, low intelligence levels, and limited retrieval methods in existing technologies, significantly improving the automation level of audio and video content processing and the user experience. Summary of the Invention

[0005] To achieve the above objectives, this application provides the following technical solution:

[0006] An AI-based platform for automatic generation and analysis of audio and video content, characterized in that it includes:

[0007] The request analysis module, in response to receiving a content generation request from a user, parses the content generation request to extract user intent information and target content description information;

[0008] The material generation module, based on the user intent information and the target content description information, calls original audio and video materials associated with the target content description information from a preset audio and video material library, or calls a pre-trained audio and video generation engine based on the user intent information to generate synthetic audio and video materials that match the target content description information;

[0009] The feature extraction module inputs the original audio-visual material or the synthesized audio-visual material into the content analysis engine, and performs multi-dimensional feature extraction operations on the original audio-visual material or the synthesized audio-visual material through the content analysis engine. The multi-dimensional feature extraction operations include audio feature extraction, visual feature extraction and text feature extraction.

[0010] The content recognition module performs content recognition on the original audio and video materials or the synthesized audio and video materials based on the extracted multi-dimensional features, and generates content recognition results, which include scene category information, object attribute information and sentiment information.

[0011] The editing module performs automated editing processing on the original audio and video materials or synthesized audio and video materials based on the content recognition results, generating target audio and video content that conforms to the user intent information and target content description information;

[0012] The associated storage module associates and stores the target audio and video content and the corresponding content recognition results, and establishes a content retrieval index based on the multi-dimensional features.

[0013] Furthermore, the request analysis module includes:

[0014] Perform semantic parsing on the content generation request to identify keywords and sentence structures in the content generation request;

[0015] Based on the keywords and sentence structure, the user intent type is determined, which includes content creation intent, content modification intent, and content analysis intent.

[0016] When the user intent type is content creation intent, the theme information, style information, and duration information are extracted from the content generation request as the target content description information; when the user intent type is content modification intent, the identification information of the content to be modified and the modification operation instruction are extracted from the content generation request; when the user intent type is content analysis intent, the identification information of the content to be analyzed and the analysis dimension information are extracted from the content generation request.

[0017] The user intent type and the corresponding extracted information are encapsulated to generate structured user intent information and target content description information.

[0018] Furthermore, the material generation module includes:

[0019] When the user intent type is content creation intent, a scene template matching the theme information is retrieved from the preset scene template library based on the theme information in the target content description information.

[0020] Based on the style information in the target content description information, retrieve the set of visual elements and the set of audio elements that match the style information from the preset material style library;

[0021] The scene template and the set of visual elements are fused together to generate a background image sequence; the set of audio elements is combined according to a preset arrangement rule to generate a background audio track.

[0022] Based on the duration information in the target content description information, the background image sequence and background audio track are synchronously cropped or expanded to generate duration-matched synthetic audio and video materials.

[0023] The synthesized audio and video materials are packaged according to a preset encoding format to generate audio and video files with a specified resolution and bitrate.

[0024] Furthermore, the material generation module includes:

[0025] When the user intent type is content modification intent or content analysis intent, extract the identification information of the content to be modified or the content to be analyzed from the content generation request;

[0026] Based on the identification information, access the metadata index table in the audio and video material library to query the storage location information corresponding to the identification information;

[0027] Based on the storage location information, the corresponding original audio and video material files are read from the distributed storage system;

[0028] The original audio and video material files are read and their format is checked to determine the compatibility between the encoding format of the original audio and video material files and the current processing environment.

[0029] When it is detected that the encoding format of the original audio and video material file is incompatible with the current processing environment, the format conversion tool is invoked to convert the original audio and video material file into a preset standard processing format.

[0030] The original audio and video material files after format conversion are loaded into a memory buffer, and a unique processing identifier is assigned to the original audio and video material files. The processing identifier is used to track the processing status of the original audio and video material files in subsequent processing steps.

[0031] Furthermore, the feature extraction module includes:

[0032] The original audio and video materials or the synthesized audio and video materials are subjected to audio and video stream separation processing to obtain audio data streams and video data streams respectively;

[0033] The audio data stream is segmented into frames, dividing the continuous audio signal into several audio frames. Time-domain feature extraction and frequency-domain feature extraction are performed on each audio frame. The time-domain features include short-time energy and zero-crossing rate, and the frequency-domain features include Mel frequency cepstral coefficients and spectral centroid.

[0034] The video data stream is subjected to frame extraction processing, and keyframe images are extracted from the video data stream at preset time intervals.

[0035] Image feature extraction is performed on each keyframe image, including edge feature extraction, texture feature extraction, and color histogram feature extraction.

[0036] The speech portion of the audio data stream is processed by speech recognition, the speech signal is converted into text data, and text feature extraction is performed on the text data, including word frequency feature extraction and syntactic structure feature extraction.

[0037] The extracted audio features, image features, and text features are correlated according to the timestamp alignment method to generate a multi-dimensional feature vector set.

[0038] Furthermore, the content recognition module includes:

[0039] Based on the multi-dimensional feature vector set, the audio portion of the original audio and video material or the synthesized audio and video material is classified into scenes to identify audio scene types, which include music scenes, dialogue scenes, ambient sound scenes and silent scenes.

[0040] Based on the image features in the multi-dimensional feature vector set, object detection is performed on the video portion of the original audio and video material or the synthesized audio and video material to identify the object categories and object locations appearing in the video frame. The object categories include human objects, object objects, and background objects.

[0041] Based on the object detection results, scene switching detection is performed on the video portion to identify scene boundaries in the video and generate scene segmentation information;

[0042] Based on the text features in the multi-dimensional feature vector set, sentiment analysis is performed on the speech content in the audio data stream to determine the sentiment tendency value of the speech content.

[0043] The content recognition result is generated by integrating the audio scene type, object detection result, scene segmentation information and sentiment value.

[0044] Furthermore, the editing module includes:

[0045] Based on the scene segmentation information in the content recognition results, determine the scene switching points of the original audio and video material or the synthesized audio and video material;

[0046] Based on the object detection results in the content recognition results, key objects in the original audio and video materials or synthesized audio and video materials are identified, and the motion trajectory of the key objects in the picture is determined;

[0047] Based on the modification operation instructions in the user intent information, determine the type of editing operation to be performed, which includes cropping, splicing, adding special effects, and adding subtitles.

[0048] When the editing operation type is a cropping operation, the cropping start point and cropping end point are set according to the scene switching point, and cropping processing is performed on the original audio and video material or the synthesized audio and video material.

[0049] When the editing operation type is a special effect addition operation, the special effect addition position is determined according to the object detection result, the special effect type is determined according to the emotional tendency value, and a specified type of visual or audio special effect is superimposed at the special effect addition position.

[0050] When the editing operation type is a subtitle addition operation, the speech content in the audio data stream is transcribed to generate subtitle text, the subtitle display timeline is determined according to the scene segmentation information, and the subtitle text is bound to the subtitle display timeline to generate a subtitle track file;

[0051] The edited audio and video data is combined with the subtitle track file and packaged to generate the target audio and video content.

[0052] Furthermore, the associated storage module includes:

[0053] Assign a unique content identifier to the target audio and video content, and write the target audio and video content into a distributed file system according to a preset storage structure;

[0054] The content recognition results are converted into structured metadata, which includes scene category tags, object tags, sentiment tags, and time sequence tags.

[0055] Establish the association between the content identifier and the structured metadata, and store the association in a relational database;

[0056] Based on the multi-dimensional feature vector set, a vector index structure is constructed, which is used to support content retrieval based on similarity calculation.

[0057] Based on the structured metadata, an inverted index structure is constructed, which is used to support content retrieval based on keyword matching;

[0058] The vector index structure and the inverted index structure are merged to generate a hybrid retrieval index, which supports composite retrieval conditions based on both vector similarity and keyword matching.

[0059] Furthermore, the platform also includes a content retrieval module, comprising:

[0060] Receive a content retrieval request input by the user, the content retrieval request including text query conditions or example audio / video files;

[0061] When the content retrieval request includes text query conditions, semantic parsing is performed on the text query conditions to extract query keywords;

[0062] Based on the query keywords, the inverted index structure is invoked to perform keyword matching and retrieval to obtain the first candidate content list;

[0063] The text query conditions are vectorized to generate query feature vectors;

[0064] Based on the query feature vector, the vector index structure is invoked to perform similarity calculation and retrieval to obtain a second candidate content list;

[0065] The first candidate content list and the second candidate content list are merged and sorted, and the final content retrieval result is generated and returned to the user based on the fusion and sorting result.

[0066] Furthermore, it also includes:

[0067] When the content retrieval request includes text query conditions, the example audio and video stream separation process is performed on the example audio and video file to obtain the example audio data stream and the example video data stream respectively.

[0068] Audio feature extraction is performed on the example audio data stream to generate an example audio feature vector;

[0069] Keyframe extraction and image feature extraction are performed on the example video data stream to generate a set of example image feature vectors;

[0070] The example audio feature vector and the example image feature vector set are fused to generate an example comprehensive feature vector;

[0071] Based on the example comprehensive feature vector, the vector index structure is called to perform similarity calculation retrieval, and the similarity distance between the example comprehensive feature vector and each content feature vector in the index is calculated;

[0072] Based on the similarity distance calculation results, candidate content with a similarity distance less than a preset threshold is selected and a list of similar content is generated.

[0073] The candidate content in the similar content list is sorted from smallest to largest according to the similarity distance, and the final content retrieval results are generated and returned to the user.

[0074] This invention discloses an AI-based platform for automatic audio and video content generation and analysis, belonging to the field of audio and video processing technology. It receives user content generation requests and performs semantic parsing to identify user intent types and target content descriptions. Based on the intent, it calls a pre-set material library generation engine to synthesize new materials. The obtained audio and video materials undergo multi-dimensional feature extraction, and the results are input into a content recognition module for scene classification, object detection, and sentiment analysis to generate recognition results. Based on the recognition results and user intent, the materials are automatically edited to generate the target audio and video content. A vector index is constructed based on multi-dimensional features, and an inverted index is constructed based on structured metadata, forming a hybrid retrieval index to support subsequent semantic retrieval. This invention achieves integrated processing across the entire process from user intent understanding, content generation, multi-modal analysis, intelligent editing to semantic retrieval, significantly improving the automation level and retrieval accuracy of audio and video content processing. Attached Figure Description

[0075] Figure 1 The diagram shows the structural modules of an AI-based audio and video content automatic generation and analysis platform claimed in this embodiment of the invention.

[0076] Figure 2 A flowchart illustrating the workflow of an AI-based automatic audio and video content generation and analysis platform claimed in this embodiment of the invention.

[0077] Figure 3 The third workflow diagram of an AI-based audio and video content automatic generation and analysis platform claimed in the embodiments of the present invention;

[0078] Figure 4 The fourth workflow diagram of an AI-based audio and video content automatic generation and analysis platform claimed in the embodiments of the present invention;

[0079] Figure 5 The fifth workflow diagram is for an AI-based audio and video content automatic generation and analysis platform claimed in this embodiment of the invention. Detailed Implementation

[0080] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0081] The terms "first," "second," and "third" in this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of those features. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. All directional indications in the embodiments of this application, such as up, down, left, right, front, back, etc., are only used to explain the relative positional relationships and movements between components in a specific orientation as shown in the accompanying drawings. If the specific orientation changes, the directional indication will change accordingly. Furthermore, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, platform, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, platforms, products, or devices.

[0082] References to embodiments herein mean that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0083] According to a first embodiment of the present invention, the present invention claims protection for an artificial intelligence-based platform for automatic generation and analysis of audio and video content, referring to... Figure 1 and Figure 2 ,include:

[0084] The request analysis module, in response to receiving a content generation request from a user, parses the content generation request to extract user intent information and target content description information;

[0085] The material generation module, based on the user intent information and the target content description information, calls original audio and video materials associated with the target content description information from a preset audio and video material library, or calls a pre-trained audio and video generation engine based on the user intent information to generate synthetic audio and video materials that match the target content description information;

[0086] The feature extraction module inputs the original audio-visual material or the synthesized audio-visual material into the content analysis engine, and performs multi-dimensional feature extraction operations on the original audio-visual material or the synthesized audio-visual material through the content analysis engine. The multi-dimensional feature extraction operations include audio feature extraction, visual feature extraction and text feature extraction.

[0087] The content recognition module performs content recognition on the original audio and video materials or the synthesized audio and video materials based on the extracted multi-dimensional features, and generates content recognition results, which include scene category information, object attribute information and sentiment information.

[0088] The editing module performs automated editing processing on the original audio and video materials or synthesized audio and video materials based on the content recognition results, generating target audio and video content that conforms to the user intent information and target content description information;

[0089] The associated storage module associates and stores the target audio and video content and the corresponding content recognition results, and establishes a content retrieval index based on the multi-dimensional features.

[0090] In this embodiment, a content generation request sent by a user terminal is received through an application programming interface. The content generation request includes request text or request voice data.

[0091] Semantic parsing is performed on request text or request voice data to identify keyword groups and sentence intents in content generation requests. Based on keyword groups and sentence intents, the user intent type is determined, which includes content creation intent, content modification intent, and content analysis intent. Target content description information corresponding to the user intent type is then extracted.

[0092] If the user's intent is content creation, then based on the theme identifier and style identifier in the target content description information, the original audio and video material segments that match the theme identifier and style identifier are retrieved from the preset audio and video material library, or the target content description information is input into the pre-trained audio and video generation engine, and the audio and video generation engine automatically synthesizes new audio and video materials based on the theme identifier and style identifier.

[0093] The obtained original audio and video materials or synthesized audio and video materials are loaded into the memory buffer, and the original audio and video materials or synthesized audio and video materials are decoded to separate the audio stream data, video stream data and subtitle stream data;

[0094] The audio stream data is processed by frame segmentation and spectrum conversion to extract audio feature vectors; the video stream data is processed by key frame extraction; image feature vectors are generated for each key frame; and the subtitle stream data is processed by text parsing to extract text feature vectors.

[0095] The audio feature vector, image feature vector, and text feature vector are associated according to the timestamp alignment rule to form a multi-dimensional feature vector set;

[0096] The multidimensional feature vector set is input into the content recognition module. The content recognition module performs scene classification, object detection and sentiment analysis on the original audio and video materials or the synthesized audio and video materials to generate content recognition results. The content recognition results include a list of scene switching points, object location annotation information and sentiment tendency score.

[0097] Based on the content recognition results and user intent type, automated editing operations are performed on the original or synthesized audio and video materials. The automated editing operations include cropping, splicing, overlaying special effects or embedding subtitles to generate the target audio and video content.

[0098] Assign a unique identifier to the target audio and video content, encapsulate the target audio and video content into an audio and video file according to a preset encoding format and bitrate, and store the audio and video file in a specified storage path in a distributed file system;

[0099] The content recognition results are converted into structured metadata, a mapping relationship is established between unique identifiers and structured metadata, and the mapping relationship is stored in the metadata database;

[0100] A vector index is built based on a multidimensional feature vector set, and an inverted index is built based on structured metadata. The vector index and the inverted index are combined into a hybrid retrieval index, which is used to support subsequent content retrieval based on text or examples.

[0101] Furthermore, referring to Figure 3 The request analysis module includes:

[0102] Perform semantic parsing on the content generation request to identify keywords and sentence structures in the content generation request;

[0103] Based on the keywords and sentence structure, the user intent type is determined, which includes content creation intent, content modification intent, and content analysis intent.

[0104] When the user intent type is content creation intent, the theme information, style information, and duration information are extracted from the content generation request as the target content description information; when the user intent type is content modification intent, the identification information of the content to be modified and the modification operation instruction are extracted from the content generation request; when the user intent type is content analysis intent, the identification information of the content to be analyzed and the analysis dimension information are extracted from the content generation request.

[0105] The user intent type and the corresponding extracted information are encapsulated to generate structured user intent information and target content description information.

[0106] In this embodiment, if the content generation request contains request text, the request text is segmented to remove stop words and extract nouns, verbs and adjectives as keyword groups; if the content generation request contains request voice data, the request voice data is first converted into text through a speech recognition engine, and then the converted text is segmented and keyword groups are extracted.

[0107] Part-of-speech tagging and dependency parsing are performed on keyword phrases to identify core verbs and core nouns in the request text. The user intent type is determined based on the semantic category of the core verbs, which includes creation verbs, modification verbs, and analysis verbs.

[0108] If the core verb is a creation verb, the user intent type is determined to be content creation intent. The noun phrases representing the theme are extracted from the request text as theme information, the adjective phrases representing the style are extracted as style information, and the time numbers representing the duration are extracted as duration information. The theme information, style information and duration information are combined into target content description information.

[0109] If the core verb is a modification verb, the user intent type is determined to be content modification intent. The identifier representing the content to be modified and the verb phrase representing the modification operation type are extracted from the request text. The modification operation type includes cropping, splicing, adding effects or adding subtitles. The identifier of the content to be modified and the modification operation type are combined into target content description information.

[0110] If the core verb is an analysis verb, the user intent type is determined to be content analysis intent. Identifiers representing the content to be analyzed and noun phrases representing the analysis dimensions are extracted from the request text. The analysis dimensions include scenario analysis, object analysis, or sentiment analysis. The identifiers of the content to be analyzed and the analysis dimensions are combined into target content description information.

[0111] User intent types and target content description information are encapsulated according to a preset data structure to generate a structured user intent data package. The structured user intent data package contains an intent type field and a description information field. The description information field further contains subfields to store specific themes, styles, durations, identifiers, operation types, or analysis dimensions.

[0112] Furthermore, referring to Figure 4 The material generation module includes:

[0113] When the user intent type is content creation intent, a scene template matching the theme information is retrieved from the preset scene template library based on the theme information in the target content description information.

[0114] Based on the style information in the target content description information, retrieve the set of visual elements and the set of audio elements that match the style information from the preset material style library;

[0115] The scene template and the set of visual elements are fused together to generate a background image sequence; the set of audio elements is combined according to a preset arrangement rule to generate a background audio track.

[0116] Based on the duration information in the target content description information, the background image sequence and background audio track are synchronously cropped or expanded to generate duration-matched synthetic audio and video materials.

[0117] The synthesized audio and video materials are packaged according to a preset encoding format to generate audio and video files with a specified resolution and bitrate.

[0118] In this embodiment, when the user's intent type is content creation intent, the target content description information is first parsed to obtain theme information, style information and duration information. The theme information includes specific object names or scene names, and the style information includes specific visual style descriptive words or auditory style descriptive words.

[0119] Access the index table of the preset audio and video material library. The audio and video material library pre-stores a large number of audio and video clips that have been labeled with theme tags and style tags. Each audio and video clip is manually or automatically labeled when it is added to the library. The labeling content includes theme tags, style tags and duration tags.

[0120] Generate theme query conditions based on theme information, generate style query conditions based on style information, perform a joint query in the index table, and retrieve all audio and video clips whose theme tags match the theme information and whose style tags match the style information.

[0121] From the retrieved audio and video clips, select several clips with the closest duration information. If there are multiple clips, sort them according to the similarity of duration and select one or more clips with the closest duration as candidate original audio and video material clips.

[0122] If no candidate original audio or video material fragments are obtained through the above search, or if the user explicitly specifies the generation method, the audio or video generation engine call process will be triggered.

[0123] When the audio and video generation engine call process is triggered, the theme information, style information and duration information are converted into the input parameter format acceptable to the audio and video generation engine. The audio and video generation engine contains a pre-trained generative adversarial network or variational autoencoder, which is used to generate new audio and video content based on the input parameters.

[0124] Send a generation request to the audio and video generation engine and wait for the audio and video generation engine to return the generation result. The generation result includes the synthesized audio and video data stream and its corresponding metadata. The metadata contains parameter information used in the generation process.

[0125] Receive the synthesized audio and video data stream returned by the audio and video generation engine, and perform integrity verification on the synthesized audio and video data stream. After the verification is passed, use the synthesized audio and video data stream as the synthesized audio and video material.

[0126] If both original audio and video clips and synthesized audio and video clips are obtained, their durations are compared based on the duration information. The one with the most suitable duration is selected as the final audio and video clip, or the two are combined and spliced ​​together to obtain clips that meet the duration requirements.

[0127] Furthermore, referring to Figure 5 The material generation module includes:

[0128] When the user intent type is content modification intent or content analysis intent, extract the identification information of the content to be modified or the content to be analyzed from the content generation request;

[0129] Based on the identification information, access the metadata index table in the audio and video material library to query the storage location information corresponding to the identification information;

[0130] Based on the storage location information, the corresponding original audio and video material files are read from the distributed storage system;

[0131] The original audio and video material files are read and their format is checked to determine the compatibility between the encoding format of the original audio and video material files and the current processing environment.

[0132] When it is detected that the encoding format of the original audio and video material file is incompatible with the current processing environment, the format conversion tool is invoked to convert the original audio and video material file into a preset standard processing format.

[0133] The original audio and video material files after format conversion are loaded into a memory buffer, and a unique processing identifier is assigned to the original audio and video material files. The processing identifier is used to track the processing status of the original audio and video material files in subsequent processing steps.

[0134] In this embodiment, when the user intent type is content modification intent or content analysis intent, the target content description information is parsed and the identifier of the content to be modified or analyzed is extracted. The identifier is the unique storage identifier of the audio and video material in the system or the user-defined name identifier.

[0135] Depending on the type of the identifier, if it is a unique storage identifier, the distributed file system's metadata service will be queried directly based on the identifier to obtain the storage location information corresponding to the identifier. The storage location information includes the storage node address where the file is located, the file path, and the file name.

[0136] If the identifier is a user-defined name identifier, the user configuration file or material mapping table is accessed first to convert the user-defined name identifier into a unique storage identifier within the system. Then, the metadata service of the distributed file system is queried based on the converted unique storage identifier to obtain the storage location information.

[0137] Based on the obtained storage location information, a read request is sent to the corresponding storage node. The read request contains the file path and file name, and the binary data stream of the original audio and video material file is read from the storage node.

[0138] The program performs format detection on the read raw audio and video material files, reads the file header information, and parses out the file's encoding format, container format, resolution, bitrate, and duration information. The parsed encoding format is then compared with a preset list of standard processing formats, which includes commonly used audio and video processing formats such as H.264 encoded MP4 files or AAC encoded audio files.

[0139] If the encoding format of the original audio and video material file is not detected to be in the standard processing format list, the format conversion tool is invoked, and appropriate transcoding parameters are selected according to the encoding format and container format. The transcoding process is then started to convert the original audio and video material file into the standard processing format.

[0140] During the transcoding process, the transcoding progress is monitored in real time. If the transcoding fails, an error log is recorded and a failure message is returned. If the transcoding is successful, a temporary file after transcoding is obtained.

[0141] The original audio and video material files, after conversion or directly read, are loaded into the memory buffer. A temporary processing identifier is assigned to the material, which contains a timestamp and a random number. This identifier is used to uniquely identify the processing instance of the material in subsequent processing steps. At the same time, the processing identifier is bound to user intent information so as to associate user requests during editing or analysis.

[0142] Furthermore, the feature extraction module includes:

[0143] The original audio and video materials or the synthesized audio and video materials are subjected to audio and video stream separation processing to obtain audio data streams and video data streams respectively;

[0144] The audio data stream is segmented into frames, dividing the continuous audio signal into several audio frames. Time-domain feature extraction and frequency-domain feature extraction are performed on each audio frame. The time-domain features include short-time energy and zero-crossing rate, and the frequency-domain features include Mel frequency cepstral coefficients and spectral centroid.

[0145] The video data stream is subjected to frame extraction processing, and keyframe images are extracted from the video data stream at preset time intervals.

[0146] Image feature extraction is performed on each keyframe image, including edge feature extraction, texture feature extraction, and color histogram feature extraction.

[0147] The speech portion of the audio data stream is processed by speech recognition, the speech signal is converted into text data, and text feature extraction is performed on the text data, including word frequency feature extraction and syntactic structure feature extraction.

[0148] The extracted audio features, image features, and text features are correlated according to the timestamp alignment method to generate a multi-dimensional feature vector set.

[0149] In this embodiment, the separated audio stream data is resampled according to a preset sampling rate to ensure that the audio data sampling rate is uniform. Then, the resampled audio data is pre-emphasized to enhance the high-frequency part. Then, the audio data is divided into continuous audio frames, each audio frame having a length of 20 to 40 milliseconds, with 50% overlap between frames.

[0150] A Hamming window function is applied to each audio frame to reduce spectral leakage. A fast Fourier transform is then performed on the windowed audio frames to convert the time-domain signal into a frequency-domain signal, thus obtaining the spectrum of each audio frame.

[0151] The spectrum of each audio frame is filtered by a Mel filter bank, and the spectrum is mapped onto the Mel scale to obtain the Mel spectrum. The logarithm of the Mel spectrum is then taken and a discrete cosine transform is performed to obtain the Mel frequency cepstral coefficients. The Mel frequency cepstral coefficients of each audio frame are combined into a one-dimensional vector as the audio feature of that frame.

[0152] The audio features of all audio frames are arranged in chronological order to form an audio feature sequence. Then, a time-dimensional pooling operation is performed on the audio feature sequence to compress the features of the entire audio stream into a fixed-dimensional audio feature vector.

[0153] The separated video stream data is decoded to obtain a continuous video frame sequence. Scene switching detection is performed on the video frame sequence, and the histogram difference between adjacent frames is calculated. If the difference exceeds the threshold, it is determined to be a scene switching point. Key frames are selected near the scene switching point, and key frames are extracted evenly at fixed time intervals. The key frames obtained by the two methods are merged and deduplicated to obtain a key frame set.

[0154] For each keyframe, size normalization is performed, the image is scaled to a preset width and height, and the normalized image is color space converted from RGB color space to HSV color space. The color histograms of the three HSV channels are extracted as color features.

[0155] An edge detection operator is applied to each keyframe to extract edge information from the image and generate an edge intensity map. The edge intensity map is then divided into blocks for statistical analysis to obtain an edge direction histogram as a texture feature.

[0156] The color and texture features of each keyframe are concatenated into a one-dimensional vector, which is used as the image feature vector of that keyframe. Average pooling is then performed on the image feature vectors of all keyframes to obtain the image feature vector of the entire video stream.

[0157] The separated subtitle stream data is parsed. If the subtitle stream is a graphic subtitle, it is converted into text through optical character recognition. If the subtitle stream is a text subtitle, the text content is extracted directly. The extracted text content is segmented into sentences, and each sentence is segmented and tagged with words and parts of speech. Stop words are removed, and nouns, verbs and adjectives are retained as keywords.

[0158] The frequency of keywords in each sentence is counted to construct a word frequency vector. The word frequency vectors of each sentence are arranged in chronological order, and the word frequency vectors of all sentences are weighted and averaged to obtain the text feature vector of the entire subtitle stream.

[0159] The audio feature vector, image feature vector, and text feature vector are aligned according to the timeline of the original audio and video materials. If a certain feature is missing, it is filled with a zero vector. The three feature vectors are concatenated into a comprehensive feature vector as a multidimensional feature vector set.

[0160] Furthermore, the content recognition module includes:

[0161] Based on the multi-dimensional feature vector set, the audio portion of the original audio and video material or the synthesized audio and video material is classified into scenes to identify audio scene types, which include music scenes, dialogue scenes, ambient sound scenes and silent scenes.

[0162] Based on the image features in the multi-dimensional feature vector set, object detection is performed on the video portion of the original audio and video material or the synthesized audio and video material to identify the object categories and object locations appearing in the video frame. The object categories include human objects, object objects, and background objects.

[0163] Based on the object detection results, scene switching detection is performed on the video portion to identify scene boundaries in the video and generate scene segmentation information;

[0164] Based on the text features in the multi-dimensional feature vector set, sentiment analysis is performed on the speech content in the audio data stream to determine the sentiment tendency value of the speech content.

[0165] The content recognition result is generated by integrating the audio scene type, object detection result, scene segmentation information and sentiment value.

[0166] In this embodiment, the audio feature vectors in the multidimensional feature vector set are input to a pre-trained audio scene classifier. The audio scene classifier determines the scene type of each audio frame based on the audio features. The scene types include music scene, dialogue scene, ambient sound scene and silent scene. The classifier outputs the scene type label and confidence score of each audio frame.

[0167] Based on the timestamp of the audio frames, consecutive audio frames of the same scene type are merged into audio scene segments. The start time, end time and scene type of each audio scene segment are recorded to generate an audio scene segment list.

[0168] The image feature vectors in the multidimensional feature vector set are input into the pre-trained object detector. The object detector performs object detection on each keyframe and identifies the category and location of each object appearing in the keyframe. The object categories include people, animals, vehicles, and buildings, and the location is represented by the coordinates of the bounding box.

[0169] The detected objects are tracked, and the same object in different keyframes is associated to form the motion trajectory of the object. The category of each object, the trajectory start time, the trajectory end time and the position information in each keyframe are recorded, and a list of object detection results is generated.

[0170] The text feature vectors in the multidimensional feature vector set are input into a pre-trained sentiment analyzer. The sentiment analyzer judges the sentiment tendency of the subtitle text and outputs the sentiment tendency score of each sentence. The sentiment tendency score includes positive score, negative score and neutral score, and the score range is 0 to 1.

[0171] The sentiment score of each sentence is associated with the time period corresponding to the sentence to form a sentiment time series. The sentiment time series is smoothed to obtain the sentiment change curve of the entire material, and the overall sentiment score is calculated.

[0172] The audio scene segmentation list, object detection result list, and emotion time series are integrated and aligned according to the timeline to generate structured content recognition results. The content recognition results are organized in JSON format, including metadata header and data body. The metadata header records the material identifier and recognition time, while the data body contains scene labels, object annotations, and emotion scores for multiple time segments.

[0173] Furthermore, the editing module includes:

[0174] Based on the scene segmentation information in the content recognition results, determine the scene switching points of the original audio and video material or the synthesized audio and video material;

[0175] Based on the object detection results in the content recognition results, key objects in the original audio and video materials or synthesized audio and video materials are identified, and the motion trajectory of the key objects in the picture is determined;

[0176] Based on the modification operation instructions in the user intent information, determine the type of editing operation to be performed, which includes cropping, splicing, adding special effects, and adding subtitles.

[0177] When the editing operation type is a cropping operation, the cropping start point and cropping end point are set according to the scene switching point, and cropping processing is performed on the original audio and video material or the synthesized audio and video material.

[0178] When the editing operation type is a special effect addition operation, the special effect addition position is determined according to the object detection result, the special effect type is determined according to the emotional tendency value, and a specified type of visual or audio special effect is superimposed at the special effect addition position.

[0179] When the editing operation type is a subtitle addition operation, the speech content in the audio data stream is transcribed to generate subtitle text, the subtitle display timeline is determined according to the scene segmentation information, and the subtitle text is bound to the subtitle display timeline to generate a subtitle track file;

[0180] The edited audio and video data is combined with the subtitle track file and packaged to generate the target audio and video content.

[0181] In this embodiment, if the user intent type is content modification intent, the modification operation type is extracted from the target content description information, and the corresponding editing operation is performed according to the modification operation type.

[0182] If the operation type is changed to cropping, the cropping start time and cropping end time are extracted from the target content description information. The cropping start time and cropping end time are corrected according to the scene switching point list in the content recognition result. The cropping start time and the cropping end time are adjusted to the closest scene switching point to ensure that the cropped segment is divided by the scene boundary.

[0183] Based on the corrected start and end times of the cropping, the corresponding audio and video data segments are extracted from the original audio and video materials to generate a cropped temporary file;

[0184] If the operation type is changed to splicing operation, the identifiers of multiple materials to be spliced ​​and their splicing order are extracted from the target content description information. The corresponding material file is called according to each identifier, each material is decoded, and the audio and video data of each material are written into a new container file in the splicing order. Smooth transition processing is performed at the splicing point to prevent audio and video jumps.

[0185] If the operation type is changed to special effect overlay operation, the special effect type and the special effect application position are extracted from the target content description information. The special effect type includes visual special effects such as filters and transition effects, or audio special effects such as echo and reverb. Based on the object detection results in the content recognition results, the specific coordinates or time period corresponding to the special effect application position are determined, and the special effect is overlaid at the specified position to generate a temporary file with special effects.

[0186] If the operation type is changed to subtitle embedding operation, the sentiment score is extracted from the content recognition result, the font color and display style of the subtitle are determined according to the sentiment score, speech recognition is performed on the speech part in the audio stream, subtitle text is generated, the subtitle text is synchronized with the video stream according to the timestamp, the subtitle track is embedded into the video file, and a temporary file with subtitles is generated.

[0187] If the user's intent type is content creation intent, then based on the theme and style information in the target content description information, the system will automatically select an effect template that matches the theme and style from the preset effects library, apply the effect template to the original audio and video materials or the synthesized audio and video materials, and generate complete audio and video content with a unified style.

[0188] If the user's intent type is content analysis intent, the original audio and video materials are directly copied without editing, but the content recognition results are associated with the original materials as additional data.

[0189] After all editing operations are completed, the generated temporary file will be encoded according to the preset output format. The encoding parameters include video resolution, video bitrate, audio sampling rate, and audio bitrate. After encoding, the final target audio and video content file will be generated.

[0190] Furthermore, the associated storage module includes:

[0191] Assign a unique content identifier to the target audio and video content, and write the target audio and video content into a distributed file system according to a preset storage structure;

[0192] The content recognition results are converted into structured metadata, which includes scene category tags, object tags, sentiment tags, and time sequence tags.

[0193] Establish the association between the content identifier and the structured metadata, and store the association in a relational database;

[0194] Based on the multi-dimensional feature vector set, a vector index structure is constructed, which is used to support content retrieval based on similarity calculation.

[0195] Based on the structured metadata, an inverted index structure is constructed, which is used to support content retrieval based on keyword matching;

[0196] The vector index structure and the inverted index structure are merged to generate a hybrid retrieval index, which supports composite retrieval conditions based on both vector similarity and keyword matching.

[0197] In this embodiment, a globally unique identifier is generated, which is composed of a timestamp, machine ID, process ID and random number to ensure that it will not be duplicated in a distributed environment;

[0198] Select the corresponding encoder according to the preset encoding format. The encoding format includes one of H.264, H.265, and VP9. Set the bit rate parameter of the encoder according to the preset bit rate. The bit rate is divided into three levels: low bit rate, medium bit rate, and high bit rate according to the application scenario.

[0199] The edited audio and video data stream is input into the encoder, which compresses and encodes the audio and video data according to the specified encoding format and bit rate to generate an audio and video basic stream that meets the format requirements.

[0200] Based on the preset container format, select the corresponding multiplexer. The container format includes one of MP4, MKV, and AVI. Input the encoded audio and video base streams, as well as possible subtitle and metadata streams, into the multiplexer. The multiplexer encapsulates these streams into a single audio and video file.

[0201] During the encapsulation process, file header information, including file type, encoding format, resolution, frame rate, bit rate, and duration, is written to the file, and index information is also written to facilitate quick location later.

[0202] After encapsulation, the hash value of the file is calculated and used as the basis for file integrity verification. The hash value is then stored together with the file.

[0203] The generated audio and video files are uploaded to the distributed file system. Before uploading, the storage path is calculated based on the file's unique identifier. The storage path is organized according to a hierarchical structure. For example, a directory is created based on the first few digits of the identifier, and the files are stored in the corresponding directory.

[0204] During the upload process, the file is transmitted in chunks, and each data block is written to multiple storage nodes simultaneously to achieve redundant backup. After the upload is completed, the storage node information and offset of each data block are recorded.

[0205] After uploading, register the file's metadata with the metadata service, including a unique identifier, storage path, file size, hash value, and upload time, for quick retrieval and access later.

[0206] Furthermore, the platform also includes a content retrieval module, comprising:

[0207] Receive a content retrieval request input by the user, the content retrieval request including text query conditions or example audio / video files;

[0208] When the content retrieval request includes text query conditions, semantic parsing is performed on the text query conditions to extract query keywords;

[0209] Based on the query keywords, the inverted index structure is invoked to perform keyword matching and retrieval to obtain the first candidate content list;

[0210] The text query conditions are vectorized to generate query feature vectors;

[0211] Based on the query feature vector, the vector index structure is invoked to perform similarity calculation and retrieval to obtain a second candidate content list;

[0212] The first candidate content list and the second candidate content list are merged and sorted, and the final content retrieval result is generated and returned to the user based on the fusion and sorting result.

[0213] In this embodiment, a content retrieval request input by the user is received through a user interface. The content retrieval request includes retrieval conditions, which may be text query terms or uploaded data of example audio and video files.

[0214] If the search criteria are text query terms, then the text query terms are preprocessed, including word segmentation, stop word removal, and stemming, to obtain a list of query keywords;

[0215] Based on the list of query keywords, the inverted index structure is accessed. Each keyword in the inverted index structure corresponds to a list of audio and video file identifiers containing that keyword and their weights. The file identifier lists corresponding to each keyword are merged, and the matching score of each file identifier is calculated. The matching score is calculated based on the keyword occurrence frequency and inverse document frequency.

[0216] File identifiers with matching scores higher than the first threshold are selected as the first candidate set and sorted from highest to lowest score.

[0217] Simultaneously, semantic encoding is performed on the text query terms, converting the query terms into query feature vectors. Semantic encoding is achieved through a pre-trained word vector model, which converts each word in the query terms into a word vector, and then performs a weighted average of the word vectors to obtain the query feature vector.

[0218] The query feature vector is input into the vector index structure. The vector index structure uses an approximate nearest neighbor search algorithm to calculate the similarity distance between the query feature vector and the comprehensive feature vector of each audio and video file in the index. The similarity distance uses cosine similarity or Euclidean distance to obtain the similarity score of each file.

[0219] File identifiers with similarity scores higher than the second threshold are selected as the second candidate set and sorted from high to low according to their similarity scores.

[0220] The first and second candidate sets are merged, and a weighted summation method is used to calculate the comprehensive score of each file. The weights are pre-set according to the application scenario. The file identifiers with comprehensive scores higher than the third threshold are used as the final search results and sorted according to the comprehensive scores.

[0221] If the search criteria is an example audio or video file, then the same feature extraction operation is performed on the example audio or video file to obtain the example audio feature vector, example image feature vector, and example text feature vector, and the three are concatenated into an example comprehensive feature vector;

[0222] Input the example's comprehensive feature vector into the vector index structure, perform a similarity search, and obtain a list of audio and video file identifiers most similar to the example and their similarity scores;

[0223] File identifiers with similarity scores higher than the fourth threshold are used as the final search results and sorted according to their similarity scores;

[0224] Based on the file identifier in the final search results, the metadata and storage path of the corresponding audio and video files are read from the distributed file system, a search result list containing file name, duration, resolution, and thumbnail is generated, and the search result list is returned to the user interface for display.

[0225] Furthermore, it also includes:

[0226] When the content retrieval request includes text query conditions, the example audio and video stream separation process is performed on the example audio and video file to obtain the example audio data stream and the example video data stream respectively.

[0227] Audio feature extraction is performed on the example audio data stream to generate an example audio feature vector;

[0228] Keyframe extraction and image feature extraction are performed on the example video data stream to generate a set of example image feature vectors;

[0229] The example audio feature vector and the example image feature vector set are fused to generate an example comprehensive feature vector;

[0230] Based on the example comprehensive feature vector, the vector index structure is called to perform similarity calculation retrieval, and the similarity distance between the example comprehensive feature vector and each content feature vector in the index is calculated;

[0231] Based on the similarity distance calculation results, candidate content with a similarity distance less than a preset threshold is selected and a list of similar content is generated.

[0232] The candidate content in the similar content list is sorted from smallest to largest according to the similarity distance, and the final content retrieval results are generated and returned to the user.

[0233] In this embodiment, the example audio and video file uploaded by the user is received, the example audio and video file is saved to a temporary storage area, and a temporary file identifier is assigned.

[0234] The example audio and video files are separated into audio and video streams, and a demultiplexer is used to extract audio stream data, video stream data, and possible subtitle stream data from the container file;

[0235] If the example audio / video file contains subtitle stream data, then the subtitle stream data is parsed to extract the subtitle text content, the subtitle text is segmented and keywords are extracted to generate word frequency vectors, and all word frequency vectors are averaged and pooled to obtain the example text feature vector.

[0236] If the example audio / video file does not contain subtitle stream data, then the example text feature vector is set to an all-zero vector;

[0237] Audio features are extracted from the extracted audio stream data. According to the described platform, the audio stream is segmented into frames, windowed, Fourier transformed, filtered by the Mel filter bank, and discrete cosine transformed to obtain Mel frequency cepstral coefficients. The Mel frequency cepstral coefficients of all frames are then averaged in the time dimension to obtain example audio feature vectors.

[0238] Video feature extraction is performed on the extracted video stream data. First, the video stream is decoded to obtain a video frame sequence. Keyframes are extracted at fixed intervals. Each keyframe is normalized in size and converted in color space. Color histograms are extracted, and edge direction histograms are extracted at the same time. Color features and texture features are concatenated to form the image feature vector of each keyframe. Average pooling is performed on the image feature vectors of all keyframes to obtain the example image feature vector.

[0239] The example audio feature vector, example image feature vector, and example text feature vector are concatenated according to a preset dimension. If the dimension of a feature vector is insufficient, it is padded with zeros. If the dimension is excessive, it is truncated. Finally, a fixed-dimensional example composite feature vector is generated.

[0240] The example composite feature vector is stored in memory and associated with a temporary file identifier for later retrieval in the similarity search step;

[0241] After the similarity search is complete, delete the example audio and video files from the temporary storage area to free up storage space.

[0242] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and platforms can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.

[0243] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. The above are merely embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made based on the description and drawings of this application, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

[0244] The specific embodiments of the invention have been described in detail above, but they are only examples, and this application is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications or substitutions to the invention are also within the scope of this application. Therefore, all equivalent changes, modifications, and improvements made without departing from the spirit and principles of this application should be covered within the scope of this application.

Claims

1. An AI-based platform for automatic generation and analysis of audio and video content, characterized in that, include: The request analysis module, in response to receiving a content generation request from a user, parses the content generation request to extract user intent information and target content description information; The material generation module, based on the user intent information and the target content description information, calls original audio and video materials associated with the target content description information from a preset audio and video material library, or calls a pre-trained audio and video generation engine based on the user intent information to generate synthetic audio and video materials that match the target content description information; The feature extraction module inputs the original audio-visual material or the synthesized audio-visual material into the content analysis engine, and performs multi-dimensional feature extraction operations on the original audio-visual material or the synthesized audio-visual material through the content analysis engine. The multi-dimensional feature extraction operations include audio feature extraction, visual feature extraction and text feature extraction. The content recognition module performs content recognition on the original audio and video materials or the synthesized audio and video materials based on the extracted multi-dimensional features, and generates content recognition results, which include scene category information, object attribute information and sentiment information. The editing module performs automated editing processing on the original audio and video materials or synthesized audio and video materials based on the content recognition results, generating target audio and video content that conforms to the user intent information and target content description information; The associated storage module associates and stores the target audio and video content and the corresponding content recognition results, and establishes a content retrieval index based on the multi-dimensional features.

2. The AI-based audio and video content automatic generation and analysis platform according to claim 1, characterized in that, The request analysis module includes: Perform semantic parsing on the content generation request to identify keywords and sentence structures in the content generation request; Based on the keywords and sentence structure, the user intent type is determined, which includes content creation intent, content modification intent, and content analysis intent. When the user intent type is content creation intent, the theme information, style information, and duration information are extracted from the content generation request as the target content description information; when the user intent type is content modification intent, the identification information of the content to be modified and the modification operation instruction are extracted from the content generation request; when the user intent type is content analysis intent, the identification information of the content to be analyzed and the analysis dimension information are extracted from the content generation request. The user intent type and the corresponding extracted information are encapsulated to generate structured user intent information and target content description information.

3. The AI-based audio and video content automatic generation and analysis platform according to claim 2, characterized in that, The material generation module includes: When the user intent type is content creation intent, a scene template matching the theme information is retrieved from the preset scene template library based on the theme information in the target content description information. Based on the style information in the target content description information, retrieve the set of visual elements and the set of audio elements that match the style information from the preset material style library; The scene template and the set of visual elements are fused together to generate a background image sequence; the set of audio elements is combined according to a preset arrangement rule to generate a background audio track. Based on the duration information in the target content description information, the background image sequence and background audio track are synchronously cropped or expanded to generate duration-matched synthetic audio and video materials. The synthesized audio and video materials are packaged according to a preset encoding format to generate audio and video files with a specified resolution and bitrate.

4. The AI-based audio and video content automatic generation and analysis platform according to claim 2, characterized in that, The material generation module includes: When the user intent type is content modification intent or content analysis intent, extract the identification information of the content to be modified or the content to be analyzed from the content generation request; Based on the identification information, access the metadata index table in the audio and video material library to query the storage location information corresponding to the identification information; Based on the storage location information, the corresponding original audio and video material files are read from the distributed storage system; The original audio and video material files are read and their format is checked to determine the compatibility between the encoding format of the original audio and video material files and the current processing environment. When it is detected that the encoding format of the original audio and video material file is incompatible with the current processing environment, the format conversion tool is invoked to convert the original audio and video material file into a preset standard processing format. The original audio and video material files after format conversion are loaded into a memory buffer, and a unique processing identifier is assigned to the original audio and video material files. The processing identifier is used to track the processing status of the original audio and video material files in subsequent processing steps.

5. The AI-based audio and video content automatic generation and analysis platform according to claim 1, characterized in that, The feature extraction module includes: The original audio and video materials or the synthesized audio and video materials are subjected to audio and video stream separation processing to obtain audio data streams and video data streams respectively; The audio data stream is segmented into frames, dividing the continuous audio signal into several audio frames. Time-domain feature extraction and frequency-domain feature extraction are performed on each audio frame. The time-domain features include short-time energy and zero-crossing rate, and the frequency-domain features include Mel frequency cepstral coefficients and spectral centroid. The video data stream is subjected to frame extraction processing, and keyframe images are extracted from the video data stream at preset time intervals. Image feature extraction is performed on each keyframe image, including edge feature extraction, texture feature extraction, and color histogram feature extraction. The speech portion of the audio data stream is processed by speech recognition, the speech signal is converted into text data, and text feature extraction is performed on the text data, including word frequency feature extraction and syntactic structure feature extraction. The extracted audio features, image features, and text features are correlated according to the timestamp alignment method to generate a multi-dimensional feature vector set.

6. The AI-based audio and video content automatic generation and analysis platform according to claim 5, characterized in that, The content recognition module includes: Based on the multi-dimensional feature vector set, the audio portion of the original audio and video material or the synthesized audio and video material is classified into scenes to identify audio scene types, which include music scenes, dialogue scenes, ambient sound scenes and silent scenes. Based on the image features in the multi-dimensional feature vector set, object detection is performed on the video portion of the original audio and video material or the synthesized audio and video material to identify the object categories and object locations appearing in the video frame. The object categories include human objects, object objects, and background objects. Based on the object detection results, scene switching detection is performed on the video portion to identify scene boundaries in the video and generate scene segmentation information; Based on the text features in the multi-dimensional feature vector set, sentiment analysis is performed on the speech content in the audio data stream to determine the sentiment tendency value of the speech content. The content recognition result is generated by integrating the audio scene type, object detection result, scene segmentation information and sentiment value.

7. The AI-based audio and video content automatic generation and analysis platform according to claim 1, characterized in that, The editing module includes: Based on the scene segmentation information in the content recognition results, determine the scene switching points of the original audio and video material or the synthesized audio and video material; Based on the object detection results in the content recognition results, key objects in the original audio and video materials or synthesized audio and video materials are identified, and the motion trajectory of the key objects in the picture is determined; Based on the modification operation instructions in the user intent information, determine the type of editing operation to be performed, which includes cropping, splicing, adding special effects, and adding subtitles. When the editing operation type is a cropping operation, the cropping start point and cropping end point are set according to the scene switching point, and cropping processing is performed on the original audio and video material or the synthesized audio and video material. When the editing operation type is a special effect addition operation, the special effect addition position is determined according to the object detection result, the special effect type is determined according to the emotional tendency value, and a specified type of visual or audio special effect is superimposed at the special effect addition position. When the editing operation type is a subtitle addition operation, the speech content in the audio data stream is transcribed to generate subtitle text, the subtitle display timeline is determined according to the scene segmentation information, and the subtitle text is bound to the subtitle display timeline to generate a subtitle track file; The edited audio and video data is combined with the subtitle track file and packaged to generate the target audio and video content.

8. The AI-based audio and video content automatic generation and analysis platform according to claim 1, characterized in that, The associated storage module includes: Assign a unique content identifier to the target audio and video content, and write the target audio and video content into a distributed file system according to a preset storage structure; The content recognition results are converted into structured metadata, which includes scene category tags, object tags, sentiment tags, and time sequence tags. Establish the association between the content identifier and the structured metadata, and store the association in a relational database; Based on the multi-dimensional feature vector set, a vector index structure is constructed, which is used to support content retrieval based on similarity calculation. Based on the structured metadata, an inverted index structure is constructed, which is used to support content retrieval based on keyword matching; The vector index structure and the inverted index structure are merged to generate a hybrid retrieval index, which supports composite retrieval conditions based on both vector similarity and keyword matching.

9. The AI-based audio and video content automatic generation and analysis platform according to claim 8, characterized in that, The platform also includes a content retrieval module, comprising: Receive a content retrieval request input by the user, the content retrieval request including text query conditions or example audio / video files; When the content retrieval request includes text query conditions, semantic parsing is performed on the text query conditions to extract query keywords; Based on the query keywords, the inverted index structure is invoked to perform keyword matching and retrieval to obtain the first candidate content list; The text query conditions are vectorized to generate query feature vectors; Based on the query feature vector, the vector index structure is invoked to perform similarity calculation and retrieval to obtain a second candidate content list; The first candidate content list and the second candidate content list are merged and sorted, and the final content retrieval result is generated and returned to the user based on the fusion and sorting result.

10. The AI-based audio and video content automatic generation and analysis platform as described in claim 9, characterized in that, Also includes: When the content retrieval request includes text query conditions, the example audio and video stream separation process is performed on the example audio and video file to obtain the example audio data stream and the example video data stream respectively. Audio feature extraction is performed on the example audio data stream to generate an example audio feature vector; Keyframe extraction and image feature extraction are performed on the example video data stream to generate a set of example image feature vectors; The example audio feature vector and the example image feature vector set are fused to generate an example comprehensive feature vector; Based on the example comprehensive feature vector, the vector index structure is called to perform similarity calculation retrieval, and the similarity distance between the example comprehensive feature vector and each content feature vector in the index is calculated; Based on the similarity distance calculation results, candidate content with a similarity distance less than a preset threshold is selected and a list of similar content is generated. The candidate content in the similar content list is sorted from smallest to largest according to the similarity distance, and the final content retrieval results are generated and returned to the user.