An industrialized digital asset production system and method for micro short dramas
Through the industrial micro-short drama digital asset production system, micro-short audio and video files are automatically processed, and the full process of industrial production from generation to distribution is solved, which solves the problems of long production cycle and high cost of micro-short video digital asset production, and improves production efficiency and asset quality.
Patent Information
- Application Number
- CN202410324229.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-03-21
AI Technical Summary
In the prior art, the digital asset production cycle of micro-short videos is long and requires a lot of manual participation, which leads to high time and economic costs and is unable to achieve large-scale matrix production.
It provides an industrial micro-short drama digital asset production system, including an enterprise content input, an industrial task scheduling system and an independent functional module. Through an automated process, it realizes full-process industrial production from preprocessing of micro-short audio and video files, voice recognition, keyword extraction, natural language generation, text to image generation, digital asset editing, material production, and blockchain automation deployment.
It significantly improves the production efficiency of digital assets and reduces the complexity of manual intervention and operation. The generated digital assets are of high quality and innovative, have unique characteristics and collection value, and solves the limitations of traditional methods.
Smart Images

Figure CN118396263B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital asset industrial production, and particularly to an industrialized micro short drama digital asset production system and method. Background Art
[0002] The definition of digital assets can be understood from different perspectives and fields, and different definitions reflect the diversity of digital assets and their applications in different fields. Digital assets not only include traditional media files such as videos and audios, but also include emerging encrypted currencies and assets generated by blockchain technology.
[0003] Digital assets being included in the financial statements means that an enterprise's data resources are regarded as an asset and incorporated into the financial statements. Currently, the accounting treatment scope, applicable standards, and related disclosure requirements for an enterprise's data resources are being gradually improved, which means that the enterprise needs to recognize and measure the data resources according to regulations and regard them as one of the enterprise's assets. The purpose of digital assets being included in the statements is to promote and standardize data-related enterprises in implementing accounting standards, ensure that economic activities of data are accurately reflected, and at the same time promote innovative research in the accounting field and efficient disclosure of information related to data resources. It provides important accounting information support for regulatory authorities to improve the governance system of the digital economy and strengthen macro management.
[0004] In the field of digital assetization, there are currently few digital assetization cases for micro short audio-visual works, especially those with content continuity and more than dozens or hundreds of episodes. The production method is still the traditional digital asset production method, with the following characteristics: 1. Manual processing of signed audio-visual works; 2. Manually hand-drawing a small number (not exceeding 5) of image works; 3. Planning and discussing implementation plans and writing smart contracts; 4. Deploying digital assets to a specified blockchain network; 5. Establishing a special operation team for the works.
[0005] The above digital asset development, production, operation, and maintenance process requires a large amount of manual participation to achieve the purpose of shaping the value and maintaining the market value of digital assets, resulting in a digital assetization production cycle of 2 to 3 months for each micro short video, and it cannot be carried out in parallel. The time cost and economic cost that need to be invested are much greater than the product life cycle of micro short videos (generally within 1 month). Therefore, although micro short videos belong to virtual digital products and are the most suitable type for digital assetization, large-scale matrix implementation cannot be carried out.
[0006] Therefore, how to provide an industrialized micro short drama digital asset production system and method is an urgent problem to be solved currently. Summary of the Invention
[0007] Embodiments of the present invention provide an industrialized micro short drama digital asset production system and method to solve the above technical problems existing in the prior art.
[0008] To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary part is not a general review, nor is it intended to identify key / important constituent elements or delineate the scope of protection of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the detailed description that follows.
[0009] According to a first aspect of an embodiment of the present invention, an industrialized micro short drama digital asset production system is provided.
[0010] In one embodiment, the industrialized micro short drama digital asset production system includes:
[0011] An enterprise content input end for enterprise users to upload micro short audio and video files through an input interface;
[0012] An industrialized task scheduling system for transmitting the micro short audio and video files to independent functional modules and invoking the corresponding independent functional modules to execute independent micro short drama processing tasks;
[0013] Independent functional modules for executing the processing tasks of the micro short audio and video files and realizing the full-process industrialized production of digital assets from generation, production, distribution to basic operation based on the task processing results;
[0014] A digital asset dashboard for displaying the output information of digital assets based on an information output interface.
[0015] In one embodiment, the independent functional modules include: a micro short audio and video preprocessing and storage module, an automatic speech recognition module, a keyword extraction module, a natural language generation module, a text-to-image generation module, a digital asset editing module, a material production module, a blockchain automatic deployment module, and a timed task distribution module. Among them,
[0016] The micro short audio and video preprocessing and storage module is used to convert micro short audio and video files in different formats into a unified encoding format and perform key frame extraction and audio-video separation of the micro short audio and video files;
[0017] The automatic speech recognition module is used to extract the speech features of the audio content in the micro short audio and video files, convert them into text data in text form based on the speech features, and mark time information;
[0018] The keyword extraction module is used to extract keywords from the text data;
[0019] The natural language generation module is used to generate a description copy of the micro short audio and video files based on the keywords in combination with a complete line script file;
[0020] The text-to-image generation module is used to convert the description text into an image;
[0021] The digital asset editing module is used for enterprise users to customize personalized information of digital assets, wherein the personalized information includes smart contract type, minting quantity and activation date;
[0022] The material production module is used to create promotional materials for digital assets and process copyright certification materials;
[0023] The blockchain automated deployment module is used to upload the digital assets to the blockchain network and automatically deploy the corresponding smart contracts;
[0024] The scheduled task distribution module is used to automatically publish the promotional materials of the digital assets at a preset time.
[0025] In one embodiment, the automatic speech recognition module includes: a speech signal processing module, a speech feature extraction module, a speech text recognition module, a context recognition module and a time marking module, wherein:
[0026] The speech signal processing module is used to eliminate noise interference in the audio content by using a filter to obtain a standardized speech signal;
[0027] The speech feature extraction module is used to display the speech attribute of the speech signal by extracting the speech feature in the speech signal, wherein the speech feature is a Mel frequency cepstral coefficient;
[0028] The speech-to-text recognition module is used to convert the speech features into text data in a text form based on a machine learning algorithm;
[0029] The context recognition module is used to recognize context information in the speech signal;
[0030] The time marking module is used to align the text data with the time series of the voice signal based on a time alignment algorithm to generate a dialogue script file containing time information.
[0031] In one embodiment, the keyword extraction module includes: a text cleaning module, an automatic extraction module, a context understanding module, a complex text processing module, a multilingual adaptation module and a learning adaptation module, wherein:
[0032] The text cleaning module is used to remove irrelevant information in the script file;
[0033] The automatic extraction module is used to extract keywords or key phrases in the script file;
[0034] The context understanding module is used to sort out the context of the story based on the lines script file;
[0035] The complex text processing module is used to identify the professional knowledge in the lines script file based on the professional field knowledge model, where the professional knowledge includes industry-specific vocabulary, academic papers, reports, and articles;
[0036] The multilingual adaptation module is used to extract keywords from the lines script files in different languages;
[0037] The learning adaptation module is used to improve the adaptability of the recurrent neural network model through model learning based on the input quantity of the micro short video files.
[0038] In one embodiment, the natural language generation module includes: a story element generation module, a keyword conversion module, a scene transition module, a style main body preset module, and an adaptation customization module, where
[0039] The story element generation module is used to generate story elements based on the extracted keywords, where the story elements include characters, backgrounds, and events;
[0040] The keyword conversion module is used to generate a continuous story description based on the complete content of the lines script file;
[0041] The scene transition module is used to generate picture prompts for constructing the story scenes based on the complete content of the lines script file;
[0042] The style main body preset module is used to adjust the generated story content based on the preset style model;
[0043] The adaptation customization module is used to customize the specific requirements of the story according to the promotional materials uploaded by the enterprise and implant the specified items in the story scenes.
[0044] In one embodiment, the text-to-image generation module includes: an image generation module, a reference image selection module, a content implantation module, a visual format adaptation module, and a continuous text processing module, where
[0045] The image generation module is used to generate corresponding images based on the picture prompts;
[0046] The reference image selection module is used to select the key frame screenshots corresponding to the picture prompts as reference images;
[0047] The content implantation module is used to extract specific content or specific styles based on the picture format materials specified by the enterprise and implant them into the images;
[0048] The visual format adaptation module is used to pre-configure different visual forms to adapt to different application scenarios;
[0049] The continuous text processing module is used to process the continuous text content in the lines script file.
[0050] In one embodiment, the digital asset editing module includes: an intelligent contract customization module, a minting quantity customization module, an activation date customization module, a source file customization module, a solution configuration module, and a blockchain configuration module, where
[0051] The intelligent contract customization module is used to obtain the contract type selected by the enterprise user and customize the intelligent contract type based on the parameters provided by the enterprise user;
[0052] The minting quantity customization module is used to set the limit of the minting quantity in the intelligent contract and customize the minting quantity based on the quantity requirement set by the enterprise user;
[0053] The activation date customization module is used to obtain the activation date of the intelligent contract input by the enterprise user and set a timer when deploying the intelligent contract;
[0054] The source file customization module is used to obtain the source file of the asset uploaded by the enterprise user and store and manage the source file uploaded by the enterprise user based on the file management method;
[0055] The solution configuration module is used to pre-define configuration templates covering different blockchain configurations and asset types and select the corresponding template according to the needs of the enterprise user.
[0056] In one embodiment, the material production module includes: a promotion material generation module, a copyright certification material submission module, a batch generation combination module, a 3D modeling conversion module, and a personalized customization module, where
[0057] The promotion material generation module is used to automatically generate promotion materials for digital assets based on a preset template, where the promotion materials include images, posters, text descriptions, promotional copy, interactive web pages, digital display web pages, and social media promotion;
[0058] The copyright certification material submission module is used to automatically process the materials and documents required for the copyright certification of digital assets;
[0059] The batch generation combination module is used to batch generate unique images for creating exclusive digital assets for micro short video files;
[0060] The 3D modeling conversion module is used to convert two-dimensional images into 3D modeling resources through a 3D engine;
[0061] The personalized customization module is used to adjust the style and format of promotional materials based on enterprise requirements.
[0062] In one embodiment, the blockchain automated deployment module includes: a file compression and upload module, a meta-file submission module, a smart contract deployment module, and a contract matrix generation module, where
[0063] The file compression and upload module is used to compress the original file corresponding to the digital asset and upload it to the InterPlanetary File System;
[0064] The meta-file submission module is used to submit the source file for generating the digital asset to the InterPlanetary File System;
[0065] The smart contract deployment module is used to automatically and batch-deploy the smart contracts required for the digital asset;
[0066] The contract matrix generation module is used to monitor the deployment status of all smart contracts and associate the digital assets of multiple short micro-video files to form a smart contract matrix.
[0067] According to the second aspect of the embodiments of the present invention, an industrialized short micro-drama digital asset production method is provided.
[0068] In one embodiment, the industrialized short micro-drama digital asset production method includes:
[0069] Enterprise users upload short micro-video files through the input interface;
[0070] Transmit the short micro-video files to the independent function modules, and call the corresponding independent function modules to execute independent short micro-drama processing tasks;
[0071] Execute the processing tasks of the short micro-video files, and based on the task processing results, realize the full-process industrialized production of digital assets from generation, production, distribution to basic operation;
[0072] Based on the information output interface, display the output information of the digital assets.
[0073] According to the third aspect of the embodiments of the present invention, a computer device is provided.
[0074] In some embodiments, the computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0075] According to the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided.
[0076] In one embodiment, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0077] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0078] 1. By integrating multiple functional modules and an industrial task scheduling system, the present invention realizes the full-process automation of micro-short drama digital assets from production to distribution, significantly improving production efficiency, reducing manual intervention, and lowering operation complexity; at the same time, through advanced technologies such as deep learning, natural language processing, and text-to-image generation, it can generate higher-quality and more innovative digital assets; and the digital assets have the unique feature of being independent and having collection value, and nearly 20,000 different works can be generated for every about 10 episodes of the drama, effectively solving the limitations of traditional methods; the present invention shows significant beneficial effects at the technical, economic, and social levels, especially having important value in improving the production efficiency of digital assets, reducing costs, enhancing digital innovation capabilities, and expanding the market coverage.
[0079] 2. The present invention effectively reduces labor costs and time costs, while increasing the market value of digital assets. Through an industrialized production method, enterprises can quickly respond to market demands and more efficiently carry out the matrix layout of digital assets; in the industrialized production of micro-short drama digital assets, it fills the gap in the existing technology, especially in the digital assetization of micro-short dramas with content continuity.
[0080] 3. The automatic speech recognition module of the present invention supports multiple languages and dialects, enabling the digital assetization system to adapt to global applications and expanding the market coverage; by reducing the use of physical media and promoting digital products, it helps to reduce environmental pollution.
[0081] 4. Through the construction of an intelligent contract matrix and decentralized storage, the management flexibility and efficiency of digital assets are enhanced, while the data security and persistence are improved; at the same time, the use of a timed task dispatcher allows enterprises to automatically publish relevant materials on multiple media channels, improving the operation efficiency and publicity effect.
[0082] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and, together with the specification, used to explain the principles of the present invention.
[0084] Figure 1It is a system block diagram of an industrialized micro short drama digital asset production system shown according to an exemplary embodiment;
[0085] Figure 2 It is a business process flowchart of an industrialized micro short drama digital asset production system shown according to an exemplary embodiment;
[0086] Figure 3 It is an illustration of the information flow of each module of an industrialized micro short drama digital asset production system shown according to an exemplary embodiment;
[0087] Figure 4 A flowchart of an industrialized micro short drama digital asset production method shown according to an exemplary embodiment;
[0088] Figure 5 It is a schematic structural diagram of a computer device shown according to an exemplary embodiment. Detailed implementation manners
[0089] Figure 1 An embodiment of an industrialized micro short drama digital asset production system of the present invention is shown. In this alternative embodiment, the industrialized micro short drama digital asset production system includes:
[0090] An enterprise content input terminal 101, configured to enable enterprise users to upload micro short audio and video files through an input interface;
[0091] An industrialized task scheduling system 103, configured to transmit the micro short audio and video files to independent functional modules and call the corresponding independent functional modules to execute independent micro short drama processing tasks;
[0092] Independent functional modules 105, configured to execute the processing tasks of the micro short audio and video files, and based on the task processing results, realize the full-process industrialized production of digital assets from generation, production, distribution to basic operation;
[0093] A digital asset dashboard 107, configured to display the output information of the digital assets based on an information output interface.
[0094] In this alternative embodiment, the independent functional module 105 includes: a micro short audio-video preprocessing and storage module (not shown in the figure), an automatic speech recognition module (not shown in the figure), a keyword extraction module (not shown in the figure), a natural language generation module (not shown in the figure), a text-to-image generation module (not shown in the figure), a digital asset editing module (not shown in the figure), a material production module (not shown in the figure), a blockchain automatic deployment module (not shown in the figure), and a timing task distribution module (not shown in the figure). Among them, the micro short audio-video preprocessing and storage module is used to convert micro short audio-video files in different formats into a unified coding format, and perform key frame extraction and audio-video separation of the micro short audio-video files; the automatic speech recognition module is used to extract the speech features of the audio content in the micro short audio-video files, and based on the speech features, convert them into text data in text form and mark time information; the keyword extraction module is used to extract keywords from the text data; the natural language generation module is used to generate a description copy of the micro short audio-video file based on the keywords and in combination with a complete line script file; the text-to-image generation module is used to convert the description copy into an image; the digital asset editing module is used for enterprise users to customize personalized information of digital assets, where the personalized information includes the type of smart contract, the number of minting, and the activation date; the material production module is used to create promotion materials for digital assets and process copyright certification materials; the blockchain automatic deployment module is used to upload the digital assets to the blockchain network and automatically deploy corresponding smart contracts; the timing task distribution module is used to automatically publish the promotion materials of the digital assets at a preset time.
[0095] In this alternative embodiment, the automatic speech recognition module includes: a speech signal processing module (not shown in the figure), a speech feature extraction module (not shown in the figure), a speech text recognition module (not shown in the figure), a context recognition module (not shown in the figure), and a time marking module (not shown in the figure). Among them, the speech signal processing module is used to eliminate noise interference in the audio content using a filter to obtain a standardized speech signal; the speech feature extraction module is used to display the speech attributes of the speech signal by extracting the speech features in the speech signal, where the speech features are Mel frequency cepstral coefficients; the speech text recognition module is used to convert the speech features into text data in text form based on a machine learning algorithm; the context recognition module is used to recognize the context information in the speech signal; the time marking module is used to align the time series of the text data with the speech signal based on a time alignment algorithm to generate a line script file containing time information.
[0096] In this optional embodiment, the keyword extraction module includes: a text cleaning module (not shown in the figure), an automatic extraction module (not shown in the figure), a context understanding module (not shown in the figure), a complex text processing module (not shown in the figure), a multilingual adaptation module (not shown in the figure) and a learning adaptation module (not shown in the figure), wherein the text cleaning module is used to remove irrelevant information in the dialogue script file; the automatic extraction module is used to extract keywords or key phrases in the dialogue script file; the context understanding module is used to sort out the context of the story based on the dialogue script file; the complex text processing module is used to identify the professional knowledge in the dialogue script file based on the professional domain knowledge model, wherein the professional knowledge includes industry professional vocabulary, academic papers, reports and articles; the multilingual adaptation module is used to extract keywords in dialogue script files of different languages; the learning adaptation module is used to improve the adaptability of the recurrent neural network model through model learning based on the input number of the micro-short audio and video files.
[0097] In this optional embodiment, the natural language generation module includes: a story element generation module (not shown in the figure), a keyword conversion module (not shown in the figure), a screen switching module (not shown in the figure), a style main body preset module (not shown in the figure) and an adaptive customization module (not shown in the figure), wherein the story element generation module is used to generate story elements based on the extracted keywords, wherein the story elements include characters, backgrounds and events; the keyword conversion module is used to generate a continuous story description based on the complete content of the line script file; the screen switching module is used to generate picture prompts for constituting story pictures based on the complete content of the line script file; the style main body preset module is used to adjust the generated story content based on a preset style model; the adaptive customization module is used to customize the specific needs of the story based on the promotional materials uploaded by the enterprise, and implant designated objects in the story scene.
[0098] In this optional embodiment, the text-to-image generation module includes: an image generation module (not shown in the figure), a reference image selection module (not shown in the figure), a content implantation module (not shown in the figure), a visual format adaptation module (not shown in the figure) and a continuous text processing module (not shown in the figure), wherein the image generation module is used to generate a corresponding image based on a picture prompt; the reference image selection module is used to select a key frame screenshot corresponding to the picture prompt as a reference image; the content implantation module is used to extract specific content or a specific style based on the picture format material specified by the enterprise, and implant it into the image; the visual format adaptation module is used to pre-configure different visual forms to adapt to different application scenarios; and the continuous text processing module is used to process text content with continuity in the dialogue script file.
[0099] In this alternative embodiment, the digital asset editing module includes: a smart contract customization module (not shown in the figure), a minting quantity customization module (not shown in the figure), an activation date customization module (not shown in the figure), a source file customization module (not shown in the figure), a solution configuration module (not shown in the figure), and a blockchain configuration module (not shown in the figure). Among them, the smart contract customization module is used to obtain the contract type selected by the enterprise user and customize the smart contract type based on the parameters provided by the enterprise user; the minting quantity customization module is used to set the limit of the minting quantity in the smart contract and customize the minting quantity based on the quantity requirements set by the enterprise user; the activation date customization module is used to obtain the activation date of the smart contract input by the enterprise user and set a timer when deploying the smart contract. The source file customization module is used to obtain the source file of the asset uploaded by the enterprise user and store and manage the source file uploaded by the enterprise user based on the file management method; the solution configuration module is used to pre-define configuration templates covering different blockchain configurations and asset types and select the corresponding template according to the needs of the enterprise user.
[0100] In this alternative embodiment, the material production module includes: a promotion material generation module (not shown in the figure), a copyright certification material submission module (not shown in the figure), a batch generation combination module (not shown in the figure), a 3D modeling conversion module (not shown in the figure), and a personalized customization module (not shown in the figure). Among them, the promotion material generation module is used to automatically generate promotion materials for digital assets based on a preset template, where the promotion materials include images, posters, text descriptions, promotional copy, interactive web pages, digital display web pages, and social media promotion; the copyright certification material submission module is used to automatically process the materials and documents required for applying for the copyright certification of digital assets; the batch generation combination module is used to batch generate unique images for creating exclusive digital assets for short micro audio-visual files; the 3D modeling conversion module is used to convert two-dimensional images into 3D modeling resources through a 3D engine; the personalized customization module is used to adjust the style and format of the promotion materials based on the needs of the enterprise.
[0101] In this alternative embodiment, the blockchain automated deployment module includes: a file compression and upload module (not shown in the figure), a meta-file submission module (not shown in the figure), a smart contract deployment module (not shown in the figure), and a contract matrix generation module (not shown in the figure). Among them, the file compression and upload module is used to compress and upload the original file corresponding to the digital asset to the InterPlanetary File System; the meta-file submission module is used to submit the source file for generating the digital asset to the InterPlanetary File System; the smart contract deployment module is used to automatically and batch-deploy the smart contracts required for the digital asset; the contract matrix generation module is used to monitor the deployment status of all smart contracts and associate the digital assets of multiple micro short audio-visual files to form a smart contract matrix.
[0102] In specific applications, the present invention, in order to match the unique commercial characteristics of "short", "frequent", and "fast" of micro short audio-visual works, enables enterprises to quickly digitalize micro short audio-visual works and provides corresponding digital asset value shaping and market value maintenance operation solutions. Further, through an industrialized production method, it provides feasibility for the enterprise digital asset matrix layout.
[0103] Enterprise users only need to upload micro short audio-visuals to the system background, and then the entire process related to the production, production, distribution, and basic operation of digital assets is completed.
[0104] In order to better understand the above technical solutions of the present invention, the above technical solutions of the present invention will be described in detail from a principle perspective as follows:
[0105] The industrialized micro short drama digital asset production system adopted to solve the technical problem can be divided into four parts: the enterprise content input end, the industrialized task scheduling system, each independent functional module, and the digital asset dashboard. Among them, the enterprise content input end and the digital asset dashboard are the user information input and output interfaces of the system. The core parts are the task scheduling system and each independent functional module. Through the scheduling system, the input information is transmitted to the corresponding module, and the production result is then imported into other modules, so as to iteratively, repeatedly, promote, and automatically complete all digital asset production processes.
[0106] The specific situations of the module functions and technical features are as follows:
[0107] 1. Micro short audio-visual preprocessing and storage module PS
[0108] 1.1) Module functions: Since the video encoding formats uploaded may not be unified, this module can convert video files in different formats into a unified encoding format, which helps to ensure the compatibility and consistency of video files in subsequent processing. Extracting key frames from videos is very useful for subsequent audio-visual analysis and processing (e.g., thumbnail generation, content recognition, etc.). Separating audio and video streams for processing is crucial for applications that only require audio or video. Batch processing of a large number of short audio-visual files: especially for large batches of files that need to be segmented into dozens or hundreds of individual episodes, an effective processing and storage solution is provided.
[0109] 1.2) Technical features: Integrating a multi-process FFmpeg video encoding processor, which is a widely used open-source multimedia processing tool that supports the conversion, processing, and streaming of multiple audio-visual formats. Through multi-process technology, multiple audio-visual source files can be processed simultaneously, significantly improving the processing efficiency. Traditionally, the time ratio for audio-visual processing is 1:1, while by adopting multi-process technology, this ratio can be increased to 1:60, that is, it only takes 1 second to process 60 seconds of audio-visual resources. The application of multi-process technology greatly improves the utilization rate of server resources, enabling more data to be processed under the same hardware configuration and reducing the hardware cost. The system adopts cloud computing and distributed processing, and is designed to be easily scalable to handle larger-scale audio-visual data processing, file storage, and more complex processing tasks in order to further improve the processing efficiency and reduce costs.
[0110] This module is a comprehensive solution aimed at improving the efficiency of audio-visual processing while ensuring data consistency and quality, providing materials for subsequent processing and analysis.
[0111] 1.3) Implementation steps: Read the header information of the source file uploaded by the user through the program to obtain information such as the original video type, encoding, and resolution; according to the original video encoding, for example, if the original video is in MOV format. Batch processing implementation: When the user uploads the original files in batches, place all the files in the same folder and write a script in the folder. After automatic execution, all files will complete operations such as unified encoding and compression.
[0112] 2. Automatic Speech Recognition Module ASR
[0113] 2.1) Module functions: Audio signal processing. The ASR module first processes the original speech signal to remove noise and standardize the speech data, that is, preprocess the original speech signal to reduce noise and improve the effect of subsequent processing. Then extract key features from the processed speech signal, which reflect important attributes of the speech, such as pitch, volume, and intonation. Extract features useful for speech recognition from the processed speech signal, such as Mel Frequency Cepstral Coefficients (MFCC).
[0114] Among them, the present invention adopts Mel scale (converting frequency into a scale consistent with human ear perception) and cepstrum analysis (the inverse Fourier transform of the logarithm of the spectrum of a signal), etc.
[0115] The Mel scale is a non-linear frequency measure based on the human ear's perception of pitch (tone). It is designed to better simulate the perception of sounds with different frequencies by the human auditory system. Cepstrum analysis has a wide range of applications in fields such as speech recognition, speaker recognition, and music analysis. Especially in speech recognition, cepstral coefficients (such as MFCC) are important features used to represent the characteristics of speech signals. Both the Mel scale and cepstrum analysis are important concepts in speech signal processing, which help the model better understand and process speech signals, thereby improving the accuracy of speech recognition.
[0116] 2.2) Technical features: The ASR module is based on deep learning and uses architectures of recurrent neural network (RNN) and convolutional neural network (CNN) to process time-series data and extract features. The ASR module is pre-trained with a large amount of speech and text data to improve its recognition accuracy. The micro short videos may include multiple languages and dialects, and even some works are produced and distributed in different regions and countries. The digital assetization system can adapt to the localization of languages, aesthetic standards, etc. in these regions, which is crucial for global applications. The source video of the micro short video is the finished video after shooting and editing. In order to maintain high recognition accuracy in various scene environments, the ASR module also adds special noise filtering and echo cancellation technologies.
[0117] 2.3) Implementation method: The ASR module uses machine learning algorithms to convert the extracted features into text, recognize words and phrases and their order, for converting the extracted features into text information. Based on deep learning theory, especially models for processing sequence data such as recurrent neural network, long short-term memory network, gated recurrent unit, etc. There is also a combination of an acoustic model (mapping audio signals to phonemes or vocabulary) and a language model (predicting the probability of a given word sequence). Then, through the understanding of the context of the conversation, the recognition accuracy is improved, especially in the case of homophones or complex grammar structures. It is used to understand the context information in speech and improve the recognition accuracy, based on natural language processing theory, including word embedding, context embedding, syntactic analysis, semantic analysis, etc. Finally, a line text with time stamps is generated to ensure the accurate transmission of information and the subsequent module to track specific content. The recognized text is aligned with the time in the original speech signal, which is convenient for subsequent processing or display. It is necessary to use time alignment algorithms, such as dynamic time warping, hidden Markov model, etc., to align the recognized text sequence with the time sequence in the original speech signal.
[0118] 3. Keyword extraction module KI
[0119] 3.1) Module Function: Process the content of the dialogue script file containing time information to remove irrelevant information; automatically identify and extract keywords or phrases from the dialogue script text, which can accurately reflect the core theme or key points of the current episode of the micro short video; the deep learning model can obtain the processed dialogue scripts of all video episodes of the micro short video, understand the context of the episodes, so as to improve the accuracy and relevance of keyword extraction; additionally attach a professional domain knowledge model to enable the module to effectively process long texts with complex structures and semantics, such as industry-specific vocabulary, academic papers, reports or articles involved in micro short videos; perform keyword extraction for texts in different languages to adapt to a multilingual environment; as more micro short video data is input, the model continues to learn and adapt, thereby continuously improving its performance.
[0120] 3.2) Technical Features: Recurrent neural networks, including long short-term memory networks (LSTM) and gated recurrent units, can effectively process sequential data and capture long-distance dependencies; convolutional neural networks, mainly used to capture local features in dialogue text, especially in dialogue text classification and keyword extraction; Transformer architecture, which improves the ability to understand the context of dialogue text through self-attention mechanism; word embedding technology, which converts words into dense vectors to capture semantic and syntactic relationships; attention mechanism, which helps the model focus on the key parts of dialogue text and improve the accuracy of keyword extraction; transfer learning, which uses a model pre-trained on a large dataset and adapts to a specific keyword extraction task through fine-tuning; data preprocessing and optimization, including word segmentation, removal of time markers, stop words, part-of-speech tagging, etc., to optimize the quality of input data.
[0121] 3.3) Implementation method: Before the keyword extraction module KI starts extracting keywords, it needs to preprocess the dialogue script. This includes text cleaning, deleting useless symbols, numbers, and stop words (such as common but uninformative words like "of", "and", "in", etc.). Use a word segmentation toolkit to split the sentence into words or lexical units. Finally, restore the words to their basic forms, for example, restore "running" to "run". After preprocessing, features need to be extracted from the text as the input of the model. Use the Word2Vec pre-trained word embedding model to convert the text into a numerical form; select to calculate the TFIDF (term frequency-inverse document frequency) of the words as features. Then select the deep learning model LSTM. Select the dialogue script that requires labeled data, that is, the dialogue script where it is known which words are keywords, and then use the training data to train the model. Apply the trained model to the dialogue script to predict whether each word is a keyword. According to the prediction output of the model, extract the words marked as keywords. Sort according to the keyword importance scores given by the model, select the most relevant words as the final keywords, and then merge the keywords with similar meanings.
[0122] 4. Natural Language Generation Module NLG
[0123] 4.1) Module functions: Story element generation, the extraction module gives keywords to create story elements such as characters, backgrounds, events, etc.; conversion from keywords to stories, the extraction module gives keywords and generates a coherent and attractive story description based on the content of the complete dialogue script (padding words) for the description of digital asset works; conversion from keywords to images, the extraction module gives keywords and generates a PROMPT (image prompt) for constructing images based on the content of the complete dialogue script (padding words) for subsequent generation of digital production images; diverse styles and themes, capable of adjusting the generated content according to preset different style models (such as fantasy, science fiction, history, etc.); adaptability and customizability, if there are additional requirements, it can be customized according to other promotional materials (pictures, photos, videos, etc.) uploaded by the enterprise according to specific requirements, adjusting the length, complexity, and language style of the story, and implanting specified items in the scene, etc.
[0124] 4.2) Technical features: Deep learning model. In addition to the self-built big data model, it can flexibly and independently choose to use large models with the Transformer architecture. The pre-set knowledge base (padding words) model can understand the context and interrelationships of the input keywords to ensure that the generated story content is logically coherent and meaningful. Use advanced algorithms to stimulate creativity and imagination to generate unique and attractive stories. Then, through text analysis, semantic understanding, and language generation, to produce smooth and natural story descriptions. Finally, optimize the quality and relevance of story generation by continuously learning user feedback.
[0125] 4.3) Implementation methods: Story element generation, based on narrative theory and creative thinking. A set of predefined story plot templates are defined, and each template describes a specific type of story structure. Define the characters in the story and their attributes, goals, and behaviors, and then generate the story plot based on the interactions of these characters. Keyword-based story generation involves information retrieval and knowledge representation. Map keywords to story elements (such as characters, events, locations), and then construct a story based on these elements. Use keywords to fill in the predefined story templates to generate stories related to the keywords. Converting keywords into visual images involves computer vision and image generation. Use GANs to generate images based on keywords, and then retrieve images related to the keywords from the image database. Use deep learning models to learn the conversion between different text styles, and customize the content according to the user's preferences and historical behaviors. At the same time, reinforcement learning methods can be used to enable the system to continuously adjust and optimize the generation strategy according to user feedback.
[0126] 5. Text-to-Image Generation Module TIG
[0127] 5.1) Module functions: Automatically generate corresponding images according to the digital asset picture PROMPT; Select the key frame screenshot corresponding to the digital asset picture PROMPT as the reference image, and generate based on the content, style, or elements of this image; Allow enterprises to specify picture format materials, extract specific content or styles from them, and implant them into the newly generated images, which is suitable for commercial promotion; Generate pictures in various forms such as cards, drawings, posters, etc. according to the preset selection to adapt to different application scenarios; Be able to process continuous text content to ensure that the generated images are relevant and consistent with the text content.
[0128] 5.2) Technical features: Deep learning and generative adversarial networks (GANs), use deep learning models, especially GANs, to generate high-quality images. Use style transfer technology in the reference image generation and content implantation functions to apply specific visual styles or elements to the new images. Combine the processing of text and visual information to better understand and generate images that meet the requirements. Provide a variety of customization options, such as image types, styles, formats, etc., to meet the needs of different users. Most popular models can be selected, and formats such as SD are compatible. The style direction of the generated images can be flexibly controlled.
[0129] 5.3) Implementation methods: First, a generative adversarial network, which consists of a generator and a discriminator. The generator attempts to generate realistic images, while the discriminator attempts to distinguish between real images and generated images. Through this adversarial process, the generator learns to generate more and more realistic images. In text-to-image generation, the generator usually receives the embedded representation of the text description as input and generates an image that matches the description. The discriminator then evaluates whether the generated image conforms to the text description.
[0130] In text-to-image generation, the text description is used as a condition to guide the generator to generate an image that matches the description. At the same time, the discriminator not only evaluates the authenticity of the image but also assesses whether the image meets the text condition. In text-to-image generation, RNNs (especially LSTM or GRU) can be used to process the text description, extract the key information in the text, and convert it into an embedding representation that can be used to guide image generation. The introduction of the attention mechanism allows the model to focus on the most relevant parts when processing the input, thereby improving the performance and interpretability of the model. In text-to-image generation, the attention mechanism can be used to make the generator focus on the words in the text description that are most relevant to the generated image part, thereby improving the detail quality of the image and its correspondence with the text. To ensure that the generated image is semantically consistent with the text description, a semantic consistency loss can be introduced as part of the training objective.
[0131] Generally speaking, text-to-image generation involves theories and technologies in multiple fields, including deep learning (especially generative adversarial networks), natural language processing, computer vision, etc. By combining these theories and methods, an artificial intelligence system can convert a text description into a corresponding image.
[0132] 6. Digital Asset Editing Module (Digital Asset Editor) SGE
[0133] 6.1) Module Functions: Customize the type of smart contract corresponding to each asset; customize the minting quantity of each asset; customize the activation date of the smart contract for each asset; source files required for customization requirements; flexibly preset multiple solutions through templates; provide blockchain configuration information for the production of other modules: what kind of contract to use, MINT quantity, relationships between contracts, contract timed activation configuration, automated release configuration of promotional materials, etc.
[0134] Example: If assets A, B, and C are configured for first-level issuance, the upper limit of the "asset points" smart contract for minting needs to be controlled at 900 copies. Assets E, D, F, G, and H are obtained for free from A, B, and C in the plot. The upper limit of the "asset blind box" smart contract for minting needs to be controlled at 12,000 copies, and the corresponding relationship between E, D, F, G, and H and the airdrop of holding A, B, and C is set. At the same time, assets I, J, and K require three to four of A, B, C, D, E, F, G, and H to participate in the synthesis in the first to third stages. The upper limit of the "asset synthesis" smart contract for minting needs to be controlled to be consistent with the lowest value of the publicized assets participating in the synthesis, and so on. At the same time, at what time point should the assets be released and activated through channels (it is necessary to automatically calculate the quantity and time nodes of each channel to avoid conflicts: such as the asset has been released but the contract has not been activated and cannot be used, and the total quantity exceeds the total minting quantity after multiple channels are released, resulting in some channels being unable to mint, etc.). The above information will be provided to other modules as important configuration information, including the quantity of assets to be generated, the contract types corresponding to the assets, the minting quantity of each asset, the associations between contracts, and so on.
[0135] 6.2) Technical features: It includes multiple self-developed smart contracts such as "asset points", "asset synthesis", "asset blind box", "asset points auction", and "asset points consumption", which can meet various needs of digital asset empowerment and can configure various strategies suitable for the operation of digital assets in short videos to achieve the purpose of shaping and enhancing the value of digital assets. Using the NLP engine, the contract type, minting quantity, release date, etc. are automatically set randomly according to previous strategies. For example, when modifying the minting quantity of a certain digital asset, the minting quantity, release date, type, etc. of other associated digital collections will be automatically changed. This avoids a large amount of calculation work when setting strategies.
[0136] 6.3) Implementation method: A smart contract is a self-executing piece of code on the blockchain used to define and enforce the terms of a contract. Different types of assets may require different types of smart contracts to manage their specific rules and behaviors. It is necessary to pre-define different types of smart contract templates, such as "asset points", "asset mystery boxes", "asset synthesis", etc. Allow users to select the contract type and provide necessary parameters (such as the minting limit) to customize and deploy the contract. Set a limit on the minting quantity in the smart contract to ensure that the minting quantity of the asset does not exceed the preset limit. And provide an interface for users to input the minting quantity of each asset and apply these settings when deploying the contract. Implement a timer mechanism in the smart contract to make the contract automatically activate on a specified date (the contract activation date refers to the time point when the smart contract starts to take effect. Setting the activation date can control the release time and availability of the asset). Then, customize the source files required by the requirements, provide an interface for users to upload the source files of the assets, and store and manage the uploaded source files through a file management system. By means of pre-defining a series of configuration templates covering different blockchain configurations and asset types, provide an interface for users to select templates and fill in specific parameters, such as contract type, minting quantity, etc.
[0137] Specifically, the following technology stacks and methods are used to implement the above functions: Front-end interface, use the front-end framework UNIAPP to build the user interface, and provide forms and input boxes for users to customize asset and contract settings. Smart contract development, use the smart contract language (Solidity) to write different types of contract templates and deploy the contract according to the parameters input by users. Blockchain interaction, use a blockchain client library (such as web3.js) to interact with the blockchain, deploy and activate smart contracts. Back-end service, build a back-end service to handle user requests, manage file uploads and storage, and coordinate blockchain interactions. The implementation of the digital asset editor module involves multiple fields such as blockchain technology, smart contract development, and front-end interface design, and these technologies and methods need to be comprehensively applied to meet the customization needs of users.
[0138] 7. Material Production Module MPP
[0139] 7.1) Module Functions: Promotion Material Generation: Automatically create various materials for NFT promotion according to preset templates, including but not limited to: images and posters; text descriptions and promotional copy; interactive web pages and digital display HTML; social media promotion content. Submission of Copyright Certification Materials: Automatically process the materials and documents required for copyright certification applications (usually a combined large file of all generated image materials), so that digital assets are legally protected. Batch Image Generation and Element Combination: Combine multiple elements (such as picture elements, text descriptions, etc.) to batch generate unique images for creating digital assets exclusive to micro short videos. 3D Modeling Conversion: When the enterprise needs it, convert 2D pictures into 3D modeling resources through the integrated 3D engine API for augmented reality (AR) or virtual reality (VR) applications. Personalized Customization: Allow enterprises to adjust the style and format of the generated materials according to specific needs based on preset templates.
[0140] 7.2) Technical Features: Automatically generate a large number of asset images based on the deliverables of other modules. Generate OBJ modeling files that are generally supported by 3D engines through code, facilitating the extended development of subsequent digital asset AR / VR applications and enhancing the asset value. Preset relevant material generation templates according to the general digital asset operation plan, and the module will generate the materials required for relevant operations based on the templates in combination with micro short drama frame screenshots, asset image elements, etc.
[0141] 7.3) Implementation Method: All generated resources will be automatically uploaded to IPFS for storage. IPFS (InterPlanetary File System). It is a peer-to-peer distributed file system designed to connect all computing devices to build a unified file system.
[0142] 8. Blockchain Automated Deployment Module BAD
[0143] 8.1) Module Functions: Batch compress and upload basic files such as images, audio and video, and code corresponding to digital assets to IPFS (InterPlanetary File System). Generate meta files of digital assets according to the configuration and submit them to IPFS. Automatically batch deploy the smart contracts required for digital assets according to the selected network (private / consortium chain, Ethereum mainnet, Polygon, BSC, etc.). Monitor all contract deployment situations and call the owner method to associate multiple micro short video digital assets to form a smart contract matrix.
[0144] 8.2) Technical Features: Tightly integrated with IPFS (InterPlanetary File System) to ensure that all uploaded digital asset files (such as images, audio / video, code) are stored in a decentralized manner, improving data security and persistence, and ensuring that users' digital assets still exist permanently under any circumstances (including when the operating entity no longer conducts operation and maintenance work). The module can automatically generate the meta-file (META) of digital assets according to the preset configuration and automatically submit it to IPFS, greatly simplifying the digital asset creation process. It supports multiple blockchain networks, including private / consortium blockchains, Ethereum mainnet, Polygon, BSC, etc., and can automatically deploy the smart contracts required for digital assets according to the user's selection. It can monitor the deployment status of all smart contracts in real time and can associate multiple digital assets by calling the owner method to form a smart contract matrix, improving the flexibility and efficiency of contract management. By associating multiple micro short video digital assets, a comprehensive smart contract matrix is formed to provide a more complex and functional digital asset management solution.
[0145] 9. Scheduled Task Distribution Module (Scheduled Task Dispatcher) STD
[0146] 9.1) Functional Modules: Integrate media channels, which can be enabled by the account information pre-filled by enterprises, and automatically publish relevant materials (graphic texts, audio / video, etc.) according to the scheduled time; integrate the asset issuance platform, and automatically publish the smart contract address and activate the corresponding smart contract in coordination with the operation strategy.
[0147] 9.2) Technical Features: Multi-platform integration ability, the ability to integrate with a variety of social media and network platforms, supporting cross-platform content publishing and management. Automated content publishing, which can automatically publish content according to the schedule preset by enterprises, including various formats of media materials such as graphic texts, audio / video, etc., effectively realizing the scheduled release of content. Account management and security, providing secure account management functions, allowing enterprises to safely fill in and store social media account information, and ensuring the security of account data. Integration with the asset issuance platform, integrating with various asset issuance platforms (such as digital asset markets), supporting the automatic release and activation of smart contracts, and facilitating coordination with the operation strategy of enterprises. Smart contract management, automatically managing the release and activation process of smart contracts, ensuring that the contracts go live on time and are synchronized with the asset issuance strategy. Flexible scheduling and planning functions, providing a user-friendly interface that enables users to easily set and adjust the publishing plan, including time, content, and target platforms. Efficient task processing mechanism, adopting advanced task scheduling algorithms to ensure the timely and efficient execution of scheduled tasks.
[0148] In summary, the simplified flowchart of the present invention is as Figure 2 shown, including the following steps:
[0149] STEP1. Content Input: The enterprise inputs short video and audio content into the system.
[0150] STEP2. Social Gamification Editor: Confirm digital asset production information.
[0151] STEP3. Preprocessing and Storage: Uniformly encode, compress, and extract key frames from the uploaded content.
[0152] STEP4. Speech Recognition: Extract audio from the video and convert it into text.
[0153] STEP5. Keyword Extraction: Extract keywords from the text to highlight the theme of the video content.
[0154] STEP6. Text-to-Image Generation: Generate corresponding images based on the extracted keywords or text content.
[0155] STEP7. Material Production: Generate materials for promotion, such as images, posters, and promotional copy.
[0156] STEP8. Blockchain Deployment: Upload digital assets and related information to the blockchain.
[0157] STEP9. Content Distribution: Publish and promote digital assets through various channels.
[0158] After the user confirmation in STEP2, STEP3 - 9 are automatically completed without the need for enterprise users to operate again.
[0159] For the background interface of short video information entry, the enterprise can enter the name, pre-release date, and upload the packaged short video source file at the terminal.
[0160] For the confirmation page of short video media information release plan, the system automatically applies the prefabricated plan to generate a media channel release plan for the short video.
[0161] For the short video digital asset editor, the system automatically applies the prefabricated plan to generate a digital asset configuration plan for the short video. After user confirmation, the material production pipeline can be automatically delivered according to the configuration for material production and digital asset deployment.
[0162] As Figure 3 shown in the information flow diagram of each module, the system architecture and module functions can be summarized as follows:
[0163] 1. Enterprise Content Input End: Enterprise users upload short video and audio content through this interface, such as pre-produced episode clips, including but not limited to video files, audio clips, or related script texts.
[0164] 2. Preprocessing and Storage Module (PS): This module automatically encodes and compresses the uploaded videos for unified format to optimize storage and subsequent processing. It also extracts key frames and separates audio and video to provide the necessary raw data for subsequent processing steps.
[0165] 3. Automatic Speech Recognition Module (ASR): Using advanced deep learning techniques, this module converts the audio content in the video into text and marks the time information, providing a basis for keyword extraction and content understanding.
[0166] 4. Keyword Extraction Module (KI): Based on the text provided by the ASR module, this module automatically identifies and extracts keywords or phrases to capture the core content of each short video episode.
[0167] 5. Natural Language Generation Module (NLG): According to the keywords extracted by the KI module, the NLG module generates attractive story descriptions or plot introductions for the copywriting of digital assets.
[0168] 6. Text-to-Image Generation Module (TIG): This module converts the story descriptions generated by the NLG module into specific images to provide a visual presentation for digital assets.
[0169] 7. Smart Digital Asset Editing Module (SGE): At this stage, enterprise users can customize the type, quantity, and activation date of the smart contracts of digital assets to increase the interactivity and collectible value of digital assets.
[0170] 8. Integrated Production Platform Module (IPP): Automatically creates various materials for NFT promotion, including images, posters, promotional copy, etc., and processes copyright certification materials.
[0171] 9. Blockchain Automatic Deployment Module (BAD): This module is responsible for uploading the generated digital assets to blockchain networks such as Ethereum or BSC, etc., and automatically deploys the corresponding smart contracts.
[0172] 10. Scheduled Task Distribution Module (STD): Used to automatically publish promotional materials on major social media and digital asset platforms at a predetermined time.
[0173] The following is a specific embodiment:
[0174] Suppose an enterprise wants to convert its new series of short micro videos into digital assets for promotion and sales. The enterprise can access the system page of short micro video management, input basic information such as "short micro video name" and "expected release time", and upload the packaged audio and video source files. After the upload is completed, the system will generate a digital asset production and operation plan using the default policy, and the planned arrangements can be viewed and modified through the provided preview function. After confirmation, the task scheduling system will automatically execute the following processes: The PS module preprocesses the video, including encoding unification and key frame extraction. The ASR module then converts the audio into text and marks the time. The KI module extracts keywords from the text, and the NLG module generates attractive story descriptions based on these keywords. The TIG module converts these descriptions into images. The SGE module allows the enterprise to customize the relevant parameters of the smart contract. The IPP module generates all the required promotion materials. Finally, the BAD module uploads these assets to the blockchain, and the STD module releases the promotion content at the scheduled time. According to the general situation of the short video series, after 2 to 3 hours, the digital assets of the entire short micro video series are completed, and the full network deployment and basic operation work are automatically executed.
[0175] Through this implementation method, the enterprise can quickly, efficiently, and batch convert short micro drama series into attractive digital assets, while reducing manual participation and improving production efficiency. The implementation of this system not only improves the market value of the content but also provides the enterprise with new business models and revenue sources.
[0176] Figure 4 An embodiment of an industrialized short micro drama digital asset production method of the present invention is shown.
[0177] In this alternative embodiment, the industrialized short micro drama digital asset production method includes:
[0178] Step S701: The enterprise user uploads short micro video files through the input interface;
[0179] Step S703: Transmit the short micro video files to the independent functional modules, and call the corresponding independent functional modules to execute independent short micro drama processing tasks;
[0180] Step S705: Execute the processing tasks of the short micro video files, and based on the task processing results, achieve the full-process industrialized production of digital assets from generation, production, distribution to basic operation;
[0181] Step S707: Based on the information output interface, display the output information of the digital assets.
[0182] In one embodiment, a computer device is provided. This computer device can be a server, and its internal structure diagram can be as Figure 5As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store static information and dynamic information data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0183] Those skilled in the art can understand that Figure 5 the structure shown in is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0184] In addition, the present invention also provides a computer device, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, it implements the steps in the above method embodiments.
[0185] Furthermore, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0186] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory, magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory or dynamic random access memory, etc.
[0187] The present invention is not limited to the structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. An industrialized short drama digital asset production system, characterized in that: include: Enterprise content input terminal, used for enterprise users to upload short audio and video files through the input interface; Industrialized task scheduling system, used to transmit short audio and video files to independent functional modules, and call corresponding independent functional modules to perform independent short drama processing tasks; Independent functional module, used to perform the processing task of short audio and video files, and based on the task processing results, realize the full process industrial production of digital assets from generation, production, distribution to basic operation; Digital asset dashboard, used to display the output information of digital assets based on the information output interface; Wherein, the independent functional modules include: A micro-short audio and video preprocessing storage module is used to convert micro-short audio and video files of different formats into a unified encoding format, and perform key frame extraction and audio and video separation of the micro-short audio and video files; An automatic speech recognition module is used to extract the speech features of the audio content in the short audio and video file, convert the speech features into text data in text form based on the speech features, and mark the time information; A keyword extraction module, used to extract keywords from the text data; A natural language generation module, used to generate a description text of the short audio and video file based on the keywords and in combination with a complete script file; A text-to-image generation module, used to convert the description text into an image; The digital asset editing module is used for enterprise users to customize the personalized information of digital assets, where the personalized information includes the smart contract type, minting quantity and activation date; The material production module is used to create promotional materials for digital assets and process copyright certification materials; A blockchain automated deployment module, used to upload the digital assets to the blockchain network and automatically deploy the corresponding smart contracts; The scheduled task distribution module is used to automatically publish the promotional materials of the digital assets at a preset time. The promotional materials include images, posters, text descriptions, promotional copy, interactive web pages, digital display web pages and social media promotion.
2. The industrialized short drama digital asset production system according to claim 1 is characterized in that: The automatic speech recognition module includes: a speech signal processing module, a speech feature extraction module, a speech text recognition module, a context recognition module and a time marking module, wherein: The speech signal processing module is used to eliminate noise interference in the audio content by using a filter to obtain a standardized speech signal; The speech feature extraction module is used to display the speech attribute of the speech signal by extracting the speech feature in the speech signal, wherein the speech feature is a Mel frequency cepstral coefficient; The speech-to-text recognition module is used to convert the speech features into text data in a text form based on a machine learning algorithm; The context recognition module is used to recognize context information in the speech signal; The time marking module is used to align the text data with the time series of the voice signal based on a time alignment algorithm to generate a dialogue script file containing time information.
3. The industrialized short drama digital asset production system according to claim 2 is characterized in that: The keyword extraction module includes: a text cleaning module, an automatic extraction module, a context understanding module, a complex text processing module, a multi-language adaptation module and a learning adaptation module, wherein: The text cleaning module is used to remove irrelevant information in the script file; The automatic extraction module is used to extract keywords or key phrases in the script file; The context understanding module is used to sort out the context of the story based on the line script file; The complex text processing module is used to identify professional knowledge in the script file based on a professional domain knowledge model, wherein the professional knowledge includes industry professional vocabulary, academic papers, reports and articles; The multi-language adaptation module is used to extract keywords from script files in different languages; The learning adaptation module is used to improve the adaptability of the recurrent neural network model through model learning based on the input number of the micro-short audio and video files.
4. The industrialized short drama digital asset production system according to claim 3 is characterized in that: The natural language generation module includes: a story element generation module, a keyword conversion module, a screen switching module, a style main body preset module and an adaptation customization module, wherein: The story element generation module is used to generate story elements based on the extracted keywords, wherein the story elements include characters, backgrounds and events; The keyword conversion module is used to generate a continuous story description based on the complete content of the line script file; The screen switching module is used to generate picture prompts for constituting story screens based on the complete content of the line script file; The style main body preset module is used to adjust the generated story content based on the preset style model; The adaptive customization module is used to customize the specific needs of the story based on the promotional materials uploaded by the enterprise, and to embed designated objects in the story scenes.
5. The industrialized short drama digital asset production system according to claim 4 is characterized in that: The text-to-image generation module includes: an image generation module, a reference image selection module, a content implantation module, a visual format adaptation module and a continuous text processing module, wherein: The image generation module is used to generate a corresponding image based on the picture prompt; The reference image selection module is used to select the key frame screenshot corresponding to the picture prompt as the reference image; The content implantation module is used to extract specific content or specific style based on the picture format material specified by the enterprise and implant it into the image; The visual format adaptation module is used to pre-configure different visual formats to adapt to different application scenarios; The continuous text processing module is used to process the text content with continuity in the dialogue script file.
6. The industrialized short drama digital asset production system according to claim 5 is characterized in that: The digital asset editing module includes: a smart contract customization module, a casting quantity customization module, an activation date customization module, a source file customization module, a solution configuration module and a blockchain configuration module, wherein: The smart contract customization module is used to obtain the contract type selected by the enterprise user and customize the smart contract type based on the parameters provided by the enterprise user; The casting quantity customization module is used to set the casting quantity limit in the smart contract and customize the casting quantity based on the quantity requirements set by the enterprise user; The activation date customization module is used to obtain the activation date of the smart contract input by the enterprise user and set a timer when deploying the smart contract; The source file customization module is used to obtain the source files of assets uploaded by enterprise users, and store and manage the source files uploaded by enterprise users based on the file management method; The solution configuration module is used to predefine configuration templates covering different blockchain configurations and asset types, and select corresponding templates according to enterprise user needs.
7. The industrialized short drama digital asset production system according to claim 6 is characterized in that: The material production module includes: a promotion material generation module, a copyright certification material submission module, a batch generation combination module, a three-dimensional modeling conversion module and a personalized customization module, wherein: The promotion material generation module is used to automatically generate promotion materials for digital assets based on a preset template; The copyright certification material submission module is used to automatically process the materials and documents required for the application of digital asset copyright certification; The batch generation combination module is used to batch generate unique images for creating exclusive digital assets for micro-short audio and video files; The three-dimensional modeling conversion module is used to convert the two-dimensional image into a three-dimensional modeling resource through a 3D engine; The personalized customization module is used to adjust the style and format of promotional materials based on enterprise needs.
8. The industrialized short drama digital asset production system according to claim 7 is characterized in that: The blockchain automatic deployment module includes: a file compression upload module, a metafile submission module, a smart contract deployment module and a contract matrix generation module, wherein: The file compression and upload module is used to compress the original file corresponding to the digital asset and upload it to the Interstellar File System; The metafile submission module is used to submit the source file for generating digital assets to the Interstellar File System; The smart contract deployment module is used to automatically deploy the smart contracts required for digital assets in batches; The contract matrix generation module is used to monitor the deployment status of all smart contracts and associate the digital assets of multiple short audio and video files to form a smart contract matrix.
9. An industrialized short drama digital asset production method, using the industrialized short drama digital asset production system according to any one of claims 1 to 8, characterized in that: The method includes: Enterprise users upload short audio and video files through the input interface; Transfer the short audio and video files to the independent function module, and call the corresponding independent function module to perform independent short drama processing tasks; Execute the processing task of short audio and video files, and based on the task processing results, realize the full process industrial production of digital assets from generation, production, distribution to basic operations; Based on the information output interface, the output information of digital assets is displayed.
Citation Information
Patent Citations
All-media content production means management platform applied to financial system
CN115665108A