Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

150 results about "Content production" patented technology

Digital human video generation method based on multi-modal large model

The invention belongs to the technical field of virtual person generation, and particularly relates to a digital person video generation method based on a multi-modal large model, and the method comprises the following steps: 1, constructing a multi-modal data system; 2, multi-modal large model training and adaptation are carried out; 3, constructing a digital human three-dimensional model; step 4, performing semantic analysis and modal mapping; 5, generating a time sequence action and a mouth shape; step 6, building and rendering a virtual scene; step 7, audio and video synchronous rendering and synthesis; step 8, quality optimization and defect repair; and step 9, performing user interaction and iterative optimization. Through technical innovation and engineering, the core pain point in digital human video generation is solved, efficient, vivid and customizable content production capacity is provided for virtual anchors, intelligent customer service, enterprise training and other scenes, and the AI digital human technology is promoted to be applied to large-scale business from experiments.
Owner:ZHE JIANG YAN HUANG KE JI YOU XIAN GONG SI

AIGC content generation method and system based on multi-modal fusion

The invention relates to the technical field of AIGC content generation, discloses an AIGC content generation method and system based on multi-modal fusion, and aims to solve the problems of decentralization, low efficiency and insufficient originality of a traditional content generation tool. Multi-modal data such as texts, images, videos and audios are integrated, user intentions are analyzed in combination with intelligent retrieval and a domain knowledge base, automatic generation from multi-modal input to high-quality creative content is achieved, a cross-modal collaborative generation technology is adopted, semantic features are dynamically aligned, and logically coherent content is generated. The content emotional value is enhanced through an emotional analysis and dynamic optimization strategy, the homogenization bottleneck is broken through, meanwhile, an automatic quality evaluation and format adaptation mechanism is integrated, deep application of scenes such as text travel, advertisement, e-commerce and interactive network television service is supported, marketing copywriting, short videos and cross-platform distribution schemes can be efficiently generated, and the market competitiveness is improved. And the content production efficiency and the creativity transmission are obviously improved.
Owner:HANGZHOU WANDIAN TECHNOLOGY CO LTD

Video generation control method and computer readable storage medium

The invention provides a video generation control method and a computer readable storage medium, and is applied to the technical field of computers. According to the method, the structural defect of a split script is solved through a large split model, the consistency of split image features is ensured through image generation based on reference image feature constraint, audio and picture synchronization is realized in combination with emotional speech synthesis and dynamic mouth shape alignment, the video fidelity is improved through a split video cue word guiding multi-subject generation algorithm, and the user experience is improved. Therefore, the technical problems of missing emotion logic, poor consistency of split image features, difficulty in audio synchronization, insufficient controllability of video generation and the like in film and television creation in the existing AIGC technology are effectively solved, and the efficiency and the quality of video content production are improved.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

Video and voice automatic translation method based on pre-training model

The invention belongs to the technical field of speech translation, and particularly relates to a video speech automatic translation method based on a pre-training model, and the method comprises the following steps: 1, preprocessing video and audio data; step 2, voice recognition and language detection; step 3, machine translation and text post-processing; step 4, speech synthesis and audio mixing; step 5, synchronizing video processing and subtitles; step 6, quality control and multi-dimensional evaluation; 7, carrying out model iteration and data closed loop; and step 8, system deployment and engineering implementation. Through deep fusion of efficient transfer learning of the pre-training model and the multi-modal technology, a high-precision, low-cost and easy-to-expand video speech translation solution is constructed, the time and labor cost of globalized content production is greatly reduced, the cross-language communication efficiency is improved, immersive multi-language experience is provided, and the method is suitable for popularization and application. And a data-driven continuous optimization mechanism is established, so that the system performance is improved along with the increase of the use scale.
Owner:ZHE JIANG YAN HUANG KE JI YOU XIAN GONG SI

Automatic lecturer video generation method based on AI speech synthesis and animation driving

The invention discloses a lecturer video automatic generation method based on AI speech synthesis and animation driving. The method comprises the following steps: performing structured analysis on a PPT or a text script through an improved interior point method and an incremental shortest path algorithm; performing semantic grouping by applying a full-dynamic parallel single-link clustering algorithm and generating an enhanced script with an expressive mark; a CosyVoice technology is combined with a low-rank approximation method to generate a high-quality voice data stream; establishing a mapping relation between contents and action expressions through semantic analysis, and generating a complete action expression instruction set; and driving the digital human model by using the msueTalk technology, and generating a final lecturer teaching video through a parallel rendering algorithm. According to the invention, the method achieves the efficient and automatic generation of the education video, remarkably improves the content production efficiency, reduces the production cost, and guarantees the specialty and expressive force of the teaching video.
Owner:SHENZHEN XUEYOU TECHNOLOGY CO LTD

Multi-mode convergence media content auxiliary creation method based on AI technology

The invention discloses a multi-mode convergence media content auxiliary creation method based on an AI technology, and the method comprises the steps: extracting key features according to a creation demand text inputted by a user, carrying out the retrieval in a knowledge base through employing a knowledge graph retrieval mode according to the key features, obtaining a creation material, and generating a first draft; carrying out cross-modal feature mapping on the first draft by adopting a multi-modal alignment model, aligning text-image-video embedded vectors through contrastive learning, and dynamically adjusting the correlation of multi-modal contents by utilizing an attention mechanism to obtain a multi-modal content packet; duplicate checking is carried out on the multi-modal content packet in multiple modes, and the multi-modal content packet after duplicate checking is sent to a user for manual editing. Through multi-mode processing methods such as AI auxiliary writing, AI illustration and video generation and whole-network duplicate checking and propagation value evaluation, the content production efficiency and quality are improved, the propagation effect is optimized, and the original content is protected.
Owner:广西日报社

House content generation method and device based on artificial intelligence, and readable storage medium

The embodiment of the invention discloses a house image content generation method based on artificial intelligence. The method comprises the steps of obtaining a basic image, wherein the basic image is an image shot in a house; processing the basic image based on a pre-trained cue word acquisition model to obtain a basic cue word; inputting the basic cue word and a preset cue word into an image generator to obtain a target image which is a realistic image; generating matched text content based on the target image; and the matched text content and the target image are associated and published to a target platform. Based on the scheme, on one hand, through full-process automatic design, links from image acquisition to publishing are linked, manual intervention is remarkably reduced, the content generation efficiency is improved, and the method is suitable for large-scale content production scenes of decoration companies, self-media and the like; on the other hand, by means of cue word design, it is ensured that the generated target image is highly close to a real house scene, meanwhile, the image-text collaborative generation mechanism guarantees consistency of content themes, and credibility and transmissibility are enhanced.
Owner:SHENZHEN BINCENT TECH

Digital human generation and streaming transmission method and system

The invention relates to the technical field of image generation, and particularly provides a digital human generation and streaming transmission method and system, and the method and system are based on an ER-NERF algorithm, and comprise the following steps: S1, carrying out the depth estimation and optical radiation field reconstruction of multi-view image data, so as to generate a realistic 2.5 D digital human model; s2, in a data transmission stage, the system pushes the preprocessed model parameters and texture information to a user terminal in a streaming media format in real time; and S3, at a user terminal, the system performs real-time analysis and reconstruction on the received data stream through a decoding module, and dynamic rendering and interactive display of the digital human are completed. Compared with the prior art, the method has the advantages that the terminal computing pressure can be reduced, cross-platform real-time interaction and efficient rendering are ensured, and the application requirements of virtual content production, remote interaction and immersive experience are met.
Owner:浪潮智慧城市科技有限公司 +1

Digital human live broadcast method and system

The invention relates to a digital human live broadcast method and system. The method comprises the following steps: acquiring a live broadcast room visual effect picture and a digital human live broadcast audio; generating digital human live broadcast content under the driving of the digital human live broadcast audio by taking the visual effect picture of the live broadcast room as a visual manuscript through an I2V model; in the process of playing the digital human live broadcast content, if the current interaction behavior of the user in the live broadcast room is monitored, generating digital human interaction content responding to the current interaction behavior; selecting an insertion frame from the key frames of the digital human live broadcast content so as to insert the digital human interaction content at the position of the insertion frame; the key frame is a video frame corresponding to a statement demarcation point in the digital human live broadcast content. According to the scheme provided by the invention, the contradiction between long content production period, high implementation cost and slow real-time interaction feedback in digital human live broadcast in related technologies can be solved.
Owner:HANGZHOU TEKAN TECHNOLOGY CO LTD

User interfaces for color and lighting adjustments for an immersive content production system

In some implementations, a computing device in communication with an immersive content generation system may generate a first set of user interface elements configured to receive a first selection of a shape of a virtual stage light. In addition, the device may generate a second set of user interface elements configured to receive a second selection of an image for the virtual stage light. Also, the device may generate a third set of user interface elements configured to receive a third selection of a position and an orientation of the virtual stage light. Further, the generate a fourth set of user interface elements configured to receive a fourth selection of a color for the virtual stage light. Numerous other aspects are described.
Owner:LUCASFILM ENTERTAINMENT COMPANY LTD

System and method for streamlining content production using artificial intelligence

A system and method for appraising film screenplays and offering predictive insights on their marketability and potential success. The artificial intelligence model and machine learning algorithms consider various production elements and generate management and production outcomes on an expedited timeline. The web services infrastructure implemented creates an enhanced ecosystem for creative development, production, and talent management by minimizing obstacles such as the inaccessibility of information and data about both national and global development trends. Moreover, the system can summarize creative data and create loglines, character breakdowns, titles, and summaries. An A.I generated a sociological analysis that provides a multi-faceted and comprehensive outlook on industry-related research, such as a cost-benefit production analysis.
Owner:WOOD ORLANDO +2

Protection method and device for medical big language model

The invention provides a medical large language model protection method and device, and relates to the technical field of artificial intelligence, and the method comprises the steps: carrying out the sensitive word detection and intention detection of user input; under the condition that the user input does not contain sensitive words and has no bad intention, whether the user input is related to the medical field or not is judged based on an AI model; and under the condition that the user input is related to the medical field, inputting the user input into the medical big language model for content production. According to the invention, the specialty and effectiveness of the medical big language model are improved.
Owner:Artificial Intelligence and Robotics Innovation Center of Hong Kong Institute of Innovation, Chinese Academy of Sciences

Immersive content production and simulation system for motion vehicle

The invention discloses an immersive content production and simulation system for a motion vehicle, and aims to solve the problems that virtual vision and physical motion are difficult to match, the content production cost is high and the interactivity is insufficient. According to the system, a six-degree-of-freedom track of a target is planned through a content making simulation operation end, and a moving carrier is driven to execute; and meanwhile, real motion data of the carrier is collected, and comparison and iterative correction are carried out on the real motion data and the target trajectory until errors converge. The system also uses artificial intelligence to generate medium-long-shot and close-shot three-dimensional asset fusion. Through closed-loop correction driven by physical data, high-precision synchronization of motion and pictures is realized, the development period is remarkably shortened, the content cost is reduced, and the personalized interactive experience is improved.
Owner:THE BEST SYNC ADVERTISING CO LTD GUANGZHOU

Text-to-action generation method and system based on fine-grained representation of body parts

The invention discloses a text-to-action generation method and system based on body part fine-grained representation, and belongs to the field of human body action generation. Body part-level fine-grained discretization modeling is performed on human body actions to obtain an action encoder and an action decoder which can stably reconstruct a continuous action sequence; the method comprises the following steps: taking a part fine-grained discrete action space as a producible space, introducing video potential representation as space-time dynamic priori, establishing an alignment modeling mechanism of the video potential representation and the action discrete space, and training to obtain a conditional generation network, so as to realize stable prediction and iterative completion of an action discrete index sequence in a reasoning stage; and a continuous human body action sequence is reconstructed through an action decoder, and finally a human body action sequence which is consistent with text semantics, richer in details and more coherent in time sequence is generated. According to the method, the technical threshold and data dependence of action content production can be reduced, and the controllability and generalization ability of action generation are improved.
Owner:ZHEJIANG UNIV

Video image processing method for reducing 3D content manufacturing cost with assistance of AI

The invention relates to the technical field of computer vision and three-dimensional modeling processing, and discloses an AI-assisted video image processing method for reducing 3D content manufacturing cost, which comprises the following steps: S1, acquiring a target video sequence, and carrying out image enhancement processing and frame sampling on the video sequence to obtain an image frame sequence; s2, performing multi-model semantic recognition processing on each image frame to obtain a plurality of semantic regions and corresponding semantic tags in the image frames; and S3, performing semantic weight scoring on the semantic region, and dividing the semantic region into a high-priority region, a middle-priority region and a low-priority region according to a scoring result. According to the method, content in the video image is subjected to classification processing, fine modeling of a high-semantic region and rapid generation of a low-semantic region by introducing a semantic recognition and region priority modeling mechanism, and through the design, invalid calculation tasks can be reduced on the basis of keeping the integrity of key content.
Owner:ZHENGZHOU UNIV

Artificial intelligence content production real-time risk control method and apparatus

The present application provides an artificial intelligence content production real-time risk control method and apparatus. The method comprises: step S1, receiving input data needing to be used for synthesizing a digital avatar video; step S2, determining the security and compliance of the input data according to a preset risk control requirement, and if the input data is secure and compliant, releasing a risk control lock applied to a digital avatar model; step S3, synthesizing the digital avatar video on the basis of the input data, and by means of the digital avatar model, performing compliance approval on the synthesized video located in a cache; and step S4, performing disk storage on a final video that has passed compliance approval. In the present application, real-time risk control management can be performed on video content generated by digital avatars, thereby preventing the malicious use of the digital avatars.
Owner:BEIJING FENGPING INTELLIGENT TECHNOLOGY CO LTD

Intelligent digital human system based on multi-modal deep learning and automatic content generation

The invention relates to the technical field of digital people, in particular to an intelligent digital people system based on multi-modal deep learning and automatic content generation, which comprises a front-end interaction module, a rear-end processing module, a multi-modal data management module, an AI synthesis engine and a distributed storage module. The system realizes full-process intelligent operation from model training to content publishing through dynamic voiceprint extraction, audio-driven video generation and multi-modal data fusion. According to the method, the defects in personalized training, automatic content generation and multi-user cooperation management in the prior art can be overcome, the naturalness, flexibility and user experience of voice synthesis and video driving are remarkably improved, and an efficient and intelligent solution is provided for digital human content production.
Owner:奚澜卜

Content production system, content production method, and computer program

Provided is a content production system for producing content. This content production system is provided with a feature amount extraction unit that extracts a feature amount of a user from information related to the user, and a content generation unit that causes a machine learning model for generating content to generate content. In response to an input instruction or input information received by an input unit, a control unit controls the content generation unit to generate content reflecting the characteristics of the user on the basis of reference content for producing content and the feature amount of the user, and causes an output unit to output the generated content.
Owner:SONY GROUP CORP

Multi-language TTS real-time synthesis method based on deep learning

The invention discloses a multi-language TTS real-time synthesis method based on deep learning. According to the method, through a deep neural network model, high-quality real-time conversion from a multi-language text to voice is realized. The method comprises the following steps: firstly, constructing a multi-language acoustic feature library and a pronunciation rule library, and extracting text semantic features by adopting an end-to-end neural network architecture; and then, an improved attention mechanism is utilized to realize accurate synthesis of voice rhythm and pronunciation, and naturalness and coherence of multi-language pronunciation are ensured. Meanwhile, a lightweight reasoning engine is designed, calculation resource allocation is optimized, and the real-time synthesis requirement is met. Compared with a traditional TTS method, the method has the advantages that the fluency and naturalness of multilingual speech synthesis are remarkably improved, the calculation delay is greatly reduced, and efficient and reliable technical support is provided for intelligent speech interaction and multilingual content production.
Owner:GUANGZHOU BAIRUI NETWORK TECH CO LTD

A content management system for brand marketing and a marketing method using it

PendingKR1020260113437AEvaluation resultThe Internet
The present invention relates to a brand marketing content management system and a marketing method using the same, which connects advertisers who want to advertise in an internet environment with creators who freely produce content in an internet environment, thereby enabling advertisers to conduct effective advertising and creators to produce desired content and receive sponsorship. The system is characterized by comprising: a customer management unit that receives content production requests from advertisers; a request matching unit that matches creators among previously recruited creators to the category of the content production request and delivers the content production request to the creator; a content evaluation unit that evaluates the advertising content produced by the creator based on established evaluation criteria; and a settlement unit that delivers sponsorship funds corresponding to the evaluation results of the advertising content to the creator.
Owner:TAEJIN METAL CO LTD

Graphical User Interface for Teaching Activities on Electronic Devices

1. Name of the product in this design: Graphical User Interface for Teaching Activities in Electronic Equipment. 2. Purpose of this design: An electronic device. 3. The key design features of this product are its graphical user interface. 4. The image or photograph that best illustrates the design's key features: the front view. 5. Purpose of the graphical user interface: Used as intelligent learning educational software. 6. Human-computer interaction method of graphical user interface: The main view is the main interface of Zhixue Education Software. Clicking the "I want to prepare lessons" module will take you to the main interface of lesson preparation, as shown in Figure 1, where you can create courseware. Click on the special feature "Classroom Activities" to enter the interface change state diagram 2. The page displays 6 activity question types, and the activity theme is displayed on the card below each question type. After selecting any card from "Unlimited Categories", click the OK button below to enter the interface change state diagram 3, where you can create the content for the category options. You can click the help button below for instructions on filling in the information. Once you have finished filling in the information, click the "Complete" button below to enter the interface change state shown in Figure 4. Click the "Subject Tools" icon to bring up a drop-down menu. Select the "Geography" icon to see three geography tools, as shown in Figure 5. Select all tools to enter the interface change state shown in Figure 6, where you can continue to operate the buttons on the relevant pages.
Owner:HISENSE COMML DISPLAY CO LTD

Three-dimensional content real-time generation method for multi-mode AI and emotional intention recognition

The invention relates to the technical field of man-machine interaction, and discloses a multi-modal AI and emotional intention recognition three-dimensional content real-time generation method, which integrates a multi-modal AI collaborative generation module, a dynamic emotional intention recognition module, a three-dimensional content real-time generation module and a holographic software and hardware collaborative module, the method comprises the following steps: realizing an immersive holographic interaction and multi-modal AI collaborative generation module, constructing a unified semantic space by adopting CLIP + VATT, analyzing an instruction by adopting LLM to generate parameters, generating 2D content by adopting Diffusion, converting NeRF into a textured 3D model in real time, and realizing sketch / semantic driven parameterization generation by combining ControlNet and LoRA. According to the scheme of the invention, the breakthrough efficiency can be improved, and the AI driven automatic generation technology is realized; a traditional 3D content production process which needs to be completed in several days can be compressed to be completed in several minutes by means of collaborative optimization of models such as Point-E, NeRF and Diffusion and combining ControlNet structured control and LoRA low-rank fine tuning, and the modeling efficiency is integrally improved.
Owner:BESTTONE HOLDING

Panoramic image generation method and device based on voice driving and electronic equipment

The invention provides a panoramic image generation method and device based on voice driving and electronic equipment, and the method comprises the steps: converting a voice instruction inputted by a user into text description information, and generating visual description information, corresponding to three different regions, of the text description information through employing a natural language model which is finely adjusted in advance; based on a preset diffusion model, converting the first visual description information corresponding to the middle area into a first image with a preset size; based on a preset diffusion model, the first image and the second visual description information corresponding to the left side area and the right side area respectively, generating second images of preset sizes corresponding to the left side area and the right side area respectively; and splicing and fusing the first image and the two second images, and converting the spliced and fused image into an equirectangular panoramic image for the VR head-mounted display device to render and display. According to the invention, a complete link from voice intention to roaming panorama is opened, and the VR content production efficiency and the immersion of a user in browsing the VR content are improved.
Owner:CHERY AUTOMOBILE CO LTD

End-network coordinated ubiquitous network congestion control method

The present application relates to the technical field of network communication congestion control, in particular to an end-network cooperative ubiquitous network congestion control method. The method is applied to a ubiquitous network comprising end nodes, intermediate nodes and content production nodes. The end nodes generate and send interest packets. The intermediate nodes obtain queue usage state, pending interest table occupation state, available bandwidth state and interest packet satisfaction state when the data packets return, generate a first congestion score or a second congestion score, and generate a congestion marking probability in combination with the occupation contribution of the target interest flow in the pending interest table to mark the return data packets with congestion. The intermediate nodes perform interest flow splitting or rate shaping according to the congestion level, and the end nodes adjust the subsequent interest packet sending window according to the congestion score information and the congestion marking result. The present application is suitable for high dynamic and multi-path ubiquitous network congestion control.
Owner:CHANGCHUN UNIV OF SCI & TECH

High dynamic range content distribution display method and video processing apparatus

The application relates to the technical field of digital film content production, distribution and screening, and discloses a high dynamic range content distribution display method and a video processing device. A standard dynamic range master is used as the only main asset, physical baseline information is generated, candidate residual information is extracted, semantic gating decision is performed, auxiliary data containing the physical baseline information and selectively containing conditional gain information is output, and a high dynamic range signal is reconstructed according to the corresponding mode selected according to the auxiliary data, so that the single master is adapted to different display capability screening terminals, and the problems of high cost and poor consistency of the existing multi-version distribution mode are solved. Meanwhile, the semantic gating mechanism is used to eliminate the AI illusion risk, the adaptive auxiliary data structure is used to reduce the transmission bandwidth, and the distribution efficiency, screening safety and content verifiability are considered.
Owner:CHINA RES INST OF FILM SCI & TECH

Short video automatic generation method and system based on semantic comprehension

The invention relates to a short video automatic generation method and system based on semantic understanding, and belongs to the technical field of intelligent content generation. The method comprises the following steps: acquiring multi-modal source data input by a user, and extracting a structured semantic feature vector; obtaining a dynamic semantic hypergraph through conflict detection and dynamic expansion of associated nodes; executing a Markov decision process on the dynamic semantic hypergraph to generate a platform instruction set; generating a video element sequence by applying a combination principle, presetting a conflict element pair, and outputting an enhanced video template; scheduling a layered material library according to the enhanced video template; and performing collaborative rendering on the layered material data flow through a heterogeneous computing architecture, and outputting a final generated video. And generating a quality feedback signal based on the behavior data of the user for the finally generated video, and dynamically updating the semantic deconstruction model and the association node of the dynamic semantic hypergraph. According to the invention, high automation and intelligence of short video content production are realized.
Owner:GOLDEN TIMES CULTURE COMM

3D video content production device for hologram device and method for driving the device

Regarding a stereoscopic image content production apparatus for a hologram device and a method for driving the apparatus, the stereoscopic image content production apparatus for a hologram device of an embodiment may include a communication interface unit that communicates with a user terminal device owned by a purchaser of a hologram device that embodies hologram content, and a control unit that provides web or app services to enable the user terminal device to create, edit, or render hologram content, and that predicts demand for products and content for the hologram device based on user experience data from the use of the hologram device and hologram content, and provides custom-made products and content.
Owner:ATECHNET CO LTD

Graphical User Interface for Media Content Production of an Electronic Device

1. Name of the design product: Graphical user interface for media content production of an electronic device. 2. Use of the design product: For an electronic device. 3. Design key points of the design product: Lies in the graphical user interface. 4. Picture or photo that best shows the design key points: The front view of Design 1. 5. Design 1 is designated as the basic design. 6. Use of the graphical user interface: The interface of this design product is the media content production function interface of an application software. The front view of Design 1 is the display interface for media content production. Users can click the black button in the interface to add media content production materials. The interaction methods of Design 2 to Design 8 are the same as those of Design 1. The grey color blocks in the interface are replaceable pictures or videos. The cross in this design interface represents text and / or numbers and / or letters and / or symbols.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Systems and methods for secure storyline media

Disclosed are systems and methods of a novel framework that automatically and dynamically detects security events at a location and generates curated media content based therefrom. The disclosed framework can curate, modify and / or create animated media files from the captured security footage. Animating the captured content from security cameras can offer a creative and effective way to address privacy concerns while still providing valuable security insights that are based on, but not limited to, dynamic visualization, privacy protection, analytical efficiency, legal and ethical compliance, resource management and the like. By integrating animated content into security systems, organizations can maintain effective surveillance and monitoring capabilities while addressing privacy concerns and protecting individual identities.
Owner:RESIDEO LLC