Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

117 results about "Content production" patented technology

Digital human video generation method based on multi-modal large model

The invention belongs to the technical field of virtual person generation, and particularly relates to a digital person video generation method based on a multi-modal large model, and the method comprises the following steps: 1, constructing a multi-modal data system; 2, multi-modal large model training and adaptation are carried out; 3, constructing a digital human three-dimensional model; step 4, performing semantic analysis and modal mapping; 5, generating a time sequence action and a mouth shape; step 6, building and rendering a virtual scene; step 7, audio and video synchronous rendering and synthesis; step 8, quality optimization and defect repair; and step 9, performing user interaction and iterative optimization. Through technical innovation and engineering, the core pain point in digital human video generation is solved, efficient, vivid and customizable content production capacity is provided for virtual anchors, intelligent customer service, enterprise training and other scenes, and the AI digital human technology is promoted to be applied to large-scale business from experiments.
Owner:ZHE JIANG YAN HUANG KE JI YOU XIAN GONG SI

AIGC content generation method and system based on multi-modal fusion

The invention relates to the technical field of AIGC content generation, discloses an AIGC content generation method and system based on multi-modal fusion, and aims to solve the problems of decentralization, low efficiency and insufficient originality of a traditional content generation tool. Multi-modal data such as texts, images, videos and audios are integrated, user intentions are analyzed in combination with intelligent retrieval and a domain knowledge base, automatic generation from multi-modal input to high-quality creative content is achieved, a cross-modal collaborative generation technology is adopted, semantic features are dynamically aligned, and logically coherent content is generated. The content emotional value is enhanced through an emotional analysis and dynamic optimization strategy, the homogenization bottleneck is broken through, meanwhile, an automatic quality evaluation and format adaptation mechanism is integrated, deep application of scenes such as text travel, advertisement, e-commerce and interactive network television service is supported, marketing copywriting, short videos and cross-platform distribution schemes can be efficiently generated, and the market competitiveness is improved. And the content production efficiency and the creativity transmission are obviously improved.
Owner:HANGZHOU WANDIAN TECHNOLOGY CO LTD

Automatic lecturer video generation method based on AI speech synthesis and animation driving

The invention discloses a lecturer video automatic generation method based on AI speech synthesis and animation driving. The method comprises the following steps: performing structured analysis on a PPT or a text script through an improved interior point method and an incremental shortest path algorithm; performing semantic grouping by applying a full-dynamic parallel single-link clustering algorithm and generating an enhanced script with an expressive mark; a CosyVoice technology is combined with a low-rank approximation method to generate a high-quality voice data stream; establishing a mapping relation between contents and action expressions through semantic analysis, and generating a complete action expression instruction set; and driving the digital human model by using the msueTalk technology, and generating a final lecturer teaching video through a parallel rendering algorithm. According to the invention, the method achieves the efficient and automatic generation of the education video, remarkably improves the content production efficiency, reduces the production cost, and guarantees the specialty and expressive force of the teaching video.
Owner:SHENZHEN XUEYOU TECHNOLOGY CO LTD

Multi-mode convergence media content auxiliary creation method based on AI technology

The invention discloses a multi-mode convergence media content auxiliary creation method based on an AI technology, and the method comprises the steps: extracting key features according to a creation demand text inputted by a user, carrying out the retrieval in a knowledge base through employing a knowledge graph retrieval mode according to the key features, obtaining a creation material, and generating a first draft; carrying out cross-modal feature mapping on the first draft by adopting a multi-modal alignment model, aligning text-image-video embedded vectors through contrastive learning, and dynamically adjusting the correlation of multi-modal contents by utilizing an attention mechanism to obtain a multi-modal content packet; duplicate checking is carried out on the multi-modal content packet in multiple modes, and the multi-modal content packet after duplicate checking is sent to a user for manual editing. Through multi-mode processing methods such as AI auxiliary writing, AI illustration and video generation and whole-network duplicate checking and propagation value evaluation, the content production efficiency and quality are improved, the propagation effect is optimized, and the original content is protected.
Owner:广西日报社

House content generation method and device based on artificial intelligence, and readable storage medium

The embodiment of the invention discloses a house image content generation method based on artificial intelligence. The method comprises the steps of obtaining a basic image, wherein the basic image is an image shot in a house; processing the basic image based on a pre-trained cue word acquisition model to obtain a basic cue word; inputting the basic cue word and a preset cue word into an image generator to obtain a target image which is a realistic image; generating matched text content based on the target image; and the matched text content and the target image are associated and published to a target platform. Based on the scheme, on one hand, through full-process automatic design, links from image acquisition to publishing are linked, manual intervention is remarkably reduced, the content generation efficiency is improved, and the method is suitable for large-scale content production scenes of decoration companies, self-media and the like; on the other hand, by means of cue word design, it is ensured that the generated target image is highly close to a real house scene, meanwhile, the image-text collaborative generation mechanism guarantees consistency of content themes, and credibility and transmissibility are enhanced.
Owner:SHENZHEN BINCENT TECH

Digital human live broadcast method and system

The invention relates to a digital human live broadcast method and system. The method comprises the following steps: acquiring a live broadcast room visual effect picture and a digital human live broadcast audio; generating digital human live broadcast content under the driving of the digital human live broadcast audio by taking the visual effect picture of the live broadcast room as a visual manuscript through an I2V model; in the process of playing the digital human live broadcast content, if the current interaction behavior of the user in the live broadcast room is monitored, generating digital human interaction content responding to the current interaction behavior; selecting an insertion frame from the key frames of the digital human live broadcast content so as to insert the digital human interaction content at the position of the insertion frame; the key frame is a video frame corresponding to a statement demarcation point in the digital human live broadcast content. According to the scheme provided by the invention, the contradiction between long content production period, high implementation cost and slow real-time interaction feedback in digital human live broadcast in related technologies can be solved.
Owner:HANGZHOU TEKAN TECHNOLOGY CO LTD

Immersive content production and simulation system for motion vehicle

The invention discloses an immersive content production and simulation system for a motion vehicle, and aims to solve the problems that virtual vision and physical motion are difficult to match, the content production cost is high and the interactivity is insufficient. According to the system, a six-degree-of-freedom track of a target is planned through a content making simulation operation end, and a moving carrier is driven to execute; and meanwhile, real motion data of the carrier is collected, and comparison and iterative correction are carried out on the real motion data and the target trajectory until errors converge. The system also uses artificial intelligence to generate medium-long-shot and close-shot three-dimensional asset fusion. Through closed-loop correction driven by physical data, high-precision synchronization of motion and pictures is realized, the development period is remarkably shortened, the content cost is reduced, and the personalized interactive experience is improved.
Owner:THE BEST SYNC ADVERTISING CO LTD GUANGZHOU

Text-to-action generation method and system based on fine-grained representation of body parts

The invention discloses a text-to-action generation method and system based on body part fine-grained representation, and belongs to the field of human body action generation. Body part-level fine-grained discretization modeling is performed on human body actions to obtain an action encoder and an action decoder which can stably reconstruct a continuous action sequence; the method comprises the following steps: taking a part fine-grained discrete action space as a producible space, introducing video potential representation as space-time dynamic priori, establishing an alignment modeling mechanism of the video potential representation and the action discrete space, and training to obtain a conditional generation network, so as to realize stable prediction and iterative completion of an action discrete index sequence in a reasoning stage; and a continuous human body action sequence is reconstructed through an action decoder, and finally a human body action sequence which is consistent with text semantics, richer in details and more coherent in time sequence is generated. According to the method, the technical threshold and data dependence of action content production can be reduced, and the controllability and generalization ability of action generation are improved.
Owner:ZHEJIANG UNIV

Video image processing method for reducing 3D content manufacturing cost with assistance of AI

The invention relates to the technical field of computer vision and three-dimensional modeling processing, and discloses an AI-assisted video image processing method for reducing 3D content manufacturing cost, which comprises the following steps: S1, acquiring a target video sequence, and carrying out image enhancement processing and frame sampling on the video sequence to obtain an image frame sequence; s2, performing multi-model semantic recognition processing on each image frame to obtain a plurality of semantic regions and corresponding semantic tags in the image frames; and S3, performing semantic weight scoring on the semantic region, and dividing the semantic region into a high-priority region, a middle-priority region and a low-priority region according to a scoring result. According to the method, content in the video image is subjected to classification processing, fine modeling of a high-semantic region and rapid generation of a low-semantic region by introducing a semantic recognition and region priority modeling mechanism, and through the design, invalid calculation tasks can be reduced on the basis of keeping the integrity of key content.
Owner:ZHENGZHOU UNIV

Intelligent digital human system based on multi-modal deep learning and automatic content generation

The invention relates to the technical field of digital people, in particular to an intelligent digital people system based on multi-modal deep learning and automatic content generation, which comprises a front-end interaction module, a rear-end processing module, a multi-modal data management module, an AI synthesis engine and a distributed storage module. The system realizes full-process intelligent operation from model training to content publishing through dynamic voiceprint extraction, audio-driven video generation and multi-modal data fusion. According to the method, the defects in personalized training, automatic content generation and multi-user cooperation management in the prior art can be overcome, the naturalness, flexibility and user experience of voice synthesis and video driving are remarkably improved, and an efficient and intelligent solution is provided for digital human content production.
Owner:奚澜卜

Content production system, content production method, and computer program

Provided is a content production system for producing content. This content production system is provided with a feature amount extraction unit that extracts a feature amount of a user from information related to the user, and a content generation unit that causes a machine learning model for generating content to generate content. In response to an input instruction or input information received by an input unit, a control unit controls the content generation unit to generate content reflecting the characteristics of the user on the basis of reference content for producing content and the feature amount of the user, and causes an output unit to output the generated content.
Owner:SONY GROUP CORP

Multi-language TTS real-time synthesis method based on deep learning

The invention discloses a multi-language TTS real-time synthesis method based on deep learning. According to the method, through a deep neural network model, high-quality real-time conversion from a multi-language text to voice is realized. The method comprises the following steps: firstly, constructing a multi-language acoustic feature library and a pronunciation rule library, and extracting text semantic features by adopting an end-to-end neural network architecture; and then, an improved attention mechanism is utilized to realize accurate synthesis of voice rhythm and pronunciation, and naturalness and coherence of multi-language pronunciation are ensured. Meanwhile, a lightweight reasoning engine is designed, calculation resource allocation is optimized, and the real-time synthesis requirement is met. Compared with a traditional TTS method, the method has the advantages that the fluency and naturalness of multilingual speech synthesis are remarkably improved, the calculation delay is greatly reduced, and efficient and reliable technical support is provided for intelligent speech interaction and multilingual content production.
Owner:GUANGZHOU BAIRUI NETWORK TECH CO LTD

A content management system for brand marketing and a marketing method using it

PendingKR1020260113437AEvaluation resultThe Internet
The present invention relates to a brand marketing content management system and a marketing method using the same, which connects advertisers who want to advertise in an internet environment with creators who freely produce content in an internet environment, thereby enabling advertisers to conduct effective advertising and creators to produce desired content and receive sponsorship. The system is characterized by comprising: a customer management unit that receives content production requests from advertisers; a request matching unit that matches creators among previously recruited creators to the category of the content production request and delivers the content production request to the creator; a content evaluation unit that evaluates the advertising content produced by the creator based on established evaluation criteria; and a settlement unit that delivers sponsorship funds corresponding to the evaluation results of the advertising content to the creator.
Owner:TAEJIN METAL CO LTD

Graphical User Interface for Teaching Activities on Electronic Devices

1. Name of the product in this design: Graphical User Interface for Teaching Activities in Electronic Equipment. 2. Purpose of this design: An electronic device. 3. The key design features of this product are its graphical user interface. 4. The image or photograph that best illustrates the design's key features: the front view. 5. Purpose of the graphical user interface: Used as intelligent learning educational software. 6. Human-computer interaction method of graphical user interface: The main view is the main interface of Zhixue Education Software. Clicking the "I want to prepare lessons" module will take you to the main interface of lesson preparation, as shown in Figure 1, where you can create courseware. Click on the special feature "Classroom Activities" to enter the interface change state diagram 2. The page displays 6 activity question types, and the activity theme is displayed on the card below each question type. After selecting any card from "Unlimited Categories", click the OK button below to enter the interface change state diagram 3, where you can create the content for the category options. You can click the help button below for instructions on filling in the information. Once you have finished filling in the information, click the "Complete" button below to enter the interface change state shown in Figure 4. Click the "Subject Tools" icon to bring up a drop-down menu. Select the "Geography" icon to see three geography tools, as shown in Figure 5. Select all tools to enter the interface change state shown in Figure 6, where you can continue to operate the buttons on the relevant pages.
Owner:HISENSE COMML DISPLAY CO LTD

Three-dimensional content real-time generation method for multi-mode AI and emotional intention recognition

The invention relates to the technical field of man-machine interaction, and discloses a multi-modal AI and emotional intention recognition three-dimensional content real-time generation method, which integrates a multi-modal AI collaborative generation module, a dynamic emotional intention recognition module, a three-dimensional content real-time generation module and a holographic software and hardware collaborative module, the method comprises the following steps: realizing an immersive holographic interaction and multi-modal AI collaborative generation module, constructing a unified semantic space by adopting CLIP + VATT, analyzing an instruction by adopting LLM to generate parameters, generating 2D content by adopting Diffusion, converting NeRF into a textured 3D model in real time, and realizing sketch / semantic driven parameterization generation by combining ControlNet and LoRA. According to the scheme of the invention, the breakthrough efficiency can be improved, and the AI driven automatic generation technology is realized; a traditional 3D content production process which needs to be completed in several days can be compressed to be completed in several minutes by means of collaborative optimization of models such as Point-E, NeRF and Diffusion and combining ControlNet structured control and LoRA low-rank fine tuning, and the modeling efficiency is integrally improved.
Owner:BESTTONE HOLDING

Panoramic image generation method and device based on voice driving and electronic equipment

The invention provides a panoramic image generation method and device based on voice driving and electronic equipment, and the method comprises the steps: converting a voice instruction inputted by a user into text description information, and generating visual description information, corresponding to three different regions, of the text description information through employing a natural language model which is finely adjusted in advance; based on a preset diffusion model, converting the first visual description information corresponding to the middle area into a first image with a preset size; based on a preset diffusion model, the first image and the second visual description information corresponding to the left side area and the right side area respectively, generating second images of preset sizes corresponding to the left side area and the right side area respectively; and splicing and fusing the first image and the two second images, and converting the spliced and fused image into an equirectangular panoramic image for the VR head-mounted display device to render and display. According to the invention, a complete link from voice intention to roaming panorama is opened, and the VR content production efficiency and the immersion of a user in browsing the VR content are improved.
Owner:CHERY AUTOMOBILE CO LTD

End-network coordinated ubiquitous network congestion control method

The present application relates to the technical field of network communication congestion control, in particular to an end-network cooperative ubiquitous network congestion control method. The method is applied to a ubiquitous network comprising end nodes, intermediate nodes and content production nodes. The end nodes generate and send interest packets. The intermediate nodes obtain queue usage state, pending interest table occupation state, available bandwidth state and interest packet satisfaction state when the data packets return, generate a first congestion score or a second congestion score, and generate a congestion marking probability in combination with the occupation contribution of the target interest flow in the pending interest table to mark the return data packets with congestion. The intermediate nodes perform interest flow splitting or rate shaping according to the congestion level, and the end nodes adjust the subsequent interest packet sending window according to the congestion score information and the congestion marking result. The present application is suitable for high dynamic and multi-path ubiquitous network congestion control.
Owner:CHANGCHUN UNIV OF SCI & TECH

High dynamic range content distribution display method and video processing apparatus

The application relates to the technical field of digital film content production, distribution and screening, and discloses a high dynamic range content distribution display method and a video processing device. A standard dynamic range master is used as the only main asset, physical baseline information is generated, candidate residual information is extracted, semantic gating decision is performed, auxiliary data containing the physical baseline information and selectively containing conditional gain information is output, and a high dynamic range signal is reconstructed according to the corresponding mode selected according to the auxiliary data, so that the single master is adapted to different display capability screening terminals, and the problems of high cost and poor consistency of the existing multi-version distribution mode are solved. Meanwhile, the semantic gating mechanism is used to eliminate the AI illusion risk, the adaptive auxiliary data structure is used to reduce the transmission bandwidth, and the distribution efficiency, screening safety and content verifiability are considered.
Owner:CHINA RES INST OF FILM SCI & TECH

Short video automatic generation method and system based on semantic comprehension

The invention relates to a short video automatic generation method and system based on semantic understanding, and belongs to the technical field of intelligent content generation. The method comprises the following steps: acquiring multi-modal source data input by a user, and extracting a structured semantic feature vector; obtaining a dynamic semantic hypergraph through conflict detection and dynamic expansion of associated nodes; executing a Markov decision process on the dynamic semantic hypergraph to generate a platform instruction set; generating a video element sequence by applying a combination principle, presetting a conflict element pair, and outputting an enhanced video template; scheduling a layered material library according to the enhanced video template; and performing collaborative rendering on the layered material data flow through a heterogeneous computing architecture, and outputting a final generated video. And generating a quality feedback signal based on the behavior data of the user for the finally generated video, and dynamically updating the semantic deconstruction model and the association node of the dynamic semantic hypergraph. According to the invention, high automation and intelligence of short video content production are realized.
Owner:GOLDEN TIMES CULTURE COMM

3D video content production device for hologram device and method for driving the device

Regarding a stereoscopic image content production apparatus for a hologram device and a method for driving the apparatus, the stereoscopic image content production apparatus for a hologram device of an embodiment may include a communication interface unit that communicates with a user terminal device owned by a purchaser of a hologram device that embodies hologram content, and a control unit that provides web or app services to enable the user terminal device to create, edit, or render hologram content, and that predicts demand for products and content for the hologram device based on user experience data from the use of the hologram device and hologram content, and provides custom-made products and content.
Owner:ATECHNET CO LTD

Systems and methods for secure storyline media

Disclosed are systems and methods of a novel framework that automatically and dynamically detects security events at a location and generates curated media content based therefrom. The disclosed framework can curate, modify and / or create animated media files from the captured security footage. Animating the captured content from security cameras can offer a creative and effective way to address privacy concerns while still providing valuable security insights that are based on, but not limited to, dynamic visualization, privacy protection, analytical efficiency, legal and ethical compliance, resource management and the like. By integrating animated content into security systems, organizations can maintain effective surveillance and monitoring capabilities while addressing privacy concerns and protecting individual identities.
Owner:RESIDEO LLC

Melody perception music reorganization method and system based on structure-texture feature decoupling

The invention discloses a melody perception music reorganization method and system based on structure-texture feature decoupling. The method comprises the steps of constructing a weak pairing data set containing the same melody skeleton and different acoustic texture audio pairs, training a comparison melody encoder based on the data set to extract pure melody structure characterization, then training a double-circulation-flux recompilation generation model, and finally completing recompilation generation of a target audio. The system comprises a data set construction module, an encoder training module, a generative model training module and a reasoning generation module, and the core of the system comprises a contrast melody encoder and a double-circulation-flux modification generative model. According to the method, the problems of feature entanglement and text control failure are solved, the melody consistency and the model generalization ability are improved, and the method can be widely applied to the fields of digital music creation, short video content production, old audio reproduction and the like.
Owner:SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Tourism culture content generation method and device based on large language model

The present application relates to the field of digital technology of tourism and travel, and discloses a tourism and culture content generation method and device based on a large language model, wherein the device comprises a tourism and culture knowledge graph global construction and dynamic updating module, a scene demand semantic analysis and generation constraint standardization definition module, a culture narrative chain intelligent construction and core content original generation module, a multi-form content adaptation and immersive expression optimization module, a tourism and culture content culture authenticity verification and compliance verification module, and a content effect quantitative evaluation and generation model iterative optimization module; the method comprises the following steps: S1, a tourism and culture knowledge graph construction stage; S2, a demand analysis and constraint definition stage; S3, a narrative chain construction and core content generation stage; S4, a content form adaptation and immersive optimization stage; S5, a content verification and compliance verification stage; and S6, an effect evaluation and model iteration stage; the present application breaks through the problem that traditional tourism and culture content production relies on artificial creation and content homogenization is serious.
Owner:HEBEI NORMAL UNIV FOR NATTIES

Method, device, and program for providing information regarding content production

Herein is proposed a method for providing information regarding content production, comprising: performing preprocessing on a script file; separating and grouping the contents of the preprocessed script file; classifying sentences which meet predetermined condition into one or more categories; storing information related to the grouping of the preprocessed script file or information on the classification of sentences into one or more categories; processing the stored information in response to requests related to the stored information; and displaying the processed information.
Owner:CJ OLIVENETWORKS

Method and device for assisting production of converged media content based on big data

The invention provides a big data-based convergence media content production auxiliary method and device. The method comprises the steps that a server side crawls public text information from a plurality of information public platforms; processing the public text information by using a text classification model to obtain a theme tag corresponding to the public text information; sending the public text information to a client corresponding to the theme tag; receiving the converged media content, extracting target text information, target video information and target audio information in the converged media content, and determining the content matching degree between the converged media content and each alternative publishing platform according to the target text information, the target video information and the target audio information; and sending the platform publishing suggestion containing the content matching degree to the client. By extracting the multi-modal information, evaluating the content matching degree between the converged media content and each alternative publishing platform and feeding back the platform publishing suggestions to the client, a creator can be helped to select a proper publishing platform, so that the suitability of the content and the platform is improved.
Owner:SHENZHEN RADIO FILM & TELEVISION GRP

Information processing system, information processing method, and information processing program

An information processing system according to the present disclosure comprises: an acquisition unit that acquires background information relating to content production and content information being edited; and a processing unit that outputs assistance information pertaining to the content production on the basis of the background information and the content information being edited.
Owner:SONY GROUP CORP

A Deep Learning-Based Real-Time Multilingual TTS Synthesis Method

This invention discloses a real-time multilingual TTS synthesis method based on deep learning. This method achieves high-quality real-time conversion of multilingual text to speech through a deep neural network model. First, a multilingual acoustic feature library and pronunciation rule library are constructed, and an end-to-end neural network architecture is used to extract semantic features from the text. Then, an improved attention mechanism is used to achieve accurate synthesis of speech prosody and pronunciation, ensuring the naturalness and coherence of multilingual pronunciation. Simultaneously, a lightweight inference engine is designed to optimize computational resource allocation and meet the requirements of real-time synthesis. Compared with traditional TTS methods, this invention significantly improves the fluency and naturalness of multilingual speech synthesis, greatly reduces computational latency, and provides efficient and reliable technical support for intelligent voice interaction and multilingual content production.
Owner:GUANGZHOU BAIRUI NETWORK TECH CO LTD

Facilitating video generation

Features described herein generally relate to content production. Particularly, the present disclosure relates to facilitating video generation. Using machine-learning models, a storyboard can be generated from an inspirational video, video attributes can be determined for the storyboard, editing scores and actions can be determined for candidate videos, candidate videos can be edited based on the editing scores and actions, and the edited candidate videos can be combined to generate a video.
Owner:10Z LLC

Intelligent doorplate dynamic management system and method based on LORA technology

The invention relates to the field of intelligent doorplate management, and particularly discloses an intelligent doorplate dynamic management system and method based on an LORA technology, which are characterized in that external heterogeneous source data such as OA, educational administration and the like are grabbed, cleaned and packaged in a standardized manner at a front end through protocol adaptation and data engine technologies, and information islands are broken to realize automatic content production; the transmission layer reconstructs network data into high-penetrability LORA radio-frequency signals for wide-area broadcasting by utilizing a cloud compiling and gateway protocol conversion mechanism; and the terminal side adopts a periodic wake-up strategy based on a CAD mechanism, and drives the electronic ink screen to refresh and immediately and automatically cut off the power supply only after the verification of the effective signaling is passed. Therefore, real-time synchronization of the physical identification and the digital service is realized without manual intervention through full-link data closed-loop circulation, and meanwhile, extremely-low-power-consumption long-term operation of the system is ensured by utilizing a bistable display characteristic.
Owner:ZHEJIANG HAIYAN POWER SYST RESOURCES ENVIRONMENTAL TECH