Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

157 results about "Visual cueing" patented technology

Visual cueing. A visual cue is a signal which your brain extracts from what you see. It indicates the state of some property around you that you are interested in perceiving. Now, only 1% of what you see actually enters through your eyes (the rest is -surprisingly correct – made up by your brain).

Remote sensing image target statistical method and system fusing large language model and visual cue driving

The invention provides a remote sensing image target statistical method and system fusing a large language model and visual prompt driving. The method comprises the following steps: acquiring a remote sensing instance segmentation image to be processed and a visual prompt thereof; inputting a to-be-processed remote sensing instance segmentation image and a visual prompt thereof into the trained remote sensing image target statistical model, and outputting a remote sensing image target statistical result; the training comprises the following steps: introducing a large language model and visual cue into an encoder architecture of a GrondingDINO model to obtain a remote sensing image target statistical model; inputting a remote sensing instance segmented image and the visual cue thereof into an encoder, and outputting an image feature, a visual cue feature and a text feature; the feature intensifier carries out fusion processing on the output of the encoder; a language-guided query selection module calculates cross-modal query according to the fusion processing result, and a cross-modal decoder obtains a target statistical result of the image based on the fusion processing result and the cross-modal query; and training by using the training data and outputting the trained model.
Owner:WUHAN UNIV

Method, device, and medium for training large scale object foundation model

Embodiments of the present disclosure provide a method, device, and medium for training a large scale object foundation model. The method comprises obtaining a training dataset comprising a plurality of subsets for a plurality of object perception tasks, wherein a sample in the training dataset comprises an image with an object, a prompt indicating the object, and labeled object perception information of the image. The method further comprises generating, by the image encoder, an image feature based on the image. The method further comprises generating, by the text encoder or the visual prompt encoder, a prompt embedding based on the prompt. The method further comprises generating, by the object decoder, object perception information of the object based on the image feature and the prompt embedding. In addition, the method further comprises training the object processing model based on the generated object perception information and the labeled object perception information.
Owner:LEMON INC(GB)

Video action recognition model training method, video action recognition method and device

The invention relates to a video action recognition model training method, a video action recognition method and a video action recognition device. The method comprises the following steps: acquiring a sample video frame image and an action category description text corresponding to the sample video frame image, inputting the sample video frame image and the action category description text into a to-be-trained recognition model, and recognizing the action category of the to-be-trained recognition model by an image encoder in the recognition model according to a preset visual prompt vector and the sample video frame image, generating a video embedding corresponding to a sample video frame image, generating a text embedding corresponding to an action category description text by a text encoder in the recognition model based on a preset text prompt vector and the action category description text, and constructing bidirectional comparison loss by taking the video embedding and the text embedding as positive sample pairs, and updating the visual prompt vector and the text prompt vector based on the bidirectional contrast loss to obtain a trained recognition model. By adopting the method, the video action recognition accuracy can be improved.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Visual language model training method, image tag prediction method, and electronic device

PCT designated stageWO2026026225A1Biological modelsFeature extractionLinguistic model
Embodiments of the present disclosure provide a visual language model training method, an image tag prediction method, and an electronic device. According to the embodiments of the present disclosure, encoding parameters respectively corresponding to an image encoder and a text encoder in a first visual language model are fixed, on the basis of initial values of visual prompts and text samples respectively corresponding to a plurality of image classification tags, a second visual language model is generated by training the first visual language model, and the visual prompts respectively corresponding to the plurality of image classification tags are mined and learned, thereby enhancing the visual representation capabilities of visual language models. Next, on the basis of the learned visual prompts, text prompts, and the text samples, a target visual language model is further trained by training the second visual language model, and adapter parameters of model adapters respectively configured for the image encoder and the text encoder are adjusted, to collaboratively optimize image feature extraction and text feature extraction. This process allows for the transfer of visual knowledge from visual prompts to text prompts.
Owner:HANGZHOU ALIBABA INT INTERNET IND CO LTD

Visual prompt multi-modal large model for multi-source remote sensing image interpretation

The invention discloses a visual prompt multi-modal large model for multi-source remote sensing image interpretation, which belongs to the technical field of crossing of remote sensing and computer vision and comprises a multi-modal content coding and integration module, a cross-domain first-stage fusion training module, a pixel level visual positioning module and a large language model. The model supports the interpretation of a remote sensing image after the remote sensing image is arbitrarily amplified and reduced, and has flexible multi-granularity vision and language interaction capability. According to the model, a large language model is used as an interface, and multi-modal content integration including multi-sensor images, visual prompts and text instructions is achieved. In addition, two types of space tasks of anaphora understanding and visual positioning are unified into a visual prompt learning framework, and comprehensive and flexible multi-granularity understanding of remote sensing data is promoted.
Owner:BEIJING INST OF TECH

Long video generation method, device and equipment, readable storage medium and program product

The invention discloses a long video generation method and device, equipment, a readable storage medium and a program product, relates to the technical field of communication, and aims to improve the quality of a generated long video. The method comprises: obtaining a video description text to be processed and video processing parameters, the video processing parameters comprising a video segment number N and a video frame number M included in each video segment, N being an integer greater than or equal to 2, and M being an integer greater than or equal to 1; dividing the video description text into text information corresponding to N time periods; taking the text information corresponding to each time period and the historical video visual prompt word corresponding to each time period as input of a video visual encoder, and running the video visual encoder to obtain a video clip corresponding to each time period; and splicing the video clips to obtain a long video. According to the embodiment of the invention, the quality of the generated long video can be improved.
Owner:CHINA MOBILE COMM LTD RES INST +1

Target navigation method and device based on active 3DGS and visual language model reasoning

The invention discloses a target navigation method and device based on active three-dimensional Gaussian spatter and visual language model reasoning, and the method comprises the steps: obtaining an RGB-D image and a pose in an unknown environment, and constructing an incremental three-dimensional Gaussian spatter map as persistent memory through active perception; generating an exploration map based on the constructed 3DGS map, extracting a leading edge point and carrying out space structure adaptive clustering; generating a guidance track, carrying out free viewpoint optimization based on the track, and rendering a leading edge point first-person view angle image containing rich information; constructing a structured visual prompt, combining with a thinking chain prompt, inputting a visual language model to carry out reasoning planning, and selecting an optimal navigation target; in the navigation process, a real-time target detector is used for screening potential targets, and a new view angle is rendered in a 3DGS space through an action decision VLM for target re-verification. The navigation success rate and efficiency are improved.
Owner:ZHEJIANG UNIV OF TECH

Explore until confident: efficient exploration for embodied question answering

A method for embodied agent exploration is described. The method includes building a semantic map of a surrounding scene based on depth information and via visual prompting of a vision language model (VLM). The method also includes utilizing conformal prediction to calibrate a question answering confidence of the VLM. The method further includes performing, by an embodied agent, scene exploration utilizing knowledge of relevant regions of the scene. The method also includes determining, by the embodied agent, when to terminate the scene exploration utilizing a calibrated question answering confidence of the VLM.
Owner:TOYOTA RESEARCH INSTITUTE INC +3

Apparatus and Method for Sensory Adjustments in Electric Vehicles

A method and apparatus enables modifying the electronic controls of EVs to mimic the sensory experience of driving a performance ICE car. The method and apparatus creates a sensory “virtual cockpit” with both electronic and mechanical enhancements for a sensory experience. By downloading and implementing the method and apparatus, one may mimic, for example, the vehicle dynamics, performance horsepower, torque curves, suspension settings, oversteer and understeer behavior, steering-wheel inputs, cabin sound, subtle cabin vibrations, and audio / visual cues via a graphical user interface. These simulations replicate, in an electric vehicle, the various aspects of an ICE vehicle to mimic the whole experience of driving various ICE performance vehicles.
Owner:LOCCISANO VINCENT

Multi-modal map enhanced retrieval method and dialogue system based on feature fusion optimization

The invention discloses a feature fusion optimization-based multi-modal map enhancement retrieval method and a dialogue system. The method comprises the following steps of: respectively carrying out pre-training and fine tuning on a visual model and a language model by utilizing a domain image and text data; constructing a knowledge graph based on the text data in the knowledge base and constructing a vector database containing associated image data; performing semantic analysis and optimization on the original query of the user by using the language model and forming a structured retrieval intention; searching related sub-graphs, text semantic vector information and associated image data based on the search intention; encoding the sub-images into knowledge contexts, inputting the knowledge contexts into a dynamic prompt generator to generate visual prompts, and extracting enhanced visual features from the associated image data through a visual model; and inputting the subgraph, the text semantic vector information and the enhanced visual features into a language model for collaborative reasoning, and generating and outputting a final answer. According to the method, deep fusion and accurate retrieval of multi-modal knowledge can be realized, and the accuracy and efficiency are remarkably improved.
Owner:ZHEJIANG UNIV

Weak supervision video anomaly detection method based on vision and text double decision

The invention belongs to the field of computer vision and image processing, and provides a weak supervision video anomaly detection method based on vision and text double decisions. According to the method, local and global time sequence modeling modules are constructed, an anomaly focusing visual prompt mechanism and a learnable text prompt vector are introduced, and fine-grained anomaly recognition is realized in combination with a dual-mode memory. In the training stage, the model extracts visual and text features on the premise of no frame-level supervision based on video-level labels, constructs a category alignment graph and optimizes category embedding. In the reasoning stage, the model dynamically updates high-confidence-degree features through a positive and negative memory mechanism, and suppression of prediction deviation between semantic proximity categories is achieved. Compared with the prior art, the method has the advantages that the abnormal behavior recognition capability and the multi-class distinguishing precision of the model are remarkably improved under the weak supervision condition, the structure is simple, deployment is easy, adaptability is high and the like, and the method is suitable for efficient video anomaly detection tasks in intelligent monitoring, behavior recognition and other scenes.
Owner:DALIAN UNIV OF TECH

Robot control method and device, equipment, medium and program product

The invention discloses a robot control method and device, equipment, a medium and a program product. A language instruction, a visual image and the initial state of a robot are obtained; determining a plurality of index values corresponding to the language instruction by utilizing the trained classifier; querying a plurality of pieces of skill prompt information corresponding to the plurality of index values in a trained skill prompt information pool; acquiring a plurality of pieces of visual prompt information corresponding to the visual image in a trained visual prompt information pool; a trained action sequence generation model is utilized to generate an action sequence based on the language instruction, the visual image, the initial state of the robot, the multiple pieces of skill prompt information and the multiple pieces of visual prompt information; and controlling the robot based on the action sequence. According to the embodiment of the invention, the robot can automatically adjust the action sequence according to different task requirements and environment changes, tasks can be executed more accurately, and the satisfaction degree of a user to robot services is improved.
Owner:CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1

A multi-modal visual understanding method based on consistent learning and mixed feature extraction

This invention relates to a multimodal visual understanding method based on consistency learning and hybrid feature extraction. It includes constructing an end-to-end fine-grained consistency learning framework, introducing a hybrid region extractor, fusing local details and global semantics to generate high-quality hybrid visual cue embeddings, combining self-reconstruction loss and latent spatial consistency loss to force the model to establish explicit alignment between the input visual cue and the output segmentation label, utilizing the geometric boundary constraints of the localization task for description generation, and simultaneously optimizing localization accuracy using the semantic depth of the description task. Furthermore, it constructs a detailed localization index expression and segmentation task to enhance the model's reasoning ability for complex long text instructions. The aim is to address the problems of feature fragmentation and insufficient accuracy in existing large models for fine-grained visual localization and description tasks. Compared with existing technologies, this invention has advantages such as high accuracy and strong generalization ability in pixel-level localization and fine-grained description.
Owner:TONGJI UNIV

Facial expression recognition method and device based on multimode attitude measurement learning, and storage medium

PendingCN121861709ARealize precise identificationSolve the problem of weak generalization abilitySemantic analysisBiological modelsBiologyMachine learning
The invention discloses a facial expression recognition method and device based on multimode attitude measurement learning, and a storage medium, relates to the technical field of facial expression recognition, and solves the technical problems of weak generalization ability, insufficient cross-modal feature fusion and easy forgetting of pre-training knowledge due to data scarcity in the existing facial expression recognition technology. The method comprises the following steps: acquiring image data containing facial expressions, preprocessing the image data, constructing a text prompt word set, and constructing triple training samples of anchor points, positive samples and negative samples; a visual cue word vector is designed to be injected into a Transform layer of a pre-trained CLIP-ViT-B / 32 model, and model trunk parameters are frozen to only train visual prompt and LayerNorm layer parameters; triple loss and classification loss combined optimization is adopted; in the reasoning process, the expression category is determined by calculating the similarity between the to-be-recognized image features and various expression text prompt features, pre-training semantic knowledge and cross-modal adaptation ability of CLIP are fully utilized, and the expression recognition precision, robustness and zero sample generalization ability are remarkably improved.
Owner:DIANSHI TECH (ZHEJIANG) CO LTD +1

Locomotive fault prompting method and device, readable storage medium and electronic equipment

The invention provides a locomotive fault prompting method and device, a medium and electronic equipment, and relates to the technical field of rail transit. The locomotive fault prompting method comprises the following steps: determining fault characteristics of a locomotive according to original fault data corresponding to a locomotive fault; determining a target fault prompting strategy in the plurality of candidate fault prompting strategies; under the condition that the target fault prompt strategy comprises visual prompt, auditory prompt and somatosensory prompt, determining a prompt mode and content of the visual prompt according to the first-level fault attribute feature, and determining a prompt mode and content of the auditory prompt according to the first-level fault attribute feature and the second-level fault attribute feature, according to the first-level fault attribute feature and the second-level fault attribute feature, determining a prompt mode and content of somatosensory prompt; and executing prompt operation based on the prompt mode and content of the visual prompt, the prompt mode and content of the auditory prompt and the prompt mode and content of the somatosensory prompt. According to the invention, the intuition and information density of locomotive fault prompt can be improved.
Owner:DATONG ELECTRIC LOCOMOTIVE OF NCR

Directive image segmentation method based on multi-mode prompt

The invention relates to a directional image segmentation method based on multi-modal prompt, which fuses natural language prompt and visual prompt information to realize accurate segmentation of a target area in an image. Specifically, the directional image segmentation model adopts an encoding-decoding structure and is composed of an image encoder, a text encoder, a pixel encoder, a visual prompt encoder and a shared mask decoder. By introducing a multi-layer embedded module and a deformable attention mechanism in a visual cue encoder, the model can fully extract space and semantic information of a cue area in a reference image. A text encoder extracts semantic features by using a pre-training language model, and performs modal fusion with visual prompts to effectively realize information complementation. In addition, the directivity image segmentation model supports segmentation tasks guided only by texts, only by vision or by the combination of the texts and the vision, and has good adaptability and task universality.
Owner:XIAMEN UNIV

Rehabilitation behavior analysis method, device and equipment based on dynamic visual cues generation

The application discloses a rehabilitation behavior analysis method, device and equipment based on dynamic visual prompt generation, comprising: collecting action video data, extracting two-dimensional skeleton point sequence data from the action video data through a pose estimator; extracting visual features and prompt features from video modal data and skeleton point modal data through a visual classification model and a skeleton classification model respectively; pre-training the visual classification model using labeled video modal data, and then performing contrast learning on the visual classification model and the skeleton classification model using unlabeled video modal data and corresponding unlabeled skeleton modal data respectively; collecting rehabilitation training video data of a patient and extracting corresponding skeleton modal data; and inputting the video data and the corresponding skeleton modal data into the trained visual classification model and the skeleton classification model respectively for classification. The application can realize more complete action representation, reduce data dependence, and improve the accuracy of rehabilitation behavior analysis.
Owner:ZHEJIANG UNIV

Audio-visual collaborative-based early parkinson disease gait intervention method, device and equipment

The application provides a gait intervention method, device and equipment for early Parkinson disease based on audio-visual cooperation. The intervention method comprises the following steps: a gait parameter calculation step; an audio-visual cooperation prompt paradigm construction step, gait rhythm characteristics are determined based on left and right foot gait cycles to generate auditory prompts synchronized with the gait rhythm characteristics, and visual prompts in the form of virtual people in a mixed reality environment are generated based on left and right foot step heights and foot position and ground detection information; a modal fusion step, the auditory prompts and the visual prompts are synchronized in the time dimension; and an audio-visual cooperation intervention step, gait intervention is performed through the synchronized auditory prompts and the visual prompts. The application can improve the gait symmetry of early Parkinson disease patients through the accurate synchronization of visual and auditory prompts in the time dimension.
Owner:INST OF SOFTWARE - CHINESE ACAD OF SCI

Model training method and device and target detection method and device

The invention discloses a model training method and device and a target detection method and device, and belongs to the field of target detection. The method comprises the following steps: acquiring a sample image, wherein the sample image comprises at least one object and a labeling category of the at least one object; performing target detection on the sample image by adopting a pre-training model to determine a prediction category of the at least one object; and training the pre-training model according to the annotation category of the at least one object and the prediction category of the at least one object to obtain a target detection model. The pre-training model comprises a visual prompt module and a target detection module, and the visual prompt module is used for carrying out weighting processing on a plurality of fusion features of the sample image according to the annotation category of the at least one object to obtain a plurality of weighted features; the target detection module is used for performing target detection according to the plurality of weighted features to determine a prediction category of the at least one object. According to the invention, the target detection accuracy of the target detection model on the image comprising the new category of objects can be improved.
Owner:BOE TECHNOLOGY GROUP CO LTD +1

Methods, devices, computer equipment and storage media for generating pet visual data

This application provides a method, apparatus, computer device, and storage medium for generating pet visual data. The method includes receiving a data generation request, which includes visual cue text and reference visual data of a target pet; invoking a pre-trained generative model based on the data generation request, wherein the generative model includes a data fusion sub-model and a visual processing sub-model; invoking the data fusion sub-model to perform data fusion based on the visual cue text and the reference visual data to obtain a visual latent vector; and invoking the visual processing sub-model to generate visual data based on the visual latent vector to obtain target visual data corresponding to the visual cue text, wherein the target visual data includes an image and / or a video of the target pet. This method can improve the consistency of pet visual data generation.
Owner:SHENZHEN LIBRO TECH CO LTD

System and method for providing gaming interfaces to test driving-related cognitive functions

A method including transmitting, to a user device of a user, a graphical user interface for display on the user device is disclosed. The graphical user interface can include one or more activation controls for activating one or more gaming interfaces. The one or more gaming interfaces can include one or more visual prompts to cause one or more user interactions with the one or more gaming interfaces by the user. The method further can include upon determining that at least one of the one or more gaming interfaces is activated, receiving, from the user device, gaming performance data associated with the one or more user interactions. The method additionally can include determining, by the one or more processors, one or more cognitive factors for the user based on the gaming performance data. The method also can include generating, by the one or more processors, an output associated with the user based at least in part on the one or more cognitive factors, as determined. Other embodiments are disclosed.
Owner:QUANATA LLC

Multi-modal instance-level understanding method and system based on visual prompt

The invention relates to the field of computer vision, and particularly discloses a multi-modal instance level understanding method and system based on visual cues, and the method comprises the steps: carrying out the space-time segmentation of a specific instance through employing an interactive instance segmentation model and a video instance tracking model, and generating a visual prompt; then, a visual encoder is used for encoding a video with a visual prompt, and a cross-modal connection module is used for mapping visual representation to a multi-modal representation space shared with a language, so that visual features are obtained; processing the input text by using a text word segmentation device to obtain corresponding text features; and finally, utilizing a large language model to uniformly model vision and language input to obtain fine-grained description or question answers about a specific instance. Compared with the prior art, the method has the advantages that accurate positioning and tracking of the specific instance in the space-time dimension are realized, and the instance-level fine-grained understanding capability of the video multi-modal model is improved.
Owner:FUDAN UNIVERSITY

Raman image fusion multi-modal fusion classification method

The invention relates to a Raman image fusion multi-modal fusion classification method, which relates to the technical field of Raman image analysis, and comprises the following steps: injecting an exclusive prompt vector P for each domain based on a visual prompt adjustment mechanism in a visual Transform, and eliminating the influence of different acquisition conditions on visual features; based on a multi-scale residual convolutional network, in cooperation with batch normalization processing, the influence of different acquisition conditions on spectral characteristics is eliminated; fusion data of visual features and spectral features are obtained based on a multi-modal fusion algorithm, and a multi-modal fusion classification model is constructed according to the fusion data, so that the advantages of images and spectrums are integrated, and the average recognition rate of data under different collection conditions is improved.
Owner:QINGDAO SINGLE CELL BIOTECH CO LTD

Visual cue-based footwork and tactical response training system for fencing

A fencing training system delivers programmable visual cues to simulate referee commands and tactical bout scenarios. It includes a transparent panel (11) with addressable LED lights (10), mounted on a tripod (12) and controlled by a microcontroller (14). Timed color and pattern sequences prompt fencers to initiate offensive, defensive, or interpretive actions. Users can customize delay intervals, cue types, and repetition frequency. The system improves timing accuracy, footwork consistency, and decision-making within the four-meter engagement zone. Configurable preparation windows and distance markers (20, 22, 24) aid spatial judgment. The system is adaptable for saber and foil fencing and may include auditory cues, motion detection, or mobile app control. Designed for portability and real-time responsiveness, it supports solo and group training. The invention may also be adapted for other sports requiring timed decision-making in confined spaces and incorporate AI-based customization using bout performance data to personalize footwork timing and cue sequences.
Owner:CHEN COOPER YE-ZOU

Matting method and electronic device

The application provides a matting method and an electronic device. The method comprises the following steps: inputting a to-be-processed image into an image encoder of an image processing model to obtain a first feature vector, inputting a visual prompt box into a prompt encoder of the image processing model to obtain a second feature vector, the visual prompt box being obtained according to a manually selected region in the to-be-processed image; inputting the first feature vector and the second feature vector into a mask decoder of the image processing model to obtain a global matting mask and a local matting mask; combining the global matting mask and the local matting mask to determine an overall mask; and cutting out an image region where a target is located from the to-be-processed image according to the overall mask. Thus, the image processing model is used to separately process the edge region to generate a local matting mask of the edge region with high accuracy, the overall mask is generated by combining the global matting mask corresponding to the foreground region, and finally, the matting is completed according to the overall mask, so that the matting effect can be improved.
Owner:HONOR DEVICE CO LTD

Systems, apparatuses, methods, and non-transitory computer-readable storage media for facilitating user interaction with a virtual keyboard

Systems, apparatuses, methods, and computer-readable storage media are disclosed for facilitating user interaction with a virtual keyboard. A computerized method comprises: receiving contextual information associated with a virtual keyboard displayed on a user interface; determining, based on the contextual information, that a visual hint display condition is met; displaying a visual hint associated with one or more keys of the virtual keyboard based on the visual hint display condition; after displaying the visual hint, receiving additional contextual information associated with the virtual keyboard; determining, based on the additional contextual information, that a visual hint dismissal condition is met; and dismissing the visual hint from the virtual keyboard based on the visual hint dismissal condition.
Owner:HUAWEI TECH CO LTD

System to facilitate parking a vehicle straight in a desired location inside a garage

An apparatus to be viewed by a driver to aid parking a car parallel to the walls of a garage, as well as the correct distance into the garage. The driver is able to park in exactly the same spot each time, at the desired distance from other cars, walls, or other stationary objects via the aid apparatus. This helps minimize banging car doors on objects upon opening, or bumping into anything in front of the vehicle, while maximizing use of space within the garage. The apparatus is installed in a fixed position in front of the driver at eye level as they sit comfortably in a driver's seat of the vehicle. As the driver drives their car into their garage to park, they are guided by visual cues from the apparatus to park the car in the same footprint on the floor each time.
Owner:POWERS ROBERT WILLIAM

Dynamic adaptation of a given assistant output based on a given persona assigned to the automated assistant.

It provides dynamic adaptation of a given assistant output based on a given persona assigned to the automated assistant. [Solution] The method generates and then adapts a given assistant output based on a given persona assigned to the automated assistant. The given assistant output is specific to the given persona and is generated without the need to later adapt the given assistant output to the given persona. The given assistant output includes a stream of text content synthesized for audible presentation to the user and a stream of visual cues used to control the client device's display and / or the automated assistant's visual representation. The given assistant output is made to reflect a given persona using a Large-Scale Language Model (LLM) or output previously generated using an LLM.
Owner:GOOGLE LLC

AIGC-based text, travel, micro and short play content generation method, electronic equipment and medium

The invention relates to an AIGC-based text travel micro and short play content generation method, electronic equipment and a medium, and the method comprises the steps: based on a first instruction sent by a user, according to a first text travel preset knowledge base, obtaining a primary text corresponding to the first instruction; based on a preset language model, according to the primary text, obtaining a split script; based on the lines, according to an emotion model, obtaining an emotion type of the lines; inputting the visual cue word into a text video model, and obtaining a video file of each sub-lens through the text video model; according to the roles, the corresponding lines and the emotion types of the lines, obtaining an audio file of each role based on a text sound generation model; and based on a preset fusion strategy, fusing the video file and all the audio files to obtain the content of the text travel, micro and short play. According to the method and the device, the movie generation cost is reduced, and the generation period is shortened.
Owner:CHONGQING INST OF ENG