Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

194 results about "Text alignment" patented technology

Data analysis problem generation method based on image input and large model combination

The invention discloses a data analysis problem generation method based on image input and large model combination, and the method comprises the following steps: S1, obtaining original image data, and carrying out the preprocessing of the original image data; s2, constructing a visual language model based on a Qwen-VL architecture, and performing bidirectional alignment to generate a visual feature vector; s3, constructing a cue word template, and fusing the cue word template through a cross attention mechanism to generate structured text description; s4, constructing a large language model based on a Transform architecture, and performing supervision and fine tuning by adopting a LoRA method to generate a data analysis problem candidate sequence; s5, semantic consistency detection and structural rule matching are carried out, and sequences which do not meet semantic specifications or structural constraints are removed; and S6, constructing an image-text alignment triple, and writing the image-text alignment triple into the data structure in the JSON format for coding and storage. According to the method, the image can be converted into a data analysis problem, and the text generation quality and efficiency are remarkably improved.
Owner:HUAZHONG NORMAL UNIV

Intelligent short message template generation method and system based on big language AI model

The invention relates to an intelligent short message template generation method and system based on a large language AI model, and the method comprises the steps: extracting cross-cultural communication frequency, sentence pattern preference analysis results and screen size adaptation data from social media interaction data and equipment use logs of target audiences, and generating a cultural background distribution matrix; identifying regional culture features in the culture background distribution matrix to obtain a taboo matching result; generating an expression habit preference matrix according to the updated culture background distribution matrix; generating a preliminary short message template according to the expression habit preference matrix and the symbol semantic mapping table, and outputting an adjusted template set conforming to message length control; analyzing the acceptability of the adjusted template set on the language conciseness requirement and the interactive button layout, and generating a new template set; evaluating a text alignment mode of the new template set; and screening candidate short message templates according to the information definition score. According to the invention, the accuracy and acceptability of cross-culture short message communication can be effectively improved, and the participation degree of audiences is enhanced.
Owner:GUANGDONG BOJIN INFORMATION TECHNOLOGY GROUP CO LTD

Cross-language text fusion intelligent alignment method and system

The invention relates to the technical field of cross-language information processing, and provides a cross-language text fusion intelligent alignment method and system.The cross-language text fusion intelligent alignment method comprises the steps that a hierarchical alignment model is constructed through preprocessing and label recognition of multi-coding-type texts, deep semantic feature extraction and labeling of a multi-language pre-training model and text semantic and format information analysis; transform is taken as a core, cross-language semantic association is enhanced through a multi-head attention mechanism, character-level and paragraph-level format collaboration is realized through label weight allocation and condition constraint, the format alignment accuracy is obviously improved in a multi-language mixed typesetting scene, and the multi-language mixed typesetting efficiency is improved. The analysis capability of the model on the structured information can be extended to layout relation processing of texts, images and tables, so that the comprehensive alignment efficiency in a multi-modal fusion scene is remarkably improved. The accuracy of text and label fusion is guaranteed, the problem of confusion of messy codes and labels is avoided, and high-precision cross-language text alignment from semantics to formats, from single mode to multiple modes and from semantics to formats is achieved.
Owner:SHANGHAI MEGALIN SOFTWARE TECH CO LTD

Reloading pedestrian re-identification method and system based on visual language pre-training model

The invention relates to the technical field of computer vision, in particular to a reloading pedestrian re-identification method based on a visual language pre-training model. The method comprises the following steps: a training stage: inputting an image to obtain a clothing mask image, and generating clothing irrelevant / relevant prompts; text encoder parameters are fixed, prompt weights are optimized, text prompts are input into an encoder to obtain features, and a classifier is constructed to achieve image-text alignment through cross entropy loss; using a visual encoder to extract mask pattern features, and constraining a class center Euclidean distance to realize image-image alignment; stripping clothes characteristics: extracting clothes area characteristics and corresponding text characteristics, optimizing by a classifier, and introducing orthogonal loss to decouple clothes correlation; in the reasoning stage, a query image is input into a trained image encoder to extract features, and cosine similarity ranking and result returning are calculated according to the features of the image library. According to the technical scheme, the recognition accuracy of the pedestrian re-recognition method under the condition of pedestrian clothing change can be improved.
Owner:重庆脑与智能科学中心

Text generation image diffusion model enhancement method based on multi-target preference optimization

The invention discloses a text generation image diffusion model enhancement method based on multi-target preference optimization. The method comprises the following steps: firstly, determining a plurality of reward models, and constructing a sample pair training set comprising positive and negative samples; and then, generating a loss weight of each sample pair in the sample pair training set, performing fine tuning training on the text map diffusion model by using the sample pair training set, and in the fine tuning training process, calculating a loss function value of each sample pair in combination with the loss weight of each sample pair until the training is completed, thereby obtaining an aligned text map diffusion model. According to the method provided by the invention, manual data annotation is not needed, the problems of preference inconsistency and over-optimization in a multi-reward scene are effectively solved, and the image quality, the text alignment capability and the multi-target optimization performance of the text-to-image generation model are remarkably improved. The method is superior to an existing optimization method under single-reward and multi-reward setting, shows higher generation quality and robustness, and can be seamlessly applied to various picture generation models.
Owner:ZHEJIANG UNIV

Landslide classification method and system based on visual language model and cross attention mechanism

The invention provides a landslide classification method and system based on a visual language model and a cross attention mechanism. The method and the system specifically comprise the following steps: data preprocessing: carrying out Canny edge detection on an RGB image, and calculating terrain attributes such as a gradient and a slope direction for a DEM (Digital Elevation Model); feature extraction: capturing local features by adopting a reflection filling convolution layer and multi-scale residual connection; a visual language model is introduced, wherein semantic enhancement features are extracted through image-text alignment by means of the visual language model; cross self-attention fusion: capturing a global context through self-attention, and focusing heterogenous data complementary information by cross attention; and classifying and outputting: outputting a result by using global average pooling and a linear classifier. Through the visual language model and the cross self-attention mechanism, the landslide recognition capability under the complex terrain is effectively improved, an efficient and reliable technical means is provided for geological disaster monitoring, and the method can be widely applied to the fields of landslide recognition, risk assessment and the like.
Owner:福州海洋研究院 +3

Image-text processing method and device for marking compression framework

The invention discloses an image-text processing method and device for marking a compression framework. The method comprises the following steps of: extracting visual features; a visual mark screening processing step; a text feature extraction step; and a multi-modal fusion and model processing step. The method has the beneficial effects that the inference efficiency of the MLLMs is remarkably improved under the condition that the visual mark compression frame does not need additional training; through global and local information fusion of the DVTS module and text guide supplement of the TGVC module, the number of visual marks is greatly reduced, meanwhile, key visual information is reserved, and visual-text alignment is enhanced; experiments show that in various image and video benchmark tests, compared with an existing method, the framework has the advantages that the calculation cost is greatly reduced, the model performance is maintained and even improved, and the framework has remarkable technical advantages and application potential.
Owner:ZHEJIANG YOULU ROBOT TECH CO LTD

Training-free consistent text-to-video generation

Embodiments of the present disclosure relate to training-free consistent text-to-image generation. A pre-trained text-to-image diffusion model is leveraged to generate images depicting a consistent subject for diverse prompts describing scenes. Inputs to the model are a text description of at least one subject with prompts (scene text descriptions) describing scenes, where each prompt is associated with a different generated image and the text description is used for all images that depict the subject. Internal activations (intermediate data) computed by the model during generation of the different images are shared for generation of the different images. A subject-driven shared attention block and correspondence-based feature injection are incorporated into the model to promote subject consistency within each image and / or between images. Additionally, layout diversity is encouraged while maintaining subject consistency. The model achieves state-of-the-art performance on subject consistency and text alignment, without requiring any optimization and naturally extends to multi-subject scenarios.
Owner:NVIDIA CORP

Intelligent decision-making method based on multi-modal large model and related equipment

The invention provides an intelligent decision-making method based on a multi-modal large model and related equipment, which are applied to an artificial intelligence technology, and the method comprises the steps: collecting a current field image of a to-be-decided region, and carrying out the preprocessing of the current field image, and obtaining a target field image; inputting the target field image into a pre-trained vision-text alignment model, and performing prediction by the vision-text alignment model by using the target field image to obtain a disease and pest diagnosis scheme corresponding to the to-be-decided region; dividing the to-be-decided region into a plurality of partitions, and processing current data of each mode in current multi-mode data of the partitions to obtain a space-time sequence of each mode; inputting the space-time sequence of each modal of the subarea into a pre-trained irrigation decision model, and performing prediction by the irrigation decision model according to the space-time sequence of each modal to obtain irrigation data of the subarea; the irrigation strategy of each subarea is generated according to the disease and pest diagnosis scheme and the irrigation data of each subarea. The cost can be reduced, and the generation efficiency and accuracy of the irrigation strategy can be improved.
Owner:INSPUR ENTERPRISE CLOUD TECHNOLOGY (SHANDONG) CO LTD

Cross-modal image-text retrieval method and device based on pulse fusion

The invention discloses a cross-modal image-text retrieval method based on pulse fusion, and belongs to the technical field of image-text retrieval. The method comprises the following steps: inputting a target image set into a target detection network to obtain an image floating point code, and inputting a target text set into a word segmentation device to obtain a word floating point code; inputting the image floating point code and the word floating point code into a pulse encoder to obtain a first image pulse code and a first text pulse code; inputting the first image pulse code and the first text pulse code into a pulse cross attention fusion module to obtain a second image pulse code and a second text pulse code; respectively carrying out weighted accumulation and average pooling on the second image pulse code and the second text pulse code to obtain an image floating point feature vector set and a text floating point feature vector set, and calculating cosine similarity to obtain an image-text alignment result; and the text retrieval result of each image in the target image set is obtained based on the image-text alignment result, so that the accuracy and efficiency of image-text retrieval are improved.
Owner:WUHAN UNIV OF TECH

Multi-modal data pairing method and system based on deep learning

The invention provides a multi-modal data pairing method and system based on deep learning, and relates to the technical field of data processing, and the method comprises the steps: obtaining a video multi-frame sequence and a target text, and respectively extracting an overlapped frame group set and a standardized text sequence; performing spatio-temporal feature extraction and text dependency relationship coding to obtain a video time sequence vector sequence and a text vector sequence; executing cross-modal alignment search, and constructing a monotonic matching path set; calculating a semantic and action entity relationship consistency score of the paired elements on the path to obtain a comprehensive score; and determining an alignment relationship between the video and the text based on the optimal path. According to the method, accurate matching of the video and the text is realized, and the cross-modal retrieval efficiency is improved.
Owner:BEIJING YIZHUANG INTELLIGENT CITY RES INST GRP CO LTD

Open vocabulary semantic segmentation method and device based on feature interaction and multi-modal data fusion

The invention provides an open vocabulary semantic segmentation method and device based on feature interaction and multi-modal data fusion. An SAM encoder of freezing parameters is adopted to extract RGB image and Mask image features in parallel, and boundary information enhancement is carried out through a feature fusion module. A CLIP image encoder of freezing parameters is used for carrying out multi-layer feature extraction on an RGB image, and multi-scale semantic information optimization is carried out through a feature enhancement module. And fusing image features extracted by the SAM branch and the CLIP branch through a feature interaction module to realize cross-network feature complementation. In a prediction stage, a category integration strategy combined with a temperature scaling operation is introduced to optimize prediction effects of known and unknown categories. According to the method, the advantages of SAM in the aspect of image boundary extraction and the image-text alignment capability of CLIP are combined, and the generalization capability and robustness of the model in an open vocabulary semantic segmentation task are effectively improved through multi-modal feature fusion and interaction.
Owner:BEIJING TECH & BUSINESS UNIV

Speech generation based on sparse speech-text alignment

Embodiments of the disclosure provide a solution for speech generation. A method includes: determining, based on a target text, a plurality of phoneme feature representations corresponding to a sequence of phonemes in the target text and respective phoneme durations for the plurality of phoneme feature representations; extending the plurality of phoneme feature representations based on the respective phoneme durations, to obtain an extended sequence of phoneme feature representations; masking at least one phoneme feature representation in the extended sequence of phoneme feature representations, to obtain a sequence of masked phoneme feature representations; and generating a target speech corresponding to the target text at least based on the sequence of masked phoneme feature representations.
Owner:BYTEDANCE TECHNOLOGY LTD +1

Vision-language model training method and device and related equipment

The invention provides a visual-language model training method and device and related equipment, and the method comprises the steps: constructing a training data set comprising an image-text alignment task, a text-driven visual positioning task and a plain text reasoning task through determining a visual encoder and a heterogeneous language model which form a multi-modal heterogeneous recognition model; carrying out dimension dynamic alignment adaptation on the heterogeneous language model, enabling a parameter architecture of the heterogeneous language model to be matched with a visual feature dimension output by a visual encoder, and carrying out supervision and fine adjustment on the visual-language model by adopting a freezing-unfreezing two-stage training strategy based on a training data set; and performing parameter training on a cross-modal alignment module connected with the visual encoder and the heterogeneous language model so as to obtain a trained visual-language model. According to the training method, time consumed by model training is saved, and the convergence speed is increased. Through dimension dynamic alignment and hierarchical weight mapping, the pre-training language ability is reserved to the maximum extent, the plain text task performance loss is reduced, and disastrous forgetting of the model is avoided.
Owner:AISINO CORPORATION

Image-text alignment model training method and device, computer equipment and storage medium

The invention discloses an image-text alignment model training method and device, computer equipment and a storage medium, and the method comprises the steps: carrying out the coding of input image data and text data, and obtaining an image block feature sequence and a text token feature sequence; performing bidirectional feature interaction on the image block feature sequence and the text token feature sequence to obtain fused image features and fused text prompt features; inputting the fused image features, the fused text prompt features and a preset output token into an SAM mask decoder to generate an image mask; filtering the fused image features based on the image mask, extracting image global features related to the text description, and obtaining text global features; constructing a joint loss function; and performing end-to-end training on the model based on the joint loss function until the model converges. According to the method, the limitation of coarse-grained alignment of a traditional model is effectively solved, and the pertinence and matching precision of image-text features are improved.
Owner:SHENZHEN TVT DIGITAL TECH CO LTD

End-to-end automatic driving long tail identification method based on comparative learning pre-training

The invention relates to the technical field of automatic driving end-to-end perception, in particular to an end-to-end automatic driving long tail recognition method based on comparative learning pre-training, and the method comprises the steps: firstly generating synthetic image data with long tail distribution characteristics through a conditional diffusion model; a fine-grained scene classifier is adopted to carry out systematic arrangement and semantic annotation on the generated samples, and a structured multi-modal image-text alignment data set is constructed; and finally, fusing the enhanced data set with the original training set, and optimizing a vision-language joint embedding space through a multi-task contrast loss function to realize parameter updating of the pre-training model. According to the method, a closed-loop optimization mechanism of a generative data enhancement and contrast learning framework is creatively established, the problem of data scarcity in a long-tail distribution scene is effectively relieved, and the cross-modal representation capability and downstream task generalization performance of the model on low resource categories are remarkably improved.
Owner:JIANGSU UNIV

Multi-modal fine-grained semantic alignment method and device based on graph neural network

The invention discloses a multi-modal fine-grained semantic alignment method and device based on a graph neural network, and relates to the field of multi-modal deep learning. Firstly, deep feature extraction is performed on input multi-modal original data, and then word-level text features and local image features are constructed into a cross-modal graph structure. And performing weighted aggregation on node neighborhood information of the cross-modal graph structure through the graph attention network. And finally, carrying out weighted fusion on the text alignment features and the image alignment features. According to the cross-modal feature fusion method, the word-level text features and the local image features are uniformly abstracted into the graph structure nodes for refined alignment, a more accurate cross-modal semantic corresponding relation can be captured, the heterogeneity problem in expression modes and semantic structures is effectively relieved, and the accuracy and reliability of cross-modal feature fusion are improved. The graph attention network can adaptively adjust the weight distribution of information propagation, highlights the effect of key features in the alignment process, and ensures that the model makes full use of important semantic relationships.
Owner:ZHENGZHOU NORMAL UNIV +1

Bidding and cross-bidding early warning and blocking method and system based on multi-modal large model

The invention relates to the technical field of artificial intelligence, in particular to a multi-modal large model-based bidding and cross-bidding early warning and blocking method and system, and the method comprises the following steps: receiving and preprocessing various qualification files, including business licenses, industry licenses, auditing reports and bidding files, submitted by suppliers; applying a document layout analysis model based on joint training of a visual coding network and a text alignment network; the method has the beneficial effects that a scanned certificate, an audit report and a bidding file are automatically analyzed into structured semantic fragments; performing same semantic space alignment on the fragments, supplier basic information, historical records and real-time captured data of an authoritative authentication website; any contradiction or gap among different sources of key information such as registration capital, permission items, financial indexes and the like can be quickly found, and interpretable risk scores and processing suggestions can be given; on the premise of ensuring the private security of the data, the qualification auditing efficiency and accuracy are improved, and the workload of manual checking is reduced.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

Vision and text alignment method and system

The invention discloses a vision and text alignment method and system, and belongs to the technical field of artificial intelligence and multi-modal semantic understanding. In order to solve the problem of insufficient vision and language deep fusion in the existing multi-mode question answering, the method mainly comprises the following steps: mapping visual features to a self-attention input space of a language model through a perceptron network, and introducing a fusion attention mechanism into each layer of decoder of the language model to realize layer-by-layer interaction processing of vision and texts. According to the method, deep alignment and fusion of visual information and text semantics can be realized, and the understanding and generation capabilities of a multi-modal question-answering system are improved.
Owner:INST OF SOFTWARE - CHINESE ACAD OF SCI

Cross-modal image-text pedestrian retrieval method based on generative model

The invention relates to a cross-modal retrieval technology, in particular to a cross-modal image-text pedestrian retrieval method based on a generative model, which comprises the following steps: acquiring text description information and a corresponding original pedestrian image; calling a diffusion generation model to generate an intermediate image based on the text description; original image features and text features are extracted through an image encoder and a text encoder respectively, and semantic-fused text feature representation is constructed; calculating the similarity between the image feature representation and the text feature representation to obtain an image-text matching score; introducing and generating an intermediate image as a semantic bridge, and realizing multi-modal fusion among the image, the text and the intermediate image through a cross attention mechanism to obtain fused image and text features; and using the fusion features to train a recognition model, and finally outputting a pedestrian image retrieval result corresponding to the text description. According to the method, the image-text alignment precision can be improved, the robustness of a scene with incomplete text description is enhanced, and the recognition accuracy in a cross-modal retrieval task is improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Multi-modal power sample feature migration method and system based on dual cross-modal information decoupling, electronic equipment and storage medium

The invention discloses a multi-modal power sample feature migration method and system based on dual cross-modal information decoupling, electronic equipment and a storage medium, which are applied to the technical field of smart power grids, and comprise the following steps: carrying out hierarchical alignment of visual features and language features by adopting a fine-grained image-text alignment strategy, and capturing distinctive multi-modal representation; based on a double-flow attention mechanism, performing enhanced feature extraction on the multi-modal data; wherein the double-flow attention mechanism comprises an image self-adaptive attention mechanism and a text guide attention mechanism; the method comprises the following steps: performing cross-modal information decoupling processing on multi-modal data, separating to obtain semantic features representing general semantics and modal specific features representing modal features, and performing fine-grained alignment recombination processing to obtain cross-modal fusion features; and migrating the deeply fused multi-modal power sample features to a downstream task of the power system. According to the method, the understanding of the multi-modal features and the further improvement of the migration capability and the efficient operation of the power system tasks are realized.
Owner:STATE GRID INFORMATION & TELECOMM BRANCH +2

Cross-modal semantic alignment method, device and equipment based on fine-grained semantic decoupling

The invention provides a cross-modal semantic alignment method, device and equipment based on fine-grained semantic decoupling, and the method comprises the steps: inputting a to-be-aligned image into a pre-training image encoder of a cross-modal semantic alignment model, and obtaining a global semantic feature, meanwhile, a semantic feature extraction module is used in a high-level encoding block of the image encoder and based on a preset query vector, high-level visual feature extraction is conducted, visual fine-grained semantic features are obtained and input into a semantic feature decoupling module to be decoupled, and fine-grained semantic features of attributes, objects and combinations are obtained; inputting prompt words corresponding to attributes, objects and combinations in the to-be-aligned image into a text encoder for text feature extraction to obtain soft prompt text features of the attributes, the objects and the combinations; and inputting the fine-grained semantic features and the soft prompt text features corresponding to the global semantic features, the attributes, the objects and the combinations into a loss calculation module for image-text alignment, and performing linear processing on an alignment result to obtain an alignment prediction result.
Owner:GUIZHOU UNIV

Image text alignment method based on multi-modal large language model

PendingCN121542772AText alignmentFeature vector
The invention relates to the technical field of image texts, in particular to an image text alignment method based on a multi-modal large language model, which comprises the following steps of: segmenting an image into local areas and encoding the local areas into visual feature vectors by adopting a pre-trained visual encoder; visual and text features are mapped to the same dimension space through a learning projection layer, an InfoNCE loss function is adopted to maximize the similarity of a positive sample pair, dynamic interaction between text query and an image area is achieved through a cross attention mechanism, language module parameters are frozen to pre-train a visual encoder, an overall module is combined to be fine-tuned, and text query is completed. An error case is analyzed, and a verified and improved alignment module is obtained through a data enhancement optimization module. The method solves the problems that a traditional image text alignment method is incomplete in manual feature extraction and low in precision, an existing deep learning model is limited in cross-modal semantic understanding ability, and the generalization ability is poor when facing long-tail data.
Owner:GUANGDONG UNIVERSITY OF FOREIGN STUDIES

Illusion phenomenon mitigation method and system of visual language large model, terminal and medium

The invention discloses an illusion phenomenon mitigation method and system for a visual language large model, a terminal and a medium, and the method comprises the steps: obtaining input image-text data, processing the image-text data to obtain a text token and an image token, and aligning the text token with the image token to obtain an input sequence of the image-text data; inputting the input sequence into the visual language large model, calculating attention distribution through a cross-modal attention mechanism, and obtaining initial vocabulary probability distribution of the current time step based on the attention distribution; and based on the initial vocabulary probability distribution and the attention distribution, performing multi-scale enhancement and fusion reasoning to obtain target vocabulary probability distribution of the current time step. According to the method, through key technologies such as image-text alignment, attention guidance, multi-scale enhancement and fusion reasoning, the understanding and generation capability of the visual language model in a complex image-text scene is effectively improved, and the illusion problem possibly occurring in the cross-modal reasoning process of the large visual language model is relieved.
Owner:SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY

Time sequence prediction method and device for performing multi-level text alignment by using large model

The invention provides a time sequence prediction method and device for performing multi-level text alignment by using a large model, and belongs to the technical field of time sequence prediction of a rail transit system based on the large model. Comprising the following steps of splitting multivariable time sequence input into a plurality of univariate time sequences according to feature dimensions, performing additive decomposition on each univariate time sequence, and performing fragmentation processing on each decomposed time sequence component; embedding the fragments into a text embedding space of a pre-training language model and aligning the fragments with the text embedding space; combining the structured prompt with the aligned time sequence representation to form input of a large model; and feeding the input combining the prompt and the alignment representation into the frozen large language model, obtaining an output representation of the model, and mapping the output representation into a final prediction result through a linear projection layer. According to the method, the time sequence data and the natural language modality are effectively aligned and fused, and the prediction accuracy and interpretability are remarkably improved.
Owner:CRRC CHANGCHUN RAILWAY VEHICLES CO LTD

Weak supervision video anomaly detection method and system based on prompt learning

The invention provides a weak supervision video anomaly detection method and system based on prompt learning, and belongs to the technical field of abnormal event detection based on computer vision, and the method comprises the steps: obtaining to-be-processed video data; and processing the acquired to-be-processed video data by using a pre-trained anomaly detection model to obtain a specific classification result of the abnormal events in the video. According to the invention, a video local and global adaptive time modeling module is introduced to capture local and global dependency relationships at the same time, and the relationship between the demand of detailed time modeling and the calculation efficiency is balanced; by utilizing an external knowledge base, the distinguishing capability of the model on different categories is improved; according to the method, a text-video comparison loss function is designed, the similarity of a correctly matched text-video pair is enhanced, the similarity of wrong matching is reduced, and too high similarity of a negative sample is effectively inhibited, so that the distinguishing capability of the model is improved, the matching of the text and the video is more accurate, and the video and text alignment capability of the model is enhanced.
Owner:BEIJING JIAOTONG UNIV

Text-driven video generation method based on single video fine tuning

The invention relates to the technical field of video generation and machine learning, and provides a text-driven video generation method based on single video fine tuning, which comprises the following steps: encoding an input video into potential feature representation through a pre-training encoder, and constructing input features; generating a video frame by adopting a video generation network based on a diffusion model, and optimizing the generation process of the video frame by adopting a pre-trained CLIP model and a pre-trained VGG model; and based on the dynamic loss weight adjustment strategy, adaptively balancing the weight of the semantic and texture constraint, and optimizing the training process of the CLIP model and the VGG model. According to the method provided by the invention, a multi-scale space adapter, semantic features based on CLIP and texture features based on VGG are combined. Through a progressive training strategy and dynamic loss weight balance, the quality and consistency of generated frames are improved, text alignment, time coherence and generalization ability are remarkably improved, and the development of the text to video generation field is promoted.
Owner:CHONGQING UNIV OF TECH

Open vocabulary action recognition method and system based on heterogeneous skeleton

The invention discloses a heterogeneous skeleton-based open vocabulary action recognition method and system. The method comprises the following steps of: constructing a heterogeneous open vocabulary skeleton data set; unifying the skeleton representation in the heterogeneous open vocabulary skeleton data set, and defining a unified space structure by establishing the maximum joint number and member number; a skeleton motion encoder model based on a Transform architecture is constructed, the skeleton motion encoder model comprises feature embedding, space-time encoding and a projection layer used for cross-modal alignment, and global time features, global space features and global visual features output by space-time encoding are subjected to cross-modal alignment through three parallel projection networks of the projection layer. Mapping to a semantic space aligned with text embedding from the pre-trained language model; training loss is constructed based on a multi-granularity motion-text alignment strategy including global instance alignment, stream specific alignment and fine granularity alignment, and a motion encoder model is trained. A wide range of experiments on a popular benchmark with heterogeneous skeleton data prove the effectiveness and generalization ability of the proposed method.
Owner:SOUTHEAST UNIV

Speech synthesis method and system for controllable latent variable modeling based on semantic distillation

The invention relates to the technical field of speech synthesis, and particularly discloses a speech synthesis method and system for controllable latent variable modeling based on semantic distillation, and the method comprises the steps: converting a Mel spectrum into continuous latent variable distribution through a speech coding module, generating continuous latent variables through re-parameterization sampling, introducing a self-supervised model for semantic distillation, and carrying out the semantic distillation. According to the method, alignment of latent variables and semantic features is constrained through marginal cosine similarity and distance matrix structure loss, a text encoder maps a phoneme sequence into latent variable distribution, time sequence alignment of a text and the latent variables is achieved in combination with monotonic alignment search, and a decoder reconstructs the latent variables into a Mel spectrum. According to the method, waveform synthesis through a vocoder and total loss function joint optimization reconstruction, KL divergence, distillation, text alignment and confrontation loss are carried out, discrete information loss is avoided through continuous latent variable modeling, semantic consistency and text alignment efficiency are enhanced, the naturalness, coherence and real-time performance of synthesized voice are improved, and the method is suitable for scenes such as voice assistants and virtual anchors.
Owner:BEIJING TIMES RUILANG TECH CO LTD

Machine-learned text alignment prediction for providing an augmented-reality translation interface

Systems and methods for providing an augmented-reality translation interface can include obtaining an image, processing the image to generate an image representation and one or more paragraph bounding boxes, and processing the image representation and the one or more paragraph bounding boxes with a machine-learned alignment classification model to generate one or more alignment classifications. In parallel or in series, text from the image can be determined and translated. The image, the translated text, and the one or more alignment classifications can then be processed to generate and provide the translation in an augmented-reality interface.
Owner:GOOGLE LLC