Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

52 results about "Dynamic text" patented technology

Dynamic text is text on a map layout that changes based on the current properties of the project, map frame, map, and so on. When you insert a piece of dynamic text, it automatically displays the current value of its respective property. When that property is updated, the dynamic text automatically updates.

Personalized and dynamic text to speech voice cloning using incompletely trained text to speech models

Systems and methods are provided for machine learning models configured as zero-shot personalized text-to-speech models which comprise a feature extractor, a speaker encoder, and a text-to-speech module. The feature extractor is configured to extract acoustic features and prosodic features from new target reference speech associated with the new target speaker. The speaker encoder is configured to generate a speaker embedding corresponding to the new target speaker based on the acoustic features extracted from the new target reference speech. The text-to-speech module is configured to generate the personalized voice corresponding for the new target speaker based on the speaker embedding and the prosodic features extracted from the new target reference speech without applying the text-to-speech module on new labeled training data associated with the new target speaker.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Hydrofracture question-answering system and method based on cross-language retrieval enhanced generation

The invention discloses a system and a method for enhancing generation of hydrofracture questions and answers based on cross-language retrieval. The knowledge question-answering system aiming at the hydraulic fracturing field is developed by utilizing a cross-language retrieval enhancement generation technology, so that a user can obtain an optimal answer and a corresponding source from a multi-language knowledge base by matching only by using Chinese retrieval; the complex process that a user needs to find useful information from numerous and jumbled papers and reports is changed. According to the system, a dynamic text vector library framework is adopted, multi-language and complex document formats are supported, the analysis function of tables and formulas is integrated, documents in the hydraulic fracturing field are stored in the same knowledge base, the knowledge base is continuously updated along with document expansion in the professional field, a user can inquire related problems in the field in the system, and the user experience is improved. By integrating and updating a large amount of document information, required professional knowledge in the field can be efficiently obtained.
Owner:ZHEJIANG UNIV

Multi-degraded image restoration method based on semantic guidance

The invention discloses a multi-degraded image restoration method based on semantic guidance. According to the method, for degraded images captured by a vehicle-mounted camera under severe weather conditions, potential features of the images are extracted by adopting a trunk network based on Transform, multi-modal semantic information is extracted in combination with a CLIP visual language model, and semantic guidance is provided for different degradation types through a dynamic text prompt generation mechanism. A self-adaptive feature fusion module is designed, channel attention and space attention mechanisms are combined to realize effective integration of multi-modal features, and a degradation feature extraction and fusion module is introduced to enhance the generalization ability of the model. According to the semantic guidance multi-type image recovery network SGIRN provided by the invention, the global modeling capability of the Transform and the cross-modal representation capability of the CLIP visual language model are combined, so that high-quality recovery of various weather degradation types is realized.
Owner:SHENYANG INST OF COMPUTING TECH CO LTD THE CHINESE ACAD OF SCI

Thermal runaway prediction method and device of power battery, storage medium and electronic equipment

The invention provides a thermal runaway prediction method and device for a power battery, a storage medium and electronic equipment, and the method comprises the steps: obtaining a time sequence of a plurality of operation parameters of the power battery in a first historical time period, and determining a plurality of embedded matrixes according to the time sequence, the plurality of operation parameters comprising the voltage and temperature of the power battery; obtaining a static text matrix and a dynamic text matrix of the power battery; and obtaining a fusion data matrix according to the embedded matrix, the static text matrix and the dynamic text matrix, and inputting the fusion data matrix into a target large language model to predict an operation parameter prediction value of the power battery in a future time period so as to predict the thermal runaway state of the power battery. The operation parameter prediction value comprises a voltage prediction value and a temperature prediction value of the power battery. According to the method, through the steps of time sequence data reprogramming, cross-modal information fusion, large language model intelligent analysis and the like, all-around perception and risk assessment of the running state of the battery are achieved.
Owner:WEICHAI POWER CO LTD

Allergic rhinitis diagnosis and allergen tracing system based on dynamic text guidance

The invention discloses an allergic rhinitis diagnosis and allergen traceability system based on dynamic text guidance, and belongs to the technical field of intelligent medical treatment. According to the method, the problems that in the prior art, patient description is inaccurate, allergens are various in variety and have differences, and allergens are difficult to accurately determine and trace are solved, the initial situational sub-graph is verified by introducing periodic features of the patient, the causal contribution degree of the initial situational sub-graph is evaluated through a Bayesian reasoning algorithm, and a mode association result library is formed; according to the method, the crossing of the allergic rhinitis diagnosis and tracing from the traditional static judgment to the dynamic accurate inference is realized, and a personalized and scientific allergen avoidance scheme is provided for patients; through similarity retrieval, environmental information injection verification and calculation of a mode goodness-of-fit score, under the condition that patient symptoms are in multi-factor mixing, a mixed causal graph is created through a graph fusion technology, and the mode goodness-of-fit score is calculated again, so that the accuracy of obtaining an allergen traceability result in practical application is ensured.
Owner:THE THIRD MEDICAL CENT OF THE CHINESE PEOPLES LIBERATION ARMY GENERAL HOSPITAL

Industrial dynamic tracking semantic analysis method and system based on artificial intelligence

The invention provides an industry dynamic tracking semantic analysis method and system based on artificial intelligence, and the method comprises the steps: firstly obtaining an industry dynamic text set which comprises a plurality of industry dynamic text units with source identifiers, constructing source semantic hierarchical mapping based on the source identifiers, and obtaining a hierarchical association relationship between the text units; calling the pre-training model to generate a semantic association rule set, constructing a dynamic semantic tracking link according to the semantic association rule set, extracting a core semantic development direction to determine a core tracking dimension, and integrating semantic information according to the dimension to generate an industry dynamic analysis report, so that the industry dynamic semantics can be efficiently and accurately tracked; and a comprehensive and accurate industry analysis result is provided for a user.
Owner:BEIXI INTELLIGENT FUTURE (SHANGHAI) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Wild animal posture estimation method based on dynamic prompt and semi-supervised multi-mode learning

The invention discloses a wild animal attitude estimation method based on dynamic prompt and semi-supervised multi-mode learning, and relates to the technical field of computer vision and animal behavior analysis, and the method comprises the steps: firstly carrying out the motion state analysis of an input video, so as to extract parameters such as the centroid velocity, the acceleration and the steering angle of a moving material; the method comprises the following steps: generating a dynamic text prompt containing directional semantics, constructing a cross-modal fusion time sequence convolutional network, fusing visual features and text prompt features through a space alignment mask mechanism and a cross attention module, capturing time sequence dependence by using a bidirectional ConvGRU network to realize attitude estimation, and finally, based on a semi-supervised multi-task learning framework, carrying out dynamic text prompt processing on the basis of the semi-supervised multi-task learning framework. A joint loss function including supervision loss, time difference loss, attitude PCA reconstruction loss, multi-modal alignment loss and motion consistency loss is adopted, a loss weight is adaptively adjusted in combination with prediction uncertainty to optimize network parameters, a smooth 2D / 3D animal skeleton is finally output, real-time reasoning can be realized on edge equipment, and the real-time reasoning efficiency is improved. Ecological analysis software is compatible, high precision is still kept under the condition of low-label data, and the performance is remarkably improved especially in an animal sharp turning scene.
Owner:BEIJING FORESTRY UNIVERSITY

Dynamic text message processing implementing endpoint communication channel selection

The present disclosure relates generally to providing a concierge service to handle a wide variety of topics and user intents via a text messaging interface. The concierge service can be part of a connection management system that can dynamically manage and facilitate natural language conversations between a user making a request or providing an instruction and one or more endpoints for the purposes of fulfilling the request or instruction.
Owner:LIVEPERSON INC

A Chromosomal Abnormality Detection Method Based on a Multimodal Large Model

This invention relates to the field of chromosome abnormality recognition technology, specifically to a chromosome abnormality detection method based on a multimodal large model. The method includes: constructing an image dataset and a text dataset; fusing image features and text features to obtain multimodal fusion features; assigning anomaly scores to image blocks belonging to band regions based on anomaly scoring rules formulated from the multimodal fusion features, and comprehensively processing the anomaly scores of all image blocks corresponding to the chromosome to determine whether the chromosome image is abnormal; and decoding and generating natural language text that meets the requirements of chromosome abnormality detection based on the multimodal fusion feature representation and anomaly scoring rules. This invention achieves accurate chromosome abnormality detection and band location positioning through multimodal fusion and dynamic text generation mechanisms, generating interpretable natural language text descriptions, and improving the practicality and interpretability of the detection results.
Owner:笑纳科技(苏州)有限公司

Method and system for enhancing reasoning stability of large model in text scene

The invention discloses a large model reasoning stability enhancement method and system in a text scene, and relates to the technical field of knowledge enhancement deep learning, and the method comprises the steps: building a structured text knowledge graph based on a text knowledge base, and generating a dynamic text knowledge embedding matrix through a graph attention network; inserting a text knowledge gating cross attention module into a decoder selection layer of the pre-trained large model, taking the text knowledge gating cross attention module as an external knowledge source, obtaining a knowledge enhanced hidden state after gating fusion, and constructing a transformation model; semantic equivalent perturbation is carried out on an input text to obtain a perturbation sample, the perturbation sample is input into the transformation model in parallel to obtain extraction probability distribution, and divergence and gradient direction consistency loss between two distributions are calculated; and combining cross entropy loss and gradient direction consistency loss to train and transform the model, and updating parameters to convergence to obtain a final large model. According to the method, the fact consistency of output can be improved, common optimization of knowledge guidance and stability constraint is realized, and the result is accurate and reliable.
Owner:DIGITAL HEALTH CHINA TECHNOLOGIES CO LTD

A lesion segmentation method of adaptive dynamic text prompt

The present application relates to the technical field of image processing, and more particularly to a lesion segmentation method of adaptive dynamic text prompt, comprising: acquiring a lesion image, and generating an adaptive dynamic text prompt based on the lesion image; constructing a multi-modal enhanced fusion text prompt adapter based on a FiLM global channel recalibration module, a spatial cross-attention interaction module and a zero initialization gate nonlinear integration module cascade; the global channel recalibration module of FiLM utilizes a feature linear modulation mechanism to perform channel-level weighting on a visual feature map; before the feature enters spatial interaction, according to the text semantic enhancement band response, the activation of the background noise channel is inhibited. The present application overcomes the problems of insufficient utilization of multi-modal prompts, dependence on artificial static labeling of text prompts, serious dependence on prior information in the reasoning stage and difficulty in multi-modal feature fusion in a low-contrast environment in the existing medical image segmentation technology.
Owner:CHANGZHOU UNIV

A text guidance and cross-attention auxiliary clothes-changing pedestrian re-identification method and system

PendingCN122369057APattern recognitionData set
This invention discloses a text-guided and cross-attention-assisted method and system for re-identifying pedestrians changing clothes. The method includes: preparing a dataset for re-identifying pedestrians changing clothes; resampling using a diversity identity sampler, prioritizing samples with high diversity to construct training batches; building a cross-modal feature representation model based on contrastive language-image pre-training, generating dynamic text prompts, and performing visual and text cross-modal alignment training; extracting key biometric features using pedestrian parsing technology, and constructing a residual-cross-attention mechanism to fuse RGB global features with biometric features; performing a second-stage optimization training by combining identity loss, triplet loss, and adaptive salient feature loss; inputting a query image into the trained model to obtain a list of pedestrian targets ranked by similarity. This invention significantly improves the recognition accuracy and model generalization performance in solving problems such as uneven data distribution and lack of semantic information in clothing-changing scenarios.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Character dynamic effect data generation method and device, storage medium and electronic equipment

The invention relates to a text dynamic effect data generation method and device, a storage medium and electronic equipment. The method comprises the following steps: constructing a copywriting library and a background library; according to the character dynamic effect template, randomly selecting target character content from a copywriting library, and generating a corresponding black-matrix character dynamic effect video; fusing the black-matrix character dynamic effect video with a target background image randomly selected from a background library by adopting a grid division strategy to generate a target dynamic effect video; and extracting a first frame of mask image of the target dynamic effect video to obtain a character mask marking image of the target dynamic effect video, and generating a training data set of the target character content in combination with the copywriting index information of the target dynamic effect video. The technical problems that a large amount of diversified dynamic character and background pairing data in a real scene is lacked, and high-quality character dynamic effect training samples are difficult to automatically generate are solved.
Owner:BEIJING QIYI CENTURY SCI & TECH CO LTD

Industrial marine meteorological disaster forecasting method based on deep learning

The invention relates to the technical field of deep learning and ocean engineering meteorology, and particularly provides an industrial ocean meteorological disaster forecasting method based on deep learning. The method comprises the following steps: acquiring real-time meteorological data, and constructing a dynamic text prompt set; performing semantic-physical feature joint modeling on the dynamic text prompt set by using a multi-modal fusion expert module to obtain a real disaster sample; constructing a dream learning framework, entering a dream stage, and obtaining a virtual disaster dream sample; in an awakening stage, high-dimensional potential representation is obtained through a real disaster sample and a virtual disaster dream sample; according to the high-dimensional potential representation, a forecast result is output, a full-closed-loop industrial meteorological disaster prediction system is formed, and the method improves the accuracy, robustness and interpretability of meteorological disaster prediction in a marine industrial scene.
Owner:SHANDONG UNIV

User interfaces for indicating time

User interfaces for indicating time, displaying a user interface based on a day of the week, user interfaces that include a dynamic text string, user interfaces that include a customizable border complication, user interfaces that include a watch hand that changes color at predetermined times, and user interfaces that include a simulated lighting visual effect.
Owner:APPLE INC

User interfaces for indicating time

User interfaces for indicating time, displaying a user interface based on a day of the week, user interfaces that include a dynamic text string, user interfaces that include a customizable border complication, user interfaces that include a watch hand that changes color at predetermined times, and user interfaces that include a simulated lighting visual effect.
Owner:APPLE INC

Asymmetric cross-modal momentum contrast learning-based galaxy form classification method and system

The invention discloses a galaxy form classification method and system based on asymmetric cross-modal momentum contrast learning, and belongs to the technical field of image classification, and the system comprises a visual encoder which extracts image features and stores the image features in a momentum feature queue; the text encoder is used for encoding the text information to generate initial text embedding; the dynamic text embedding projection layer is used for embedding the initial text into the projection to obtain a text projection; the classification head is used for predicting a galaxy form classification result; the processing module is used for calculating total loss including visual contrast loss and cross-modal alignment loss, and updating model parameters of the visual encoder, the dynamic text embedding projection layer and the classification head by using the total loss; for the total loss, the weighted weight of the visual contrast loss is greater than the cross-modal alignment loss. According to the method, asymmetric interaction of images and texts can be realized, text information is effectively fused in a visual dominant galaxy classification task, the classification precision is improved, and resource consumption is optimized.
Owner:JIANGSU HONGXIN SYST INTEGRATION

Open vocabulary object detection method, apparatus, device, and storage medium

PendingCN122313006AVisual perceptionVocabulary Object
This invention provides an open-vocabulary object detection method, apparatus, device, and storage medium, relating to the field of object detection technology. The method includes: extracting a visual feature map of the image to be detected based on an image encoder in an open-vocabulary object detection model; obtaining an alignment matrix of the visual feature map based on the positive dynamic text embedding, negative dynamic text embedding, and visual features of each target region in the visual feature map for each text category; wherein the positive and negative dynamic text embeddings for each text category are trained based on the visual features and easily confused visual features of that text category; and inputting the alignment matrix between each target region and each text category into a detection head to obtain the detection result. This invention improves the ability to distinguish easily confused categories during the detection process by establishing discriminative constraints for easily confused categories, thereby effectively improving the accuracy of open-vocabulary object detection results.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

A tender document analysis method, device, equipment and storage medium

This invention relates to the field of document parsing, providing a method, apparatus, device, and storage medium for parsing tender documents. It utilizes a large language model to dynamically segment and extract semantic features from tender documents in a knowledge base, obtaining segmented text vectors. Based on the semantic features of the segmented text blocks, a multi-dimensional index structure is established according to multiple segmented text vectors. Pre-defined parsing items are converted into query commands through a question-and-answer interface of the knowledge base. Based on the query commands, the target information is obtained using the large language model and the multi-dimensional index structure. Compared to existing technologies where manual or simple tools struggle to effectively integrate the relationships within tender documents, this invention uses a large language model to extract semantic features from the tender documents and leverages the semantic association capabilities of the large language model to intelligently match and analyze the parsed tender requirements with the knowledge base, providing strong support for accurate recommendations and decision-making assistance.
Owner:SINOTRANS +1

Code editor dynamic text rendering acceleration method based on GPU context awareness

The invention discloses a GPU context awareness-based code editor dynamic text rendering acceleration method, which comprises the following steps of: establishing a multi-dimensional semantic tag system comprising variables, functions, errors and structures, extracting semantic sub-dimensions of characters, generating character semantic tags through operation, caching the character semantic tags to a semantic tag SSBO, constructing a mapping rule of the semantic tags and rendering styles, and carrying out text rendering on the rendering styles according to the mapping rule. The CPU encodes the semantic rules into semantic rule textures; identifying structural units and layout attributes thereof, encoding the structural units into structural textures by the CPU, and uploading the structural textures to a structural layout SSBO of the GPU; the calculation shader reads two types of SSBO, respectively allocates four types of weights of semantics, interaction, structure and effect for characters, and generates a pattern weight texture after determining a fusion coefficient; the vertex shader calculates the final screen position of the character in parallel according to the structure unit to which the character belongs; and the fragment shader reads the two types of SSBO, each texture and the final screen position output by the vertex shader, hierarchical fusion is carried out according to a weight sequence, and character pixel-level rendering is completed.
Owner:北京麟卓信息科技有限公司

Dynamic prompt decoupling Transform-based skeleton human body action fine-grained recognition method

The invention relates to a skeleton human body action fine-grained recognition method based on dynamic prompt decoupling Transform. Comprising the steps of (1) data acquisition and preprocessing, (2) decoupling type visual Transform modeling, (3) dynamic text prompt generation, (4) visual-text semantic alignment, performing text-guided visual feature enhancement through a semantic adjustment module so as to bridge a modal gap between visual and text, and (5) multi-modal cooperative training. The visual enhancement features and the dynamic text features are aligned through comparative learning, semantic stability of the dynamic text features is kept through consistency learning, and action classification is completed through cross entropy loss. Through interaction of a cross-modal attention mechanism, text-guided visual feature enhancement is realized, so that a modal gap is effectively relieved, and priori knowledge of a text is converted into discrimination improvement of visual features.
Owner:NANJING UNIV OF POSTS & TELECOMM

Text generation method and system based on deep synthesis

The invention discloses a text generation method and system based on deep synthesis, and relates to the technical field of deep synthesis, and the method comprises the steps: determining multiple groups of to-be-trained data according to the priorities of multiple primary text data, corresponding text contents and a corpus database; and the text optimization event is determined based on the multiple groups of to-be-trained data, the corresponding identifiers and the dynamic text model, so that the accuracy of the text optimization event is improved. Determining a plurality of model optimization items based on the identification of the model optimization event, and triggering a deep synthesis of each model optimization item to determine a plurality of sets of created content, and marking a content quality level of each set of created content; according to the method, the corresponding content correlation coefficient is determined according to the matching of the multiple groups of created contents, the multiple text features are determined according to the content quality grade of each group of created contents, the corresponding content correlation coefficient and the application scene of the to-be-detected text data, and the corresponding final text is actively generated, so that the accuracy of the final text is improved.
Owner:BEIJING INST OF TECH ZHUHAI CAMPUS

Dynamic text watermark addition method, dynamic text watermark tracing method, dynamic text watermark processing system and apparatus, and medium

The present application relates to a dynamic text watermark addition method, a dynamic text watermark tracing method, a system, an apparatus and a medium. The addition method comprises: acquiring a query request sent by a terminal and terminal information of the terminal, parsing the query request to obtain storage information and sensitive indicator information of text data to be subjected to watermark injection, and on the basis of the storage information, matching said text data in a database; on the basis of the sensitive indicator information, determining a watermark injection ratio, and on the basis of the terminal information, obtaining watermark codes to be injected, wherein said watermark codes comprise authentic watermark codes and dummy watermark codes; and on the basis of the watermark injection ratio, obtaining a conversion function, and sending the conversion function to the database, such that the database executes the conversion function to add said watermark codes into said text data, so as to obtain text data subjected to watermark addition, and sends to the terminal the text data subjected to watermark addition.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Unmanned aerial vehicle pesticide package recovery detection and identification method based on deep learning

The invention relates to the technical field of deep learning, in particular to an unmanned aerial vehicle pesticide package recovery detection and recognition method based on deep learning, which comprises the following steps: inputting a farmland surface image to be recognized into a trained pesticide package recognition model: generating a plurality of candidate regions through a candidate region network; inputting each candidate area into RoIAlign to carry out visual feature extraction to obtain a visual feature vector; for each candidate region, generating a dynamic text prompt based on the visual feature vector of the candidate region, and performing text feature extraction on the dynamic text prompt to generate a text feature vector; fusing the visual feature vector and the text feature vector of each candidate region; generating a pesticide package identification result based on the fusion feature vector of each candidate region; and finally, generating a path planning scheme of the unmanned aerial vehicle based on a pesticide package identification result and controlling the unmanned aerial vehicle to realize pesticide package recovery. According to the invention, the accuracy and robustness of pesticide package identification during unmanned aerial vehicle pesticide package recovery can be improved.
Owner:CHONGQING UNIV OF EDUCATION

Refining training sets and parsers for large and dynamic text environments

Briefly stated, the invention is directed to retrieving a semantically matched knowledge structure. A question and answer pair is received, wherein the answer is received from a query of a search engine. A question is constraint-matched with the answer based on maximizing a plurality of constraints, wherein at least one of the plurality of the constraints is a similarity score between question and answer, wherein the constraint matching generates a matched sequence. For one or more answer sequences, a subsequence is found that are not parsed as answer slots. Query results are obtained from another search engine based on a combination of the answer or question, and the non-answer subsequence. And a KB based is refined on the query results and the constraint matching and based on a neural network training, for a further subsequent semantic matching, wherein the KB includes a dense semantic vector indication of concepts.
Owner:ONTOCORD LLC

System and method for extracting and matching dynamic meaning features

The invention provides a dynamic text feature extraction and matching system and method, the system comprises an input module, a text analysis module, a text matching module and an output module, the input module receives text data and then preprocesses the text data, and the text analysis module analyzes semantics, extracts text features and converts the text features into vectors through a natural language processing technology; the text and meaning matching module uses vector similarity calculation to find out a result with the highest Chinese meaning matching degree in the matching model, and the output module carries out structured output on the result. The method can ensure that the text-meaning matching result is suitable for content comparison of papers or long articles, identifies contents with the same essence but different expressions in papers of different languages, further links the same key points, and can be used as basic theory text contents during investigation report or new theory writing.
Owner:YUAN ENERGY PTE LTD

Method for accelerating dynamic text rendering of code editor based on GPU context awareness

This invention discloses a GPU-based context-aware code editor dynamic text rendering acceleration method. It establishes a multi-dimensional semantic tagging system including variables, functions, errors, and structures. Semantic sub-dimensions of characters are extracted and processed to generate character semantic tags, which are then cached in semantic tag SSBOs. A mapping rule between semantic tags and rendering styles is constructed, and the CPU encodes the semantic rules into semantic rule textures. Structural units and their layout attributes are identified, and the CPU encodes them into structural textures and uploads them to the GPU's structural layout SSBOs. The compute shader reads two types of SSBOs and assigns four weights—semantic, interactive, structural, and effect—to the characters. After determining the fusion coefficients, it generates style-weighted textures. The vertex shader calculates the final screen position of the character in parallel based on the structural unit to which the character belongs. The fragment shader reads the two types of SSBOs, each texture, and the final screen position output by the vertex shader, and fuses them layer by layer according to the weight order to complete pixel-level character rendering.
Owner:北京麟卓信息科技有限公司

A pedestrian re-identification method based on structured attribute perception prompt learning, medium and equipment

This invention discloses a pedestrian re-identification method, medium, and device based on structured attribute-aware prompt learning, comprising: acquiring and preprocessing pedestrian image data, constructing a software-hardware collaborative dynamic text prompt template; inputting the pedestrian image and the corresponding software-hardware collaborative text prompt into a pre-trained image and text encoder to extract visual and text features, freezing the pre-trained network parameters, optimizing the learning of the software-hardware collaborative text prompt for each pedestrian, and inputting the corresponding frozen text encoder and trainable image encoder to obtain text and image embedding vectors; using the text embedding vector as a query and the image embedding vector as a key and value, performing depth alignment of visual and text features through N cross-modal networks to obtain the final fusion features; using the trained image encoder to extract pedestrian image features from the query image and image library images, and completing pedestrian re-identification retrieval by calculating feature similarity; this invention has high accuracy and strong generalization ability.
Owner:BEIHANG UNIV