Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1921 results about "Target text" patented technology

A target text (TT) is a translated text written in the intended target language, which is the result of a translation from a given source text. According to Jeremy Munday's definition of translation, "the process of translation between two different written languages involves the changing of an original written text (the source text or ST) in the original verbal language (the source language or SL) into a written text (the target text or TT) in a different verbal language (the target language or TL)". The terms 'source text' and 'target text' are preferred over 'original' and 'translation' because they do not have the same positive vs. negative value judgment.

Cross-modal image-text analysis method for machine vision

The invention relates to the technical field of machine vision, and discloses a machine vision-oriented cross-modal image-text analysis method, which comprises the following steps of: partitioning an input image to generate an image block sequence; inputting the image block sequence into a visual converter for multi-scale feature extraction, and generating target visual features; encoding the input text to generate a target text feature; inputting the target visual features and the target text features into a deep reconstruction bottleneck network for compression alignment, and generating a cross-modal compression vector; and inputting the cross-modal compression vector into a large language model to generate cross-modal decoding information, so that cross-modal redundant information can be effectively filtered, compact shared semantic representation can be learned, the information integrity of the compression process is ensured through bidirectional reconstruction verification, cross-modal semantic alignment is realized, and the method has the advantages of high efficiency and high reliability. Omnibearing cross-modal content generation from the whole to details is achieved, and the requirements of different application scenes are met.
Owner:SHENZHEN YOULIANCHUANG WISDOM TECH CO LTD

Question answering method based on large model, and electronic device

The present application relates to the technical field of artificial intelligence, and in particular to a question answering method based on a large model, and an electronic device. The electronic device comprises a communication interface, a memory, and a processor, the memory is configured to store a computer instruction, and the processor is configured to execute a computer program to enable the electronic device to: determine, on the basis of each pre-stored text block, a target text block matched with a question to be answered; obtain, if non-text data including a picture and / or a table is pre-stored for the target text block in advance, stored summary text of the non-text data; and input the question to be answered, the target text block, and the summary text into a first large model on the basis of a preset format to obtain answer information. Since the summary text is text that summarizes all content in the non-text data, the first large model does not need to process the picture and / or the table and still can ensure that the content recorded in the picture and / or the table is taken into consideration during generation of the answer information, thereby improving the model-based question answering accuracy.
Owner:HISENSE GRP HLDG CO LTD

Depth map generation method and device based on large model, three-dimensional reconstruction method and device, electronic equipment and storage medium

The invention provides a depth map generation method and device based on a large model, a three-dimensional reconstruction method and device, electronic equipment and a storage medium, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like, can be applied to real-time road scene depth perception, environment three-dimensional reconstruction and obstacle avoidance, and can be applied to real-time road scene depth perception. And virtual and real scene fusion and other scenes can be realized. The specific implementation scheme is as follows: performing visual coding on a monocular image to obtain a coded image; inputting the coded image and the target text into a pre-trained large language model for fusion to obtain fusion features; generating global guide features based on the fusion features, wherein the global guide features comprise joint semantic information of visual features and text features; adding noise to the color image of the monocular image to obtain a noise feature sequence; de-noising the noise feature sequence under the condition of the global guide feature, and generating an implicit feature matched with the joint semantic information; a depth map is generated based on the implicit features.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Three-dimensional digital human generation method and system capable of voice interaction

The invention belongs to the technical field of three-dimensional reconstruction, and discloses a three-dimensional digital human generation method and system capable of voice interaction. According to the invention, brand new speaking audios in different languages are automatically generated according to different languages of the input target text and the sampled human voice audios; the sequential stability and detail reduction capability of three-dimensional human motion are guaranteed by using multi-model joint estimation and a sequential loss function, and facial expression details and hand postures in the image can be accurately estimated. After the high-precision three-dimensional human body model is obtained through estimation, human body action and expression generation is carried out based on voice driving, accurate synchronization of actions and expressions generated through voice is achieved, and facial expression movement and body posture movement, namely a whole-body three-dimensional human body model, conforming to brand-new speaking audio are accurately generated; and finally, rendering the whole-body three-dimensional human body model into a real digital human capable of voice interaction by using a three-dimensional neural rendering model. According to the invention, the realization of single person picture input, high-precision three-dimensional digital person generation and voice interaction is facilitated.
Owner:NANJING UNIV OF SCI & TECH

Causal interpretability and illusion suppression method, system and device for text generation

The invention relates to the technical field of financial management, and discloses a causal interpretability and illusion suppression method, system and device for text generation, and the key point of the technical scheme is that the method comprises the following steps: S1, obtaining related data according to a generation target, extracting the causal relationship between entities, and constructing a causal map; s2, extracting a causal chain related to the generated target from the causal atlas, and inputting the causal chain and the input variables into the large language model to obtain a target text; s3, performing anti-fact intervention processing on the input variable, inputting the input variable into the large language model, and recording a logic test result; s4, identifying an entity from the target text, comparing the entity with the fact data related to the generated target, and calculating an entity alignment score; and S5, according to the causal chain, the target text, the logic test result and the entity alignment score, generating an interpretation report and performing structured output, so that the output text has higher logicality and higher credibility.
Owner:JIANGSU SUNING BANK CO LTD +2

Intelligent system conflict point review system based on knowledge graph and large language model

The invention relates to the technical field of electrical digital data processing, and discloses a system conflict point intelligent review system based on a knowledge graph and a large language model, which comprises the following steps: constructing a dual-mode storage space containing an unstructured index and a structured logic graph, analyzing target text extraction features and triggering graph-based generation logic; converting the topological structure of the associated sub-atlas into a natural language instruction sequence to construct a forced logic constraint template, filling the template with a text, and inputting a pre-training language model to generate a verification result; according to the method, the discrete atlas topology is mapped into the linear logic constraint, random divergence of the generative model is restrained on the calculation principle, and precise decoupling and dynamic evolution of unstructured semantics and structured logic are achieved.
Owner:CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY

Spoken language pronunciation training correction system based on intelligent equipment

The invention belongs to the technical field of intelligent voice processing, and particularly relates to a spoken language pronunciation training and correcting system based on intelligent equipment, which acquires a user rhythm feature set including pitch change rate, accent intensity, pause duration and intonation contour by acquiring spoken language audio data sent by a user for a target text, and corrects the spoken language pronunciation training and correcting system. The method comprises the following steps: acquiring spoken language audio data, converting the spoken language audio data into a phoneme sequence aligned with target text time, generating a rhythm deviation degree report according to comparison evaluation of a user rhythm feature set and the phoneme sequence, an execution intonation mode, accent distribution, speech stream sound change and speech speed rhythm, and determining the rhythm deviation degree according to a rhythm problem type in the report in combination with user historical learning data. And forming and outputting a correction scheme including a text prompt, a targeted minimum contrast training unit and a listen-and-read simulation task, solving the problem of weak capability of correcting hyper-phoneme in a second language spoken language of an adult, and improving the authentic and fluency of the spoken language, thereby realizing efficient personalized learning.
Owner:HUNAN DIGITAL TECHNOLOGY CO LTD

Rhythm migration method and device, electronic equipment and storage medium

The invention relates to the technical field of voice processing, and provides a rhythm migration method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a decoupled rhythm feature based on a source rhythm voice, and a decoupled tone feature based on the voice of a target speaker, the decoupled rhythm feature represents the rhythm of the source rhythm voice, and the decoupled tone feature represents the tone of the target speaker; the decoupled timbre features represent the timbre of the voice of the target speaker; generating a target voice vector sequence based on the text features of the target text, the voice features of the voice of the target speaker, the decoupled rhythm features and the decoupled timbre features; and synthesizing a target audio based on the target voice vector sequence. According to the method and the device, the decoupled rhythm features and the decoupled timbre features are acquired, and the target voice is generated based on the features, so that the problem of feature mixing is effectively relieved, the timbre purity of the target speaker in cross-person rhythm migration is ensured, the expressive force of rhythm migration is improved, and the synthesized audio is more natural and vivid.
Owner:IFLYTEK CO LTD

Text segmentation method and related equipment

PendingCN121328561AMathematical modelsSemantic analysisSemantic changeSemantic variation
The invention provides a text segmentation method and related equipment. The method comprises the steps of obtaining a to-be-processed text; segmenting the to-be-processed text into ordered statement sequences to obtain an initial statement set of the to-be-processed text; wherein the ordered statement sequence comprises a plurality of statements; the semantic variation, the confusion degree variation and the information entropy variation of a first target text block are calculated when a to-be-decided statement in an initial statement set of the to-be-processed text is added into the first target text block, and the first target text block is a set of multiple statements meeting a merging condition; determining the collaboration degree of the statement to be decided and the first target text block based on the semantic variable quantity, the confusion variable quantity and the information entropy variable quantity; and partitioning the to-be-processed text based on the collaboration degree of the to-be-decided statement and the first target text block to obtain a target partitioned text of the to-be-processed text. The text segmentation quality can be improved.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Fault-tolerant processing method and system for super-long text in large model service, and storage medium

The invention relates to the technical field of large language model application, in particular to a fault-tolerant processing method and system for a super-long text in large model service and a storage medium, and the method comprises the following steps: obtaining a super-long text to be processed and a maximum Token threshold of a large model context window, completing Token processing and judging whether the super-long text exceeds the limit or not; starting a multi-level alternative scheme including rapid compression, dynamic compression and chat history compression; starting an error recovery mechanism including hierarchical exception processing, intelligent degradation and parameter verification; outputting the target text of which the Token number is compliant; the system comprises a text acquisition module, a multi-level alternative scheme execution module, an error recovery module and an output module. The storage medium stores a computer program, realizes the method during execution, can adapt to multiple models, and meets industrial-grade super-long text processing requirements; the problems of Token overrun errors and inference service instability caused by lack of fault-tolerant mechanisms and insufficient boundary processing in the prior art are solved.
Owner:POWERCHINA BEIJING ENG CORP

Cross-modal image-text retrieval method and device based on pulse fusion

The invention discloses a cross-modal image-text retrieval method based on pulse fusion, and belongs to the technical field of image-text retrieval. The method comprises the following steps: inputting a target image set into a target detection network to obtain an image floating point code, and inputting a target text set into a word segmentation device to obtain a word floating point code; inputting the image floating point code and the word floating point code into a pulse encoder to obtain a first image pulse code and a first text pulse code; inputting the first image pulse code and the first text pulse code into a pulse cross attention fusion module to obtain a second image pulse code and a second text pulse code; respectively carrying out weighted accumulation and average pooling on the second image pulse code and the second text pulse code to obtain an image floating point feature vector set and a text floating point feature vector set, and calculating cosine similarity to obtain an image-text alignment result; and the text retrieval result of each image in the target image set is obtained based on the image-text alignment result, so that the accuracy and efficiency of image-text retrieval are improved.
Owner:WUHAN UNIV OF TECH

Sign language translation model, system and method based on full-modal alignment

The invention discloses a sign language translation model, a sign language translation system and a sign language translation method based on full-modal alignment. The sign language translation method comprises the steps of extracting multi-modal features of hand, face and body postures from an input video and performing preliminary fusion; deep secondary fusion and alignment are carried out through multi-scale time sequence coding and a cross-modal collaborative attention mechanism, and a spatial-temporal feature sequence of full-modal alignment is generated; performing boundary detection and dynamic segmentation on the feature sequence by using a CTC-based sequence prediction model, and outputting a discrete sign language word sequence with a timestamp; and finally, capturing a sign language grammar structure of the sequence through a graph structure enhanced Transform encoder, inputting the sign language grammar structure into a Transform decoder integrating grammar consistency loss, and generating a target text conforming to target natural language grammar and semantic rules. According to the method, the problems of continuous sign language action adhesion and grammar structure difference are effectively solved, and the sign language translation accuracy and the natural language generation fluency are greatly improved.
Owner:SURELY ACCESSIBLE TECH (SUZHOU) CO LTD

Multi-modal data pairing method and system based on deep learning

The invention provides a multi-modal data pairing method and system based on deep learning, and relates to the technical field of data processing, and the method comprises the steps: obtaining a video multi-frame sequence and a target text, and respectively extracting an overlapped frame group set and a standardized text sequence; performing spatio-temporal feature extraction and text dependency relationship coding to obtain a video time sequence vector sequence and a text vector sequence; executing cross-modal alignment search, and constructing a monotonic matching path set; calculating a semantic and action entity relationship consistency score of the paired elements on the path to obtain a comprehensive score; and determining an alignment relationship between the video and the text based on the optimal path. According to the method, accurate matching of the video and the text is realized, and the cross-modal retrieval efficiency is improved.
Owner:BEIJING YIZHUANG INTELLIGENT CITY RES INST GRP CO LTD

Toxic text collection method and system based on retrieval enhancement generation

The invention relates to a toxic text collection method and system based on retrieval enhancement generation, and the method comprises the steps: firstly obtaining target text data through a crawling platform, manually constructing a small-scale initial data set through a cold start mode, and injecting the initial data set into a knowledge base as basic data; and performing semantic retrieval on each batch of target texts, obtaining the first k texts with semantic similarity from the knowledge base, reasoning by adopting a plurality of large language models through thinking chain reasoning in combination with the texts, generating a toxic label, and labeling the target texts to be labeled in the current batch. And injecting the labeled target text into a knowledge base, continuously collecting and labeling data in an iteration mode, and continuously optimizing the performance of a retriever through an incremental learning mechanism in the iteration process so as to realize toxic text collection. According to the method, large-scale, fine-grained and high-consistency toxic text tagging corpora can be efficiently accumulated, and the problems of high tagging cost, inconsistent quality and poor expansibility in traditional toxic text data collection are remarkably relieved.
Owner:JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS

Time sequence card generation method and device based on large language model

The invention belongs to the technical field of large language models, and discloses a time sequence card generation method and device based on a large language model, and the method comprises the steps: receiving a card creation instruction, and obtaining a card data type in the card creation instruction; extracting target text data in a dialog box of the large language model based on the card data type; converting the target text data into target card data corresponding to the card data type; generating a card title corresponding to the target card data according to the target text data; combining the target card data and the corresponding card titles to obtain time sequence cards; and putting the time sequence card into a tool window adjacent to the dialog box for displaying. According to the method, the manual operation of the user can be reduced, and the co-creation work efficiency and convenience of the large language model and the user are improved.
Owner:GUANGZHOU ZHIYONGKAIWU ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Wound knowledge graph construction method and system based on three-level fault-tolerant mechanism

The invention provides a trauma knowledge graph construction method and system based on a three-level fault-tolerant mechanism. The method comprises the following steps: acquiring a trauma condition core knowledge framework; extracting trauma condition rules based on the trauma condition core knowledge framework, and establishing a trauma condition rule base; dynamically generating a target cue word based on the trauma condition rule base by utilizing a large language model; processing the target text based on the target cue word by using a large language model to generate a candidate file; verifying the candidate file to obtain a verification file; and constructing the trauma knowledge graph based on the verification file. According to the scheme, a mode of combining ontology construction from bottom to top and entity node construction from top to bottom is adopted, various methods such as manual construction, rule extraction and a large language model are fused, the knowledge extraction precision and the map coverage degree are improved, the information processing capacity and the treatment decision-making efficiency in a complex trauma condition scene are remarkably improved, and the method is suitable for popularization and application. And construction of an intelligent medical system is supported.
Owner:THE NAVAL MEDICAL UNIV OF PLA

Text generation method and device based on large model, equipment and storage medium

The invention discloses a text generation method and device based on a large model, equipment and a storage medium, and relates to the technical field of natural language processing, and the method comprises the steps: carrying out the semantic analysis of a preset field related document based on an encoder of an initial large model, and generating a target concept set and a target embedding vector; constructing an initial weighted directed graph according to the target concept set and the target embedded vector, and expanding the initial weighted directed graph to obtain a target weighted directed graph; determining a target sub-graph meeting a preset optimal condition from the target weighted directed graph by utilizing a preset optimization algorithm, and obtaining a target structured text part set based on the target sub-graph; and performing fine tuning on the initial large model to obtain a fine-tuned large model, and converting the target structured text part set into a target text meeting a preset field text specification by using the fine-tuned large model. Content innovation can be realized while strict logic is ensured and requirements of professional fields are met.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

A rule corpus-based text specification marking method and system

The application relates to the technical field of text label marking, and provides a text specification marking method and system based on a rule corpus, which comprises the following steps: analyzing a policy and regulation document, identifying and marking condition morphemes and conclusion morphemes in the policy and regulation document, constructing a logical relationship between the two by using a large language model, and forming a rule corpus composed of structured morpheme pairs; performing semantic embedding on the corpus to generate a semantic vector library; performing multi-label coding on a verification data set based on the rule corpus, and constructing a multi-label training data set; training a deep learning classification model by taking semantic vectors as features and multi-labels as targets, so that a text specification marking model is obtained; and automatically marking target text by using the model. The application significantly improves the accuracy, interpretability and business adaptability of text marking, improves the update quality of a system label data set, and reduces the system maintenance cost.
Owner:SSE INFORMATION NETWORK LTD

Digital transformation intelligent question-answering method, device and equipment for small and medium-sized enterprises and storage medium

The invention provides a small and medium-sized enterprise digital transformation intelligent question answering method and device, equipment and a storage medium, and the method comprises the steps: carrying out the semantic recognition and entity extraction of a question input by a user, and generating a question semantic vector; retrieving a corresponding target text fragment in a preset vector database according to the question semantic vector; searching a sub-graph structure associated with the question semantic vector in a preset knowledge graph according to the target entity of the question semantic vector; reordering the text segments and the sub-graph structures based on correlation scores to obtain an ordering result set; and inputting the sorting result set into a retrieval enhancement generation model, and generating question and answer output content. Through the implementation of the scheme of the application, association matching of question semantics and enterprise knowledge can be realized by utilizing the text semantic information and the knowledge graph structure information at the same time in the question and answer generation process, the consistency of question and answer contents in logic structure and semantics is ensured, and the accuracy of digital transformation question and answer of small and medium-sized enterprises is improved.
Owner:YUNDI SMART TECH CO LTD

Text generation method and electronic equipment

The invention discloses a text generation method and electronic equipment, and relates to the technical field of deep learning, and the method comprises the steps: obtaining a plurality of candidate output texts corresponding to a text input by a user through a pre-training language model, splitting the plurality of candidate output texts into a plurality of statements, according to the method, each statement is extracted, each statement is scored to obtain the statement score of each statement, and then the target text is screened from the multiple statements according to the statement score of each statement, so that representative statements with focused semantics and rich information can be selected from the multiple statements, and trunk content extraction based on the statement scores is realized; and then the sequence-to-sequence generation model is utilized to perform optimization processing on the screened target text to obtain an optimized text, so that unified expression of the trunk content on language styles and sentence pattern structures is realized, and the accuracy of the finally generated text is improved. Therefore, the technical problem that it is difficult to ensure the accuracy of the output text by directly selecting the optimal candidate answers to determine the final output can be solved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

CNN-based call quality inspection method, apparatus and device, and storage medium

The invention belongs to the technical field of artificial intelligence, and discloses a CNN-based call quality inspection method, device and equipment and a storage medium, the CNN-based call quality inspection method comprises the following steps: inputting a word vector sequence obtained by converting a target text corresponding to a target seat call into a preset sentiment analysis model to obtain a sentiment tag corresponding to the target text; analyzing the importance degree of each word vector in the word vector sequence to obtain a keyword vector in the word vector sequence; performing named entity recognition on each word vector in the word vector sequence to obtain an optimal entity tag sequence corresponding to the word vector sequence; and splicing the audio feature, the emotion tag, the keyword vector and the optimal entity tag sequence, inputting the obtained multi-modal fusion feature into a preset evaluation model, and outputting to obtain a quality inspection result. The method and the system can be applied to business management systems of financial science and technology, medical health and the like, and solve the technical problem that efficiency and quality cannot be considered in a call quality inspection mode based on the prior art.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Voice generation method and device based on pseudo-autoregression modeling, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice generation method, device and equipment based on pseudo-autoregression modeling and a medium, and the method comprises the steps: obtaining a training sample containing a text sequence, a prompt voice segment and a target semantic token sequence; performing continuous fragment mask training on the text-to-semantic model to obtain a pseudo-autoregression trained text-to-semantic model; generating candidate speech output by using the text-to-semantic model and the initial semantic-to-acoustic model which are subjected to pseudo-autoregression training, and constructing a preference data pair; updating the semantics-to-acoustics model based on the preference data pair to obtain a preference optimized semantics-to-acoustics model; and generating target voice output based on the target text and the target prompt voice. According to the method, the time sequence modeling capability of the model is enhanced through pseudo-autoregression training, and the voice generation quality is directly optimized through the preference data pair, so that the voice alignment precision and the subjective listening feeling performance are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

A method, device, and medium for processing NOTAM text based on semantic enhancement

This invention relates to the field of text processing technology, and in particular to a method, device, and medium for processing navigational notice text based on semantic enhancement. The method includes: first, acquiring a content carrier to be processed; then, acquiring the semantic vector and glyph feature vector of the content carrier; concatenating the two types of vectors to form an enhanced text representation; extracting temporal features from the enhanced text representation to obtain temporal features containing forward and backward logical relationships within the text; acquiring the weights of words and sentences in the temporal features and performing weighting to obtain weighted word representations and weighted sentence representations; performing correction processing on the weighted representations to generate corrected text; and finally, validating the corrected text and outputting the target text. This invention can improve the accuracy and efficiency of content carrier processing.
Owner:CIVIL AVIATION UNIV OF CHINA

Token compression method and device for large model dialogue

The invention provides a Token compression method and device for large model dialogues, which can effectively reduce the number of Tokens (lexical elements) of texts input into a large model in multiple rounds of dialogues, and can also retain key information of early parts of historical dialogue texts at the same time. The method is applied to a Token compression management agent, and comprises the following steps: after obtaining a question text input by a user to a terminal, determining the number of lexical elements in an initial text (including a historical dialogue text and the question text) to be sent to a large model; if the initial text lexical element number does not meet the threshold requirement, switching the state into a lexical element compression state, interacting with the large model to obtain the lexical element number of compressed intermediate dialogue texts (parts except for the previous k rounds of dialogue texts, k is greater than or equal to 0), and keeping a semantically generated abstract text. And splicing the first k rounds of dialogue texts, abstract texts and question texts to obtain a target text. If the number of the target text lexical elements meets the threshold requirement, sending to the first large model.
Owner:ULTRAPOWER SOFTWARE

Audio synthesis method, audio synthesis model training method, apparatus, electronic device, computer-readable storage medium, and computer program product

An audio synthesis method, an audio synthesis model training method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product, which relate to artificial intelligence technology. The method includes: invoking an audio synthesis model based on language information and preset style information of a target text to perform following processing, the audio synthesis model including a prior encoder and a waveform decoder: generating audio features corresponding to the target text based on the language information and the preset style information by using the prior encoder; performing normalizing flow processing on the audio features by using the prior encoder, to obtain a hidden variable of the target text; and performing waveform decoding on the hidden variable of the target text by using the waveform decoder, to obtain a synthetic waveform conforming to an audio style described in the preset style information and corresponding to the target text.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

End-cloud cooperative data mining method, device, system and computer cluster

An end-cloud cooperative data mining method, comprising: determining a target text and a task configuration file according to a business requirement by a cloud end; encoding the target text by a text encoder to obtain a text feature; placing the text feature in the task configuration file and issuing it to a vehicle end together; encoding image data by a first picture encoder to obtain an image feature by the vehicle end; calculating a value of a similarity of the text feature and the image feature; determining a target picture according to the value of the similarity and the task configuration file; and uploading the target picture to the cloud end; wherein the first picture encoder is obtained by compressing and optimizing a second picture encoder; and the text encoder and the second picture encoder are two modules of a picture-text multimodal large model. The application is applied to an automatic driving shadow mode, can complete data mining of any interesting target text by using one large model, and does not need to design and develop detection rules for each type of interesting target. The picture encoder of the vehicle end can be continuously updated and optimized, and the model iteration efficiency is improved.
Owner:HUAWEI TECH CO LTD

UI automatic test method and test system, and storage medium

The invention provides a UI automatic test method and test system and a storage medium. The method comprises the steps that an optical character recognition result of a target text and a to-be-tested UI is obtained, and the optical character recognition result comprises a plurality of elements and element texts, coordinate frames and visual features of all the elements; obtaining the semantic similarity between the target text and each element text; obtaining attribute similarity between the target text and each element based on the visual features of each element; according to the semantic similarity and the attribute similarity, elements, corresponding to the target text, in the optical character recognition result are determined to serve as target elements; taking the position of the coordinate frame of the target element in the UI as a target position; and generating a presentation image which highlights the target position and / or the target element. Therefore, fuzzy matching based on semantics can be realized, multi-modal features including text semantics and visual features are fused, and the robustness and accuracy of UI element positioning are improved in combination with context sensing.
Owner:AIJI MICRO CONSULTING (XIAMEN) CO LTD

Text generation method, server, storage medium, and program product

The present disclosure provides a text generation method, a server, a storage medium, and a program product. According to the method of the present disclosure, on the basis of given theme information, an outline of a target text to be generated is determined, related materials of the outline are acquired, on the basis of the related materials of the outline, fine-tuning training is performed on a pre-trained model to obtain a text generation model, domain knowledge contained in a large amount of related materials of the outline can be stored in the text generation model in the form of parameters, and domain knowledge contained in a large amount of related materials is injected into the text generation model, so that the text generation model learns more domain knowledge; and quality screening is performed on the related materials of the outline to obtain a small amount of high-quality preferred materials related to the outline, the text generation model is used to generate a target text on the basis of the preferred materials, during the generation of the target text, reference is made to the domain knowledge contained in a small amount of high-quality preferred materials, and reference is also made to the domain knowledge contained in a large amount of related materials of the outline, thereby improving text generation quality.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Multimodal interleaved image-text generative model based on dynamic feature synchronizer

The present disclosure relates to the technical field of multimodal learning, and particularly relates to a multimodal interleaved image-text generative model based on a dynamic feature synchronizer. An image encoder extracts multi-resolution multi-scale feature maps from input images of interleaved image-text data; and a dynamic feature synchronizer in a multimodal large language model acquires fine-grained information, so as to determine output feature data corresponding to the interleaved image-text data and then generate a target image and / or target text associated with the interleaved image-text data.
Owner:TSINGHUA UNIVERSITY

Text summary generation model training method, text summary generation method and device

The present application relates to a text summarization model training method, a text summarization method, an apparatus, a computer device, a storage medium, and a computer program product. The method comprises: inputting a training text and a corresponding training image set into an initial text summarization model to obtain a predicted text summary, and generating a target loss based on the difference between the predicted text summary and the labeled text summary; inputting masked training data corresponding to the training text and first training data into the initial text summarization model to obtain masked predicted data, and generating a reconstruction loss based on the difference between the masked labeled data and the masked predicted data; adjusting model parameters of the initial text summarization model based on the target loss and the reconstruction loss until convergence conditions are met, thereby obtaining a target text summarization model; and using the target text summarization model to generate a text summary of the text. This method can improve the prediction accuracy of the model and the quality of the generated text summary.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1