Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

747 results about "Text generation" patented technology

Sign language translation method and system based on pre-training diffusion large language model

The invention provides a sign language translation method and system based on a pre-training diffusion large language model, and belongs to the field of sign language video translation. The method comprises the following steps: preprocessing a video containing sign language actions to obtain a sign language video frame sequence, inputting the sign language video frame sequence into a visual feature extraction network to extract features, and fusing to obtain a time sequence visual fusion feature sequence; giving a text cue word of a sign language translation task, constructing an initial mask sequence for a target translation position, taking the text cue word, the time sequence visual fusion feature sequence and the initial mask sequence as guide conditions, injecting the guide conditions into a diffusion language model, iteratively denoising and predicting lexical elements of a masked position in combination with a diffusion mask mechanism, and obtaining the sign language translation task. A natural language translation sequence is obtained, and sign language translation is completed; wherein when the diffusion language model is trained, through an internal feature alignment mechanism, the guiding effect of guiding conditions on text generation is optimized, so that the accuracy, coherence and robustness of long text translation are improved, and the actual requirements of a barrier-free public service scene are better met.
Owner:ZHEJIANG UNIV

Original script-oriented AI autonomous plot structure adaptive generation system

PendingCN121233763ASemantic analysisBiological modelsNeural oscillationAlgorithm
The invention discloses an original script-oriented AI autonomous plot structure adaptive generation system, and relates to the technical field of creation assistance, and the system specifically comprises the following modules: a structure extraction analysis module, an emotional role analysis module, a plot inference module, a scene generation optimization module, a conservation target generation module, an adaptive control module, and a constraint punishment module. According to the method, a multi-level narrative structure and a causal relationship graph are constructed to form an emotion vector and a trajectory curve, an optimal causal path is generated based on a graph neural network, a graph convolution / attention mechanism and a graph generation algorithm, and emotion toning and conservation target driven text generation are performed on scene and dialogue levels. Structural consistency and emotional arcs are optimized in real time in combination with self-adaptive control, plot path weighting and rewriting triggering are performed by utilizing a multi-dimensional emotional space and a neural oscillator network, intelligent structured management of a script is realized, and script creation efficiency and quality are improved.
Owner:GOLDEN TIMES CULTURE COMM

Semantic perception black box large language model training data auditing method and system

The invention relates to the technical field of training data auditing, in particular to a semantic perception black box large language model training data auditing method and system.The method comprises the following steps that content is returned based on a text generation interface, lexical elements and candidate content are analyzed for multi-round sampling distribution, and a convergence index is determined; and calculating semantic paths and weights of the real content and the candidate content, and analyzing differences between the real content and the candidate content and the reference content to obtain an affiliation adaptation judgment result. According to the method, through multi-round sampling behavior fluctuation tracking, section feature dynamic extraction, stability change judgment and sequence-level weight aggregation, fine separation of training data attribution signals is achieved, a multi-level judgment system is constructed for complex expression and diversified output, judgment accuracy is optimized through a tension comparison signal group and an attribution adaptation judgment mechanism, and the judgment accuracy is improved. And a sensitive and adaptive training data member auditing strategy is formed, so that the risk hidden danger omission is effectively prevented, and the model data security boundary control is enhanced.
Owner:NANKAI UNIV

Large-model-oriented multi-modal retrieval enhancement generation method and large-model-oriented multi-modal retrieval enhancement generation device

The invention provides a multi-modal retrieval enhancement generation method and device oriented to a large model, and the method comprises the steps: coding multi-modal information, uniformly mapping all modal codes to a shared representation space, obtaining each modal vector in the shared representation space, and obtaining a multi-modal retrieval enhancement generation result based on query information and context information inputted by a user; identifying a task intention of the modal vectors sharing the representation space, outputting a modal weight vector, dynamically selecting a fusion mode based on the task intention, fusing the modal vectors based on a gating fusion network and the fusion mode, and dynamically adjusting the proportion of different modal vectors in fusion according to the modal weight vector. And in response to the semantic consistency conflict of each modal vector, selecting a target conflict mediation strategy based on a strategy selection rule, carrying out conflict mediation on each modal vector based on the target conflict mediation strategy, generating a retrieval answer based on each fused modal vector and context information, and sending the retrieval answer to a server. And executing a text generation task based on the retrieval answer by adopting an auxiliary model.
Owner:CHINA LIFE INSURANCE CO LTD SHANGHAI DATA CENT

SMPL-X action-to-text generation method based on global and local feature fusion

ActiveCN121502730ASemantic analysisBiological modelsAlgorithmAction semantics
The invention provides a global and local feature fusion SMPL-X action-to-text generation method, and belongs to the field of artificial intelligence. The method comprises the following steps: preprocessing and coding an input SMPL-X action sequence into a double-flow action feature; through a cross-modal mapping module, the double-flow action features are mapped to a pre-trained large language model through independent projection branches, and global conditions and local action prefix embedding are obtained; through a text generation module, a decoder of a pre-trained large language model is used as a trunk network, text cue word embedding is extracted based on a text instruction given by a user, local action prefix embedding and text cue word embedding are spliced and then input into the decoder, and global conditions are injected into each layer of the decoder through a cross attention mechanism. And generating a description text in an autoregression mode. According to the method, the description text which is consistent with action semantics and has sufficient details can be stably and accurately generated, and the generation stability and the cross-scene applicability are improved when disturbance exists in the action sequence.
Owner:ZHEJIANG UNIV

Watermark generation detection method based on stochastic smoothing testible robust large language model

The invention discloses a watermark generation detection method based on a stochastic smoothing testible robust large language model. The method comprises the following steps: training a watermark generator according to a lexical sequence data set; according to the prompt word text data set, the watermark generator and the large language model, generating texts with embedded watermarks and texts without embedded watermarks, and training a watermark detector; and converting a text into which a preset watermark is to be embedded, inputting the converted text into the trained watermark generator, outputting the text into which the watermark is embedded, processing the text into which the watermark is embedded by adopting a random smoothing-based mode, and inputting the processed text into the trained watermark detector for detection, thereby realizing watermark generation detection. According to the method, the watermark can still be reliably detected after the text is tampered, the anti-attack ability of the watermark is improved by introducing the random smoothing technology, the watermark detector has the cross-model generalization ability through the code conversion and training strategy, and efficient watermark embedding and detection can be achieved while the text generation quality is kept.
Owner:ZHEJIANG UNIV

Large model-based legal text generation proofreading method and system

The invention provides a legal text generation proofreading method and system based on a large model, and relates to the technical field of computer information processing. The method comprises the following steps: firstly, acquiring an input case text, and matching a corresponding target legal instrument template based on a preset legal instrument template library; secondly, calling a pre-trained large language model according to a format requirement, and generating a first-edition legal document; based on a pre-stored legal knowledge graph, performing internal logic and legal basis consistency verification on the first-edition legal document to obtain a consistency verification result; and finally, according to a consistency verification result, locating conflicts and errors existing in the first-edition legal document so as to generate a proofreading report containing targeted modification suggestions. According to the technical scheme provided by the invention, the legal document with a standard format can be automatically generated, logic and legal reference errors in the legal document can be intelligently detected, and specific modification suggestions are provided, so that the document quality and the generation efficiency are improved.
Owner:BEIJING NEW ORANGE TECH CO LTD

Generative interaction method and device for medium-free aerial imaging

The invention discloses a generative interaction method and device for medium-free aerial imaging, and relates to the technical field of man-machine interaction for medium-free aerial imaging. The method comprises the steps of obtaining a reference description text, hand motion data and hand posture data corresponding to a target user, wherein the hand posture data is used for indicating a final posture of hand motion; inputting the hand motion data and the hand posture data into a motion text generation model, outputting a motion text, inputting the reference description text and the motion text into a text fusion model, and outputting a target description text; the target description text is input into a model generation network, a three-dimensional model is output, and the three-dimensional model comprises at least one entity model and at least one environment scene model corresponding to the entity model; and mapping the three-dimensional model to a medium-free aerial imaging device to obtain an aerial suspension image. The problems of low aerial imaging efficiency, lack of dynamic interaction capability and the like of the existing aerial imaging scheme are solved.
Owner:HUIZHI WORLD (HANGZHOU) TECHNOLOGY CO LTD +1

Generating or modifying text using a digital assistant and / or language model

The present disclosure generally relates to generating and / or modifying text using a digital assistant. This disclosure encompasses a system for requesting information before generating text using models based on the type of information needed. This disclosure further encompasses a system for a remote model requesting additional information before generating text. This disclosure encompasses user interfaces for text generation / modification. This disclosure encompasses a system for proofreading text using a model. This disclosure encompasses a system for generating a response to a received communication with operations for updating the generated response.
Owner:APPLE INC

Customer service information generation method and system based on multi-modal intention recognition

The invention discloses a customer service information generation method and system based on multi-modal intention recognition, and the method comprises the steps: obtaining customer service session data, and carrying out the heterogeneous data classification operation of the session data; performing feature extraction operation on each piece of parting modal information; establishing a semantic bridging matrix among the modal features, performing hierarchical attention routing, and outputting an intention label; a knowledge graph is injected, the knowledge graph comprises a static knowledge base, a real-time service flow and a user historical portrait, and three-source dynamic knowledge is output; constructing a decision tree based on the intention label, the three-source dynamic knowledge and the space-time identifier; and executing a corresponding text generation operation, a visual mark generation operation or a voice synthesis operation based on a response action corresponding to the leaf node of the decision tree, and generating customer service information. According to the method and the system, the buyer intention recognition accuracy of the multi-modal session data can be greatly improved, and the response delay time of generating the customer service information is shortened.
Owner:深圳乐搏科技有限公司

Multi-modal large model illusion detection and suppression method based on attention time sequence difference

A multi-modal large model illusion detection and inhibition method based on attention time sequence difference comprises the following steps: inputting text lexical elements of cue words and visual lexical elements of images into a multi-modal large model, and obtaining an internal attention graph of the decoding stage of the multi-modal large model; then calculating the attention proportion of the visual lexical units at the current generation moment, and making a difference between the attention proportion and the proportion at the previous moment; and if the difference value exceeds a set threshold value, determining that the lexical elements are visual related lexical elements. When the visual related lexical elements are recognized, performing secondary forward propagation of primary visual enhancement to obtain more accurate output; and if not, directly entering the next step of generation. According to the method, visual related lexical elements in text generation are recognized and refined through attention time sequence difference, two-time forward propagation is adopted, the visual attention of second-time forward propagation is enhanced based on a visual attention graph of first-time forward propagation, and illusion can be recognized and corrected on the premise that the language expression ability is not reduced.
Owner:HANGZHOU DIANZI UNIV

Image context based text generation

Methods, systems, and storage media for generating contextually relevant text from image descriptions and user intent are disclosed. Exemplary implementations may: receive an image and a user-defined intent for text output; analyze the received image to generate a contextual description of the image; generate a query based on the contextual description of the image and the user-defined intent; and generate the text output based on the query.
Owner:SHUTTERSTOCK

Method for generating image description text based on large model

The invention discloses a method for generating an image description text based on a large model, which relates to the technical field of image processing, and comprises the following steps: an image preprocessing step: dividing a target level through semantic segmentation and extracting key visual information by adopting hybrid denoising and self-adaptive normalization; a feature extraction step: fusing the multi-scale visual features and the semantic features, and generating a high-dimensional fusion feature vector through cross-modal alignment; a large model initialization and adaptation step: loading the pre-training model and performing incremental fine tuning, and dynamically adjusting the Prompt template; a text generation step: generating candidate texts through logic constraint and beam search; and a text optimization adjustment step of outputting a final description text based on multi-dimensional evaluation and user preference iterative correction. According to the method, the semantic matching degree, logic coherence and common sense accuracy of image description are improved, multi-element scenes are adapted through dynamic adaptation and iterative optimization, and high-quality text description support is provided for high-precision and diversified scenes.
Owner:BEIJING LINGMANG TECH CULTURE CO LTD

Fine-grained image clustering method and system based on FG-CLIP and language enhancement

The invention provides a fine-grained image clustering method and system based on FG-CLIP and language enhancement, and belongs to the field of image clustering. A pre-training text generation model is utilized to generate diversified coarse-grained text descriptions, then a text fine-grained module and a text screening module are utilized to perform fine processing on a text, and two enhanced views are fused to form consistent text representation, so that the comprehensiveness of text information is enhanced. Then, a neighbor set is constructed according to the feature similarity through a random neighbor sample selection module, attention to samples in the same cluster is improved, and the selection range of positive samples is expanded; and finally, inputting the image features, the text features and neighbor samples thereof into each modal cluster-level mapping head for dimension mapping, and realizing alignment and joint training of the two modal features through random neighbor contrast loss. Under the assistance of the text information, the fine-grained clustering targets can be accurately distinguished, so that the accuracy and robustness of fine-grained image clustering are effectively improved.
Owner:UNIV OF JINAN

Text generation method and device, model training method and device and computing equipment

The embodiment of the invention provides a text generation method and device, a model training method and device and computing equipment. The text generation method comprises the steps of obtaining to-be-processed data; the to-be-processed data is input into a content generation model, initial probability distribution of all lexical elements of the target text is obtained, and the initial probability distribution is used for indicating the initial probability that the lexical elements are candidate lexical elements; obtaining a target feature vector of the target text according to the initial probability distribution and the embedding vector of each candidate lexical element; the target feature vector is input into a content generation model, updating probability distribution of all lexical elements of the target text is obtained, the target text is generated based on the updating probability distribution, and the updating probability distribution is used for indicating the updating probability that the lexical elements are candidate lexical elements. And a parallel decoding mechanism and a global feature optimization mechanism greatly improve the generation efficiency while ensuring the accuracy and continuity of the generated text.
Owner:SWEET POTATO TECHNOLOGY (SHANGHAI) CO LTD

Generative text model query system

Text generation prompts may be determined based on an input document and a text generation prompt template. The text generation prompts may include text from the input document and questions related to the text. The text generation prompts may be sent to a remote text generation modeling system, which may respond with text generation prompt response messages including novel text portions generated by a text generation model. The text generation prompt response messages may be parsed to generate answers corresponding with the questions.
Owner:CASETEXT INC

Intelligent text generation method and system based on natural language processing

The invention discloses an intelligent text generation method and system based on natural language processing, and relates to the technical field of natural language processing. The method comprises the following steps: firstly, combining a natural language processing technology and a deep learning algorithm, and carrying out multi-task identification analysis on an original text to obtain text original element information including styles, styles, emotions, entities and topics; performing fusion processing on the text generation control information input by the user and the style, the genre and the emotion in the element information to obtain a text generation control vector containing a target style, a target genre, a target emotion and a forbidden word set, and performing retrieval to obtain background knowledge and / or fact data of an original entity and a theme of the text; and finally, integrating the control vector and the retrieval result into a cue word, and importing the cue word into a large language model to output and obtain a new text, so that text contents with high quality, high correlation, high fact consistency and high controllability can be generated by deeply analyzing the original text and combining conditions specified by a user.
Owner:CHONGQING SILICON LIFE TECHNOLOGY CO LTD

Multi-modal visual position identification reordering method and system based on guidance

The invention relates to the technical field of visual position recognition, and particularly discloses a multi-modal visual position recognition reordering method and system based on guidance, and the method comprises the steps: obtaining a query image, and retrieving a plurality of candidate images based on a pre-trained visual basic model and the query image; constructing a composite multi-modal prompt object, wherein the composite multi-modal prompt object comprises an image pair formed by the query image and the current candidate image, and an instruction text used for guiding a multi-modal large language model to perform visual comparison; outputting a structured similarity judgment result, wherein the result comprises a quantitative similarity score; and sorting based on the similarity scores corresponding to all the candidate images, and determining the candidate image with the highest score as an optimal matching result. Through combination of guiding type prompt engineering and structured output, an intermediate text generation link is avoided fundamentally, and the calculation efficiency is improved while the fidelity of all original visual information is reserved.
Owner:SHENZHEN 1024 ROBOT TECHNOLOGY CO LTD

Training multi-stage malleable hybrid networks

Multi-stage hybrid network integrates relationship regularization links and explainable elements to improve alignment with human values, explainability, robustness, and efficiency. The network comprises neural components, event prediction elements, and probability models across multiple stages, with relationship constraints enforcing structured knowledge representation. Explainable elements provide interpretable rationales for decisions, enhancing transparency. Training incorporates supervised learning, human-guided refinement, semi-automated knowledge engineering, and adversarial robustness techniques. A Socratic reasoning module detects contradictions and refines outputs for logical consistency. Indexed model elements enable dynamic memory optimization for improved efficiency. Candidate outputs may be scored, verified, or selected using neural and symbolic criteria. The invention supports retry loops and configurable subsystem pipelines to improve output quality. Applications include text generation, speech recognition, translation, and decision support. By combining structured constraints, human oversight, and modular architectures, the system improves the trustworthiness, safety, and adaptability of AI systems across diverse modalities and tasks.
Owner:D5AI LLC

Text generation methods and apparatuses, storage medium devices, and program products

This specification provides text generation methods, apparatuses, and storage medium devices. One method includes the following operations. In an iteration of a plurality of iterations under a large language model (LLM): estimating a first text sequence following a current text sequence based on a speculative decoding method, forming a plurality of candidate sequences based on the current text sequence and subsequences of the first text sequence, allocating logical blocks to text units in the plurality of candidate sequences in a key-value cache, to store attention information of the text units, mapping the allocated logical blocks to physical blocks based on a first criterion, and determining, by the LLM, a newly generated text unit in the iteration by using attention information of each candidate sequence in the key-value cache, to form a current text sequence for a next iteration of the plurality of iterations.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

AI file query method based on human-computer interaction

The invention discloses an AI file query method based on human-computer interaction, which comprises the following steps: S1, collecting user input, and generating a query text; s2, intention recognition and slot extraction are performed on the query text, a slot set is generated, and standardization processing is performed; s3, utilizing the improved CoSent model to generate a query vector and a document vector, and storing the query vector and the document vector in a vector index; s4, retrieving the candidate document based on the query vector in the vector index, calculating semantic similarity and obtaining a comprehensive score in combination with a matching result; s5, carrying out reordering on the candidate documents; s6, generating a permission set based on the user identity, and executing permission filtering and field masking; and S7, inputting the document fragments passing the permission filtering and the query text into a text generation model, outputting a query result, and recording a log and user feedback. According to the invention, the accuracy and safety of file query are obviously improved.
Owner:ANHUI BOGUANG ARCHIVES TECH CO LTD

Multi-agent cooperation and dynamic feedback long text generation system and method

PendingCN121413623ASemantic analysisBiological modelsEvaluation resultContinual improvement process
The invention provides a multi-agent cooperation and dynamic feedback long text generation system and method, and the system comprises a performance agent which is used for receiving a theme or an outline inputted by a user, planning and scheduling a long text generation process, coordinating the interaction of an author agent and an evaluation agent, and dynamically adjusting a generation strategy based on an evaluation result; the author agent is used for generating text content according to an instruction of the director agent and correcting a generation result; the evaluation agent is used for performing multi-dimensional quality evaluation on the text content generated by the author agent and returning optimization suggestions; and the memory bank is used for storing the historical generation information and the global setting information and providing context support in subsequent text generation. Through the synergistic effect of the director agent, the author agent, the evaluation agent and the memory bank, a dynamic feedback closed-loop mechanism is established, so that the text generation process has continuous improvement capability, and the overall quality and stability of long text generation are improved.
Owner:BEIJING JIBU QIANLI TECHNOLOGY CO LTD

Multi-agent collaborative question and answer method and system based on Bayesian Nash equilibrium

The invention relates to the technical field of natural language processing, and provides a multi-agent collaborative question answering method and system based on Bayesian Nash equilibrium, and the method comprises the steps: obtaining a user query instruction, generating a reasoning strategy, calling a plurality of agents, enabling each agent to convert a current observation state into a belief state through a belief network according to the current observation state, action parameters are generated, the large model is controlled to execute a text generation task, each agent maintains a value estimation network, and the value estimation networks take the belief state and the action parameters as input to generate instant rewards; aggregating the results of executing the text generation task by each agent to obtain a final answer; the belief state of each agent is aggregated into a unified group belief representation through a belief encoder, comprehensive loss is generated through a centralized hybrid network, and parameters of each agent are updated. Stronger consensus achievement capability and higher problem solving quality are shown in complex reasoning and planning tasks.
Owner:DAREWAY SOFTWARE

Government affair question and answer method, device and equipment based on image-text understanding and storage medium

The invention discloses a government affair question-answering method, device and equipment based on image-text understanding and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining initial image data of a government affair question, preprocessing the initial image data to obtain to-be-processed image data, and storing the to-be-processed image data in a database; determining image information by using an optical character recognition technology, an image classification technology and a target detection technology; carrying out correlation analysis on the image information, carrying out semantic relation and logic relation matching to generate structured information, constructing a knowledge graph, matching the structured information with the knowledge graph to obtain a matching result, generating an initial government affair reply by utilizing a context sensing model and dialogue context information, and sending the initial government affair reply to a server; constructing a user portrait based on the historical behavior data of the user, and determining policy content in the knowledge graph based on the structured information to obtain government affair recommendation information; and integrating the initial government affair reply and the government affair recommendation information by using an image-text generation model to obtain a target government affair reply comprising an interactive diagram so as to improve the efficiency of government affair question answering.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Systems and methods for an indexing tree based artificial intelligence conversation agent

A method of retrieval-augmented text generation is provided to generate a response to by an artificial intelligence (AI) agent. The method incudes constructing an indexing tree reflecting sematic relationship among textual contexts of one or more documents. The indexing tree includes at least a first level of nodes representing a plurality of propositions indicating facts about at least one entity from the documents, a second level of nodes representing a plurality of proposition aggregates that aggregate one or more first-level propositions mentioning a same entity, and a third level of nodes representing a plurality of summaries that summarize one or more second-level proposition aggregates. The method also includes receiving the user query; retrieving, from the indexing tree, one or more nodes across multiple levels based on similarity comparison with the user query based at least on a similarity between embeddings of the nodes and an embedding of the user query.
Owner:SALESFORCE INC

Ultra-long text generation method and device, equipment and storage medium

The invention discloses a super-long text generation method and device, equipment and a storage medium, and the method comprises the steps: responding to a document generation instruction, and generating an initial text outline of a target text through a large model; wherein the initial document outline comprises a plurality of chapter titles; based on the keyword information of the initial document outline, the text complexity of the target text is obtained through calculation, and the text complexity is associated with the subject breadth of keywords and the expected word number of the target text; if the text complexity of the target text is greater than a preset threshold value, determining the initial text outline as a target text outline; and on the basis of the target text outline, the text content corresponding to each chapter title is generated in parallel and spliced to obtain the target text, so that the generation efficiency, the text coherence and the overall text quality of the generated super-long text can be improved, and the generation cost of the super-long text is reduced.
Owner:SHENZHEN YUEHUA EXPRESS CO LTD

Retrieval enhancement generation method and system based on long-term memory and multi-source knowledge iterative fusion

The invention discloses a retrieval enhancement generation method and system based on long-term memory and multi-source knowledge iterative fusion. The method comprises the following steps: firstly, after user inquiry is received, triggering long-term memory database retrieval, external knowledge base retrieval and large language model internal knowledge generation in parallel to construct a multi-source candidate knowledge base, and improving result correlation by adopting a self-adaptive retrieval method based on a dynamic top-k value; secondly, inputting user query and multi-source knowledge into the large language model for iterative fusion, and generating a final answer through conflict detection, fusion processing and confidence evaluation; and finally updating the long-term memory database. The method can improve the accuracy and stability of the generated answers in a multi-round interaction and complex task scene, reduces the risk of answer one-sidedness, inference incompleteness or fact deviation, enhances the multi-source knowledge utilization ability and long-term learning ability of the system, and improves the user experience. Therefore, the comprehensive performance and reliability of the system in applications such as intelligent question answering, information retrieval assistance and text generation are obviously expanded.
Owner:SCHOOL OF SOFTWARE ZHEJIANG UNIV (NINGBO) MANAGEMENT CENT (NINGBO SOFTWARE EDUCATION CENT) +3

Multi-modal document analysis method and device, electronic equipment and storage medium

The invention relates to the technical field of multi-modal document analysis, and discloses a multi-modal document analysis method and device, electronic equipment and a storage medium. The method comprises the steps that picture elements in a document are automatically recognized and extracted through a third-party library module, placeholders are generated in combination with picture positions, and the placeholders are stored in the third-party library module; and then semantic understanding and text generation are carried out by using a multi-modal large model, so that the visual content is converted into a structurable information text. The method has the advantages that by introducing the multi-modal large language model, unified analysis of multi-modal information such as texts, pictures, tables and flow charts is achieved, layout semantics and logic structures of documents are reserved, fusion analysis of image-text content is achieved, the integrity and the intelligent level of document analysis are remarkably improved, and the method is suitable for large-scale popularization and application. Compared with a traditional single-mode method, the method has the advantages that unified understanding and structured expression can be carried out on multi-mode content, and richer and more accurate structured results can be output.
Owner:GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)

Text generation-oriented dynamic self-distillation method for large language model

The invention discloses a text generation-oriented large language model dynamic self-distillation method, which comprises the following steps of: firstly, adding a task instruction prefix to input text data, automatically aligning an input text sequence to input a large language model, and introducing confidence to generate a text sequence; secondly, for a text sequence generated by the large language model, confidence-driven loss optimization is carried out on the large language model in combination with confidence loss and distillation loss; then, in a large language model self-distillation generation stage, a confidence-quality joint evaluation framework is constructed, and adaptive weight distribution is performed on a large language model generation text sequence; and finally, evaluating a self-distillation training effect based on the multi-dimensional global evaluation index in a self-distillation training period gap. According to the method, high-reliability knowledge migration and self-distillation are realized under the condition of low resources, and the text is accurately and efficiently generated.
Owner:HANGZHOU DIANZI UNIV

User session response system, method and equipment based on health data and medium

The invention relates to the technical field of artificial intelligence and medical information processing, and discloses a user session response system and method based on health data, equipment and a medium. The system comprises a session semantic recognition module, an entity information extraction module, an archive data alignment module, a health classification tree construction module, a data hierarchical matching module and a response text generation module. The method comprises the following steps: recognizing consultation semantics contained in a health consultation session of a target user; determining a target session text in the health consultation session according to the consultation semantics; performing data alignment on the medical entity information in the target session text according to the historical health file, constructing a structured health file based on a data alignment result, and generating a health state classification tree; performing classification matching on the consultation semantics by utilizing the health state classification tree to obtain a target health state classification; and generating a corresponding target response text based on the target health state classification, and pushing the target response text to the target user side. According to the invention, the healthy session response accuracy can be improved.
Owner:BEIJING NAYA MEDICAL TECH CO LTD