Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

681 results about "Text processing" patented technology

In computing, the term text processing refers to the theory and practice of automating the creation or manipulation of electronic text. Text usually refers to all the alphanumeric characters specified on the keyboard of the person engaging the practice, but in general text means the abstraction layer immediately above the standard character encoding of the target text. The term processing refers to automated (or mechanized) processing, as opposed to the same manipulation done manually.

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Method for improving long text processing efficiency and accuracy

The invention discloses a method for improving long text processing efficiency and accuracy, and relates to the technical field of natural language processing and large language models.According to the method, text word segmentation embedding, sliding block preprocessing, YaRN position code injection, dynamic sparse attention calculation, multi-level attention fusion, graded KV cache management and output generation are sequentially executed; position drift is inhibited through logarithmic scaling, and key contexts are adaptively screened according to the attention activeness, so that the attention calculation complexity is close to linearity; in million-level Token reasoning, the video memory occupation of the method is reduced, the remote dependency recall rate is improved, and the method is suitable for scenes such as document analysis, code auditing and multi-mode streaming understanding.
Owner:BEI JING JING YUE KE JI YOU XIAN GONG SI

Prompt word attack detection method and system of large language model and electronic equipment

The invention provides a cue word attack detection method and system of a large language model and electronic equipment, and relates to the technical field of text process.The method comprises the steps that a to-be-processed text block in an original text is determined; performing heuristic semantic analysis on the text block to obtain a first suspicious text block and a first risk score thereof; performing semantic analysis on the first suspicious text block and the context information thereof through a first language model to obtain a second suspicious text block and a second risk score thereof; performing structured verification of the attack intention on the second suspicious text block according to a preset meta prompt word through a second language model to obtain an attack intention verification result; the attack intention verification result comprises a third suspicious text block and a third risk score thereof; and executing a target response action on the original text based on the first risk score, the second risk score and the third risk score. According to the method, the efficient processing requirement of the long text can be considered, and the detection accuracy of complex attacks can be improved.
Owner:HANG ZHOU LING XIN SHU KE XIN XI JI SHU YOU XIAN GONG SI

Computer-aided system for multidimensional generative value assessment and applicant selection

ActiveDE202025107568U1InstrumentsData packData stream
A computer-implemented system for multidimensional generative value assessment and applicant selection, consisting of: a data collection unit configured to electronically receive applicant data consisting of structured academic records, work experience records, digital documentation, and unstructured narrative responses generated from generative self-assessment instruments and contextual interviews; a feature extraction unit coupled to the data acquisition unit, configured to apply computer-assisted text processing, semantic analysis, and token-level attribute identification to transform narrative responses and structured data into multidimensional feature vectors that represent generative indicators of innovation, mentoring, collaborative performance, resilience, social contribution, ethical consistency, and predicted institutional impact; a weighting calculation unit configured to assign weight values ​​to the extracted feature vectors based on a digital generative profile definition matrix that includes dimensions, sub-criteria, indicators, documentation requirements and importance coefficients, with the weighting being distributed across the generative dimensions defined in the digital matrix and configurable according to the institutional context; a quantitative rating unit configured to calculate a generative rating score by aggregating weighted feature vectors derived from self-assessment inputs, interview-based ratings, document analyses, and authenticity predictions, with the aggregation including normalization, nonlinearity correction, conflict handling, and artifact frequency balancing to obtain a consolidated score; a proof verification unit configured to electronically validate referenced digital evidence by performing content extraction, metadata verification, pattern matching, and cross-document correlation to determine authenticity, credibility, and contextual relevance with respect to the calculated feature vectors; a classification determination unit configured to assign a classification level to an applicant by comparing the generative assessment score with a set of system-defined calculation thresholds, including at least a lower threshold, a middle threshold and an upper threshold, the classification levels representing different generative maturity states and determining subsequent eligibility for selection; a decision generation unit configured to produce a digital output data set that includes classification level, feature aggregation summaries, evidence validation results, and recommended organizational actions, wherein the decision generation unit encodes the data set in a digitally signed, tamper-proof format and stores it on a non-volatile storage medium; and A system control unit acts as an operational interface to all other units and is configured to orchestrate data flow, scheduling, process state transitions, and event logging to ensure verifiable traceability, consistency, and auditability of the evaluation and selection processes.
Owner:BERNARDO OHIGGINS UNIVERSITY +3

Precise alignment method and system for multilingual terminologies in nuclear power field

The invention relates to the technical field of text processing, and discloses an accurate alignment method and system for multilingual terminologies in the nuclear power field, and the method comprises the steps: collecting original terminologies from a nuclear power design document, an operation manual and an international standard database, and constructing an original terminology set in the nuclear power field; encoding the original term set after term cleaning to obtain standardized term blocks; extracting semantic feature vectors according to the standardized term chunks; based on the attention scores of the semantic feature vectors in the sequence, performing weighted fusion on the semantic feature vectors to obtain enhanced semantic representation; constructing candidate term pairs in the nuclear power field according to the similarity of the term pairs in the enhanced semantic representation; performing consistency check on the candidate term pairs through entity relationships and attributes of the knowledge graph, and outputting the candidate term pairs passing the consistency check as target term pairs in the nuclear power field; according to the invention, the efficiency of accurate alignment of multilingual terminologies in the nuclear power field can be improved.
Owner:JIANGSU NUCLEAR POWER CORP

Text classification method and system based on semantic analysis

The invention relates to the technical field of text processing, in particular to a text classification method and system based on semantic analysis, and the method comprises the following steps: segmenting semantic units, constructing a direction change sequence, positioning mutation nodes, generating a consistency section, forming a convergence section, and outputting a classification result. According to the method, a continuous change sequence is formed by constructing a semantic embedding vector and calculating a direction difference, a semantic mutation point can be anchored and divided into sections by combining mutation intensity identification and local jump tracking, and a semantic closed structure and a convergence section are extracted by means of context direction consistency judgment and generic label comparison; precise recognition of a semantic relation chain is realized, semantic jump and conflict starting points can be dynamically sensed, the semantic boundary recognition capability is improved, and the understanding and classification capability of a model on semantic attribution in a complex context is enhanced on the premise of not depending on a fixed dictionary and shallow statistics. The problems that a traditional model is slow in response to an abrupt change structure and weak in semantic convergence recognition are effectively solved.
Owner:上海笑聘网络科技有限公司

A method, device, and medium for processing NOTAM text based on semantic enhancement

This invention relates to the field of text processing technology, and in particular to a method, device, and medium for processing navigational notice text based on semantic enhancement. The method includes: first, acquiring a content carrier to be processed; then, acquiring the semantic vector and glyph feature vector of the content carrier; concatenating the two types of vectors to form an enhanced text representation; extracting temporal features from the enhanced text representation to obtain temporal features containing forward and backward logical relationships within the text; acquiring the weights of words and sentences in the temporal features and performing weighting to obtain weighted word representations and weighted sentence representations; performing correction processing on the weighted representations to generate corrected text; and finally, validating the corrected text and outputting the target text. This invention can improve the accuracy and efficiency of content carrier processing.
Owner:CIVIL AVIATION UNIV OF CHINA

Non-intra-macro identifier processing method in macro reference row, electronic equipment and medium

The invention relates to the technical field of macro text processing, in particular to a non-macro identifier processing method in a macro reference row, electronic equipment and a medium, the method comprises the following steps: S1, obtaining storage information corresponding to each macro reference of the macro reference row, the storage information corresponding to the macro reference comprises a macro reference identifier, a line number in a macro expanded text of a macro reference line and an intra-macro identifier list corresponding to the line number in the macro expanded text of the macro reference line; step S2, obtaining non-intra-macro identifier positioning information and a corresponding to-be-added macro reference identifier in a macro expanded text of the macro reference line; and S3, adding the non-intra-macro identifier positioning information to an intra-macro identifier list corresponding to the corresponding to-be-added macro reference identifier. According to the method, the non-intra-macro identifier binding object after row expansion can be accurately and quickly referenced for the macro, and a debugging function is realized based on the binding object.
Owner:BEIJING NORI INTEGRATED CIRCUIT DESIGN CO LTD +2

Cross-border digital service multi-language real-time interaction and semantic error correction method and system

The invention discloses a cross-border digital service multi-language real-time interaction and semantic error correction method and system, and belongs to the technical field of text processing.The method specifically comprises the steps that when cross-border interaction begins, voice, text and auxiliary multi-mode information of a user are collected, fused semantic representation is generated, the semantic representation is input into a cross-language prediction model, and the cross-language prediction model is obtained; a target language candidate result is obtained, a reverse mapping channel is established, when semantic errors or ambiguity occurs in the target language candidate result, collaborative correction is conducted in combination with the scene rule base, user historical preferences and real-time feedback, the corrected result is processed through a predefined target context sensitive model, and the target language candidate result is obtained. Automatically adjusting the culture expression, the terminology and the compliance, and outputting to the cross-border service terminal in real time; according to the method, the target context sensitive model is introduced in the output stage, cultural expression, terminologies and law compliance are automatically adjusted, and multi-language seamless communication and high-reliability semantic transfer in cross-border digital services are achieved.
Owner:JIANGSU ZHIMENG INTELLIGENT TECH CO LTD

Self-adaptive retrieval enhancement generation system based on multi-dimensional problem features and implementation method

The invention relates to a self-adaptive retrieval enhancement generation system based on multi-dimensional problem features and an implementation method. The system comprises an RAG initialization module, a text processing module, a model initialization module, a vector processing module, a retriever module and an RAG retrieval generation module. The input end of the retriever module receives the text vector of the vector processing module, the output end of the retriever module is connected with the RAG retrieval generation module, a strategy pool of five types of retrieval strategies and the self-adaptive decision module are arranged in the retriever module, and the retriever module is used for determining retrieval strategies through keyword matching and problem length judgment and constructing RAG cue words. And the RAG retrieval generation module is used for calling the retriever module to input the obtained text segment into the large language model to generate an answer after reconstructing the user question. By adopting the method, the retrieval efficiency and the answer generation quality of the RAG system can be improved.
Owner:NAT UNIV OF DEFENSE TECH

Long text abstract generation method and device

The invention provides a long text abstract generation method and device, and belongs to the technical field of natural text processing, the method comprises the following steps: using Euclidean norm to carry out importance sorting on Tokens and carrying out compression storage according to a sparse rate, so that global key information is completely reserved and memory occupation is obviously reduced; in the decoding stage, local attention scores and global attention scores are calculated in parallel, entropy differences are mapped into fusion weights through Sigmoid by combining temperature adjusting parameters, and dynamic balance of local details and long-distance dependence is achieved. The local key value pairs and the global key value pairs are subjected to weighted integration based on the fusion weight, a continuous semantic spectrum is formed in a single decoding layer, splicing breakage caused by traditional partitioning is eliminated, the problems of input limitation and semantic splitting are effectively relieved, the context length capable of being processed by a model is expanded under the condition that the calculation amount is not remarkably increased, and the method has the advantages of being simple in structure and convenient to operate. And local and global context information is adaptively fused, so that the accuracy and continuity of the abstract are effectively improved.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719 +1

Aviation text content cleaning and labeling method, system and equipment and medium

PendingCN121543549ANatural language analysisBiological modelsDuplicate contentAviation
The invention relates to the technical field of aeronautical text data processing, and discloses an aeronautical text content cleaning and labeling method, system, device and medium wherein the method comprises: noise filtering: identifying and removing noise in an aeronautical text in combination with static cleaning and a general large model; format standardization: converting the aviation text after noise removal into a standardized format text; duplicate removal and error correction: detecting duplicate contents based on a hash algorithm, and correcting spelling errors and grammar errors based on a general large model to obtain an aviation text subjected to duplicate removal and error correction; entity identification: key entities are extracted based on the general large model, and the extracted key entities are labeled; active learning: screening high-value samples in the marked key entities based on an uncertainty query strategy; and dynamic optimization: performing verification and iterative optimization on the marking result in combination with the aviation knowledge base. According to the method, the automation level, the labeling accuracy and the system self-adaptive capability of aviation text processing can be remarkably improved.
Owner:四川腾盾科技有限公司 +1

Paper text processing method based on intelligent glasses

The invention relates to the technical field of artificial intelligence, and discloses a paper text processing method based on intelligent glasses. The method comprises the steps that intelligent glasses are awakened through an awakening word, a user can conduct image recognition through a voice instruction, and whether a textbox is complete or not is judged. If the textbox is complete, character recognition is carried out; if not, the intelligent glasses help the user to adjust through voice guidance until the textbox is completely displayed. If the recognized characters are not the language set by the user, the intelligent glasses automatically translate the recognized characters into the user language, and voice synthesis is carried out to generate a voice file. And after generation, inquiring the user whether to play the voice, and if not, encrypting and storing the voice file. The problems that current text detection precision is not high, the reading process is not natural, user control is complex, and a feedback mechanism is single are solved.
Owner:SHENZHEN SENSING FUTURE TECHNOLOGY CO LTD

Data processing method and apparatus, and target question-answering model training method and apparatus

Provided in the embodiments of the present disclosure are a data processing method and apparatus, and a target question-answering model training method and apparatus. The data processing method comprises: adjusting an initial text processing result corresponding to initial text in an initial text pair, in order to obtain an updated text processing result, and on the basis of the initial text and the updated text processing result, constructing an updated text pair; on the basis of the updated text pair, updating an initial text processing model, in order to obtain a reference text processing model; on the basis of the initial text pair, using the initial text processing model and the reference text processing model to respectively obtain predicted loss results and reference loss results that correspond to a plurality of tokens in the initial text processing result; on the basis of the predicted loss results and the reference loss results, determining a loss change result for each token; and on the basis of the loss change result for each token, executing a data processing task. Thus, token-level distinction is realized, and reasoning-focused supervised fine-tuning can be performed on the basis of the token-level distinction, thereby improving the reasoning capability of models.
Owner:ALIBABA (CHINA) CO LTD

Industrial knowledge base dynamic construction method and system based on multi-source heterogeneous data

The invention provides an industry knowledge base dynamic construction method and system based on multi-source heterogeneous data, and relates to the technical field of data processing.The method comprises the steps that target industry multi-source heterogeneous data, user historical behavior data, a natural language question text and a processing text are collected to generate a query data stream; mapping the user question into a query semantic vector, converting the user interaction behavior into a behavior feature vector, and generating a joint query vector; industry knowledge units are automatically extracted from multi-source heterogeneous data and external knowledge sources, and a dynamic knowledge topological graph is constructed; establishing a dynamic knowledge routing table, taking the joint query vector as a routing request, matching a target knowledge node, recording a query path and caching an association rule; and counting and querying path node transition probability to identify a hotspot path, mining a deep association rule, and reversely adjusting a knowledge topological graph structure, so that user intentions can be accurately matched, an industry knowledge base can be dynamically optimized, and knowledge retrieval and association efficiency can be improved.
Owner:BEIJING ALL VIEW CLOUD DATA TECH CO LTD

Document segmentation method and device based on large language model, equipment and storage medium

The invention provides a document segmentation method and device based on a large language model, equipment and a storage medium, and relates to the technical field of text processing. The method comprises the steps of inputting a to-be-segmented target document into a pre-trained large language model, and executing the following operations through the large language model: performing text layout analysis on the target document, and identifying titles and all paragraphs of each level in the target document; for each paragraph, inserting an associated title related to the paragraph in all titles into an initial position of the paragraph to obtain a corresponding target paragraph; and sorting all the target paragraphs based on the semantic similarity among all the target paragraphs, and determining a segmentation result of the target document based on all the sorted target paragraphs. By the adoption of the technical scheme, when document segmentation is carried out, semantic loss in the document segmentation process can be effectively reduced, and therefore the document segmentation effect is improved.
Owner:CHINA LIFE ASSET MANAGEMENT CO LTD

Method, system and medium for constructing a variant character dictionary of ancient chinese medical books and text alignment

The present application belongs to the technical field of natural language processing for traditional Chinese medicine ancient books, and particularly relates to a method and system for constructing a variant character dictionary and text alignment of traditional Chinese medicine ancient books, and a medium. The present application combines the recognition of variant characters and the construction of a variant character dictionary to achieve a text alignment method for traditional Chinese medicine ancient books. Specifically, the present application uses deep learning and natural language processing technology to automatically extract variant character features, significantly improving the coverage range and recognition accuracy; through dynamic programming, semantic similarity calculation and knowledge graph fusion, the multi-modal features are comprehensively considered to significantly improve the alignment accuracy. At the same time, the model can dynamically adapt to new texts and variant characters, and has stronger expansibility and adaptability; and the knowledge graph is used to optimize the alignment result, improving the accuracy and efficiency of text processing. The final generated result is the aligned text sequence, in which the variant characters are correctly recognized and mapped to standard characters. The present application has good application prospects in the digitization of traditional Chinese medicine ancient books.
Owner:CHENGDU UNIV OF TRADITIONAL CHINESE MEDICINE

Cross-language text classification and processing method and system based on deep transfer learning

The invention provides a cross-language text classification and processing method and system based on deep transfer learning, and relates to the technical field of text processing, and the method comprises the steps: extracting feature representations of a source language text and a target language text at different linguistic levels through a multi-level semantic transfer network; and determining an optimal alignment path, performing nonlinear mapping alignment to obtain fusion features, propagating category semantics by using a semantic bridging function, iteratively updating pseudo-tag confidence distribution of the target language text, and completing classification in combination with a multi-task learning model. According to the method, the problem of text classification in a cross-language scene is effectively solved, and the accuracy and efficiency of low-resource language text processing are improved.
Owner:SHANGHAI XIRUAN TECH CO LTD

Intra-macro identifier position information acquisition method, electronic equipment and medium

The invention relates to the technical field of macro text processing, in particular to an intra-macro identifier position information acquisition method, electronic equipment and a medium, and the method comprises the following steps: setting a corresponding macro virtual file for each macro reference in a chip design code original file; loading a chip design code original file, and obtaining initial position information of an identifier in each macro based on a macro virtual file corresponding to each macro reference; expanding each target macro reference in a target code line in the original file of the chip design code to obtain a target macro expanded text; obtaining a column position offset and a row position offset corresponding to each target macro reference based on the target macro expanded text and the macro initial expanded text of each target macro reference; and generating position information of the intra-macro identifier in each target macro reference in the target macro expanded text based on the column position offset and the row position offset corresponding to each target macro reference and the initial position information of the intra-macro identifier in each target macro reference.
Owner:BEIJING NORI INTEGRATED CIRCUIT DESIGN CO LTD +2

Large and small model collaborative natural language processing method, system and equipment and medium

The invention discloses a big and small model collaborative natural language processing method, system and device and a medium, which are applied to the field of language processing, and the method comprises the following steps: performing data preprocessing on to-be-processed text data to obtain word embedding vector data; obtaining the semantic complexity of the word embedding vector data and the input text length of the to-be-processed text data, and calculating the task complexity according to the semantic complexity and the input text length; splitting the text processing task into a plurality of sub-tasks according to the task complexity; according to a preset task complexity threshold, dynamically allocating each sub-task to the large model and the small model for processing to obtain a text reasoning result of each sub-task; and fusing the text reasoning results to obtain a text processing result. According to the method, by dynamically distributing the tasks to the large and small models and fusing the reasoning results of the large and small models, the reasoning efficiency and precision are effectively balanced, the accuracy and the real-time performance of the natural language processing tasks are improved, and meanwhile the dynamic adaptability of the system is enhanced.
Owner:GOSUNCN TECH GRP

Contract identification and intelligent management system and method based on MCP

The invention provides an MCP-based contract identification and intelligent management system and method. The method comprises the following steps: acquiring contract documents in various formats of a target contract, performing text processing and image processing on the contract documents, and performing multi-modal fusion to obtain intermediate data; constructing a knowledge graph, and performing clause logic conflict detection and compliance risk assessment on the target contract to obtain a contract assessment result; constructing a reinforcement learning decision model according to the knowledge graph and the contract evaluation result; based on the reinforcement learning decision model, MCP interface standard data is determined, and an adapter model is constructed; processing the contract document by using an adapter model to obtain a contract MCP data stream, reconstructing a knowledge graph and a reinforcement learning decision model for the contract MCP data stream based on MCP interface specification data, and integrating a model processing result and a model operation state; and executing the contract processing flow to obtain a contract management scheme. According to the invention, efficient and accurate contract document processing and management are realized.
Owner:BEIJING YULORE INNOVATION TECH

File processing methods, electronic device, storage medium and computer program product

The present disclosure relates to the fields of large model technology and text processing. Disclosed are file processing methods, an electronic device, a storage medium and a computer program product. A method comprises: in response to an input instruction acting on an operation interface, displaying on the operation interface a file to be processed, said file containing text to be processed of at least one modality; and, in response to a processing instruction acting on the operation interface, displaying a processing result on the operation interface, the processing result being used for representing that said file has been successfully named and stored on the basis of a target processing rule, the target processing rule being a processing rule corresponding to a target type of said file, the target type being determined by comparing said text with preset text contained in a plurality of preset files, and the types of different preset files being different. The present disclosure solves the technical problem in the prior art of relatively low file processing accuracy.
Owner:ALIBABA (CHINA) CO LTD

Text processing method and apparatus

The embodiments of the present disclosure relate to the technical field of computers. Provided are a text processing method and apparatus. The method comprises: on the basis of target text and at least two pieces of sub-text corresponding to the target text, using a text processing model to determine an initial reasoning path associated with a target entity, wherein the target entity is an entity included in the target text; on the basis of the target text, the at least two pieces of sub-text and the initial reasoning path, using the text processing model to obtain an initial text processing result; on the basis of the initial text processing result, using the text processing model to update the initial reasoning path, so as to obtain a target reasoning path; and on the basis of the target text, the at least two pieces of sub-text and the target reasoning path, using the text processing model to obtain a target text processing result. By means of obtaining an initial text processing result and using same to adjust a target reasoning path, the method significantly improves the accuracy of a target text processing result finally obtained with reference to the target reasoning path.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD +1

Entity pair guided scientific and technical literature document level relation extraction method and system

The invention provides an entity pair guided scientific and technical literature document level relation extraction method and system, and the method comprises the steps: carrying out the entity recognition of an input scientific and technical literature document, and obtaining an entity set in the scientific and technical literature document; based on an entity pair pre-screening mechanism of multiple sampling and similarity verification, screening out a candidate entity pair set from all possible entity pairs of the entity set; then generating enhanced relation description fusing corresponding entity type information and relation semantics between the entity pairs; based on a pre-constructed relation semantic knowledge base, a double-layer filtering mechanism is adopted, and a corresponding fine screening candidate relation set is retrieved for each enhanced relation description; and guiding the large language model to perform triple fact judgment by using detailed semantic description of the candidate relationship to obtain an output result. According to the method, high-precision relation extraction is ensured, and meanwhile, the calculation overhead of long text processing is remarkably reduced, so that scientific and technical literature document-level relation extraction is more accurate and efficient.
Owner:CHENGDU DOCUMENT & INFORMATION CENT OF CHINESE ACAD OF SCI

Heterogeneous medical text-oriented index name normalization and mapping method and system

The invention belongs to the field of text processing, and particularly relates to a heterogeneous medical text-oriented index name normalization and mapping method and system, which specifically comprises the following steps of: extracting a triple containing index names, measurement units and numerical values from a text by utilizing an entity recognition model, and retrieving candidate terms based on a standard term library, generating a rule matching probability by calculating literal form similarity and dimensional conversion consistency, mapping an index name to a high-dimensional vector space, separating primary and secondary feature components by using orthogonal projection, calculating a semantic matching probability through cosine similarity, and after the two probabilities are fused, if the confidence degree is in a fuzzy interval, judging that the index name is not in the fuzzy interval; and if not, further performing weighted correction in combination with the semantic association degree of the definition text, sorting according to the corrected confidence, and selecting an optimal standard term as normalized output. According to the method, the problem that index names in heterogeneous medical texts are difficult to align is solved, and the standardization degree of medical terms is improved.
Owner:WEST CHINA HOSPITAL SICHUAN UNIV

Document processing method, document question and answer method and computer program product

The invention provides a document processing method, a document question and answer method and a computer program product, and relates to the technical field of document question and answer. The document processing method comprises the steps of performing picture recognition and text processing on a to-be-processed document in a target format to obtain a plurality of text segments; wherein the to-be-processed document comprises a table and / or a picture; based on the plurality of text segments, generating multi-level abstracts and multi-level recommendation questions corresponding to the multi-level abstracts; and performing associative storage on the plurality of text segments, the multiple levels of abstracts and the multiple levels of recommendation questions to obtain a document database. The document question-answering method comprises the steps of obtaining an input question of a user; based on the input question, performing retrieval in a document database, and determining a recalled text segment and a recalled recommendation question; and generating an answer result in combination with the recalled text segment and the recalled recommendation question. According to the method and the device, characters, pictures, tables and the like existing in the document can be processed, and question and answer recall processing is carried out based on the extracted multi-level abstract and the generated recommendation question.
Owner:INNOVATION QIZHI TECH GRP CO LTD

Semantic analysis and recognition method based on artificial intelligence

The invention discloses a semantic analysis and recognition method based on artificial intelligence, particularly relates to the field of semantic analysis, and is used for solving the problems of frequent semantic offset and inconsistent translation results caused by difficulty in accurate recognition and disambiguation of cultural load words in existing cross-language text processing. The method comprises the following steps: performing text structure decomposition and culture metaphor analysis on a cross-language contrast corpus, identifying culture load words with specific culture meanings, further obtaining semantic vector sets of the culture load words in a source culture context and a target culture context, and calculating a semantic escape distance between the culture load words and the target culture context to generate an escape path; and training a context disambiguation model in combination with the escape path and the context, so that the model can output final semantic probability distribution of the culture load words in the target culture when the to-be-analyzed text is input, and generates a semantic mapping prompt in a cross-language semantic conversion process based on the probability distribution. Therefore, cross-culture accurate semantic analysis and prompt of the culture load words are realized.
Owner:XIAN DAMAI NETWORK TECH CO LTD

Text processing method and device, electronic equipment, medium and program product

The invention provides a text processing method and device, electronic equipment, a medium and a program product, and can be applied to the technical field of big data and the technical field of artificial intelligence. The method comprises the steps of obtaining a to-be-processed unstructured text; performing event abstract classification on the unstructured text to obtain a plurality of semantic event units; performing modular analysis on the semantic event unit to obtain corresponding module attribute information; and performing structured conversion and format verification based on the module attribute information to obtain target structured data corresponding to the unstructured text.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Coronal mass ejection detection method and device based on CLIP model

The invention discloses a coronal mass ejection detection method and device based on a CLIP model. The method comprises the following steps: acquiring CME historical data; preprocessing the CME historical data to obtain CME preprocessed data; and processing the CME preprocessing data by using a CLIP model to obtain CME image segmentation data and CME arrival time prediction data. The method comprises the following steps: performing feature extraction on each frame of an image by using an image processing module of a CLIP model to generate an image feature vector, encoding description of video data or related text information by using a text processing module of the CLIP model to generate a text feature vector, calculating the similarity between the image feature vector and the text feature vector, and detecting a CME image. And the CME image detection efficiency and accuracy are improved.
Owner:CHINESE PEOPLES LIBERATION ARMY UNIT 31016

Using masked text processing for information processing with documents

A method implements masked text processing for information processing with documents. The method involves receiving a document page as an image including text image data. The method further involves extracting text unit data and text location data from the image corresponding to the text image data using an optical character recognition (OCR) engine. The method further involves generating mask data for the text unit data with color data based on text type data. The method further involves producing a masked image by replacing the text image data with the mask data using the color data with the location data in the image. The method further involves transmitting the masked image to a machine learning model to execute a downstream task.
Owner:SCHLUMBERGER TECH CORP