Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

711 results about "IntraText" patented technology

IntraText is a digital library that offers an interface while meeting formal requirements. Texts are displayed in a hypertextual way, based on a Tablet PC interface. By linking words in the text, it provides Concordances, word lists, statistics and links to cited works. Most content is available under a Creative Commons license It also offers publishing services that enable similar advantages.

Text content generation method based on artificial intelligence

The invention relates to the technical field of artificial intelligence and natural language processing, and particularly discloses a text content generation method based on artificial intelligence, and the method comprises the steps: obtaining natural language text input, and extracting a semantic recognition feature vector; acquiring context state data and encoding the context state data into a state recognition feature vector; generating a fusion feature vector containing a semantic and state association relationship through fusion analysis; constructing a causal discrimination model based on the fusion features, and outputting the matching confidence of semantics and states; dynamically adjusting a generation strategy according to the confidence coefficient, if the matching degree is high, generating a standard text, otherwise, triggering an error correction mechanism to output a corrected text; and finally, performing logic consistency verification on the generated text to ensure that physical constraints, technological procedures and safety standards in the industrial field are met. According to the method, by introducing multi-level feature fusion, causal reasoning, intelligent error correction and rule verification mechanisms, context perception and safety controllability in the text generation process are achieved.
Owner:JINING POLYTECHNIC

Document processing method and system based on text content extraction

The invention relates to a document processing method and system based on text content extraction. The method comprises the steps that an original document containing text, image and format information is received, the encoding format of the document is automatically detected, character set conversion is executed, and hierarchical indexes including page numbers, paragraphs and tables are established for an unstructured document; the method comprises the following steps: synchronously processing text content and visual layout through a pre-trained visual-language model, extracting word-level and sentence-level semantic features by a text stream embedding layer, analyzing spatial distribution features of document elements by a visual encoder, and fusing text and visual features through a cross-modal attention mechanism; and loading the domain knowledge graph matched with the document type, and executing entity linking to associate the text mentions to the knowledge nodes. According to the document processing method and system based on text content extraction, through the synergistic effect of vision-text joint coding and knowledge enhancement, the accuracy of financial contract key clause recognition tasks is improved, the error rate is lower than that of industry benchmark products, and the semantic understanding precision is remarkably improved.
Owner:WIN THE BID HUIKANG TECH CO LTD

Text information structured recovery method and system based on large language model and application

The invention provides a text information structured recovery method and system based on a large language model and application. The text information structured recovery method comprises the following steps: S1, extracting original text content from a webpage or an unstructured document; s2, designing a cue word template according to different scenes and target structures, and generating cue words; s3, guiding a large language model to analyze the original text content and the cue word in the step S2, and generating a text result with a hierarchical structure; s4, analyzing a text result in the step S3, constructing a semantic structure tree, and forming a multi-layer nested structure; and S5, applying the multi-layer nested structure in the step S4 to database modeling and content indexing, compared with the defects of non-uniform structure loss, poor universality and semantic understanding intelligence deficiency in the prior art, the manual processing cost can be remarkably reduced, the data structuring efficiency and accuracy are improved, and the semantic understanding intellectuality is improved. And a stable and high-quality structured text support is provided for a large model ecological system.
Owner:SHENZHEN NAT HEALTH CULTURE COMM CO LTD

Multi-modal consistency verification and harmonization models

A system may access content comprising text content, visual content, and / or audio content. The system may perform, based on a harmonization model and / or consistency model, a harmonization check and / or a consistency check on the content. The system may recognize, based on the harmonization check and / or the consistency check, a conflict to be corrected. The system may identify a property of the content that should be changed based on the recognized conflict. The system may generate a corrective action based on the property of the content that should be changed.
Owner:REVE AI INC

Automatic video editing method based on semantic analysis

The invention relates to an automatic video editing method based on semantic analysis, and belongs to the technical field of video editing. The method comprises the following steps: extracting anchor point information of text content in a streaming media file based on a knowledge graph, establishing a support relationship, and determining a main anchor point, a sub anchor point and a argument anchor point by combining the support relationship and the anchor point information; constructing a dynamic trigger mechanism, setting a verification direction of the dynamic trigger mechanism, determining a semantic extension direction of the anchor point information according to the verification direction, and verifying whether a precondition chain supporting the anchor point information exists or not; and setting a credibility scoring strategy, determining the credibility of the anchor point information according to the logic complexity, the language confidence and the text consistency, forming a sequence set of the anchor point information, and splicing to form a logically coherent video editing text result. According to the method, automatic and intelligent editing of the video is realized based on semantic analysis, and the editing efficiency and quality are remarkably improved.
Owner:BEIJING GONGXIN INTERNET TECHNOLOGY CO LTD

System and Method for Modifying Textual Content

A system and method are provided for dynamically modifying textual content by applying a text analysis tool and selectively using a large language model (LLM). The method includes applying a text analysis tool to text having been added to a document, identifying a portion of text in the document as violating a rule in a rule set associated with the document, and providing a first input to a large language model (LLM). The first input comprises at least one prompt requesting a revision to the portion of text in the document. The method also includes receiving, from the LLM, a response to the first input, the response comprising a suggested modification to the document related to the portion of text; and providing an option to apply the suggested modification to the document.
Owner:SHOPIFY INC

Localized bidding document error checking method based on large language model

The invention relates to the technical field of text inspection and analysis, in particular to a localized bidding document error inspection method based on a large language model, which comprises the following steps of: constructing a bidding requirement knowledge graph which comprises a plurality of requirement item nodes, the requirement item node attribute comprises requirement content, a chapter to which the requirement item node attribute belongs and an importance level; establishing a preliminary mapping relationship between the requirement item node and the text content set, and generating a structured bidding file parse body; and carrying out double-round model analysis and error checking, generating an error record for the items judged to be not satisfied, and outputting a structured error list comprising the non-satisfied items and reasons of the non-satisfied items. The method improves the accuracy and coverage of error checking, and is especially suitable for checking scenes of bidding and tendering documents with numerous and jumbled contents and various formats.
Owner:SHANGHAI BELDEN PROJECT MANAGEMENT CONSULTING CO LTD

AI reading control method and system based on artificial intelligence

The invention relates to the technical field of natural language processing, in particular to an AI reading management and control method and system based on artificial intelligence, and the method comprises the following steps: carrying out text word segmentation processing on an input original text, segmenting the text into independent sentences, recognizing basic word units in each sentence, analyzing semantic adjacency relationships and syntactic structure features among vocabularies, and carrying out word segmentation processing on the word units; screening and extracting potential phrases representing paragraph meanings, and establishing a candidate semantic unit set; according to the method, through word segmentation, syntactic structure recognition and semantic adjacency analysis of the original text, potential phrases capable of representing paragraph significance are extracted, the candidate semantic unit set is constructed, and modeling of the semantic structure in the text is achieved. And associating the set with the reading fixation duration and the playback action of the user sentence by sentence to obtain a reading behavior response of each semantic fragment, and executing semantic weighting and reading behavior cross analysis according to the reading behavior response. Through the linkage mode, the semantic focus actually focused by the user at present can be recognized.
Owner:SHENZHEN JOYAR SMART MFG TECH LTD

Long text writing generation method and system based on multi-agent cooperation

The invention belongs to the field of natural language processing and artificial intelligence, and particularly relates to a long text writing generation method and system based on multi-agent cooperation. The method comprises the following steps: receiving a text theme input by a user, completing complex theme decoupling, and outputting an outline directory; the method comprises the following steps: collecting keywords according to a text topic and an outline catalog, obtaining a retrieval result by retrieving the Internet, a domain database and network resources, and preprocessing to obtain a reference with a uniform format; performing systematic evaluation on the reference literature from the field integrating degree and the logic relevance, and performing grading processing on the reference literature to obtain a high-quality reference literature; based on the text theme and the high-quality reference, long text content is generated according to a set logic structure and writing specifications. According to the method, the fragmented expression problem is effectively avoided, the academic value and reading experience of the text are remarkably improved, the inherent defects of a traditional long text generation technology are practically overcome, and a high-efficiency solution is provided for text creation in the professional field.
Owner:MILITARY SCI INFORMATION RES CENT ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY

Machine learning large language model ensemble deployment in content summarization

System and method generating a summarization of text content, performed in a machine learning neural network large language model (LLM) ensemble. The method comprises inputting text content that includes an unstructured text dataset to a trained baseline LLM. The LLM ensemble includes the trained baseline LLM, a trained classification LLM, and multiple finetuned LLMs. Generating, based on performing natural language processing tasks, a baseline summary of the text content based on the trained baseline LLM, and a classification of topics of the text context via the trained classification LLM. Generating respective finetuned LLM summaries of the text based upon inputting the text content to the multiple finetuned LLMs. Determining, based on a semantic similarity analysis, respective text semantic similarity measures across the baseline summary compared to the finetuned LLM summaries. And generating a summarization of the text content for a topic based on the similarity measures.
Owner:BHAN VAIBHAV

Document interpretation and report generation method and device, equipment and medium

The invention relates to the technical field of natural language processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a document interpretation and report generation method, device, equipment and medium, which comprises the following steps: receiving an original document set to generate a structured document object, executing optical character recognition on an image content set to generate a recognition text set, the recognition text set and the text content set are combined into a unified text sequence, element item extraction is executed based on the interpretation template parameter set to generate an interpretation element set, a retrieval enhancement context is retrieved and generated from the domain knowledge base, and the unified text sequence, the interpretation template parameter set and the retrieval enhancement context are input into a language model to generate an interpretation result. And generating report content based on the historical report template set. According to the method, automatic closed loop of document interpretation and report generation is realized through multi-modal unified processing and semantic enhanced reasoning, the efficiency is improved, and the manual dependence and compliance risk are reduced.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Document format processing method and system based on large language model, terminal and medium

The invention relates to the field of document processing, and particularly provides a document format processing method, system, terminal and medium based on a large language model.The method comprises the steps that a generation instruction containing a target document type and a source material are received, and original text content is generated through the large language model integrating domain knowledge; obtaining a matched structured format template, and analyzing the style rule into a format instruction set; identifying logic elements and hierarchical relationships thereof in the original text based on a natural language processing technology; performing association mapping on the format instruction and the logic element, and performing automatic style rendering by calling a document object model interface to generate an intermediate document with a standard format; and finally, outputting after quality verification. According to the method, an automatic process of content generation and intelligent formatting is constructed, and the document writing efficiency and normalization are improved.
Owner:浪潮智慧科技有限公司 +2

Systems, apparatuses, methods, and non-transitory computer-readable storage media for adaptive information retrieval for question-answering

Methods and systems for retrieving relevant information in response to an input question. The method includes obtaining text content related to the input question and partitioning the content into one or more paragraphs based on predefined rules. The method further involves extracting one or more evidence spans that are relevant to the input question by inputting the text content and the question into a trained language model. A semantic search is then performed on both the paragraphs and the extracted evidence spans, ranking the candidate passages based on their relevance to the input question. Each candidate passage may comprise either a paragraph or an evidence span that addresses the question. The disclosed methods and systems improve the quality and relevance of retrieved information by combining heuristic-based content partitioning with machine learning-based evidence extraction.
Owner:HUAWEI TECH CO LTD

Electronic health record automatic coding method and system

The invention provides an electronic health record automatic coding method and system, and the method comprises the steps: segmenting an electronic health record text into a plurality of text segments, and obtaining a plurality of unstructured text contents; obtaining structured key medical information based on a preset large language model and the plurality of unstructured text contents; generating a preliminary coding result based on a preset coding rule base and the structured key medical information; and obtaining a final coding result based on a preset verification model, a preset auditing model and the preliminary coding result so as to complete automatic coding of the electronic health record. According to the method, the key health information in the non-structural text is coded, and the coding result is verified, so that the accuracy of the coding result is improved.
Owner:THE THIRD AFFILIATED HOSPITAL OF SUN YAT SEN UNIV

Knowledge document-oriented method for structurally and stably presenting large language model output

The invention discloses a knowledge document-oriented method for structurally and stably presenting large language model output. The method comprises the following steps of: preprocessing a knowledge document and converting the knowledge document into a text by a system; the back-end service module adopts a prompt word engineering technology based on a preset JSON mode, constructs a structured prompt containing clear instructions and output format constraints, calls a large language model to carry out analysis and information extraction on text contents, and forces the model to generate structured JSON data following the preset mode; the back-end service provides access to the structured JSON data for the front-end application; and the front end executes templated mapping rendering according to a preset user interface component template isomorphic to the JSON mode, and fills the content of each field of the JSON into a corresponding visual component. According to the method, through the cooperation of the back-end constraint prompt word engineering and the front-end modularized template rendering, the defects that the output content format of a large language model is unstable and is difficult to be directly applied to standardized interface presentation are relieved.
Owner:BOBAN ZHIJIE (BEIJING) TECHNOLOGY CO LTD

Generating probabilistic data structures for lookup tables in computer memory for multi-token searching

Methods, systems, and non-transitory computer readable storage media are disclosed for optimizing computer memory usage for lookup lists in computer memory via probabilistic data structures. For example, the disclosed system generates a probabilistic data structure (e.g., a Bloom filter) to represent data in a lookup list including multi-token items by hashing items of the lookup list to sets of bit values in a bit vector. The disclosed system classifies text content in a digital document by utilizing a maximum number of tokens from multi-token items in the lookup list to select and compare sets of sequential tokens in the digital document to the probabilistic data structure. The disclosed system also iteratively reduces the number of tokens in sets of sequential tokens for subsequent comparisons. Furthermore, in some aspects, the disclosed system causes a computing device to modify a digital document and / or database operations based on the classifications.
Owner:ONETRUST LLC

Text content index automatic identification method based on semantics

The invention discloses a semantic-based text content index automatic identification method, which relates to the technical field of information retrieval, and comprises the following steps: initializing sparse projection and LSH signature, performing iterative optimization by using a Lagrange duality form and gradient update, adjusting hash digits, obtaining a fragment index through a k-d tree, and constructing an inverted index. According to the method, compression is performed through Delta coding, an index map is constructed based on Jaccard similarity, compression is performed through WebGraph, CSNMF is used in combination with Z-Laplacian regularization, a low-rank basis matrix and a low-rank coding matrix are generated, a compressed inverted index is reconstructed after iterative optimization, and reconstructed inverted index entries are generated. According to the method, through multi-resolution hash table initialization, joint feature optimization, local adaptive quantization and low-rank index reconstruction, the semantic expression ability and the compression effect of an index structure are improved, the index precision and efficiency are improved, and intelligent identification of index content is achieved.
Owner:BEIJING GEPU TECHNOLOGY CO LTD

File auditing method, device and equipment based on multi-modal large model

The invention provides a file auditing method, device and equipment based on a multi-modal large model, and the method comprises the steps: splitting data in a to-be-audited file, and obtaining text data in the to-be-audited file and image data in the to-be-audited file; segmenting the file to be audited to obtain a plurality of text paragraphs; based on a pre-constructed cue word template, guiding the multi-modal large model to identify texts in the image data, calling a knowledge base to verify text contents identified in the image data, guiding the multi-modal large model to determine themes of all text paragraphs, and indexing in the knowledge base according to the paragraph themes, and the plurality of text paragraphs are verified based on the index result, so that the automatic auditing process of the to-be-audited file is realized, and the file auditing efficiency and accuracy are improved. The powerful processing capability of the multi-modal large model and the rich information resources of the knowledge base are fully utilized, and powerful technical support is provided for file auditing.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

AIGC content security monitoring system and method based on dynamic reasoning and context awareness

The invention relates to the technical field of natural language processing, in particular to an AIGC content safety monitoring system and method based on dynamic reasoning and context awareness, and the system comprises a semantic graph construction module which is used for extracting entity nouns and predicate verbs in a text according to a received AIGC interaction text flow, generating a semantic concept node set, and sending the semantic concept node set to a database; and performing directed connection and hierarchical nesting on the concepts in the semantic concept node set according to a logic direction according to a subject-predicate-object dependency relationship rule. According to the method, the curvature value of the semantic track is calculated, and the similarity between the direction vector and the center of the sensitive semantic cluster is combined for double verification, so that sudden turning of an intention in a dialogue process or progressive induction to a sensitive field can be perceived, and abnormal mutation can be recognized through a curvature pulse form; therefore, hostile attack behaviors are accurately captured in real time in dynamic interaction, and the defense capability for context dependent attacks and implicit induction behaviors is improved.
Owner:XINGXUAN DIGITAL TECHNOLOGY (SHANGHAI) CO LTD

Heterogeneous document structured data extraction system and method based on multi-modal fusion

The invention discloses a heterogeneous document structured data extraction system and method based on multi-modal fusion, and the method comprises the following steps: S1, receiving a heterogeneous document, and carrying out the preprocessing; s2, extracting visual and semantic modal information based on the visual backbone network and an OCR module, and fusing spatial coordinates, text content and layout features; s3, receiving a dynamic target mode Schema, executing field-level semantic matching, and calculating semantic similarity; s4, performing consistency verification, verifying and filling the field values, and correcting the associated fields; s5, converting the format into a data file in a specified format, and reserving fields to be mapped with an original document semantic block; s6, automatically switching a cloud mode and a local mode according to a deployment environment; and S7, outputting the data file of the target field. According to the method, high-precision and low-delay structured data extraction of the heterogeneous format document can be realized, and the automatic processing efficiency and the data credibility are improved.
Owner:ANHUI HANGTIAN INFORMATION CO LTD

Biding document auditing method based on large language model

ActiveCN120611723ASemantic analysisCommerceLinguistic modelCitation frequency
The invention discloses a bidding document auditing method based on a large language model, belongs to the technical field of text auditing, and solves the problems that after text contents of bidding documents are classified according to risk levels, the continuity of the bidding document contents is broken, and the bidding document auditing efficiency is high. Therefore, the problem that resistance text vulnerabilities cannot be recognized among bidding document texts with different risk levels in the auditing process can occur. Comprising the following steps: constructing a risk factor graph, analyzing a bid invitation file, constructing a structured knowledge network containing hard terms and elastic terms, and defining dependency and conflict relationships between the terms; and dynamically grading bidding document chapter risks, and fusing four-dimensional indexes including commitment density, fuzziness, reference frequency and topology centrality. According to the method, the absolute commitment is extracted from the high-risk chapter to serve as the anchor point, the semantic association fragment of the non-high-risk area is reversely traced, the risk level of the association area is dynamically improved after the constraint effect of the semantic association fragment is verified, and the problem of antagonistic vulnerability identification fault caused by grading auditing is solved.
Owner:SHANDONG DONGHANG INTELLIGENT TECH CO LTD

A Question Answering Method and System for Structured Long Documents

The present invention relates to the technical field of large language models, and specifically provides a question-answering method for structured long documents, which includes the following steps: S1, Parse documents in different formats, and construct structured metadata of the documents according to the parsing results; S2, Divide the documents into multiple text segments, perform vectorization processing on each text segment, and store them in a dedicated vector database; S3, Construct multiple text content acquisition tools respectively for extracting text content from different parts of the documents; Design and implement a vector-based retrieval tool for finding text segments related to the user's query in the vector database; S4, Construct an Agent that includes multiple text content acquisition tools and retrieval tools, and intelligently select text content acquisition tools or retrieval tools for the user's question to obtain relevant text content required for the LLM to answer questions; S5, After obtaining the relevant text content, analyze the relevant text content through the LLM to generate a final answer.
Owner:NO 63921 UNIT OF PLA

Electronic file content anti-counterfeiting and tampering detection method adopting locked signature

The invention discloses an electronic file content anti-counterfeiting and tampering detection method adopting a locked signature, and belongs to the field of information security. According to the method, file content is automatically segmented according to semantics through a context coding self-adaptive segmentation algorithm, and each segment of content fingerprint is generated by applying a content embedding disturbance weighted digest algorithm. And constructing each segmented abstract into a directed relation chain by adopting an embedded relation chain transformation signature mechanism, and carrying out digital signature overall locking. During detection, the segmented abstract and the relation chain structure are compared again, and precise positioning and tracing of tiny tampering, segmented change and structure adjustment of the text content are achieved. And through a signature duration chain entropy analysis algorithm, performing comparison and evolution path analysis on previous version states, and outputting a detailed traceability report. According to the method, authenticity verification and tampering detection of the electronic file content are greatly improved, and the method is suitable for the fields of data anti-counterfeiting, compliance supervision, electronic evidence management and the like.
Owner:CHINA ELECTRONICS STANDARDIZATION INST

Self-media content streaming matching method and system based on dynamic semantic analysis

The invention provides a self-media content streaming matching method and system based on dynamic semantic analysis, and the method comprises the steps: carrying out the dynamic semantic analysis of the text content of a voice stream of a creator, generating a semantic theme sequence, aligning the semantic theme sequence with an emotion fluctuation curve of the voice stream of the creator, and generating an emotion semantic incidence matrix; dividing a mapping relationship between an emotional intensity numerical range in the emotional fluctuation curve and a semantic topic type in the semantic topic sequence; performing behavior association fitting on the mapping relationship based on user historical feedback behavior data, and generating a dynamic association rule between the emotion intensity numerical range and the user feedback behavior; and adjusting a preset self-media content delivery strategy in real time, and generating an adjusted delivery strategy so as to carry out self-media content delivery stream matching. According to the method, the emotion semantic association matrix is constructed, accurate space-time matching of the content theme and the emotion expression is realized, and the target of adaptively optimizing the flow casting matching according to the real-time emotion state of the creator is achieved.
Owner:SHANGHAI YUXING CULTURAL COMMUNICATION CO LTD

Document adaptive conversion method and device based on multi-modal large model and medium

The invention provides a document adaptive conversion method and device based on a multi-modal large model and a medium, and belongs to the technical field of data processing. The method comprises the following steps: inputting a multi-modal document into a pre-constructed multi-modal large model, wherein the multi-modal large model comprises a multi-modal joint framework fusing a ViT visual model and an LLM model; using a ViT visual model and an LLM model to respectively extract multi-scale visual features corresponding to non-text content and semantic features corresponding to text content, and using a cross-modal attention mechanism to bidirectionally align the multi-scale visual features and the semantic features to obtain a multi-modal document; carrying out self-adaptive blocking on the multi-modal document by using a self-adaptive blocking strategy, and modeling a relative position relationship between blocks; establishing a dynamic mapping rule from a document element to an HTML tag, and mapping the multi-modal document into a front-end interaction component in combination with a natural language instruction of a user; and dynamically rendering the front-end interaction component in an on-demand mounting and resource isolation mode. The problem of insufficient document restoration capability in the prior art can be solved.
Owner:SHANDONG INSPUR DIGITAL BUSINESS TECHNOLOGY CO LTD

Document translation method and system, computer equipment and storage medium

The invention provides a document translation method and system, computer equipment and a storage medium, and the method comprises the steps: taking a document associated with steel as a target document, then carrying out the structured fragmentation of the target document, and obtaining a more precise semantic unit; for the content of each fragment, retrieving professional terms from a preset translation term contrast library, constructing prompt words, and performing translation term matching on the text content of each fragment according to the preset translation term contrast library, so that the translation terms of the same text content are the same; inputting the term cues and the original text into a multi-stage large language model, and respectively carrying out literal translation processing, quality evaluation and optimization suggestion, so as to generate interpretation content which is strong in specialty and naturally expresses; and finally, integrating and outputting all the fragmented contents into a structured complete translation document. The translation quality of technical data in the iron and steel industry can be remarkably improved, and the problems that an existing general translation system is not uniform in terms, split in structure, stiff in expression and the like are solved.
Owner:CISDI RES & DEV CO LTD

Generation of Interactive Data Visualizations and Textual Content

A system can be used to generate interactive data visualizations and textual content. The system receives a layout of a content item. The content item can be a data visualization dashboard or an article. The system can generate a content structure tree based on the layout. The content structure tree represents hierarchical and semantic relationships between sections of the content item. The system can receive input adding content elements, such as data visualizations or textual paragraphs, to sections within the content item. The system can identify text roles for different sections based on the content structure tree and the content elements. The system can generate text suggestions for the content item based on the text roles. The system can also generate text content for a text suggestion using a large language model, display the generated text content within the content item, and iteratively refine the content item.
Owner:SALESFORCE INC

Tokenization systems and methods for redaction

A tokenization system receives a request for redaction of sensitive textual content in a document, identifies a portion of the document as the sensitive textual content, and edits the document, including replacing the sensitive textual content thus identified with tokens, each token having a token value and a pattern that identifies a start and an end of the token value. The editing produces a transformed version of the document with the tokens and without the sensitive textual content. The tokenization system may then communicates the transformed version of the document with the tokens and without the sensitive textual content to the client computing system, an automated recognition service, or a redaction plug-in to a frontend application.
Owner:OPEN TEXT CORPORATION

Multimodal duplicate checking method and system for bidding documents based on large model

The present invention belongs to the field of natural language processing and information retrieval technology. It provides a large-scale model-based multimodal bid duplication checking method and system, which performs structured parsing and multimodal feature extraction on input bids to generate feature representations of semantic blocks and non-text elements. It uses a dynamic context-aware large language model to perform deep semantic encoding on text blocks, combining hierarchical position encoding to preserve the logical relevance of long texts. It achieves efficient matching of massive semantic vectors through a hybrid retrieval architecture, and optimizes computing efficiency by combining a distributed computing framework with hardware acceleration instructions. Finally, it reconstructs structured features of non-text content such as tables and charts, achieving cross-modal semantic association analysis.
Owner:INSPUR GENERSOFT CO LTD

Arranging and / or clearing speech-to-text content without a user providing express instructions

Implementations described herein relate to an application and / or automated assistant that can identify arrangement operations to perform for arranging text during speech-to-text operations—without a user having to expressly identify the arrangement operations. In some instances, a user that is dictating a document (e.g., an email, a text message, etc.) can provide a spoken utterance to an application in order to incorporate textual content. However, in some of these instances, certain corresponding arrangements are needed for the textual content in the document. The textual content that is derived from the spoken utterance can be arranged by the application based on an intent, vocalization features, and / or contextual features associated with the spoken utterance and / or a type of the application associated with the document, without the user expressly identifying the corresponding arrangements. In this way, the application can infer content arrangement operations from a spoken utterance that only specifies the textual content.
Owner:GOOGLE LLC