Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1514 results about "Text processing" patented technology

In computing, the term text processing refers to the theory and practice of automating the creation or manipulation of electronic text. Text usually refers to all the alphanumeric characters specified on the keyboard of the person engaging the practice, but in general text means the abstraction layer immediately above the standard character encoding of the target text. The term processing refers to automated (or mechanized) processing, as opposed to the same manipulation done manually.

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Text detail map-based method for supervising end-to-end text detection and recognition

A text detail map-based method for supervising end-to-end text detection and recognition, pertaining to the field of text processing. The method comprises the following steps: given an input image containing text in any shape, processing the input image by means of two separate processing branches; designing a text attention head (TAH), and designing a feature pyramid enhancement fusion module (FPEFM); the FPEFM performing feature self-enhancement at different sizes, fusing text image local features and global text position information extracted by a TAH module, and fusing features extracted by the TAH from feature maps of different sizes; stacking a plurality of FPEFMs to continuously enhance the feature representation capability of a model and the depth of the model; and sampling the feature maps to a unified size to obtain a final enhanced feature map.
Owner:CHONGQING UNIV OF TECH

Speech recognition and natural language processing integration method and system

The invention discloses a speech recognition and natural language processing integration method and system, and the method comprises the steps: carrying out the multi-modal data fusion according to a speech signal of a user, context text information and environment sensor data, and obtaining a fused multi-modal feature vector; inputting the multi-modal feature vector into a speech recognition model based on an adaptive deep neural network, and performing speech-to-text processing to obtain text output; inputting the text output into a semantic analysis model based on a graph neural network, and performing context semantic analysis and user intention recognition to obtain semantic representation of the user intention and a confidence score of the semantic representation; and according to the semantic representation and the confidence score thereof, dynamically adjusting parameters of the speech recognition model and the semantic analysis model by using a feedback optimization technology based on reinforcement learning, and generating a model optimization strategy. According to the embodiment of the invention, the accuracy of speech recognition and the semantic comprehension capability of natural language processing can be improved.
Owner:GUANGZHOU JIUSI INTELLIGENT TECH CO LTD

Text processing method and device based on hybrid expert model, equipment and medium

The invention relates to the artificial intelligence technology, can be applied to service system platforms such as medical health and financial science and technology, and discloses a data processing method, device, equipment and medium based on a hybrid expert model.The method comprises the steps that to-be-processed data is input into the hybrid expert model, routing probability distribution is obtained, the Tsallis entropy of the routing probability distribution is calculated, and the Tsallis entropy of the routing probability distribution is calculated; generating a prediction result according to the Tsallis entropy; evaluating loss based on an auxiliary entropy loss function, constructing a subspace, and adjusting pre-training parameters in the subspace to obtain an optimized hybrid expert model; and obtaining a target prediction result according to the optimized hybrid expert model. According to the dynamic routing mechanism, experts are flexibly selected, an entropy loss function is assisted to optimize a routing decision, uncertainty is reduced, and model convergence is accelerated; and interference of new tasks on old tasks is eliminated through re-parameterization and subspace design, so that the model achieves good balance between stability and plasticity, and generalization and adaptability of model data processing are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Enterprise project text processing model directional training method and device

The embodiment of the invention provides a directional training method and device for an enterprise project text processing model, and the method and device achieve the screening of high-quality training samples through the construction of a project text scoring model and the multi-dimensional evaluation of the text coverage degree, the technical correlation degree and the structural integrity. And analyzing a text structure relationship based on the knowledge graph and the graph convolutional network, and constructing a differentiated text template library. The system performs directional training on a pre-training large model in combination with an enterprise portrait vector, and continuously optimizes model parameters and a knowledge base by adopting an incremental learning strategy. According to the method, the defects of a traditional technology in the aspects of training sample screening, knowledge structure analysis and model personalized customization are effectively overcome, and the practicability and adaptability of a project text processing model are remarkably improved.
Owner:ZHEJIANG WANCHUANG HUILI TECHNOLOGY SERVICE CO LTD

Community governance decision generation method and device based on knowledge graph

The embodiment of the invention discloses a community governance decision generation method and device based on a knowledge graph. The method comprises the following steps: collecting community governance data; performing information extraction and text processing on the community governance data by adopting a data extraction model to generate structured data; constructing a causal analysis atlas, a decision treatment atlas and a space-time correlation atlas by using the structured data; fusing the causal analysis graph, the decision governance graph and the space-time association graph to obtain a community governance knowledge graph, wherein the community governance knowledge graph supports semantic-based query and reasoning; acquiring resident appeal information, and performing risk reasoning based on a preset reasoning rule to obtain a risk reasoning result; and based on a risk reasoning result, in combination with spatio-temporal information defined by a spatio-temporal association graph, querying in the community governance knowledge graph, and generating a community governance scheme according to a query result. According to the invention, the dynamism, the accuracy and the intelligent level of community management can be improved.
Owner:BEIJING UNION UNIVERSITY +1

Engineering management report generation method and system based on natural language processing

The invention relates to the technical field of text processing, and provides an engineering management report generation method and system based on natural language processing, which are used for realizing deep analysis of construction logs and providing a comprehensive and accurate decision basis for engineering management. The method comprises the steps that a construction log text set generated in the engineering construction process is acquired, natural language processing is conducted on the construction log text set, key entities and progress events in each construction stage description unit are extracted, and a construction feature set containing entity identification information and event time sequence information is obtained; performing semantic alignment processing on the construction feature set and a preset engineering specification knowledge base to generate specification association information between the construction stage description unit and an engineering specification requirement; and performing structured integration processing on the standard associated information based on a hierarchical template matching algorithm to generate a structured report text meeting the engineering management requirements.
Owner:CMA METEOROLOGICAL OBSERVATION CENT

Knowledge base question and answer platform construction method based on large language model

The invention relates to the technical field of natural language processing, and discloses a knowledge base question and answer platform construction method based on a large language model, which comprises a knowledge acquisition module, a data preprocessing module, a text processing module, a vectorization module, a question understanding module, a mixed retrieval module, a prompt generation module, an answer generation module and an answer quality analysis module. A secondary inquiry processing module and a feedback learning module; according to the method, a semantic segmentation algorithm is combined with semantic retrieval and keyword retrieval, so that the flexibility is high; normalized prompts are constructed, input is performed according to correlation sorting, and the accuracy of answers is improved; multi-dimensional confidence evaluation is introduced, strict multi-layer security and compliance filtering is set, and the reliability of the system is ensured; the relevance of multiple rounds of dialogues is judged and complemented, so that interaction is more natural and efficient; knowledge is collected and updated in real time, a knowledge base and a retrieval strategy are continuously optimized, and a closed loop of data-application-feedback-tracing-optimization is formed.
Owner:JIANGSU INSPIRE INTERNET OF THINGS TECH CO LTD +1

Multi-algorithm collaborative intelligent PDF (Portable Document Format) document analysis system

The invention belongs to the technical field of document analysis, and particularly relates to a multi-algorithm collaborative PDF document intelligent analysis system which comprises a preprocessing module used for recognizing the type of a PDF document, performing layout correction on a scanning document PDF and uniformly adjusting the scanning document PDF into a vertical storage format; the layout analysis module is used for identifying page element categories based on a target detection algorithm, and processing element superposition and separation problems through a merging de-duplication or priority discarding strategy; and the text processing module is used for analyzing all places with characters in the PDF based on the OCR, outputting the characters and corresponding textbox coordinates, and dividing the contents of the text blocks in combination with layout analysis. The system can extract characters of the PDF of the scanned copy, distinguish element types such as titles, texts, tables and formulas, solve the problem of confusion of cross-page tables, column texts and characters with similar shapes, and can completely reserve the content structure of the document.
Owner:BEIJING HUAYUN WORLD TECH CO LTD

Method and system for retrieving DOCX document content based on keywords

The invention belongs to the technical field of text processing, and particularly relates to a method and system for retrieving DOCX document content based on keywords, which comprises the following steps: analyzing an Office Open XML structure of a DOCX document, combining with multi-dimensional features such as style names, and utilizing a title classification score model to accurately distinguish a title and a text, so that a semantic hierarchical structure of the document is effectively reserved; and secondly, a multi-level semantic extension mechanism is introduced, and a Sension-BERT, a HowNet knowledge base and a Word2Vec model are fused, so that intelligent extension of synonyms and synonyms of keywords is realized, and the recall rate and semantic understanding ability of retrieval are remarkably improved. And in addition, a BM25 model is combined with paragraph length normalization and structure position weight to calculate a correlation score, so that retrieval results are sorted more accurately and reasonably. The construction of the reverse index is combined with the position coding and compression optimization strategy, and the retrieval efficiency and the storage performance are both considered.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Method for improving long text processing efficiency and accuracy

The invention discloses a method for improving long text processing efficiency and accuracy, and relates to the technical field of natural language processing and large language models.According to the method, text word segmentation embedding, sliding block preprocessing, YaRN position code injection, dynamic sparse attention calculation, multi-level attention fusion, graded KV cache management and output generation are sequentially executed; position drift is inhibited through logarithmic scaling, and key contexts are adaptively screened according to the attention activeness, so that the attention calculation complexity is close to linearity; in million-level Token reasoning, the video memory occupation of the method is reduced, the remote dependency recall rate is improved, and the method is suitable for scenes such as document analysis, code auditing and multi-mode streaming understanding.
Owner:BEI JING JING YUE KE JI YOU XIAN GONG SI

Chinese language and literature computer networking query reading system

The invention relates to a Chinese language and literature computer networking query reading system. According to the system, a query request processing module carries out mixed feature analysis processing according to a query request of a user, and a mixed query vector fusing radicals and semantics is generated; a distributed retrieval module performs distributed retrieval in a preset space-time fragment index structure based on the mixed query vector to obtain a candidate literature data set; the preliminary retrieval optimization module is used for carrying out self-adaptive conversion processing on abnormal characters and traditional and simple characters in the candidate literature data set according to a Chinese character evolution rule and a context to obtain standardized text data; and the visual query output module performs knowledge graph enhancement processing on the standardized text data to generate a visual query result. According to the system, efficient query and visual display of Chinese language and literature works are achieved through multi-modal processing, multi-dimensional indexing and knowledge graph technologies, and the problems that a traditional query mode is single in input, limited in retrieval dimension, insufficient in text processing and the like are effectively solved.
Owner:FUJIAN AGRI VOCATIONAL & TECH COLLEGE

Text processing method and system based on large model, terminal and storage medium

The invention belongs to the technical field of text processing, and particularly relates to a text processing method and system based on a large model, a terminal and a storage medium, and the method comprises the steps: preprocessing an input text, inputting the preprocessed text into a pre-established first large model, and outputting an error list or an error-free prompt; when the first large model outputs an error list, the output error list is detected, and the preprocessed text is corrected according to the detected error list; and inputting the corrected text or the text corresponding to the error-free prompt into a pre-established second large language model, and outputting a sensitive word prompt or a sensitive word-free prompt. According to the method, the sensitive word problem possibly introduced or triggered by the error correction operation is actively recognized and processed through the serial flow of error correction and sensitive word analysis. And the second model performs sensitive word analysis on the basis of the error-corrected text, so that the real-time accuracy and safety of sensitive word recognition are remarkably improved.
Owner:山东浪潮智能生产技术有限公司

Prompt word attack detection method and system of large language model and electronic equipment

The invention provides a cue word attack detection method and system of a large language model and electronic equipment, and relates to the technical field of text process.The method comprises the steps that a to-be-processed text block in an original text is determined; performing heuristic semantic analysis on the text block to obtain a first suspicious text block and a first risk score thereof; performing semantic analysis on the first suspicious text block and the context information thereof through a first language model to obtain a second suspicious text block and a second risk score thereof; performing structured verification of the attack intention on the second suspicious text block according to a preset meta prompt word through a second language model to obtain an attack intention verification result; the attack intention verification result comprises a third suspicious text block and a third risk score thereof; and executing a target response action on the original text based on the first risk score, the second risk score and the third risk score. According to the method, the efficient processing requirement of the long text can be considered, and the detection accuracy of complex attacks can be improved.
Owner:HANG ZHOU LING XIN SHU KE XIN XI JI SHU YOU XIAN GONG SI

Bidding document information extraction method

The invention relates to the field of text processing, in particular to a bidding document information extraction method. Comprising the following steps: segmenting a bidding and tendering file into pages, and identifying the pages to obtain corresponding texts; generating complementary text description for images and tables in the page and adding the complementary text description to the tail of a text corresponding to the page to form an enhanced text block sequence; matching a label from the text block sequence according to a pre-constructed hierarchical label system, and generating a corresponding cue word template according to the label and a pre-constructed cue word template library; inputting the cue word template, the enhanced text block sequence and the context text abstract as a combination into a large language model to obtain a structured extraction result with a hierarchical relationship; and matching the extracted entity content with a local dictionary, carrying out aggregation arrangement on a result after the matching is passed, and outputting a structured data file. On the premise that the model does not need to be retrained, the illusion risk of the generated content is reduced.
Owner:SHANGHAI MECHANICAL & ELECTRICAL EQUIP TENDERING CO LTD

HTML (Hypertext Markup Language) information extraction method, device and equipment based on multi-LoRA cascade strategy and medium

The invention discloses an HTML (Hypertext Markup Language) information extraction method and device based on a multi-LoRA cascade strategy, equipment and a medium, and relates to the technical field of HTML information extraction. The information extraction method comprises the following steps: acquiring a document and inputting a large language model; and the large language model judges whether a table is included. And if the table is included, calling a table processing LoRA adapter to extract table content and convert the table content into pseudo natural language description, calling a text processing logic module to extract adjacent text contexts of the table, and then performing semantic integration to obtain first text information. And if the table is not included, calling a text processing module of the table processing LoRA adapter to extract the text content, and obtaining the first text information. And calling a key information extraction LoRA adapter, and extracting the structured key value pair from the first text information to generate JSON data. And calling a nested structured generation module to convert the JSON data into a target sequence in a multi-layer nested JSON format.
Owner:XIAMEN UNIV OF TECH

Computer-aided system for multidimensional generative value assessment and applicant selection

ActiveDE202025107568U1InstrumentsData packData stream
A computer-implemented system for multidimensional generative value assessment and applicant selection, consisting of: a data collection unit configured to electronically receive applicant data consisting of structured academic records, work experience records, digital documentation, and unstructured narrative responses generated from generative self-assessment instruments and contextual interviews; a feature extraction unit coupled to the data acquisition unit, configured to apply computer-assisted text processing, semantic analysis, and token-level attribute identification to transform narrative responses and structured data into multidimensional feature vectors that represent generative indicators of innovation, mentoring, collaborative performance, resilience, social contribution, ethical consistency, and predicted institutional impact; a weighting calculation unit configured to assign weight values ​​to the extracted feature vectors based on a digital generative profile definition matrix that includes dimensions, sub-criteria, indicators, documentation requirements and importance coefficients, with the weighting being distributed across the generative dimensions defined in the digital matrix and configurable according to the institutional context; a quantitative rating unit configured to calculate a generative rating score by aggregating weighted feature vectors derived from self-assessment inputs, interview-based ratings, document analyses, and authenticity predictions, with the aggregation including normalization, nonlinearity correction, conflict handling, and artifact frequency balancing to obtain a consolidated score; a proof verification unit configured to electronically validate referenced digital evidence by performing content extraction, metadata verification, pattern matching, and cross-document correlation to determine authenticity, credibility, and contextual relevance with respect to the calculated feature vectors; a classification determination unit configured to assign a classification level to an applicant by comparing the generative assessment score with a set of system-defined calculation thresholds, including at least a lower threshold, a middle threshold and an upper threshold, the classification levels representing different generative maturity states and determining subsequent eligibility for selection; a decision generation unit configured to produce a digital output data set that includes classification level, feature aggregation summaries, evidence validation results, and recommended organizational actions, wherein the decision generation unit encodes the data set in a digitally signed, tamper-proof format and stores it on a non-volatile storage medium; and A system control unit acts as an operational interface to all other units and is configured to orchestrate data flow, scheduling, process state transitions, and event logging to ensure verifiable traceability, consistency, and auditability of the evaluation and selection processes.
Owner:BERNARDO OHIGGINS UNIVERSITY +3

Image description generation method and system based on enhanced fine-grained information

The invention discloses an image description generation method and system based on enhanced fine-grained information. The method comprises the following steps: 1) extracting regional features and global features of an image, and carrying out weighted fusion to generate multi-view feature information; 2) combining text feature extraction and cosine similarity screening, dynamically fusing text-image features, and optimizing cross-modal features through a dynamic multi-view enhancement mechanism; (3) a hierarchical memory enhancement encoder is adopted, multi-head attention and Transform structure hierarchical coding features are combined, and detail information is reserved through a memory unit; and 4) aggregating image and text weights in two stages through a double cross-modal attention decoder to generate high-quality description. The system comprises a feature extraction module, a text processing module, an encoding module and a decoding module. Through fine-grained feature enhancement and cross-modal interaction, the problems of detail loss and semantic deviation in a traditional method are effectively solved, and the method is suitable for image semantic analysis of complex scenes.
Owner:NANJING UNIV OF SCI & TECH

Image-text processing method and device for marking compression framework

The invention discloses an image-text processing method and device for marking a compression framework. The method comprises the following steps of: extracting visual features; a visual mark screening processing step; a text feature extraction step; and a multi-modal fusion and model processing step. The method has the beneficial effects that the inference efficiency of the MLLMs is remarkably improved under the condition that the visual mark compression frame does not need additional training; through global and local information fusion of the DVTS module and text guide supplement of the TGVC module, the number of visual marks is greatly reduced, meanwhile, key visual information is reserved, and visual-text alignment is enhanced; experiments show that in various image and video benchmark tests, compared with an existing method, the framework has the advantages that the calculation cost is greatly reduced, the model performance is maintained and even improved, and the framework has remarkable technical advantages and application potential.
Owner:ZHEJIANG YOULU ROBOT TECH CO LTD

Retrieval system for bidding law document large language model

The invention relates to the field of text processing, in particular to a retrieval system for a bidding law document large language model. Comprising a document preprocessing module, a semantic dicing module, an entity and relation recognition module, a knowledge graph construction module, a multi-strategy mixed retrieval module and a retrieval result preferential module. During working, the bidding law document is preprocessed, based on document word number judgment, a chapter-level slicing strategy or a clause-level slicing strategy is adopted for cutting, a LateChunking algorithm is used for determining cutting points, semantic units are extracted, entity elements are recognized, a knowledge graph is constructed, and an optimal retrieval result is obtained through multi-strategy mixed retrieval and preferential processing. According to the method, the limitation of a traditional single retrieval mode in bidding legal chief document processing is overcome, and the provision retrieval accuracy and context coherence are improved, so that the illusion risk caused by information missing or misunderstanding of a large language model is greatly reduced, and the credibility and practicability of a legal intelligent question and answer result are enhanced.
Owner:SHANGHAI MECHANICAL & ELECTRICAL EQUIP TENDERING CO LTD

Task processing method, text processing method, automatic question-answering method, task processing model training method, information processing method based on task processing model, and cloud training platform

PCT designated stageWO2025196533A1Digital data information retrievalSemantic analysisInformation processingDuplicate content
Provided in the embodiments of the present disclosure are a task processing method, a text processing method, an automatic question-answering method, a task processing model training method, an information processing method based on a task processing model, and a cloud training platform, which are applied to the technical field of computers. The task processing method comprises: acquiring task data of a target task; and inputting the task data into a task processing model, so as to obtain a task processing result of the target task, wherein the task processing model is obtained by means of performing training on the basis of a plurality of pieces of target sample data and target sample results of the plurality of pieces of target sample data, the target sample data is obtained by means of performing screening on the basis of detection results of repeated content of a plurality of pieces of sample data, and the content of the target sample results is not repeated. Target sample data is obtained by means of performing screening on the basis of detection results of repeated content, and the content of target sample results is not repeated, such that the repetition illusion of a model is reduced without damaging the processing capability of the model itself and relying on external knowledge, thereby improving the accuracy and integrity of a task processing result.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Text processing method and apparatus based on artificial intelligence, and electronic device, computer program product and computer-readable storage medium

Provided in the present application are a text processing method and apparatus based on artificial intelligence, and an electronic device, a computer program product and a computer-readable storage medium. The present application can be applied to the field of artificial intelligence and the field of large models. The method comprises: acquiring first integrated text, wherein the first integrated text is obtained by means of correcting first original text; splicing the first integrated text and the first original text to obtain spliced text; performing multi-dimensional evaluation processing on the spliced text to obtain an evaluation result corresponding to each dimension, wherein the multi-dimensional evaluation processing comprises at least two of the following: semantic evaluation processing, grammar evaluation processing and typesetting evaluation processing; and fusing evaluation results of at least two dimensions to obtain a correction evaluation result of the first integrated text.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Internet live broadcast illegal behavior monitoring method and system based on image and voice recognition

The invention discloses a method and a system for monitoring internet live broadcast illegal behaviors based on image and voice recognition, and relates to the technical field of live broadcast monitoring. The method comprises the following steps: acquiring live stream data of each anchor and verifying data integrity; processing the live broadcast stream data to obtain preprocessed live broadcast stream data; performing feature recognition on the video frame set by using the determined image violation processing model, and outputting image violation marking information; performing text processing on the voice-to-text data by using the determined text processing model, and outputting text violation marking information; based on the image violation mark information and the text violation mark information, determining whether the anchor violates the rule or not, and giving an alarm when the anchor violates the rule; and the data acquisition parameters are corrected based on the historical violation data of the anchor, so that the problem that it is difficult to accurately acquire and process the live broadcast data to identify Internet live broadcast violation behaviors and optimize the monitoring parameters according to the historical violation data by using image and voice recognition technologies at the same time in the prior art is solved.
Owner:BEIJING FEIYING INTERACTIVE TECHNOLOGY CO LTD

Water conservancy knowledge structured extraction and verification method and device

The invention provides a water conservancy knowledge structured extraction and verification method and device, and belongs to the technical field of artificial intelligence, and the method comprises the steps: carrying out the differential text processing of different formats of files, and generating an intermediate file; classifying the intermediate file into a regulation class or a non-regulation class based on a preset rule base; performing hierarchical title identification on the regulatory files to form entry knowledge blocks, and converting table contents into HTML (Hypertext Markup Language) knowledge blocks; performing semantic segmentation on the non-regulation file to generate knowledge blocks; performing knowledge block checking and filing, and marking an abnormal alarm block; converting the table knowledge blocks into natural language description by utilizing a large model; and positioning the context of the original text of the alarm knowledge block, and performing intelligent correction through a large model. According to the method, a traditional semantic analysis model and a large language model are creatively fused, a closed-loop process of preprocessing, extraction, verification and correction is formed, the problems of structured analysis and error correction of complex texts in the water conservancy field are solved, and the knowledge processing efficiency and accuracy are remarkably improved.
Owner:长江水利委员会网络与信息中心

Fault-tolerant processing method and system for super-long text in large model service, and storage medium

The invention relates to the technical field of large language model application, in particular to a fault-tolerant processing method and system for a super-long text in large model service and a storage medium, and the method comprises the following steps: obtaining a super-long text to be processed and a maximum Token threshold of a large model context window, completing Token processing and judging whether the super-long text exceeds the limit or not; starting a multi-level alternative scheme including rapid compression, dynamic compression and chat history compression; starting an error recovery mechanism including hierarchical exception processing, intelligent degradation and parameter verification; outputting the target text of which the Token number is compliant; the system comprises a text acquisition module, a multi-level alternative scheme execution module, an error recovery module and an output module. The storage medium stores a computer program, realizes the method during execution, can adapt to multiple models, and meets industrial-grade super-long text processing requirements; the problems of Token overrun errors and inference service instability caused by lack of fault-tolerant mechanisms and insufficient boundary processing in the prior art are solved.
Owner:POWERCHINA BEIJING ENG CORP

Text processing method and system based on semantic density

The invention provides a text processing method and system based on semantic density, which are applied to the technical field of text processing, and the method comprises the following steps: obtaining a target text, and carrying out feature extraction on the target text to obtain a plurality of text features of each paragraph of the target text; calculating a semantic density score of each paragraph according to the text feature of each paragraph; classifying each paragraph according to the semantic density score of each paragraph to obtain a paragraph category of each paragraph; if the paragraph category of the paragraph is high in density, performing block processing on the paragraph according to the text type of the target text and the semantic density score of the paragraph to obtain a plurality of text blocks corresponding to the paragraph; and if the paragraph category of the paragraph is low density, processing the paragraph according to the text type of the target text, the paragraph and the semantic density score thereof, and each other paragraph and the semantic density score thereof, so that information splitting and context loss can be avoided, redundant calculation and resource waste are reduced, and the field adaptability is improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Text processing method and device, electronic equipment, storage medium and program product

The invention provides a text processing method and device, electronic equipment, a storage medium and a program product, and relates to the technical fields of artificial intelligence, natural language processing, large language models, automatic driving, intelligent traffic and the like. The method comprises the following steps: determining an initial voice recognition text and phoneme data of a target voice, and generating target prompt information based on the initial voice recognition text and the phoneme data; obtaining a target recognition text through a large language model based on the target prompt information; the large language model can correct the initial speech recognition text based on the phoneme data; as the accuracy of the phoneme data is far greater than the accuracy of the initial speech recognition text, more and more effective information can be provided for reference for the processing process of the large language model by combining the phoneme data, the large language model is helped to obtain a correct text in the processing process, and then the accuracy of text processing is improved.
Owner:GUANGZHOU TENCENT TECH CO LTD

Precise alignment method and system for multilingual terminologies in nuclear power field

The invention relates to the technical field of text processing, and discloses an accurate alignment method and system for multilingual terminologies in the nuclear power field, and the method comprises the steps: collecting original terminologies from a nuclear power design document, an operation manual and an international standard database, and constructing an original terminology set in the nuclear power field; encoding the original term set after term cleaning to obtain standardized term blocks; extracting semantic feature vectors according to the standardized term chunks; based on the attention scores of the semantic feature vectors in the sequence, performing weighted fusion on the semantic feature vectors to obtain enhanced semantic representation; constructing candidate term pairs in the nuclear power field according to the similarity of the term pairs in the enhanced semantic representation; performing consistency check on the candidate term pairs through entity relationships and attributes of the knowledge graph, and outputting the candidate term pairs passing the consistency check as target term pairs in the nuclear power field; according to the invention, the efficiency of accurate alignment of multilingual terminologies in the nuclear power field can be improved.
Owner:JIANGSU NUCLEAR POWER CORP

Video generation method and device based on multi-agent cooperation

The invention relates to the technical field of artificial intelligence content generation, and provides a video generation method and device based on multi-agent cooperation. The method comprises the following steps: selecting a proper copywriting processing agent according to text information input by a user, so as to use the copywriting processing agent to expand or condense a plurality of sub-mirror descriptions; selecting a proper image generation agent; inputting the picture description, the picture style and the mirror operation method in the mirror splitting description into the selected image generation agent to generate a mirror splitting picture; selecting a proper video generation agent; inputting the split picture and the character action and the mirror moving method in the corresponding split description into a selected video generation agent to generate a split video; and selecting a proper automatic editing agent, and performing post-production on the split video to generate a finished video. The technical problems that in the prior art, the video generation process is fragmented, and multi-element collaboration is difficult are solved, and automatic production of complex video content is achieved.
Owner:WUHAN SINAN YIYI INTELLIGENT TECHNOLOGY CO LTD

Text processing method, text processing apparatus, electronic device, and computer-readable storage medium

A text processing method includes obtaining a query text, invoking a search engine interface based on the query text to obtain a plurality of text search results corresponding to the query text, obtaining, from the plurality of text search results, a plurality of answer text segments matching the query text, determining a relevance between the query text and each of the plurality of answer text segments, determining one of the plurality of answer text segments that corresponds to a maximum relevance as a reference text of the query text, and invoking a language model based on the query text and the reference text to obtain a reply text of the query text.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD