Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

603 results about "Regular expression" patented technology

A regular expression, regex or regexp (sometimes called a rational expression) is a sequence of characters that define a search pattern. Usually such patterns are used by string searching algorithms for "find" or "find and replace" operations on strings, or for input validation. It is a technique developed in theoretical computer science and formal language theory.

Data asset classification and dynamic authority management system

The invention relates to the technical field of data security, in particular to a data asset classification and grading and authority dynamic management system, which comprises an asset parameter acquisition module, a classification feature judgment module, a grading weight calculation module and an authority dynamic adaptation module. According to the method, metadata and path parameters are automatically collected through a distributed crawler in combination with a regular expression, naming parameters and content feature parameters are separated to generate a structured parameter set, field naming similarity is quantized through a Levenshtein distance, content character distribution features are verified through chi-square verification, and a naming similarity and distribution feature dual verification mechanism is constructed. According to the method, sensitivity, access frequency and data volume weight are dynamically distributed through an entropy weight method, grading parameters are generated through linear superposition, an RBAC model is combined with a Dijkstra algorithm to verify access path topology legality, unauthorized access is blocked, and the heterogeneous data classification grading and authority control dynamic adaptive capacity is improved.
Owner:国义招标股份有限公司

Government affair data automatic classification and grading method

The invention discloses a government affair data automatic classification and grading method, and relates to the technical field of data management and information processing. According to the method, data of different sources and formats are converted into structured formats by adopting a regular expression and a pattern matching technology, and the consistency and accuracy of the data are improved by means of a machine learning algorithm, data quality detection, anomaly correction and the like; by constructing a government domain ontology and a knowledge graph and combining a natural language processing technology, semantic classification is effectively carried out on government affair data, overlapping and uncertainty between categories are solved, robustness and accuracy of classification decision are further improved through an integrated learning algorithm, high efficiency and reliability of classification are ensured, and the method is suitable for large-scale popularization and application. By establishing a data grading dynamic adjustment and update mechanism and combining real-time data monitoring, sensitivity evaluation and automatic adjustment of access control rules, the flexibility and security of data management are ensured.
Owner:SCI CITY (GUANGZHOU) INFORMATION TECH GRP CO LTD

Artificially intelligent systems and methods for financial coaching

Artificially intelligent systems and methods for financial coaching provide personalized, fiduciary-compliant financial guidance through advanced machine learning architectures with measurable performance criteria. The systems implement privacy-preserving processing pipelines that detect personally identifiable information using multi-layered pattern recognition including regular expressions for formatted data sequences, named entity recognition with confidence thresholds above 0.85, and contextual analysis algorithms. A multi-step artificial intelligence processing workflow includes automated language detection, emotional tone classification with confidence scoring, financial profile transformation using predefined templates, context-aware question rephrasing, and semantic similarity matching employing vector embeddings with financial domain vocabulary weighting applying multiplier values between 1.3-2.0. Specialized training methodologies expand datasets through mathematical transformation functions utilizing statistical standard deviations with incremental variations between 0.5-2.0. Mood-based escalation logic automatically transfers users to human advisors when emotional indicators exceed confidence thresholds above 0.8. The systems maintain response times below 5 seconds while providing regulatory compliance through curated content sources and predefined fiduciary instruction parameters.
Owner:BRIGHTPLAN LLC

Power knowledge question and answer matching method and system based on improved retrieval enhancement generation technology

The invention relates to the field of power question answering, in particular to a power knowledge question-answer matching method and system based on an improved retrieval enhancement generation technology, and the method comprises the steps: employing a training set to carry out the decomposition of a pre-constructed Embedding model, and carrying out the training of the Embedding model, and obtaining an Embedding fine tuning model; the method comprises the following steps: vectorizing a power field file by adopting an Embedding fine tuning model to obtain a power knowledge base, and constructing a regular expression rule base; the user consultation content is converted into a query vector, the query vector is matched with the regular expression rule base to obtain an accurate matching result, and meanwhile the query vector is matched with the power knowledge base to obtain a fuzzy matching result; semantic correlation sorting is carried out on the accurate matching result and the fuzzy matching result, and then resorting is carried out by adopting a resorter in combination with a service rule; and inputting the reordering result into a large language model to obtain a result corresponding to the user consultation content. The method can effectively solve the problem of inaccuracy during retrieval of a specific ID, and improves the retrieval precision of a retrieval enhancement generation system.
Owner:STATE GRID INFORMATION & TELECOMM BRANCH

Financial sensitive data desensitization method and device, equipment and storage medium

The invention discloses a financial sensitive data desensitization method and device, equipment and a storage medium, and relates to the technical field of data security, and the method comprises the steps: recognizing a sensitive data segment of to-be-desensitized financial data in an interface request, obtaining a sensitive data type and a target specific character position, and storing the sensitive data type and the target specific character position; determining to-be-encrypted data in the to-be-desensitized financial data by using a preset regular expression and based on the target specific character position; extracting a first preset number of bits of data from dynamic identification information corresponding to the to-be-desensitized financial data to obtain a first type of characters, and intercepting a second preset number of bits of data at the tail of the to-be-desensitized financial data to obtain a second type of characters; and combining and encoding the first type of characters and the second type of characters to obtain an initial character sequence, encrypting the to-be-encrypted data by using a target converted character sequence obtained by carrying out system conversion on the initial character sequence to obtain encrypted data, and desensitizing the to-be-desensitized financial data based on the encrypted data to obtain desensitized financial data. Therefore, the data desensitization efficiency can be improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Real-time natural language processing and fulfillment

A system and method of real-time feedback confirmation to solicit a virtual assistant response from an evolving semantic state of at least a portion of an utterance. A user accesses a virtual assistant on an electronic device having the system and / or method configured to capture a command, a question, and / or a fulfillment request from audio such as, the speech emitted from the speaking user. The speech may be intercepted by a speech engine configured to transcribe the speech into text that is matched with the fragment pattern's regular expression to generate a fragment and / or the speech may be processed with a machine learning model to identify fragments. The fragments are identified by a domain handler configured to update a data structure of the current semantic state of the utterance in real-time on an interface of an electronic device.
Owner:SOUNDHOUND AI IP LLC

Component identifier matching method and system based on NLP semantic segmentation and multi-level word bank

The invention discloses a component identifier matching method and system based on NLP semantic segmentation and a multi-level word library, and relates to the field of constructional engineering informatization, and the method comprises the steps: building a standard main word library and a preset compensation word library, carrying out the semantic segmentation of an input component identifier, and obtaining a segmented lexical element sequence, calculating semantic similarity and morphological similarity between the segmented lexical element sequence and entries in a standard main word bank, and performing standard term mapping based on the semantic similarity and the morphological similarity to obtain a standard term mapping result; carrying out structured data conversion on the segmented lexical element sequence which does not complete the standard term mapping by utilizing a regular expression mode to obtain structured data, and carrying out AI extension library matching on the segmented lexical element sequence which does not complete the structured data conversion to obtain a synonym matching result; and combining the standard term mapping result, the structured data and the synonym matching result to obtain component identification description. By implementing the method, the matching accuracy of the component identifier names in the engineering project can be improved.
Owner:CHINA CONSTR THIRD ENG BUREAU GRP CO LTD

Artificial intelligence assistant for network services and management

An artificial intelligence assistant analyzes network observations to generate regular expressions using an artificial intelligence model for network automation management pipelines that identify and resolve network issues. A method includes obtaining at least one instruction and information about a plurality of enterprise network assets and configuration of an enterprise network that includes the plurality of enterprise network assets and generating at least one regular expression using an artificial intelligence model based on context description of the at least one instruction and the information about the plurality of enterprise network assets and the configuration of the enterprise network. The method further includes generating at least one solution for configuring at least one network asset of the plurality of enterprise network assets based on the at least one regular expression and providing the at least one solution to cause a configuration change in the at least one network asset.
Owner:CISCO TECHNOLOGY INC

Layout check and automatic correction system based on rules

The invention provides a rule-based layout check and automatic correction system, which relates to the technical field of integrated circuits, and comprises a rule input and management module, a layout data processing module, a rule check engine, a problem analysis and visualization module, an automatic correction module and a user interaction and control interface. Through a rule template library and a complex rule analysis mechanism, an efficient rule configuration mode is provided for a user, and the user can quickly adjust parameters based on a preset template to generate a user-defined design rule, or directly upload complex rule description containing a regular expression and a mathematical formula; and the system analyzes the data into executable logic units and stores the executable logic units in a rule database. According to the mode, the complexity and the time cost of design rule configuration are remarkably reduced, so that the system can flexibly adapt to different process nodes and design requirements, the universality and the efficiency of layout inspection are improved, and convenient and diversified rule management experience is provided for designers.
Owner:上海芯无双仿真科技有限公司

Digital intelligent carrier pigeon WeChat intelligent dialogue generation method based on deep learning

The invention belongs to the technical field of dialogue generation, and relates to a digital intelligent carrier pigeon WeChat intelligent dialogue generation method based on deep learning. Dialogue data are obtained, and semantic anomaly detection is carried out through a regular expression and a BERT model in combination with a public corpus; organizing dialogue data in a mode of constructing a tree structure and introducing a three-dimensional timestamp; for emotion information of emoticons, text and visual features of expressions are combined through a hybrid encoder, a dynamic attention mechanism is adopted, weighting and memory length adjustment are carried out according to importance of historical statements, after user portrait features are introduced, user features and dialogue codes are fused, and therefore, the emotion information of the expressions is obtained. In the generation stage, through a Transform-based model and a Beam Search algorithm, a diversity penalty term is introduced in the generation process, and generation parameters are dynamically adjusted according to a user portrait, so that the diversity and security of generated dialogues are ensured, the precision and coherence of a dialogue generation model are improved, and the technical problems of context coherence and personalized generation are effectively solved.
Owner:SICHUAN KUAIDATONG TECHNOLOGY CO LTD

Phishing document de-obfuscation and feature extraction method and application thereof in attack detection

The invention discloses a phishing document de-obfuscation and feature extraction method and an application thereof in attack detection. The de-obfuscation comprises the steps of obtaining an obfuscation macro code of a phishing document, constructing a prompt engineering template by utilizing a pre-training language model, analyzing an obfuscation logic structure and generating a de-obfuscation rule and a reduction strategy; the method comprises the following steps: structuring a confused macro code into an abstract syntax tree through an analysis tool, matching a typical confusion mode based on a regular expression, and performing simulation and cell reference analysis in combination with a function to realize structure preliminary reduction, control flow semantic reduction and operation path construction; according to the unmixing rule and the abstract syntax tree, the confusion structure is converted into a readable macro statement, a macro code instruction sequence with confusion semantics removed is generated, and a semantic sequence obtained after unmixing is output. The feature extraction comprises word feature extraction, Token feature extraction, abstract syntax tree feature extraction and relation feature extraction. According to the method, the bottleneck that traditional phishing attack detection and confusion documents are difficult to recognize is broken through, and the detection accuracy of phishing document attacks is improved.
Owner:GUIZHOU UNIV

Intelligent matching system based on big data and matching method thereof

The invention relates to the technical field of policy matching, and discloses an intelligent matching system based on big data and a matching method thereof. The system comprises a data acquisition module, a condition analysis module and an intelligent matching module. The data acquisition module is used for acquiring policy text, enterprise qualification and historical case data and constructing a matching basic database; the condition analysis module is used for extracting policy dominant conditions through a regular expression, mining implicit conditions from failure cases by utilizing an association rule algorithm, and calculating condition weights by adopting a random forest algorithm; the intelligent matching module integrates dominant and implicit conditions to calculate a matching coefficient, dynamically divides risk levels and generates an early warning, the matching comprehensiveness and the early warning sensitivity are improved by mining the implicit conditions, dynamically adjusting the matching model and reinforcing a learning optimization mechanism, upgrading from text matching to risk prejudgment is achieved, and the method has the advantages of being high in practicability and easy to popularize. The enterprise declaration success rate can be improved, and the repeated declaration rate is reduced.
Owner:BEIJING LIANYUE TECHNOLOGY CO LTD

Large model training data set construction method based on table type data

The invention discloses a large model training data set construction method based on table type data. The method comprises the steps that 1, data preprocessing is carried out; step 2, constructing an automatic processing method and an extraction template based on header association; step 3, constructing a deformable template model; and step 4, binding question types and fields. According to the method, manual intervention is reduced by applying an automatic processing method based on header association and an extraction template, a question and answer template is preset, key fields are matched by using a regular expression, a table format is automatically adapted through a field identification module, and then data quality is guaranteed through a data quality evaluation and feedback mechanism; a deformable template model is provided to be combined with a dynamic adjustment mechanism, and field identification, dynamic adjustment and mapping modules cooperate to ensure accurate query; heterogeneous table data is standardized, problem types and corresponding fields are automatically bound to improve the processing efficiency, and the diversity of a data set is increased by using a data enhancement strategy based on a natural language.
Owner:NANJING AGRICULTURAL UNIVERSITY

Interaction control method and system for intelligent agent and front-end and back-end systems

The invention discloses an interaction control method and system for an intelligent agent and a front-end and back-end system, and the method achieves the automatic control of the intelligent agent on the front-end and back-end system through the integration of a large language model. The method comprises the following steps: an intelligent agent performs natural language intention recognition through a large language model, automatically selects a proper tool, extracts parameters required by the tool, converts the intention into a target task, generates a structured control instruction, maintains a complete session history and supports multiple rounds of dialogue interaction; the back end receives a user request, calls an intelligent agent to generate a structured instruction, transmits the structured instruction to the front end, and dynamically adjusts a subsequent instruction according to the received front end feedback; and the front end receives the instruction transmitted by the rear end, automatically identifies an instruction format through a regular expression, performs corresponding operation after parameter verification and feeds back a result. According to the method, unified control of the agent on the front-end UI and the back-end business logic is realized, complex multi-round dialogue and multi-step operation are supported, and the method has a wide application prospect.
Owner:ZHEJIANG LAB

Multi-modal document retrieval method and device, electronic equipment and storage medium

The embodiment of the invention discloses a multi-modal document retrieval method and device, electronic equipment and a storage medium. The similar problems that in an existing retrieval enhancement system, multi-modal data (pictures and tables) in documents are difficult to process, the content stored in a knowledge base is incomplete, the correlation between retrieval results and problems is low, non-professionals are difficult to construct the knowledge base and a question and answer system by themselves, and the like can be solved. The method comprises the steps of obtaining a to-be-processed multi-modal document; preprocessing the to-be-processed multi-modal document according to a preset regular expression to obtain a target text list; according to the target text list, a first knowledge base and a second knowledge base are created, and an index relation exists between the first knowledge base and the second knowledge base; according to the dense vector and the sparse vector corresponding to the second knowledge base, the to-be-retrieved content is retrieved, a target retrieval result is obtained, and the target retrieval result is determined according to sub-segment results obtained through retrieval of the dense vector and the sparse vector.
Owner:国家超级计算天津中心

Dynamic adaptive log analysis method and system based on local large language model

The invention belongs to the technical field of log analysis, and provides a dynamic self-adaptive log analysis method and system based on a local large language model, and the method comprises the steps: 1, distributing sampling weights according to log levels, and preferentially covering key logs; 2, guiding the locally deployed LLM to generate a regular expression for log analysis through a structured prompt; 3, analyzing the logs one by one according to the generated rules by using a traditional log analysis engine, and outputting structured logs; 4, when the unanalyzed log rate exceeds a preset threshold value, asynchronous rule generation is triggered; through the rule generation and verification module based on the local large language model, automatic generation and verification of the log analysis rule are realized, the cost of manually compiling the rule is remarkably reduced, and meanwhile, the accuracy and adaptability of the rule are improved. Therefore, the problems that a traditional log analysis method highly depends on manpower and is poor in cross-domain adaptability are solved.
Owner:EAST CHINA NORMAL UNIV

Method, electronic device and computer-readable storage medium for extracting entity relationships

A method, electronic device and computer-readable storage medium for extracting entity relationships, which relates to artificial intelligence technologies such as natural language processing, knowledge graphs, deep learning, and large language models. The method for extracting entity relationships includes: inputting a target long text into a target large language model to obtain a target keyword list based on an output result of the target large language model; inputting the target keyword list into multiple target relationship agents respectively to obtain multiple target regular expressions corresponding to different entity relationships based on output results of the multiple target relationship agents; and processing texts in a preset text set using the multiple target regular expressions to obtain entity relationship extraction results.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

LLM-based multi-source heterogeneous data intelligent fusion and adaptive processing method

The invention relates to the technical field of data processing, in particular to an LLM-based multi-source heterogeneous data intelligent fusion and self-adaptive processing method, which comprises the following steps of: receiving multi-source heterogeneous data, matching character sequences of mailbox fields and identity card number fields in records one by one by adopting a preset regular expression model, according to a preset key value pair mapping dictionary, items are uniformly converted, data records which cannot pass format verification, type conversion and value domain mapping are marked as doubt, and data records which cannot be processed or are marked as suspicious are separated. According to the method, a hierarchical processing flow is constructed, and deterministic format verification, type conversion and value domain mapping are arranged at the front end of a processing link, so that most simple data quality problems can be quickly processed with low calculation overhead. Then, numeric statistics and character string comparison are used to carry out fuzzy correction on the data, and spelling errors and numeric abnormities which cannot be covered by rules are further solved.
Owner:GUANGXI BEITOU SOFTWARE CO LTD

Occupational information assessment method based on rule judgment and large language model

The invention discloses an occupational information evaluation method based on rule judgment and a large language model, and belongs to the technical field of information processing, and the method comprises the following steps: S1, obtaining a multi-modal document uploaded by a user, converting the multi-modal document into a Markdown format, and completing a shunting operation and a converging operation; s2, extracting a named entity of the multi-modal document after the confluence operation, and constructing a knowledge graph; s3, obtaining query content, and matching most relevant regulation content by using the knowledge graph; s4, generating specific improvement suggestions by utilizing the rule model according to the most relevant regulation content; and S5, adjusting the specific improvement suggestions by using a large language model to obtain a final evaluation result. Through the OCR technology, regular expression analysis and context-associated picture analysis, unified structured processing of texts, tables and pictures is achieved, and the problem that unstructured data are difficult to integrate in a traditional method is effectively solved.
Owner:SICHUAN UNIV

Large model output content security test method and device

The invention relates to the field of large model security testing, and particularly provides a large model output content security testing method and device, and the method comprises the following steps: S1, preparing and managing a test set, a sensitive word library and a regular expression which are required by testing; s2, reading a test set, and obtaining a large model output result according to the test set and the large model interface information; s3, judging whether the output content of the large model is safe or not according to the sensitive lexicon and the regular expression; s4, extracting semantic risk features according to the output content of the large model by using the large model and the oriented Prompt, and automatically storing the semantic risk features after confidence verification; and S5, storing the information result of each request in a file. Compared with the prior art, the test time can be shortened, and the evaluation efficiency can be improved; and the security of the output content of the large model can be effectively evaluated by using a method for dynamically constructing the sensitive word bank by using the output result of the large model.
Owner:INSPUR QILU SOFTWARE IND

Method for deploying regular expression based on P4 in intelligent network card / DPU

The invention discloses a method for deploying a regular expression based on P4 in an intelligent network card / DPU. The method specifically comprises the following steps that the regular expression defined by a user is compiled into a finite-state machine; a matching table structure and a state register are defined in the P4 program and used for storing a state conversion rule of the regular expression, and a current matching state is recorded; and mapping a matching logic and a state register defined by the P4 into a hardware module of the intelligent network card / DPU, and realizing high-speed state conversion and matching by utilizing the parallel processing capability of the intelligent network card / DPU. According to the method for realizing regular expression matching based on P4 in the intelligent network card or the DPU, the regular expression is converted into the finite-state machine, and the programmability and the hardware acceleration capability of the P4 language are combined, so that the network flow detection and analysis efficiency is greatly improved, and the network flow detection and analysis efficiency is improved. And an effective solution is provided for next-generation network security and performance optimization.
Owner:ZHEJIANG RUIWEN TECH CO LTD

Structured analysis and semantic template normalization method for unmanned vehicle operation logs

The invention provides a structured analysis and semantic template normalization method for an unmanned vehicle operation log. For the problems of unstructured unmanned vehicle log data, changeable formats, complex semantics and the like, logs are converted into semantic vectors by adopting a multi-language sentence vector coding model, and semantic grouping is performed by using a small-batch K-means clustering algorithm after dimension reduction is performed by adopting an incremental principal component analysis method; extracting a variable field from a clustering result by using a regular expression and generating a standardized template; and intelligent template combination is realized by calculating cosine similarity and editing distance between the templates. The method has the following three technical characteristics: 1) the template consistency is improved by adopting a strategy of clustering first and then normalizing; 2) introducing double similarity constraints to reduce template redundancy; and 3) retaining variable sequence and semantic information to support subsequent analysis. The method is suitable for log data analysis and intelligent operation and maintenance of various complex systems.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Non-depth flow analysis and streaming matching search analysis method

PendingCN121508886ASecuring communicationSearch analyticsData pack
The invention provides a non-deep traffic analysis streaming matching search analysis method, which comprises the following steps of: constructing a deep data packet detection architecture on the basis of a data platform development kit (DPDK); aiming at the non-encrypted traffic, establishing a multi-mode recognition mechanism combining a regular expression and a feature bit stream mode; aiming at the encrypted traffic, constructing an encrypted traffic feature library according to the statistical features, the protocol features and the behavior features; real-time risk detection of network traffic is realized by adopting a streaming matching search technology; and constructing an intelligent decision and response mechanism, and integrating flow analysis and identification results. According to the non-deep traffic analysis streaming matching search analysis method provided by the invention, real-time analysis, risk identification and supervision of network non-encrypted traffic and encrypted traffic are realized by taking a data platform development kit (DPDK) as a basis and combining a high-performance streaming regular expression engine and a finite-state machine principle; the method can be widely applied to scenes of network communication supervision, data security protection, malicious traffic monitoring and the like.
Owner:BEIJING ACT TECH DEV CO LTD

PDF text extraction method based on image segmentation and OCR

The invention discloses a PDF (Portable Document Format) text extraction method based on image segmentation and OCR (Optical Character Recognition), which belongs to the technical field of image processing and comprises the following steps: S1, carrying out image processing on a PDF file to be analyzed; s2, constructing a column-dividing judgment model, and judging whether columns exist in the PDF file subjected to the imaging processing or not through the column-dividing judgment model; s3, constructing an image segmentation model, segmenting the PDF file with columns, calling an OCR (Optical Character Recognition) algorithm interface according to coordinates and a sequence to extract text information according to segmentation results, and splicing the text information according to the sequence; s4, page header and page footer filtering is carried out based on the text coordinates; s5, finally, regular expression filtering and table information filtering are carried out, and text data are obtained. According to the method, the PDF files in various formats can be processed, the method is particularly suitable for the PDF files with the column dividing condition, the recognition accuracy is high, and good applicability is achieved.
Owner:CHENGDU AIRCRAFT INDUSTRY GROUP

Method for analyzing rationality of acquisition and bid evaluation information of power equipment

The invention discloses a rationality analysis method for acquisition and bidding evaluation information of power equipment, which comprises the following steps: constructing a regular expression to perform paragraph matching on preprocessed bidding document and bidding document texts, and extracting specific parameter contents; obtaining an absolute error and a relative error based on each bidding value and each bidding value, and comparing the absolute error and the relative error with a preset threshold to obtain a numerical deviation degree grade of the bidding file relative to the bidding file; the method comprises the following steps: splitting a bid invitation file and a bidding file into character sequences, creating a two-dimensional table, analyzing to obtain an editing distance value of each text, obtaining a comprehensive matching degree of the text through accurate matching and fuzzy matching based on a preset synonym library, and further obtaining a text deviation degree grade; constructing a deviation degree comprehensive evaluation rule matrix based on the numerical deviation degree grade and the text deviation degree grade of the bidding file relative to the bidding file, and obtaining a corresponding comprehensive evaluation result; according to the invention, the efficiency and accuracy of collection and bidding evaluation of the power equipment are improved, and personal errors and compliance risks are reduced.
Owner:GUIZHOU POWER GRID CO LTD

Railway official document keyword extraction method and device and electronic equipment

The invention relates to a railway official document keyword extraction method and device and electronic equipment, and the method comprises the steps: based on a pre-constructed railway official document format rule base, extracting a key field of a fixed position from an input text through regular expression matching and position locking; a Jieba word segmentation device is used for loading a railway-specific term library for word segmentation, and a multi-word combination entity boundary is dynamically corrected through a dependency relationship rule; executing a TF-IDF algorithm on the text after word segmentation to generate an initial word weight, adjusting the weight according to the position area of the word in the official document and a preset coefficient, and performing position weighting; and combining the words of which the weights are greater than a set threshold value with the extracted key fields, and outputting a final keyword set after verification of a term library. According to the method, missing detection caused by low frequency of a traditional algorithm is avoided, splitting errors of a general word segmentation device are eliminated, the term recognition error rate is reduced, the core word sorting priority is improved, and the semantic weight of keywords is strengthened; the new term storage time is shortened, and the updating cost problem is solved.
Owner:INST OF COMPUTING TECH CHINA ACAD OF RAILWAY SCI +2

Picture type PDF document analysis method based on convolutional neural network, multi-modal model and regular expression

The invention discloses a picture type PDF document analysis method based on a convolutional neural network, a multi-modal model and a regular expression, and belongs to the technical field of artificial intelligence and text processing. The method comprises the following steps: detecting types of layout elements of a preprocessed PDF document to obtain bounding box coordinates of each layout element; performing content identification on each layout element according to the type of the layout element; using a regular rule engine and a large language model to perform structured information extraction on the identification content, and extracting to obtain a plurality of predefined first business fields corresponding to each layout element; the character recognition result and the table recognition result are combined, and the combined result serves as content needing to be extracted; and taking a proofreading result as an analysis result of the scanned PDF document. According to the method, high-precision structured extraction of paragraphs, tables, formulas and other contents in the PDF document is realized, and the method has good universality, expandability and automation capability.
Owner:MILITARY SCI INFORMATION RES CENT ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY

Automatic identification method and system for correspondence-related regulations

The present invention relates to the field of document recognition technology, and discloses a method and system for automatically identifying regulations associated with letters. The method comprises the steps of file uploading, preprocessing, text recognition, chapter recognition, content extraction, tree structure building, and result display. In the image preprocessing, the present invention has accurate grayscale, flexible noise reduction, adaptive binarization, and efficient tilt correction, thereby improving image quality and ensuring the basis for text recognition. The text recognition process is innovative, and line segmentation and character segmentation are reliable and accurate. Character recognition combined with a standard character template library is intelligent and efficient, thereby improving accuracy. Table of contents recognition is judged in multiple dimensions through regular expressions, font features, and numbering structures, and the hierarchy is scientifically and verified, and the content extraction is reasonable. The tree structure is built based on the hierarchy, combined with refined node and content pairs, with clear levels. The modules of the system work together to achieve automated processing, providing users with convenient and intuitive display of regulatory information.
Owner:NANJING ANXIA ELECTRONIC TECH CO LTD

Systems and methods for securing data based on discovered relationships

Techniques for automatically discovering and protecting sensitive data are disclosed. In some embodiments, a set of data objects is searched for data matching a first set of one or more regular expressions and for metadata matching a second set of one or more regular expressions. A confidence score is then generated for a particular data objects in the set of data objects as a function of regular expressions in the first set of one or more regular expressions that match data stored in the particular data object and regular expression in the second set of one or more regular expressions that match metadata associated with the particular data object. One or more operations may be performed to protect sensitive data stored in the particular data object based, at least in part, on the confidence score.
Owner:ORACLE INT CORP

Slurm scheduling specification integration method and system

The invention relates to the field of high-performance computing cluster resource scheduling management, and discloses a Slurm scheduling specification integration method and system, and the method comprises the following steps: 1, analyzing a heterogeneous job description file, and extracting a resource demand parameter and a dependency relationship through a regular expression rule base; 2, based on the extracted original parameters, converting the original parameters into SLURM standard parameters through a preset mapping rule; step 3, according to the converted standard parameters; 4, calling a Slurm interface command to submit a script; and step 5, monitoring the execution state of the submitted job, and triggering a re-submission process for the abnormal job with the resource overrun or dependency missing. Through a multi-level analysis architecture and a regular expression rule base, job description files in different formats are effectively compatible, the problem that analysis of a non-standardized input format by a traditional method fails is solved, unified processing of cross-platform job definition is achieved, and the manual adaptation cost is remarkably reduced.
Owner:北京月新时代科技股份有限公司