Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

98 results about "Structured text" patented technology

Structured text, abbreviated as ST or STX, is one of the five languages supported by the IEC 61131-3 standard, designed for programmable logic controllers (PLCs). It is a high level language that is block structured and syntactically resembles Pascal, on which it is based. All of the languages share IEC61131 Common Elements. The variables and function calls are defined by the common elements so different languages within the IEC 61131-3 standard can be used in the same program.

Electronic device, method, and non-transitory computer-readable storage medium for performing paste function

PCT designated stageWO2026135162A1Biological modelsInput/output processes for data processingTransformation of textDisplay device
This electronic device comprises a memory for storing instructions, a display, and at least one processor including processing circuitry. The instructions, when executed individually or collectively by the at least one processor, cause the electronic device to: when a first user command for capturing a first screen of the display is identified, acquire a captured image of the first screen and additional information related to the first screen; when a text included in the captured image is identified, convert the identified text into structured text data on the basis of the additional information; acquire paste content on the basis of the structured text data and context information about a second screen of the display; and input the paste content to an input field included in the second screen.
Owner:SAMSUNG ELECTRONICS CO LTD

Intelligent alarm method for industrial safety monitoring

The application relates to an intelligent alarm method for industrial safety monitoring and belongs to the technical field of industrial safety production monitoring, which solves the problems of insufficient scene understanding ability and high false alarm rate of a traditional AI monitoring system. The method comprises the following steps: acquiring an industrial production scene video stream in real time and extracting a key frame image; determining a scene description structured text based on the key frame image and a scene classification prompt word by using a multi-modal visual language model; acquiring a determined high-risk scene type based on the scene description structured text and a scene reasoning prompt word by using a large language model; determining a violation feature of the scene based on the determined scene type and a violation feature prompt word by using a model calling platform; determining a judgment conclusion about the violation feature based on the violation feature of the scene and the key frame image by using a multi-modal visual language model; and generating an alarm event according to the judgment conclusion. The application realizes an intelligent alarm method for industrial safety monitoring.
Owner:BEIJING ENERGY INVESTMENT HLDG +3

A strategy model training method and device, electronic equipment and storage medium

The application relates to the technical field of computers, and provides a strategy model training method and device, electronic equipment and a storage medium. The method obtains unlabeled sample data; a strategy model to be trained is used to sample the unlabeled sample data multiple times, to generate multiple candidate outputs corresponding to the unlabeled sample data; the intrinsic reward corresponding to each candidate output is calculated, and the parameters of the strategy model are updated based on the intrinsic reward corresponding to each candidate output, to obtain a target strategy model, which is used to receive a natural language query request input by a user and generate structured text response data corresponding to the natural language query request. Through multiple samplings of the unlabeled sample data by the strategy model to be trained and the calculation of the intrinsic reward of the candidate output, the reinforcement learning update of the model parameters is directly completed based on the intrinsic reward, and the strategy model is freed from the serious dependence on high-quality artificial labeled data and an external reward model in traditional large model alignment training.
Owner:北京衔远有限公司 +1

Intelligent service request submittal system

ActiveUS12682324B23d user interfaceDatabase
Techniques for aircraft manufacture and product support are described. These techniques include providing a three-dimensional (3D) user interface for an aircraft and identifying a component relating to an area of concern for the aircraft, based on input to the 3D user interface. The techniques further include generating structured text for a service request relating to the area of concern based on the identified component, and submitting the service request to facilitate resolving the area of concern for the aircraft.
Owner:THE BOEING CO

Method and apparatus for recognizing new words and first occurrence events

ActiveCN121787410BText streamHotline
The application discloses a new word and a first occurrence event recognition method and device, wherein the method comprises the following steps: obtaining hot-line work order text data according to a preset period, and preprocessing the hot-line work order text data to obtain a structured text stream; performing word segmentation and sliding window processing on the text stream to obtain an initial phrase candidate set; performing statistical access screening on the initial phrase candidate set, and performing alias merging on the candidate phrases passing the statistical access by using consistency rules of specified dimensions to obtain a word candidate set; updating a current space-time reference baseline based on the word candidate set, and recalculating historical full-amount data according to a specified fixed period to re-calibrate the space-time reference baseline; and determining whether there is a new word in the word candidate set and whether there is a first occurrence event based on the re-calibrated space-time reference baseline. Based on the unified space-time reference baseline, the new word and the first occurrence can be accurately and timely determined.
Owner:CAPINFO CO LTD

Power grid cime model attribute graph modeling method and system, computer device and medium

PendingCN122452378ALinguistic modelAlgorithm
The application relates to the technical field of artificial intelligence, in particular to a power grid CIME model attribute graph modeling method and system, computer equipment and a medium; the method comprises the following steps: analyzing a power grid CIME format file, extracting electrical equipment entities and connection relationships, and constructing a static attribute graph; fusing dynamic operation attributes in real-time operation data into the attributes of corresponding graph nodes in the static attribute graph to form a dynamic power grid attribute graph; based on a graph-text multi-dimensional semantic decoupling projection strategy, the dynamic power grid attribute graph is projected into a structured text sequence; the structured text sequence is packaged into a prompt word and input into a large language model to execute a power grid cognitive reasoning task. In this way, the technical problem of insufficient real-time reasoning capability of a large model in the application scenario of the existing technology facing large language model cognition is solved, and the accuracy, efficiency and real-time performance of the large language model for the power grid cognitive reasoning task are improved.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO LTD HANGZHOU POWER SUPPLY CO +4

A method and device for generating information for hepatocellular carcinoma risk stratification and treatment recommendations

The application provides a liver cancer risk stratification and treatment recommendation information generation method and device, the method comprises the following steps: obtaining liver cancer related structured clinical data of a target object; based on a pre-set electronic medical record narrative template, converting the structured clinical data into an electronic medical record style narrative text containing clinical semantics, embedding key clinical variable markers corresponding to the structured clinical data into the electronic medical record style narrative text, and constructing a model input sequence; inputting the model input sequence into a pre-trained large language model for inference operation, outputting a structured text stream, and the large language model is obtained through decision tree constraint based on a liver cancer diagnosis and treatment guideline and multi-objective reinforcement learning strategy training; and extracting comprehensive decision assistance information from the structured text stream by using a pre-set analysis rule, realizing high-credibility clinical assistance decision, and significantly improving the accuracy and logical consistency of the treatment recommendation information.
Owner:TSINGHUA UNIVERSITY

Intelligent bid evaluation and bidding content page code indexing method and system based on large language model fusion

PendingCN122364230ALinguistic modelDirectory structure
This application discloses a method and system for intelligent page number indexing of bidding content based on a large language model, relating to the field of computer data processing technology. The method includes: extracting structured review item information from the bidding documents; extracting text from each page of the bidding documents according to document type to generate a structured text set; extracting the title information of each page; establishing a correspondence between each title and physical page number based on their physical page number position in the bidding documents to obtain the document directory structure; inputting the structured review item information and the document directory structure into a pre-trained first large language model to establish a mapping relationship and obtain the confidence score of the mapping relationship; and obtaining the page number indexing result of the bidding content, including the physical page number range and the confidence score of the mapping relationship. This solves the technical problems of low efficiency, inaccurate page number positioning, and severe failure of traditional keyword mapping methods in existing technologies.
Owner:ANHUI HIGH QUALITY MINING TECH DEV CO LTD

Underground pipeline pipe gallery resource connection method and device based on large language model

ActiveCN122197244BOpen up the relationshipRealize processingGlobal topologyLinguistic model
The application provides a kind of underground pipeline pipe gallery resource interconnection method and device based on large language model, belongs to computer technology field, specifically includes obtaining the pipeline detection data of all underground pipelines in target area, constructs global topology knowledge graph;Extract local topology knowledge graph from global topology knowledge graph;The structured text is obtained by converting local topology knowledge graph, and the pipeline gallery resources in structured text are evaluated and filtered using a large language model, to obtain structured text with compatibility evaluation label;Using a large language model, execute logical reasoning based on the local topological relationship represented by the structured text with compatibility evaluation label, generate reasoning results that meet pipeline planning and complete path chain from start to end;Based on the reasoning result, a candidate interconnection scheme is generated. Through the processing scheme of the present application, an interconnection scheme is generated that takes into account the reuse of existing resources, cost control and safety compliance, improving the intensive utilization efficiency of urban underground space.
Owner:SHANGHAI YINGYI URBAN PLANNINGDESIGN CO LTD

Hierarchical text structuring and domain-specific information reflecting method for improving performance of information extraction based on large-scale language model

According to an embodiment of the present disclosure, a computer program stored in a computer readable storage medium is provided. The computer program causes a processor to execute the following method for improving information extraction performance based on a large-scale language model, the method comprising: a step of separating text data by sentence unit; a step of hierarchically structuring the upper and lower level relationships between the separated sentences using symbols or indentations; a step of adding prefixes for task indication and suffixes for output format specification in the hierarchically structured text, thereby generating hierarchical input that can be provided to the large-scale language model; and a step of providing the above hierarchical input to the large-scale language model to output structured data containing domain-specific information.
Owner:RENSHOT GMBH

Protocol finite state machine extraction method and system

The application provides a protocol finite state machine extraction method and system, which comprises the following steps: preprocessing an original protocol document to generate a structured text block and a vector knowledge base; performing entity recognition on the structured text block to obtain a target entity list; performing skeleton extraction on the structured text block based on a first large language model and the target entity list to generate an initial global protocol finite state machine skeleton; performing iterative structure correction on the initial global protocol finite state machine skeleton, and if the corrected initial global protocol finite state machine skeleton meets the structure constraint, regarding it as a target global protocol finite state machine skeleton; performing iterative filling of details on the target global protocol finite state machine skeleton based on the target entity list, the structured text block and the vector knowledge base to obtain an initial protocol finite state machine; and performing format conversion on the initial protocol finite state machine to obtain a target protocol finite state machine. The application can improve the accuracy and reliability of the generated model, and overcome the problems of low model processing efficiency and information dispersion.
Owner:BEIJING UNIV OF POSTS & TELECOMM

A small language-oriented audio and video subtitle optimization generation method and system

PendingCN122313949ASoutheast asiaSpoken language
This invention relates to the field of artificial intelligence technology and discloses a method and system for optimizing and generating audio and video subtitles for less commonly spoken languages. The method includes the following steps: Step 1: Audio extraction and preprocessing; Step 2: Language recognition for the less commonly spoken language; Step 3: Language-aware punctuation restoration module; Step 4: Subtitle readability-driven segmentation module; Step 5: Subtitle format encapsulation and output. This application establishes a complete technical process encompassing audio preprocessing, language-aware speech recognition, structured text restoration, subtitle segmentation optimization, and time alignment. This method significantly improves the accuracy of speech-to-text conversion in less commonly spoken language videos and enhances the structural integrity and readability of subtitles. It is particularly suitable for practical application scenarios involving less commonly spoken languages ​​in regions such as Southeast Asia, where data is scarce and languages ​​are diverse. It has broad application value and practical significance in media dissemination, educational videos, and government services.
Owner:XINGZHOU DIGITAL TECH (ZHUHAI) CO LTD

A Visual Language-Based Method for Identifying Crop Diseases and Pests

PendingCN122313285APattern recognitionDisease
This invention discloses a method for identifying crop diseases and pests based on visual language, comprising the following steps: S100, temporal image acquisition and temporal texture set construction; S200, bidirectional dynamic masking processing and mask temporal feature map generation; S300, visual language anchor point library construction and bimodal feature encoding; S400, temporal anchor point alignment and coarse identification of diseases and pests; S500, disease and pest evolution inference and confidence level determination; S600, structured text output and presentation of recognition results. This invention uses temporal image sequence acquisition and temporal texture evolution features as the core recognition basis, completely breaking through the technical limitations of traditional single-frame static recognition, effectively reducing the probability of confused recognition and misjudgment, and significantly improving the accuracy and discriminative power of disease and pest identification. Simultaneously, by guiding temporal feature alignment and disease and pest evolution confidence level determination through language anchor points, it can simultaneously output the disease and pest evolution stage and damage level, providing direct and reliable decision support for precision plant protection in the field.
Owner:HENAN ZHILIAN TIME & SPACE INFORMATION TECH CO LTD

A Host Asset Identification and Profiling Method Based on Multimodal Network Feature Fusion Enhancement

PendingCN122087691Aimprove concealmentHighly non-invasiveText processingBiological modelsEngineeringSemantic feature
This invention relates to the field of secure communication application technology, specifically to a host asset identification and profiling method based on multimodal network feature fusion enhancement. First, a data collection agent is deployed at the core node of the target network to collect multimodal data through traffic mirroring and log crawling, and a cross-modal aligned multi-source network data representation is constructed. Next, a TCN-BiLSTM model combined with an attention mechanism is used to extract deep periodic features, and a GraphSAGE model is used to generate graph embedding vectors bound to topological relationships. A pre-trained BERT model is used to obtain global protocol semantic feature vectors and communication intent distribution. Subsequently, a modality-specific mapping is used to unify dimensions, and a two-layer mechanism of intra-modal self-attention and inter-modal mutual attention is constructed to dynamically allocate weights, fusing multimodal features to extract core features of host assets. Finally, the fused features are mapped to structured text, and a CLM strategy combined with LoRA technology is used to fine-tune the LLM, generating a natural language asset profile containing role positioning, behavioral patterns, and topological relationship logic.
Owner:GUANGXI UNIVERSITY OF TECHNOLOGY

A Spatial Decision Recommendation Generation Method and System Based on LLM and Geographic Knowledge Base

PendingCN122088638AAccurately understand decision-making objectivesAccurately understand complex decision-making needsSemantic analysisInference methodsLinguistic modelSimilarity computation
A spatial decision-making suggestion generation method and system based on LLM and a geographic knowledge base, relating to the field of data processing technology, is disclosed. This method first establishes a knowledge base containing spatial geometric data of geographic entities. When a user's natural language decision query is received, a large language model is used for semantic parsing to extract the decision objective, spatial entities, and constraints. The system retrieves candidate knowledge from the knowledge base through semantic similarity calculation and evaluates the spatial compliance confidence level of each candidate based on the spatial geometric data and constraints. Subsequently, the candidate knowledge is hierarchically labeled according to confidence level to construct structured text, which is then injected into the large language model through a preset template. Finally, the system generates a spatial decision report based on the injected structured information. Implementing the technical solution provided in this application can improve the accuracy of spatial decision-making suggestions.
Owner:SHANGHAI GISINFO TECH CO LTD

A layout design rule generation method, device and medium based on a multi-modal agent technology

The application discloses a layout design rule generation method and device based on a multi-modal intelligent agent technology and a medium, relates to the field of electric digital data processing, and converts a design rule manual into a structured text format and divides the structured text format according to process levels. A visual model is used to perform semantic extraction on a schematic diagram in the manual, generate an image description text, and fuse the image description text with an original rule text. A candidate function set is screened from a function library, historical experience rule library is combined, and a function or function combination required by a current rule is decided. Detailed usage information is acquired, process level auxiliary constraints and global instructions are retrieved from the manual, code specifications are retrieved from the experience knowledge base, all knowledge is enhanced to guide rule semantic restatement, image information selection, function parameter analysis and pre-output self-checking, and layout design rule checking code is generated. The application solves the problems of low manual conversion efficiency and errors, and has the advantages of high generation efficiency, accurate results and good interpretability.
Owner:FUDAN UNIVERSITY

A large model-based human-computer voice precise interaction method

ActiveCN121905189BSpeech recognitionElectric/fluid circuitSound source locationSound sources
The present application belongs to the technical field of voice interaction, and particularly relates to a human-machine voice precise interaction method based on a large model. In view of the problems of voice source confusion, low instruction recognition rate and insufficient interaction safety caused by the complex acoustic environment in the vehicle cabin, the present application fuses reverberation features and harmonic attenuation features to construct acoustic fingerprints, and realizes high-precision sound source positioning in combination with seat occupancy state verification; performs priority filtering and semantic pre-alignment on non-main driver voice, and eliminates fragmented interference; introduces a large language model to standardize the mapping of structured text into instructions conforming to syntax, format and semantic specifications, and embeds a dynamic permission verification mechanism based on sound source position. The scheme significantly improves the voice instruction recognition accuracy and system robustness, effectively suppresses false triggering, enhances user privacy protection, provides an efficient and precise voice interaction experience for the intelligent cabin, and has outstanding technical practical value and industrial application prospects.
Owner:JINAN BOSAI NETWORK TECH CO LTD

Multimodal BEV perception consistency alignment checking method based on language structure prior

The invention discloses a multi-modal BEV perception consistency alignment checking method based on language structure prior, and relates to the technical field of automatic driving, and the method comprises the steps: obtaining multi-modal sensor data, and mapping the multi-modal sensor data to a virtual projection space; and based on the large language model, generating language structure priori in a structured text form, and mapping the language structure priori to the virtual projection space to generate a structure constraint field. In the training stage, based on a structure constraint field, weak supervision consistency loss is constructed, an evidence gating mechanism is introduced, loss weight is dynamically adjusted according to an evidence intensity score, and joint optimization is carried out on the network. In the reasoning stage, consistency checking is carried out on BEV sensing results by utilizing a structure constraint field, and confidence reestimation, hallucination suppression and open set rejection are carried out in combination with a structure conflict score and an evidence intensity score. The problems of structure default and high-confidence illusion under the weak evidence working condition are solved, and the physical consistency and reliability of the sensing result are improved.
Owner:TONGJI UNIV

A few-shot contraband image recognition method fusing language and frequency prior

This invention relates to the field of artificial intelligence technology, specifically to a method for few-shot contraband image recognition that integrates language and frequency priors. This method includes the following steps: text prompt generation: inputting the contraband category C and scene context S into a large language model to generate structured prompt text, and extracting the prompt embedding vector through a text embedding encoder. This invention combines the structured text prompt information generated by GPT-4 with task-related frequency component analysis, jointly guiding the visual-language model (CLIP) to learn more discriminative feature representations at multiple levels, thereby improving the model's classification performance and generalization ability under data-scarce conditions.
Owner:DATA SPACE RES INST

A multi-modal large model hidden danger identification method and system for the urban energy industry

This invention discloses a multimodal large-scale model hazard identification method and system for the urban energy industry, belonging to the field of automatic hazard identification technology in the energy industry. It includes S1, data preprocessing; S2, multi-dimensional low-rank parameter fine-tuning: constructing a ROI-oriented multi-branch LoRA fine-tuning framework to perform local visual feature fine-tuning on the image ROI feature set, generating global image fusion features; constructing a field-level differentiated LoRA fine-tuning module group to perform field-level independent semantic fine-tuning on the structured text field set, generating exclusive text features for each field; S3, cross-modal feature fusion and decision output. This invention's multimodal large-scale model hazard identification method and system for the urban energy industry constructs a full-link structure of pre-calibration locking of the area of ​​interest – dual-mode initial screening – refined ROI extraction – calibration and adaptation fine-tuning, deeply integrating calibration information into the entire process of data preprocessing and multi-dimensional low-rank parameter fine-tuning, improving image processing efficiency and hazard identification capabilities.
Owner:SHANDONG HETONG INFORMATION TECH CO LTD

Credit evaluation processing method and apparatus

The embodiment of the present specification provides a credit evaluation processing method and device, wherein the credit evaluation processing method comprises: in the credit evaluation processing process, inputting a to-be-evaluated image into a text extraction model for text extraction to obtain structured text, performing identity verification on user identity information contained in the structured text, and performing similarity verification on entity fields contained in the structured text and preset entity fields, if the verification passes, inputting the to-be-evaluated image and prompt text into a large language model for verification processing, obtaining a verification result, and synchronizing the structured text and the verification result to a credit evaluation platform.
Owner:CHONGQING ANT CONSUMER FINANCE CO LTD

Open-vocabulary substation equipment segmentation and inspection system based on multi-modal prompt learning

ActiveCN122156832BData packVisual technology
This invention relates to the fields of intelligent inspection of power equipment and computer vision technology, specifically to an open-vocabulary substation equipment segmentation and inspection system based on multimodal cue learning. The system includes: an anchoring modeling module, which acquires structured text of the target scene as an initial text prefix input to a text encoder and acquires sample images as visual anchors; a collaborative adaptation module, which acquires a real-time image stream, extracts multi-scale visual features through a visual encoder, and inputs them into a visual-language orthogonal coupling projection layer to obtain target cue parameters; a fitting segmentation module, which calculates the inner product of the visual cue feature tensor and the text cue feature tensor to generate a cross-modal attention heatmap and outputs a pixel-level segmentation mask; and a closed-loop deployment module, which generates state recognition results based on the pixel-level segmentation mask and serializes and stores the target cue parameters as a parameter data package with attribute labels. This invention can achieve open-vocabulary segmentation under small sample conditions and reduce the risk of catastrophic forgetting.
Owner:SHENZHEN LAIDA SIWEI INFORMATION TECH CO LTD

Object-oriented code processing method and device, electronic equipment and storage medium

This application provides an object-oriented code processing method, apparatus, electronic device, and storage medium, relating to the field of programmable logic controller (PLC) technology. The object-oriented code processing method provided by this application includes: converting an acquired structured text language program into code to generate a code program, and compiling the code program to obtain an object file and a first variable library; given that the acquired PLC's operating mode is incremental update mode and the PLC has a second variable library, using a pre-built carriage merging algorithm to compare and analyze the second variable library with the first variable library to obtain a memory block processing file to be updated; and sending the object file, the first variable library, and the memory block processing file to be updated to the PLC, enabling the PLC to run the code program based on the object file, the first variable library, and the memory block processing file to be updated. This application can reduce program complexity and achieve incremental update operation.
Owner:NR ELECTRIC CO LTD +1

A problem-based generation education domain knowledge base search optimization method and device

The application discloses an education field knowledge base search optimization method based on question generation, which comprises the following steps: firstly, obtaining the education field text by analyzing the education knowledge base direction field corpus; obtaining the semantic model by using the education field text to pre-train the language model migration learning; designing the fixed question and answer pair template based on the existing structured text information in the knowledge base to obtain the knowledge base question and answer pair; training the question generation model by using the knowledge base question and answer pair data and the Chinese open source question and answer pair data, deploying the question generation reasoning service; generating the question and answer pair to expand the knowledge base; simultaneously coding the entity node text in the knowledge base structured information and the question text in the question and answer pair by using the semantic model to construct the vector library, and performing the semantic similarity calculation after the user query input; recalling the best result in the online semantic matching; and the application greatly improves the recall rate of the returned result of the user search behavior, improves the learning efficiency and improves the user experience.
Owner:ZHEJIANG LAB

Supply chain business document evaluation method and device, electronic equipment and storage medium

PendingCN122347323AInformation processingCompliance analysis
The present disclosure relates to the technical field of data evaluation, and provides a supply chain business file evaluation method and device, electronic equipment and storage medium. The method comprises: performing multi-modal information extraction processing on a supply chain business file to be evaluated to obtain structured text, structured data and image feature information; performing multi-dimensional compliance analysis processing on the three to obtain a clause completeness evaluation value, a contract risk evaluation value, a technical parameter clarity evaluation value and a graphic-text consistency evaluation value; performing weighted quantification processing on the clause completeness evaluation value, the contract risk evaluation value, the technical parameter clarity evaluation value and the graphic-text consistency evaluation value based on a preset weight value to obtain a target evaluation value; and performing evaluation suggestion generation processing based on the target evaluation value and a preset evaluation threshold to obtain an evaluation suggestion text, thereby enhancing the forward-looking and accuracy of risk identification, improving the objectivity and traceability of the evaluation result, and enhancing the consistency of multi-modal information processing.
Owner:BEIJING QDING INTERCONNECTION TECHNOLOGY CO LTD

A method for testing a multi-mode master-slave control system and a laser device

PendingCN122239671AElectric testing/monitoringSystem level testingTarget text
This application discloses a testing method and laser equipment for a multi-mode master-slave control system, relating to the field of laser testing technology. It obtains structured text representing a hierarchical testing system, including a first text representing board-level characteristic testing rules and a second text representing system-level functional testing rules. Based on the first and second texts, board-level characteristic tests and system-level functional tests are performed on the master control board, slave control board, and beam combiner board, respectively, yielding board-level test results and system-level test results. Based on the board-level and system-level test results, a target text representing a comprehensive test report is generated. This achieves comprehensive testing and verification of the hardware foundation and system functional coordination of each board in the multi-mode master-slave control system, accurately identifying hardware defects and potential collaborative problems on each board. It provides a standardized solution to ensure the overall quality and reliability of the multi-mode master-slave control system, significantly improving the testing efficiency and reliability of the multi-mode master-slave control system.
Owner:WUHAN RAYCUS FIBER LASER TECHNOLOGY CO LTD

Document verification method and device, electronic equipment and storage medium

The application relates to the technical field of text processing, and discloses a document verification method and device, electronic equipment and a storage medium, the method comprising the following steps: text extraction is performed on a plurality of to-be-verified documents to obtain a structured text set; the document structure of the structured text set is analyzed, and a corresponding regular expression set is generated based on the document structure; the to-be-verified documents and standard verification documents are matched based on the regular expression set, and a semantic similarity score and a fuzzy matching score are obtained respectively; a comprehensive score is calculated based on the semantic similarity score and the fuzzy matching score, and the similarity relationship between the to-be-verified documents and the standard verification documents is verified based on the comprehensive score. The application significantly improves the work efficiency and verification accuracy in a large-scale document verification scenario, and solves the problem that the document content cannot be efficiently and accurately verified to conform to a specific standard.
Owner:ZHEJIANG CHINT INSTR & METER

Content extraction method, apparatus, device, storage medium, and product

The application relates to the technical field of illegal website detection, and discloses a content extraction method, device and equipment, a storage medium and a product, which comprise the following steps: in response to a content extraction request of a target page, loading the target page based on a preset browser instance to obtain a rendered target page; performing text recognition on a visual content image of the rendered target page to generate a text recognition result; and determining structured text content of the target page according to the text recognition result. The application shifts the entry point of content extraction from a document object model structure to a visual presentation level, presents the page to a final visual state through browser dynamic rendering, and then extracts text from the corresponding image. As long as the content is visually visible in the browser, it can be effectively extracted, and all interference contents injected through a visually invisible mode are naturally excluded because they do not exist in the visual content image, so that the purity and accuracy of the extracted content are significantly improved.
Owner:BEIJING HONGTENG INTELLIGENT TECH CO LTD