Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

247 results about "Content extraction" patented technology

Content extraction is the task of separating boilerplate such as comments, navigation bars, social media links, ads, etc, from the main body of text of an article formatted as HTML. The main content typically accounts for only a small portion of a page’s source code (highlighted in red in the image below).

Report generation method and system based on multi-agent architecture

The invention provides a report generation method and system based on a multi-agent architecture, and the method comprises the steps: firstly receiving and analyzing a report generation instruction, obtaining a theme demand, a framework specification and an initial reference material, then starting a multi-agent cooperation framework, generating an agent task distribution table, and generating an agent task distribution table; and the resource retrieval agent executes network resource directional retrieval according to the task allocation table to generate an associated resource set, and the content extraction agent performs structured conversion on the associated resource set and the initial reference material to obtain a structured content unit with chapter codes. The method comprises the following steps: performing module classification and logic series connection on a structured content unit by a report integration agent in combination with a framework specification to generate a report first draft, and finally performing content verification and optimization on the report first draft by a multi-agent collaborative framework to generate a final report text conforming to the framework specification, thereby realizing automation, intellectualization and high efficiency of report generation. And report quality is improved.
Owner:JIEHELIX (SHANGHAI) MEDICAL TECH CO LTD

Document content extraction method and system based on multimodal model collaboration, terminal and medium

The invention belongs to the technical field of document content extraction, and particularly discloses a document content extraction method and system based on multimodal model collaboration, a terminal and a medium. Comprising the following steps: identifying the type of an input to-be-processed document, and judging the document type; on the basis of the type identification result, calling a multi-modal model to analyze the document content, and outputting space coordinates, visual features and semantic features of document elements; generating a content sequence according with a reading habit through a semantic sequence reconstruction algorithm; paragraph boundary detection, paragraph recombination and semantic association modeling of charts and texts are completed based on the multilayer attention network and the graph neural network; grammar error correction, format optimization and title hierarchy generation are carried out by using a large language model and a hierarchical classification network; and converting the identification result into a structured output file. According to the method, the processing requirements of different types of documents can be considered, and high-precision analysis and efficient output are realized under the scenes of complex layouts, multiple languages and formula tables.
Owner:TUOSI (SHANDONG) INFORMATION TECHNOLOGY CO LTD

Enterprise data link treatment and value management method and system

The invention discloses an enterprise data link management and value management method and system, and the method comprises the steps: carrying out the real-time capturing of multi-source heterogeneous data of all business systems in an enterprise through a distributed data collection engine, and recognizing the types of original data in different formats through a preset data source adapter; if structured data is detected, a relational database connector is adopted for extraction, and if the structured data is recognized as unstructured data, a document analysis module is started for content extraction, and an initial data set containing metadata tags is obtained; performing format conversion and field mapping on the initial data set according to a pre-established data standardization rule base, eliminating duplicate records and abnormal values through a data cleaning algorithm, performing automatic evaluation on a data quality grade by adopting a naive Bayes classifier, if a data quality score is lower than a preset threshold value, triggering a data recovery process, and if the data quality score is lower than the preset threshold value, performing data recovery. And standardized data meeting a unified standard is obtained. The normativity and value utilization efficiency of data management are effectively improved.
Owner:FRIENDSHIP INT ENG CONSULTING CO LTD

File uploading attack interception method based on semantic entropy enhancement

The invention provides a file uploading attack interception method based on semantic entropy enhancement, and aims at overcoming the defects of an existing file uploading security protection technology in the face of complex attacks. The method specifically comprises the steps that S1, file format analysis and content extraction are conducted, hidden scripts are mined through nested content recognition, and intermediate representation is generated through grammar cleaning and coding specifications; s2, constructing an abstract syntax tree and semantic entropy calculation, tracking a pollution chain, analyzing a high-risk function, identifying a high-entropy character string, modeling and controlling flow complexity, and generating a semantic entropy vector; s3, dynamic scoring and decision making are carried out, and accurate judgment is carried out in combination with white list perception, feature comparison, multi-modal model scoring, adaptive threshold and sandbox observation; and S4, carrying out real-time interception and feature synchronization, blocking malicious file landing, generating an attack log and synchronizing an attack fingerprint. The method takes the semantic entropy vector as a core, breaks through the limitation of static features, remarkably improves the recognition rate of complex attacks, reduces missed judgment, and guarantees the safety of Web applications.
Owner:CHINA LIFE INSURANCE CO LTD

Full-text retrieval method and system fusing various types of documents

The invention provides a full-text retrieval method and system fusing various types of documents, and relates to the technical field of information retrieval, and the method comprises the following steps: obtaining document representation through document content extraction and structure recognition, generating a cross-modal semantic vector by using word embedding and nonlinear transformation, constructing a hierarchical index and a cross-document association graph, and obtaining a full-text retrieval result; the basic correlation score is calculated after the query request is received, and the comprehensive score of the candidate content segments is calculated based on the association graph to determine the optimal retrieval result, so that unified representation and retrieval of heterogeneous documents are realized, the cross-document retrieval precision and relevance are improved, and the processing capability of a retrieval system on complex queries is enhanced.
Owner:BEIJING CHANGFA TECH CO LTD

Knowledge extraction method and system based on semantic consistency evaluation and hybrid verifiable reward

The invention belongs to the technical field of artificial intelligence, and discloses a knowledge extraction method and system based on semantic consistency evaluation and hybrid verifiable reward, and the method comprises the steps: constructing a training data set with evidence labeling; based on the training data set, reinforcement learning training is carried out on a pre-trained large language model, a group strategy optimization GRPO algorithm is adopted, and model output is evaluated by using a mixed reward function; and on the basis of an evaluation result of the mixed reward function, updating parameters of a large language model so as to generate structured knowledge which is correct in format, accurate in content and provided with verifiable evidence. According to the method, in reinforcement learning training, effective decoupling format, content and credibility evaluation is realized, and refined feedback is provided for the model; when the model outputs knowledge, traceable original text evidence is provided for the model, so that the credibility and the interpretability of the model are enhanced, the accuracy of outputting the JSON format by the model is improved, and the accuracy and the integrity of the model in the aspect of content extraction are enhanced.
Owner:SHENZHEN WANGLIAN ANRUI NETWORK TECH CO LTD

Dynamic user portrait generation system and method based on intention-to-problem model

The invention discloses a dynamic user portrait generation system and method based on an intention-to-problem model, relates to the technical field of dynamic user portrait generation, and is used for solving the problem that dynamically changing user requirements cannot be adapted. Semantic features are extracted in combination with current dialogue content and matched with label embedding vectors, target labels are recognized, question texts are generated, the system extracts text, voice and visual multi-modal features after dynamically scoring based on dialogue rounds and semantic consistency and judging question insertion opportunities and user responses, response vectors are generated through fusion, and the question insertion opportunities and the user responses are obtained. Constructing a semantic tag vector by combining knowledge graph reasoning, finishing adaptive scoring between user response and a property tag, triggering incremental updating of the property tag when a scoring value is higher than a threshold value, counting the number of updating times, executing weight reconstruction if the number of updating times exceeds a total reconstruction threshold value, and optimizing a tag weight by combining a time attenuation mechanism; and the user portrait accuracy is optimized.
Owner:JI YIFENG (SUZHOU) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Bill content extraction method and device, equipment and storage medium

The invention discloses a bill content extraction method and device, equipment and a storage medium, and relates to the technical field of data processing, and the method comprises the steps: receiving a to-be-extracted bill file and extraction elements uploaded by a user; preprocessing the to-be-extracted bill file according to the file type to obtain a to-be-extracted image-text; according to the extraction elements and the bill type template, generating adaptive cue words; and inputting the cue word and the to-be-extracted image-text into a preset visual language model to obtain bill content corresponding to the to-be-extracted bill file. According to the method, the to-be-extracted bill file is preprocessed according to the file type, so that the complex bill layout is simplified, and the to-be-extracted image-text is obtained; according to the method and the device, the extracted elements are extracted, then the flexibly defined cue words capable of being matched with the document features are generated according to the extracted elements and the bill type template, so that the situation that a complex bill file is difficult to adapt by depending on a fixed template can be avoided, finally, the adaptive bill content can be extracted based on the cue words, and the extraction accuracy of the bill content is improved.
Owner:CHINA MERCHANTS BANK

Intelligent routing and unified adaptation method for large language model

The invention discloses an intelligent routing and unified adaptation method for a large language model, and the method comprises the steps: defining all access details of the model through a declarative configuration file, and achieving the zero-code access of the large language model without writing any code for a newly-added model; a completely consistent calling interface is provided for an upstream application, the isomerism of all downstream large language models is shielded, and a unified request and response abstraction layer is constructed; through a strategy engine and a JSON path technology, complex streaming response including content thinking is precisely processed, and intelligent analysis and content extraction are carried out; dynamic configuration and intelligent strategy hot update are supported; intelligent routing of the model is realized, and an optimal large language model instance is dynamically selected; and meanwhile, enterprise-level governance capability is provided, governance functions such as fusing, degradation, current limiting and monitoring are integrated, and stability guarantee is provided for model calling. Therefore, the maintainability, the expandability and the user experience consistency of the system are comprehensively improved.
Owner:NANJING INFORMATION HIGH-SPEED RAILWAY RES INST OF SCI AND TECH

Intelligent table filling method and system based on deep semantic analysis

The invention relates to the technical field of information extraction, in particular to an intelligent table filling method and system based on deep semantic analysis. The intelligent table filling method based on deep semantic analysis comprises the following steps: performing content extraction and format standardization on an input document to form a standard document text; a large language model is adopted, deep semantic analysis is executed on a standard document text, an entity of the document is recognized, and corresponding semantics are analyzed; analyzing semantics, position coordinates and logic partitions of fields in the target Excel template; establishing a corresponding semantic mapping relationship based on the semantic similarity between the entity of the document and the field of the target Excel template in the vector space; and based on the semantic mapping relationship, executing a hierarchical filling operation according to a semantic similarity threshold. According to the intelligent table filling and reporting method and system based on deep semantic analysis, the accuracy of automatic filling and reporting is improved through format verification, logic consistency check and integrity verification.
Owner:BEIJING YUNJIANXIN TECH CO LTD

Large model output content security test method and device

The invention relates to the field of large model security testing, and particularly provides a large model output content security testing method and device, and the method comprises the following steps: S1, preparing and managing a test set, a sensitive word library and a regular expression which are required by testing; s2, reading a test set, and obtaining a large model output result according to the test set and the large model interface information; s3, judging whether the output content of the large model is safe or not according to the sensitive lexicon and the regular expression; s4, extracting semantic risk features according to the output content of the large model by using the large model and the oriented Prompt, and automatically storing the semantic risk features after confidence verification; and S5, storing the information result of each request in a file. Compared with the prior art, the test time can be shortened, and the evaluation efficiency can be improved; and the security of the output content of the large model can be effectively evaluated by using a method for dynamically constructing the sensitive word bank by using the output result of the large model.
Owner:INSPUR QILU SOFTWARE IND

Logic rule-based relative support and confidence for semi-structured document content extraction

One method includes extracting word-elements, each corresponding to a respective element of a ground truth cell-item array from an annotated document, applying logic rules to the extracted word-elements so that the applicability, or not, of each logic rule to each element of the ground truth cell-item array is determined. Based on the applying of the logic rules, metrics are obtained that indicate, for each word-element of the annotated document, the applicability of the logic rules, and the frequency with which applicable logic rules is satisfied. A first aggregation process is performed that aggregates the metrics across a group of unstructured, and annotated, documents, and a second aggregation process is performed that aggregates the metrics regarding a model-generated cell item array that was created based on the group of annotated documents. Finally, respective outcomes of the first and second aggregation processes are compared so as to identify logic rules of interest.
Owner:DELL PROD LP

Text generation method and electronic equipment

The invention discloses a text generation method and electronic equipment, and relates to the technical field of deep learning, and the method comprises the steps: obtaining a plurality of candidate output texts corresponding to a text input by a user through a pre-training language model, splitting the plurality of candidate output texts into a plurality of statements, according to the method, each statement is extracted, each statement is scored to obtain the statement score of each statement, and then the target text is screened from the multiple statements according to the statement score of each statement, so that representative statements with focused semantics and rich information can be selected from the multiple statements, and trunk content extraction based on the statement scores is realized; and then the sequence-to-sequence generation model is utilized to perform optimization processing on the screened target text to obtain an optimized text, so that unified expression of the trunk content on language styles and sentence pattern structures is realized, and the accuracy of the finally generated text is improved. Therefore, the technical problem that it is difficult to ensure the accuracy of the output text by directly selecting the optimal candidate answers to determine the final output can be solved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Paper archive digital quality self-checking system based on AI detection

The invention discloses a paper archive digital quality self-inspection system based on AI detection, which relates to the technical field of archive management and comprises an image acquisition module, a content extraction module, a threshold adjustment module, a quality detection module, a result evaluation module, a feedback correction module and a data management module. The image acquisition module is responsible for acquiring images of paper archives by using a scanner or a digital camera and performing preliminary image preprocessing. According to the AI detection paper archive digital quality self-inspection system, unified execution of quality inspection standards is ensured through a rule engine, subjective differences of quality standard understanding of different quality inspection personnel are eliminated, and a detection threshold value is dynamically adjusted according to the age, the material and the storage state of the archive so as to adapt to quality inspection requirements of different archives; and particularly, the difficulty in identifying handwritings, faded handwritings, special symbols and the like in historical archives is optimized, so that the comprehensiveness and the efficiency of detection are greatly improved.
Owner:SUZHOU JIAHONG INFORMATION TECHNOLOGY CO LTD

Semantic anchor point definition method based on guide word and topological structure

The invention provides a semantic anchor point definition method based on a guide word and a topological structure. The semantic anchor point definition method is used for solving the problem that key fields in structured bills with the same semantic content but large format difference are difficult to position. The method comprises the following steps: constructing a semantic field and guide word dictionary; carrying out OCR identification on the bill image; identifying a matched guide word block in the text block; establishing a spatial adjacency relationship between the guide words; and finally generating an anchor point structure triple containing the semantic field (S), the adjacent structure (N) and the semantic content extraction direction (D). According to the method, the structured representation for expressing the semantic layout of the bills can be constructed, good cross-format adaptability is achieved, stable anchor point support is provided for follow-up format matching and field content extraction, and the method is suitable for processing scenes of multi-type heterogeneous bills such as customs production place certificates, customs declarations and invoices.
Owner:HUNAN UNIV

Image sensitive content auditing method and system

The invention discloses an image sensitive content auditing method and system, and the method comprises the steps: obtaining the original description of a to-be-audited image, and the original description comprises the text summarization of the visual content of the image and the text content in the image; extracting an entity area in the to-be-audited image, retrieving related external background information based on the to-be-audited image and the entity area, eliminating interference information by using a label of the entity area, and generating a retrieval enhancement result; semantic correlation screening based on sensitive topics is carried out on the retrieval enhancement results, the screened retrieval enhancement results are fused into the original description in an iteration mode, and final fusion description is generated; and constructing a training data set containing positive and negative judgment sample pairs based on the fusion description, performing fine tuning on the visual language model, and performing sensitive content auditing on a new image by using the fine-tuned visual language model. The method can effectively improve the recognition accuracy and reliability of the sensitive content of the image.
Owner:HANGZHOU YUNSHEN TECH CO LTD

Large model memory enhancement method based on time perception consistency feedback and optimization

A large model memory enhancement method based on time perception consistency feedback and optimization comprises the following steps: 1, data preprocessing: obtaining a historical interaction text generated by each agent in a target scene, and processing the text by using a preprocessing module; 2, attribute mining: extracting agent personal attributes, site attributes, logic attributes and core arguments from each generated content; and 3, keyword management and memory storage: maintaining an independent keyword historical file and a memory library file for each agent, and constructing an evolvable memory which changes along with time. 4, performing memory retrieval and context construction; 5, performing multi-dimensional consistency evaluation, and generating a comprehensive score; and 6, self-adaptive optimization is carried out, and self-feedback and self-evolution of model behaviors are realized. By means of the method, the memory ability and historical consistency of the multi-agent large language model can be effectively enhanced.
Owner:LIAONING UNIVERSITY

Intelligent marketing decision analysis method and system for real-time competitive product strategy

The invention provides an intelligent marketing decision analysis method and system for a real-time competitive product strategy, and the method comprises the steps: collecting the marketing content of a competitive product published on a social media platform, extracting an account number, an arriving person portrait feature, a content structure label, a platform identifier and publishing time, and constructing a standardized strategy behavior vector; performing time sequence alignment on the strategy behavior vectors and our marketing indexes, and identifying competing product strategy behavior subsets which have significant influence on the our indexes; based on the subset, generating a structured response strategy containing a recommended person type, a content structure label, a delivery platform and a time period; and performing multi-dimensional matching on the marketing resource library and the content template library, and outputting an executable marketing resource combination and scheduling instruction. According to the method, a closed loop from competitive product behavior perception and causal influence identification to response strategy automatic generation and landing execution is realized, and the real-time performance, the accuracy and the automation level of marketing response are improved.
Owner:GUANGZHOU YUNZHIDACHUANG TECH CO LTD

Demand specification generation method and system based on multi-modal understanding

ActiveCN121525656AText processingKnowledge based modelsLinguistic modelSpecification document
The invention provides a demand specification generation method and system based on multi-modal understanding, and relates to the field of artificial intelligence, and the method comprises the steps: carrying out the content understanding operation of data of each modal through employing a corresponding content extraction tool, so as to generate a demand representation document after multi-modal unification; establishing a demand knowledge base according to an existing template base and a corresponding rule base, matching and comparing the demand knowledge base with the demand representation document, checking the integrity of the demand representation document, and generating a question for seeking a decision; creating an overall framework of the standard document according to the standard standard template, and performing content filling on the overall framework according to the demand representation document and an answer to the question seeking the decision to generate the standard document; and collecting feedback information for the standard document, analyzing the feedback information by using a large language model, and converting the feedback information into executable modification suggestions for the standard document. According to the method, the automation degree, the integrity and the accuracy of demand specification generation are improved.
Owner:ZHIJIA ARTIFICIAL INTELLIGENCE TECH (TIANJIN) CO LTD

Big model-based document content customized extraction method and system

The invention relates to a document content customized extraction method and system based on a large model. The method comprises the following steps: inputting a document, and dividing the document into a text document or an image document; formulating a structured cue word according to the document and user requirements; calling a large model to process the document, and performing content extraction based on the structured cue word to generate a preliminary extraction result; comparing the preliminary extraction result with the extraction target, updating the structured cue word for iterative optimization, and recalling the large model until the extraction result meets a precision threshold value; and confirming that the extraction result meeting the precision threshold meets an extraction target and storing the extraction result as a final extraction result, and storing the final extraction result as an output format. According to the method, efficient extraction of multi-modal document content information is achieved through large model calling, customized extraction is achieved through structured cue words, the output precision is improved through an interactive process, a large amount of time cost is saved for a user, and remarkable achievement benefits are brought.
Owner:INNOVATION ACAD FOR MICROSATELLITES OF CAS +1

Intelligent image batch generation method and device, equipment and storage medium

The invention discloses an intelligent image batch generation method and device, equipment and a storage medium, and the method comprises the steps: receiving an Excel data file uploaded by a user, carrying out the table structure analysis and data content extraction of the Excel data file, and obtaining target data; obtaining a design template selected by a user, and establishing association between the target data and elements in the design template by adopting a binding algorithm to obtain a binding relation mapping table; an AI image processing function is called to process the bound target data to obtain processed data, and the AI image processing function comprises at least one of AI sectional drawing, AI image sharpening, AI document generation and AI text drawing; based on the binding relation mapping table, applying the processing data to the bound design template to obtain an initial image; and performing image rendering on the initial image through a rendering engine to generate a target image. According to the method, various AI tools are integrated, so that the processing efficiency is remarkably improved.
Owner:BEIJING WEZONET NETWORK TECH CO LTD

Generative interface for multi-platform content

Embodiments described herein relate to systems and methods for automatically generating content for a generative answer interface of a collaboration platform. The system receives a natural language user input identifying corresponding blocks of text or snippets using a content extraction service. A prompt is generated using the blocks of text and is used to obtain a generative response. The generative response and links to corresponding content are displayed in the generative answer interface and can be inserted into content of the collaboration platform. The systems and methods described use a network architecture that includes a prompt generation service and a set of one or more purpose-configured large language model instances (LLMs) and / or other trained classifiers or natural language processors used to provide generative responses for content collaboration platforms.
Owner:ATLASSIAN PTY LTD

Multi-source document management method and device based on knowledge construction and fusion storage

The invention discloses a multi-source document management method and device based on knowledge construction and fusion storage, and relates to the technical field of document management, and the method comprises the steps: receiving a multi-source document to an object storage system, and recognizing the document type; performing document analysis and structure extraction according to the document type; extracting pictures in the image-text mixed content, uploading the pictures to an object storage system, generating a mapping dictionary, and inserting picture marks in a text part of the mapping dictionary; carrying out structured processing on contents of the table key value pairs to extract summaries and abstracts; performing standardization processing and semantic slicing on the processed image-text mixed content, table key value pair content and / or plain text content to generate knowledge fragments; constructing a knowledge extracting questions from knowledge fragments; the knowledge fragments are stored in a first index, and the questions and IDs of the corresponding knowledge fragments are stored in a second index; generating vectors for each knowledge fragment and question, embedding and writing the vectors into a vector database; and performing document management based on the knowledge base. The document management efficiency is improved.
Owner:DIGITAL CHINA SYST INTEGRATION SERVICE

PDF document intelligent identification and content extraction method based on deep learning

The invention discloses a PDF document intelligent identification and content extraction method based on deep learning, and relates to the technical field of artificial intelligence, deep learning, computer vision and document image processing, and the method comprises the steps: obtaining a positioning table region of each table in a PDF whole page image; obtaining a basic grid structure; cells with cross-row or cross-column structures are obtained; performing consistency detection and repair on the cells with the cross-row or cross-column structure by using a structure verification network to obtain a repaired table structure; and performing text recognition on each logic cell in the repaired table structure, and binding row and column position information corresponding to each logic cell to obtain table content which can be output in a preset structured format. According to the method, various forms of PDF tables such as scanners and pictures can be effectively processed, different table styles, fonts and backgrounds are adapted, the requirement on the quality of the input image is reduced, and high-precision table recognition and content extraction are ensured.
Owner:ZHONGSHAOXUAN TECHNOLOGY GROUP CO LTD

Method for identifying whether authenticity of news is changed or not and related equipment

The invention provides a method and related equipment for identifying whether the authenticity of news is changed or not, and relates to the technical field of news authenticity identification. Each piece of news in an obtained target news set is judged to obtain clarified news; extracting clarified data from the clarified news by utilizing the constructed content extraction cue word; taking the clarified data as retrieval content and / or a text of the clarified news to generate a retrieval vector, and retrieving news related to the retrieval content and / or the retrieval vector from the historical news to obtain a candidate news set; analyzing the content consistency and content conflict between all news in the candidate news set and the clarified news, judging whether news corresponding to the clarified news exists in the candidate news set or not, if the news corresponding to the clarified news exists in the candidate news set, considering that the authenticity of the news corresponding to the clarified news is changed, and otherwise, judging that the authenticity of the news corresponding to the clarified news is changed. Otherwise, considering that no news with changed authenticity exists in the candidate news set; and the authenticity identification capability of the news is improved.
Owner:HUNAN INST OF INFORMATION TECH

Automatic electronic chart document review method and device based on multi-modal large model, medium and equipment

The invention discloses an automatic electronic chart document review method and device based on a multi-modal large model, a medium and equipment, and the method comprises the steps: constructing a review rule expression and a knowledge graph corresponding to each key entity, generating a document review key point expression for recognizing the key entity based on the knowledge graph, constructing a rule set based on the corresponding relation of the knowledge graph, the review rule expression and the document review key point expression; performing content extraction on the obtained to-be-examined document to obtain to-be-examined text data; extracting a key entity from the to-be-examined text data; and performing review verification on the key entity based on the review rule expression to obtain a review result. According to the method and the device, the review rule expression and the knowledge graph corresponding to the required key entity can be constructed through the pre-acquired rule document, so that the method and the device can be suitable for review requirements of different fields and different types of documents, and have better flexibility.
Owner:SHENZHEN TIANHAI CHENGUANG TECH CO LTD

Internet webpage content feature extraction method based on artificial intelligence

The invention discloses an Internet webpage content feature extraction method based on artificial intelligence, which comprises the following steps: S1, acquiring a webpage HTML source file and a rendering image, and preprocessing the webpage HTML source file and the rendering image; s2, performing semantic classification on DOM nodes, encoding the DOM nodes into three types of identifiers, and constructing a node label sequence; s3, performing time sequence synchronization on the node attribute vector, the node tag sequence and the visual area set, and performing block-level slicing; s4, inputting the block-level slices into a gated recursive attention network, extracting multi-modal joint representation, and executing attention aggregation; s5, performing structure alignment on the feature fusion sequence, calculating a cross-node consistency distance, and screening a target slice set with high structure cohesion; and S6, mapping the target slice set to a content feature space, and generating a content feature tag set. According to the method, the structural accuracy, the semantic integrity and the multi-modal fusion precision of webpage content extraction are improved.
Owner:NANJING YUANPENG SOFTWARE TECHNOLOGY CO LTD

Code checking method and device, electronic equipment and storage medium

The embodiment of the invention provides a code checking method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence application, and the method comprises the steps: obtaining a code file and a project document, and carrying out the content extraction of the project document through a literary and physical agent, the method comprises the steps of obtaining function requirement information and programming specification requirements for a code file, conducting code analysis on the code file through a programming agent according to the function requirement information and the programming specification requirements, extracting a first code block needing to be modified from the code file, and outputting a code suggestion for the first code block, the code suggestions are evaluated through the review intelligent agent and the literary and sports intelligent agent, the evaluation result for the code suggestions is obtained, and the review report file corresponding to the evaluation result is output, so that the comprehensiveness and accuracy of inspection are improved, the inspection efficiency is effectively improved, developers can visually and effectively learn based on the review report file, and the development efficiency is improved. And the programming capability is improved.
Owner:LEAYUN TECH CO LTD OF ZHUHAI +1

Multi-modal mixed document OCR (Optical Character Recognition) and structured extraction method

The invention relates to the technical field of content extraction, in particular to a multi-modal hybrid document OCR (Optical Character Recognition) and structured extraction method, which comprises the following steps of: acquiring an image text region bounding box and classifying a style, extracting a font or stroke sequence to generate a character positioning structure, dividing paragraph and sentence groups to classify semantic fields, and calculating a field matching relationship to generate structural mapping. According to the method, logic mapping is constructed through character two-dimensional coordinate sorting and paragraph contours, complete reconstruction of a page structure is enhanced, semantic field categories are extracted by using syntactic density of sentence paragraph division and inter-paragraph features, the accuracy of field classification is improved, and the method has the advantages of being simple in structure, convenient to operate and high in practicability. A field mapping relation is established through Jaccard similarity and part-of-speech consistency analysis between a head word and a standard field keyword, a field path index and a structure node link are clarified, field semantic affiliation and structure position output are unified, and the document structure reduction degree and field extraction accuracy are improved.
Owner:HANGZHOU JINGSHENG HANGXING TECH CO LTD

Content extraction and delivery method based on multi-mode quotient single video

The invention relates to the technical field of big data analysis, and particularly discloses a content extraction and delivery method based on a multi-modal business order video, and the method comprises the steps: calling a preset domain name-brand mapping library to output a preliminary brand identifier or extract to-be-verified brand information, and obtaining directional multi-modal features; determining a core brand based on the initial brand identifier or the directional multi-modal feature, the to-be-verified brand information and a brand vector library, and performing standardization processing to obtain a standardized brand identifier; based on the standardized brand identifier, in combination with the characteristics of a platform to which the business order video belongs and target user portrait data, calling a delivery strategy library to generate a directional delivery strategy adapted to a brand type, pushing the business order video to a target user group according to the directional delivery strategy, and monitoring a delivery effect index in real time to obtain monitoring data; adjusting the current multi-dimensional weight according to the monitoring data, and generating an effect brief report; the brand identification accuracy in the business bill video is improved, and the delivery strategy can be continuously optimized.
Owner:SHANGHAI XINBANG INFORMATION TECH CO LTD