Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

910 results about "Annotation" patented technology

An annotation is extra information associated with a particular point in a document or other piece of information. It can be a note that includes a comment or explanation. Annotations are sometimes presented in the margin of book pages.

Heterogeneous document set-oriented cross-modal semantic alignment and logic consistency verification system

The invention relates to document verification, in particular to a heterogeneous document set-oriented cross-modal semantic alignment and logic consistency verification system, which is used for heterogeneous document input, supports multi-format document input and comprises multi-modal elements including texts, pictures, tables and charts. The document analysis module is used for carrying out structured extraction on document contents; extracting multi-modal elements, identifying and classifying various elements in the document, and establishing position and type labels of a foundation; the knowledge graph construction module is used for uniformly modeling heterogeneous elements into a multi-modal knowledge graph; the graph neural network semantic alignment module is used for realizing accurate cross-modal semantic alignment by using a specially designed graph neural network based on the multi-modal knowledge graph; the hybrid consistency verification engine is used for performing logic consistency verification in combination with a symbol logic verification mechanism and a semantic consistency verification mechanism; according to the method, the defect that accurate cross-modal semantic alignment and logic consistency verification are difficult to carry out on professional documents with multi-modal elements can be effectively overcome.
Owner:ANHUI GAOSHAN TECH CO LTD

Knowledge question and answer library agent construction method and system

The invention provides a knowledge question and answer library agent construction method and system. Efficient knowledge management and question and answer are achieved through cooperation of multiple agents. According to the system, firstly, a multi-modal information extraction agent is constructed, and heterogeneous data such as texts, images and tables are converted into structured vectors and stored; meanwhile, the knowledge graph is dynamically constructed and continuously optimized by the self-adaptive knowledge graph construction agent, and a new relationship is derived through combination of symbolic logic and a graph neural network, so that an evolvable knowledge network is formed. In the question and answer stage, a query analysis agent deeply analyzes the intention of a user and generates sub-queries; retrieving the vector library and the knowledge graph in parallel by the retrieval enhancement generation agent; and the reasoning and synthesizing agent integrates multi-source information and generates an accurate answer with a complete source label through a large language model. Dynamic knowledge management, precise semantic analysis and system self-evolution are achieved, and the method is particularly suitable for professional field scenes needing high-reliability questions and answers.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Penetration test automation method and device based on large language model and ATTCK framework

The invention discloses a method based on a large language model and ATTamp; the invention discloses a CK framework penetration test automation method and device, and the method comprises the steps: firstly carrying out the structural analysis of multi-source input information and tool output, and guaranteeing that key fields are not discarded; then combining a retrieval enhancement generation technology and a network security knowledge base to provide domain knowledge support for the large language model, so as to generate a model with ATTamp; a penetration test task tree marked by CK tactics, technologies and sub-technologies; on the basis, an optimal tool is automatically selected through a tool resource library and a multi-dimensional screening mechanism, an execution instruction is generated, and finally an execution result is returned to the input analysis module to form a self-adaptive optimization test closed loop. According to the method, semantic fidelity compression and standardization processing of long information can be realized aiming at the problems of large output format difference, more information redundancy and the like of different penetration testing tools, and efficient, explainable and auditory technical support can be provided for automatic penetration testing in a complex network environment.
Owner:GUANGZHOU UNIVERSITY

Fine-tuning system for large language models trained for open-ended domain-specific tasks

There are provided systems and methods for a fine-tuning system for large language models trained for open-ended domain-specific tasks. An online transaction processor or other service provider may provide computing services and platforms to entities, which may include chatbots, information retrieval systems, question-and-answer systems, and the like. To provide better LLM training and fine-tuning, which may improve LLM performance in answering users' questions in an automated manner, the service provider may implement a fine-tuning system that may utilize automated annotations of training data, such as query and response pairs. An LLM may be prompted to determine an annotation to such pairs, and the annotations may be used to label the training data. A fine-tuning system and operations may then be implemented to fine-tune the LLMs using different processes including question-answering, retrieval augmented generation, or a continuous fine-tuning based on a size of the training data.
Owner:PAYPAL INC

Digital archive intelligent processing method, storage medium and system

The invention relates to a digital archive intelligent processing method, a storage medium and a system, which are suitable for multi-source heterogeneous archive management scenes such as colleges and universities. The method comprises the following steps of: classifying structured and unstructured data such as paper archive scanning pieces and database views, and extracting metadata and entity information by adopting a scanning and OCR (Optical Character Recognition) technology; a complex table and document content are analyzed through a model, semantic analysis (entity recognition, relation extraction and event abstract) is achieved in combination with a language model of a Transform architecture, and a structured report containing a data abstract, an entity relation graph and abnormal annotations is generated. The system is internally provided with a parameter template automatic generation module, supports cross-page content continuous restoration and sensitive data encryption desensitization, and realizes safe sharing through an API interface. The method solves the problems of low efficiency, difficulty in multi-source data fusion and the like of traditional archive processing, improves the automation level and data value mining capability of archive management, and is suitable for intelligent upgrading of complex archive scenes.
Owner:CHINA AGRI UNIV

Power system equipment image anomaly detection and quality diagnosis method based on multi-modal visual language model

The invention discloses an electric power system equipment image anomaly detection and quality diagnosis method based on a multi-modal visual language model. The method comprises the following steps: constructing a large-scale multi-modal data set comprising an electrical equipment image, an object detection annotation, a pairing question and answer knowledge base and an official supervision document, constructing a basic diagnosis model based on a visual language model, and carrying out instruction tuning; carrying out post-training on the model by adopting group relative strategy optimized reinforcement learning, and generating an interpretable step-by-step diagnostic reasoning chain; in the reasoning process, related knowledge is dynamically retrieved based on a retrieval enhancement generation technology of a graph structure, and the accuracy and compliance of a diagnosis decision are enhanced; and finally, generating a diagnosis report containing the exception type, the root cause and the decision suggestion. Compared with a traditional method, the method solves the three problems of data scarcity, opaque reasoning and knowledge isolation in the field of electric power detection, and has the remarkable advantages that the diagnosis process can be explained, complex multi-step reasoning is supported, and domain knowledge can be dynamically integrated.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL +1

Cross-domain equipment fault diagnosis method and system based on cooperation of large and small models

The invention provides a cross-domain equipment fault diagnosis method and system based on large and small model cooperation, and relates to the technical field of equipment fault diagnosis. According to the method, the causal field generalization structure is introduced into the small model, explicit decomposition is carried out on the stable causal law and the field specific difference, and meanwhile, the causal field generalization structure is corrected by using the large model, so that the small model can automatically identify and retain the causal relationship which is universally applicable to each device and each field; therefore, the influence of inter-domain distribution difference is effectively eliminated. Theoretical analysis shows that the generalization error of the model mainly depends on the accuracy of the stable causal item, and the structure can minimize error drift caused by distribution drift. Therefore, the robustness of health state evaluation and fault prediction can be remarkably improved in a cross-domain scene, and the fault diagnosis model can still keep the prediction capability close to the training domain level under the condition of no target domain annotation data.
Owner:HEFEI UNIV OF TECH

Intelligent cell type annotation method based on key marker gene

The invention discloses a key marker gene-based intelligent cell type annotation method, which comprises the following steps of: constructing a static knowledge base by using known marker genes in a reference database, and endowing the marker genes with cell specific weights by using a TF-IDF method, so that the annotation accuracy and interpretability are improved. Meanwhile, under the condition that static matching is insufficient, the literature is understood through a large language model, mark information is extracted, dynamic completion of the knowledge base is achieved, the defect that updating of a traditional knowledge base is lagged is overcome, and good adaptability and expansibility are achieved. Besides, static and dynamic matching scores are fused in the annotation process, so that more robust cell type identification is realized, annotation requirements of multi-tissue, multi-species and novel cell states are adapted, high-precision and extensible cell type annotation can be realized in a scene with insufficient reference knowledge or a fuzzy sample, and the annotation efficiency is improved. And the method has good universality and practicability.
Owner:ZHEJIANG UNIV +1

Interface verification method and device based on dynamic rule, medium and program product

The embodiment of the invention provides an interface verification method and device based on a dynamic rule, a medium and a program product, and relates to the technical field of parameter verification. The method comprises the following steps: before calling interface processing, acquiring a user-defined annotation associated with an interface; determining a to-be-verified field and corresponding rule index information based on the user-defined annotation; acquiring a corresponding target verification rule from a rule cache list based on the rule index information; wherein the rule cache list is loaded and cached on the basis of a dynamic rule base under the condition of compiling initialization or rule change; and verifying the corresponding field to be verified based on the target verification rule. According to the embodiment of the invention, a cache rule-oriented verification mode is adopted to dynamically adapt to the business change of the rule base, so that the interface verification efficiency in a dynamic business scene can be effectively improved, and the maintenance cost can be reduced.
Owner:BEIJING TOPSEC NETWORK SECURITY TECH +2

Continually evaluating and modifying artificial intelligence assistant

The present disclosure relates to systems, non-transitory computer-readable media, and methods for generating modifications to an LLM based artificial intelligence assistant based on classifying the severity of errors and focusing the modifications on resolving high-severity errors. In particular, the disclosed systems receive prompts via an artificial intelligence assistant graphical user interface and generate responses to the prompts using the LLM based artificial intelligence assistant. Further, the disclosed systems determine errors in the responses using an annotation tool to generate annotated errors and an error analysis mechanism to generate indications of the errors based on the annotated errors. Additionally, the disclosed systems classify the errors as one of high-severity, mid-severity, or low-severity. Moreover, the disclosed systems generate modifications to components of the LLM based artificial intelligence assistant based on the high-severity errors.
Owner:ADOBE INC

Text data extraction method, system and equipment based on multi-modal fusion and self-evolution learning and medium

PendingCN121390035ASemantic analysisText processingLearning machineEvolutionary learning
The invention relates to the technical field of text data processing, and discloses a text data extraction method, system, equipment and medium based on multi-modal fusion and self-evolution learning, which comprises the following steps of: performing feature extraction and spatial alignment on a printed text, a handwritten annotation and a dynamic table of a mixed format document to obtain a semantic feature of an image-text table, and inputting the semantic feature into a dynamic analysis layer; analyzing metaphor expressions and synonymous heterogeneous fields through field extraction and a context semantic reasoning mechanism, and outputting structured data; performing grammar compliance verification by adopting a regularization engine, and performing comparison verification through a federal learning mechanism; and inputting the verified data into the reinforcement learning model, updating the analysis rule and the model parameters through strategy iteration, and feeding back the updated analysis rule and model parameters to the dynamic analysis layer to complete closed-loop optimization. According to the method, the processing precision and efficiency of the complex document are greatly improved, the manual intervention requirement is remarkably reduced, and meanwhile, the privacy protection and compliance requirements are met.
Owner:YUNNAN ELECTRIC POWER TESTING & RES INST (GRP) CO LTD +1

Defect detection data screening method based on Pontryagin maximum principle

The invention relates to the field of computer vision algorithms, in particular to a defect detection data screening method based on a Pontryagin maximum principle, which comprises the following steps of: training a training data set and testing an evaluation data set to obtain a model reference index; based on a preset proxy data set, respectively calculating definition, labeling integrity and distribution deviation degree indexes to obtain an initial quality weight vector; iteratively updating the basic model parameters and reversely iteratively updating the target vector to obtain a sample quality score; a training scoring device scores and sorts full samples of the training data set; and dividing the sorted full samples into a plurality of candidate screening intervals, and screening high-quality data to train a final model. According to the method, the entropy weight method is adopted to weight the three dimensions to obtain the initial quality weight, so that the initial weight of the sample can reflect the own basic quality difference, the problem that the traditional uniform weight ignores the sample quality difference is avoided, and the accuracy of the sample quality score is further improved.
Owner:苏州深视信息科技有限公司

Automatic generation method of health science popularization article

The invention discloses an automatic generation method for health science popularization articles, and belongs to the crossing field of computer technology and health science popularization. The method comprises the following steps that S1, health authority information is collected through a multi-source crawler, and an original material pool is obtained through Sension-BERT vectorization duplicate removal and BioBERT medical entity labeling; s2, classifying nine types of health themes by using a RoBERTa fine tuning model; s3, multiple AI model supplementary materials are made into a structured package; s4, predetermining optimal selection questions in combination with AI three-dimensional scoring and editing; s5, calling the multi-dimensional knowledge base to generate a professional knowledge packet; s6, generating an outline and performing three-dimensional five-score system auditing of knowledge point fullness and the like; s7, the LLM expands and writes the text and marks knowledge sources; s8, performing double-stage auditing; s9, anthropomorphic draft moistening; s10, intelligently illustrating the picture; s11, when the matching degree is smaller than a preset threshold value, automatically updating the knowledge base; and S12, integrating and outputting. According to the method, the health science popularization article is generated, the manual participation time is shortened to be within 30 minutes, the medical error rate is reduced to be below 5%, the daily output of a 10-person team is improved by 3-5 times, the annual human cost is reduced by 60% or above, and the large-scale and high-quality science popularization requirements are met.
Owner:GUANGZHOU FAMILY DOCTOR ONLINE HEALTH MANAGEMENT SERVICE CO LTD

Address data matching method and related equipment

The embodiment of the invention provides an address data matching method and related equipment, and belongs to the technical field of geographic information services. The method comprises the following steps: constructing an address annotation corpus according to input address information data and a preset address database; the method comprises the following steps: generating a geographic information embedding vector according to a preset geographic information knowledge graph, performing address element analysis in combination with an address annotation corpus to obtain an address element sequence so as to construct a dictionary tree, and performing similarity screening through a spatial hierarchical matching algorithm to obtain a similar address set; generating an address embedding vector matrix through a preset word embedding vector model, and performing feature extraction through a preset semantic feature extraction model to obtain semantic-level similar features; according to input address information data, multi-dimensional character similarity matching is carried out to obtain character-level similar features, then weighted fusion is carried out in combination with semantic-level similar features, and target matching address data is determined according to a weighted fusion result. According to the embodiment of the invention, the address data matching accuracy and efficiency can be improved.
Owner:CHINA TELECOM CORP LTD

File positioning management method and system based on artificial intelligence

The invention provides a file positioning management method and system based on artificial intelligence, and relates to the technical field of artificial intelligence. According to the method, files are collected from multiple sources and subjected to standardization processing, an element set is generated in combination with multi-modal analysis of texts, images, audios, videos and tables, cross-modal alignment is achieved through a semantic representation model, hierarchical indexes of semantics, keywords and relations are constructed, and a unique traceability identifier is generated; in the query stage, intention recognition and joint retrieval are carried out, a result subjected to permission verification and traceability information labeling is output, online optimization and incremental reconstruction are executed based on user feedback, and comprehensiveness, accuracy, traceability and self-adaptive optimization of file positioning are achieved.
Owner:ZUNYI NORMAL COLLEGE

Domain-specific processing and information management using machine learning and artificial intelligence models

Systems and techniques are provided for automatically analyzing and processing domain-specific image artifacts and document images. A process can include obtaining a plurality of document images comprising visual representations of structured text. An OCR-free machine learning model can be trained to automatically extract text data values from different types or classes of document image, based on using a corresponding region of interest (ROI) template corresponding to the structure of the document image type for at least initial rounds of annotations and training. The extracted information included in an inference prediction of the trained OCR-free machine learning model can be reviewed and validated or corrected correspondingly before being written to a database for use by one or more downstream analytical tasks.
Owner:32HEALTH INC

Large language model auxiliary vulnerability detection method and system based on abstract syntax tree decomposition and annotation enhancement

The invention discloses a vulnerability detection method and system based on abstract syntax tree decomposition and large language model assistance. The vulnerability detection method and system are used for solving the problem that an existing pre-training model is insufficient in detection accuracy under complex code logic and multiple execution paths. The method comprises the following steps: firstly, analyzing a code snippet into an abstract syntax tree, splitting the abstract syntax tree into a plurality of sub-trees through an improved decomposition algorithm, and combining each sub-tree with a natural language annotation generated by a large language model to form an abstract sub-tree with the annotation; then, a semantic aggregator based on Transform is used for modeling the relation between the sub-trees, features are fused to a target vulnerability vector, and finally, vulnerabilities are predicted through a classifier. Based on the technical scheme, the vulnerability detection accuracy is effectively improved, and the performance of the vulnerability detection model is greatly improved.
Owner:HUNAN UNIV OF SCI & TECH SANYA RES INST

Multi-mode self-supervision abnormal mode detection method and system

The invention relates to the technical field of artificial intelligence, discloses a multi-modal self-supervision abnormal mode detection method and system, and aims to solve the problems that in the prior art, a large amount of annotated data is relied on, modal fusion is insufficient, the anomaly discrimination ability is weak, and the dynamic environment is difficult to adapt. The method comprises the following steps: synchronously acquiring videos, audios, sensor time sequences and log text data, and carrying out time alignment; extracting spatial-temporal characteristics of each modal through a modal specific encoder; constructing a contrast learning task under a label-free condition, generating positive and negative sample pairs by utilizing data enhancement, and driving model learning discriminative representation; and cross-modal feature alignment and dynamic weighted fusion are realized by adopting an attention mechanism, and joint representation is generated. According to the scheme, efficient anomaly detection without annotation data is realized, the multi-modal fusion representation capability is remarkably improved, the false alarm rate is reduced, the environmental adaptability is enhanced, and the real-time monitoring requirement is met.
Owner:SHANGHAI SHENTONG YUANENG TECHNOLOGY CO LTD

Multi-modal large model reasoning method and system based on self-driven feedback and symbol collaboration

The invention discloses a multi-modal large model reasoning method and system based on self-driven feedback and symbol collaboration, and the method comprises the steps: carrying out the structural representation of multi-modal information through a knowledge graph, carrying out the entity recognition and relation extraction in combination with a large model, constructing a unified knowledge graph, and generating a knowledge ternary set; defining a symbol logic expression, constructing a diversified symbol logic rule by using the knowledge ternary set, and calculating a symbol consistency award of a reasoning path; constructing a symbol-human feedback collaborative reward mechanism to obtain a mixed reward function; in the process of interacting with the multi-modal environment, sampling a group of outputs for specific tasks, and constructing an intra-group relative reward optimization strategy network objective function in a multi-task scene in combination with a mixed reward function; the system interacts with the environment to realize autonomous evolution cycle to generate a training sample, and iterative cycle realizes self-driven feedback without a large amount of manual annotation data; and the logicality, the interpretability and the autonomous evolution ability of the multi-modal large model in a complex reasoning task are promoted.
Owner:XI AN JIAOTONG UNIV

Strategy model training method and device, strategy generation method and device, storage medium and terminal

The embodiment of the invention discloses a strategy model training method and device, a strategy generation method and device, a storage medium and a terminal. Firstly, a set of sample answers are generated for a target question through a preset large language model to serve as references. And during training, inputting an answer generated by the current strategy model and a sample answer into the evaluation model for comparison, and outputting a good and bad sorting result. The ranking is quantized as a reward value whose size is positively correlated with the degree to which the current answer is superior to the sample answer. And finally, the system adjusts strategy model parameters according to the reward value and guides the strategy model parameters to be continuously optimized. According to the method, scoring according to standard answers is replaced with relative sorting, and the problem that open questions lack clear judgment standards is solved. The method has the beneficial effects that the training data cost and labeling dependence are remarkably reduced, so that the model can realize stable and autonomous efficiency improvement in the vertical field, and a self-driven benign evolution cycle is formed.
Owner:RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD

Multi-dimension-based defect detection labeling quality automatic evaluation method and system

The invention relates to the field of defect detection, in particular to a multi-dimension-based defect detection labeling quality automatic evaluation method and system, and the method comprises the following steps: constructing an initial domain knowledge base, initializing a severity weight, a dynamic reliability weight and a multi-dimensional smoothing coefficient, and loading a pre-training defect detection model; inputting a defect sample batch to be evaluated, iteratively training the defect detection model, and calculating a multi-dimensional original evaluation index; obtaining an evaluation dimension set through dimension reduction and standardization processing, and generating a dynamic tracking result of each dimension index by adopting an index moving average algorithm in combination with a smoothing coefficient; in combination with the severity weight and the dynamic reliability weight, a comprehensive mark quality score is obtained through weighted fusion, and suspected error mark samples are screened out; and performing iterative training on the defect detection model, and outputting a final result. According to the invention, continuous optimization of evaluation parameters is realized through a man-machine cooperative feedback closed loop, and the practicability and stability of the technical scheme are further enhanced.
Owner:苏州深视信息科技有限公司

Convolutional neural network-based building plane element recognition model construction method

The invention relates to the technical field of image recognition, in particular to a building plane element recognition model construction method based on a convolutional neural network, and the method comprises the following steps: obtaining an input APN sample; elements in the APN sample are labeled, and a labeled file is generated; converting the annotation file into a format required by YOLO, and normalizing the annotation file; performing data enhancement on the APN sample, expanding the training data volume, and outputting to obtain an element detection result and a feature vector; installing dependency and carrying out data configuration; setting command line model parameters and customizing training scripts; and carrying out model training and model reasoning. According to the method, the plane layout is abstracted based on the machine learning method, and the vector data type has high flexibility and deformation capability, so that the vector output which keeps the original plane image form unchanged can be converted into various objects according to the purpose of a user. And classifying and identifying the plane elements by using the GNN.
Owner:BEIJING UNIV OF CIVIL ENG & ARCHITECTURE

Multi-modal sample data generation method and device, electronic equipment and storage medium

The invention relates to the technical field of computers, in particular to a multi-modal sample data generation method and device, electronic equipment and a storage medium, and the method comprises the following steps: for a target webpage, obtaining a webpage screenshot and element information of UI elements in the webpage; generating annotation information based on the element information of the UI element; taking the annotation information and a preset task cue word as text input, taking the webpage screenshot as picture input, and inputting the text input and the picture input into a large language model to obtain description information which is output by the large language model and is used for describing the webpage screenshot; and taking the webpage screenshot and the description information as sample data for training the multi-modal large language model. According to the embodiment of the invention, the generation efficiency of the sample data of the multi-modal large language model can be improved.
Owner:MOORE THREADS TECH CO LTD

Medical instruction data generation and use method based on multi-modal knowledge graph

The invention discloses a medical instruction data generation and use method based on a multi-modal knowledge graph. Firstly, medical images and text information in a public data set are collected, and high-quality image-text pair data are screened out from the medical images and the text information; secondly, performing entity recognition on the medical text and retrieving related triple knowledge from the multi-modal knowledge graph to form knowledge enhanced medical image-text data; secondly, inviting a medical expert to label real scene question and answer pairs based on the knowledge-enhanced medical image-text data, and taking the real scene question and answer pairs as a seed instruction data set to guide a large-scale visual language model to generate diversified instruction question and answer data; and finally, using the medical image understanding and question-answering ability of the supervised fine-tuning training model, and performing reinforcement learning supervised by a knowledge graph triple path. According to the method, the problem of scarcity of medical instruction data can be effectively relieved, so that the trained model has higher accuracy and interpretability in a medical question and answer task.
Owner:EAST CHINA UNIV OF SCI & TECH

Picture sharing method and device based on AI platform, electronic equipment and storage medium

The invention relates to the technical field of picture sharing, and discloses a picture sharing method and device based on an AI platform, electronic equipment and a storage medium, and the method comprises the steps: obtaining a generated data stream of a picture uploaded by a first user, converting the generated data stream into a structured JSON file, and enabling a built-in cross-platform compatible tag to support multi-terminal sharing; based on the second user creation preference portraits and the key features, non-key parameters are screened through an AI semantic matching model, an adaptive priority adjustable component is generated, and a difference thermodynamic diagram and a new picture are generated after user adjustment; the enhanced copyright metadata containing original author identification, content fingerprints and the like are embedded into the JSON parameter stream; and calculating a hierarchical hash value and carrying out block chain evidence storage during each parameter iteration, calculating contribution degree distribution earnings based on a multi-dimensional index, generating a corresponding NFT copyright certificate, and supporting on-chain reverse tracing. According to the method and the device, the problem of low efficiency in secondary creation parameter adjustment caused by lack of standardized structured packaging and dependency relationship labeling in parameter generation can be solved.
Owner:URBAN PLANNING & DESIGN INST OF SHENZHEN UPDIS

Classifier multi-iteration-based subject corpus labeling method and system

The invention discloses a subject corpus labeling method and system based on classifier multi-round iteration, and belongs to the field of natural language processing and machine learning. The method comprises the steps that a multi-level subject classification system is constructed, multiple rounds of iterative reasoning are conducted on a target subject based on a Fasttext classifier and a seed set, positive and negative samples are dynamically generated in each round of iteration, and classification accuracy is gradually improved; performing multi-dimensional scoring on the positive samples by using the large model to screen high-quality positive samples, and extracting keywords to optimize classification boundaries; a subject problem is generated through a WebQA method, and a low-recall-rate subject corpus is retrieved and supplemented; training an information filter to identify, filter and reject low-quality contents such as advertisements and garbage; and finally, efficient and accurate subject labeling is realized through a classifier and filter series connection process. According to the method, the problems of low efficiency, large resource consumption and subject understanding deviation of traditional labeling are solved, and the method is suitable for efficient and accurate subject labeling scenes of large-scale text data.
Owner:ZHEJIANG LAB

Oral diagnosis and treatment knowledge graph dynamic construction system based on deep learning

The invention discloses an oral diagnosis and treatment knowledge graph dynamic construction system based on deep learning, and relates to the cross technical field of artificial intelligence and medical information technology, and the system comprises a data preprocessing and labeling module, an entity recognition stage module, a relationship classification stage module, a graph construction and application module and a local database. A structured annotation file is generated through data cleaning and expert annotation, the adaptability and quality of a data source are guaranteed, the problems that traditional preprocessing is insufficient in specialty and disordered in format are solved, a BERT-LSTM-CRF model is adopted, BERT is used for obtaining depth vector representation, Bi-LSTM is used for capturing long-distance dependency features, CRF is used for optimizing a tag sequence, the problems that a universal model is fuzzy in recognition and low in recall rate are solved, and the method is suitable for large-scale popularization and application. According to the method, semantic coding is carried out on the original sentences embedded with the special symbol marking entity pairs, then the entity pair relation is judged through a full connection layer and a Softmax classifier, accurate recognition of the deep clinical logic relation is achieved, and the problem that only shallow layer correlation is extracted traditionally is solved.
Owner:AFFILIATED STOMATOLOGICAL HOSPITAL OF NANJING MEDICAL UNIV

Domain relation extraction method and system based on large language model

The invention relates to the technical field of artificial intelligence, and provides a domain relation extraction method and system based on a large language model. The method comprises the following steps: identifying entity information in a field text set, and obtaining a labeled entity set with type labels; on the basis of entity pairs in the labeled entity set and predefined relation types, a question set is constructed by utilizing a judgment question generation algorithm, and a domain relation judgment question set is obtained; reconstructing the structured domain data into training data through a dialogue format conversion algorithm, and training a large language model based on the training data by adopting a QLoRA quantization fine tuning algorithm to obtain a domain fine tuning model; a double-layer retrieval algorithm is applied to retrieve and obtain related information from the knowledge graph, the domain relation judgment question set and the related information are combined and input into the domain fine tuning model for reasoning, and a domain relation triple is obtained. According to the method, high-precision and interpretable domain relation extraction is realized, and the accuracy and robustness of the model in the vertical domain are improved.
Owner:ZHONGJINKE INFORMATION TECH CO LTD +1

Mould engineering drawing labeling system, method and equipment, storage medium and program product

The invention provides a mold engineering drawing labeling system, method and equipment, a storage medium and a program product, and belongs to the field of computers. The system comprises an input analysis module used for obtaining three-dimensional model information, two-dimensional engineering drawing information and mapping information of a target mold; packaging the three-dimensional model information, the two-dimensional engineering drawing information and the mapping information into structured data based on a set data format; the semantic understanding module is used for performing semantic understanding based on the structured data by utilizing a large language model to obtain feature functions and annotation requirements; the dynamic annotation engine module is used for generating an annotation strategy based on the feature function and the annotation requirement; and the parameterized annotation generation module is used for generating an annotation instruction code based on an annotation strategy by utilizing a large language model, and writing the annotation instruction code into the target mold engineering drawing. The method is at least used for solving the problems that in an existing method, labeling efficiency is low, deep semantics cannot be understood, and consequently certain limitation exists in feature recognition application.
Owner:LENS TECH CHANGSHA

Fully-weathered argillaceous siltstone elastic-plastic damage constitutive model establishing system

The invention discloses a fully-weathered argillaceous siltstone elastic-plastic damage constitutive model establishing system, and relates to the technical field of elastic-plastic damage constitutive models.The system comprises a parameter collecting and labeling module used for obtaining a data set and conducting classification labeling and test condition information; the parameter processing and matching module is used for preprocessing the parameters to extract a feature parameter set and performing information association to supplement boundary condition data; the damage elastic-plastic feature module is used for presetting a database and a weight distribution rule and performing weighted fusion to form a damage elastic-plastic coupling feature parameter set; the constitutive model construction module is used for constructing a constitutive model, performing iterative training on the constitutive model and calculating damage parameters and a prediction result; and the model verification report module is used for verifying the damage parameters and the prediction result and judging the effectiveness of the constitutive model according to the verification result. According to the method, the multi-dimensional core characteristic parameter set is constructed by obtaining mechanical property test parameters and the like, and the limitation of single test data is broken through.
Owner:EAST CHINA JIAOTONG UNIVERSITY