Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

703 results about "Text graph" patented technology

In natural language processing (NLP), a text graph is a graph representation of a text item (document, passage or sentence). It is typically created as a preprocessing step to support NLP tasks such as text condensation term disambiguation (topic-based) text summarization, relation extraction and textual entailment.

Multi-agent dynamic arrangement method based on multi-modal analysis and adaptive retrieval

The invention discloses a multi-agent dynamic arrangement method based on multi-modal analysis and adaptive retrieval, and relates to the technical field of artificial intelligence and information retrieval. Comprising the steps of S1, converting a text, an image, structured data and voice content input by a user into a unified multi-mode semantic representation, S2, converting the unified multi-mode semantic representation into a specific execution process, and S3, automatically scheduling a reasoning agent, a knowledge obtaining agent and an execution agent according to DAG nodes, task elements and available resources, and obtaining the task elements and the execution agent according to the reasoning agent, the knowledge obtaining agent and the execution agent. S4, after task process construction and agent arrangement are completed, dynamic retrieval, evidence convergence and strategy optimization are carried out on information requirements related to a user task, so that a reasoning agent obtains complete knowledge support with consistent context, and S5, knowledge evidence is combined with a task process, so that the task process is completed. The method comprises the following steps: step S6, implementing problem solving, strategy generation and task closed-loop execution through a reasoning agent, step S6, performing actual operation on a target task by an execution agent according to an executable instruction sequence output by the reasoning agent, and outputting a result, and step S7, performing result verification according to an output result returned by the execution agent, and the correctness, integrity and consistency of an output result are examined through rule verification, model evaluation and evidence alignment.
Owner:INSPUR GROUP CO LTD +1

Artificial intelligence-powered large-scale content generator

An AI-powered content generation system that creates consistent, coherent, and engaging multi-modal content by integrating multiple specialized AI components. The system analyzes user input, identifies key elements, and maintains continuity throughout the generation process. It incorporates a feedback loop to learn and adapt based on user preferences, enabling personalized content experiences. The modular architecture allows for seamless integration of AI components focusing on text, images, audio, and interactive elements. The system ensures consistency across modalities and over extended periods, while managing rights, licenses, and royalties using blockchain technology. This advanced platform revolutionizes content creation, consumption, and management in the digital age.
Owner:QOMPLX INC

Multi-modal fusion intelligent question answering and knowledge retrieval method and system

The invention discloses a multi-modal fused intelligent question answering and knowledge retrieval method and system, and the method comprises the steps: building a multi-modal data index model oriented to a heterogeneous knowledge source, carrying out the feature mapping of text, image, table, chart, audio and video contents through a unified semantic embedding space, and generating a cross-modal index set; after a query request is received, performing semantic matching and structure matching on the cross-modal index set by using a multi-channel retriever to obtain candidate evidence fragments; and based on an evidence granularity decomposition strategy, performing minimum evidence unit division on text statements, table units, chart data points and multimedia frame contents in the candidate evidence fragments, and establishing a semantic consistency graph among the units. According to the method, high-credibility traceable generation of question and answer results is realized through multi-modal fusion and space-time consistency constraint, and the retrieval precision and interpretation transparency in a complex knowledge scene are remarkably improved.
Owner:NANJING CHUANGLIAN INTELLIGENT SOFT INFORMATION TECH CO LTD

Expert question and answer technical method, system and equipment based on local geological knowledge graph semantic reasoning

The invention discloses an expert question and answer technical method, system and equipment based on local geological knowledge graph semantic reasoning, and the method comprises the following steps: collecting and integrating multi-source heterogeneous geological data, including texts, images and remote sensing, and constructing a unified geological knowledge base covering multi-modal data; on the basis of a pre-training language model, semantic analysis is performed on a geological text, and'entity-relationship-entity 'structured knowledge is automatically generated by utilizing a triple extraction module. The method has field breadth and multidisciplinary fusion, is different from a knowledge graph technology focusing on single fields of mineralogy, geophysics and the like in the prior art, innovatively constructs a large-scale comprehensive knowledge graph system covering the whole geological disciplinary, supports cross-field knowledge extraction and reasoning, and is high in practicability. And complex interdisciplinary geological problems can be handled.
Owner:JIANGSU PROVINCIAL GEOLOGICAL BUREAU BIG DATA CENTER

Multi-modal data driven general report generation method and system based on large model

The invention discloses a multi-modal data driven general report generation method and system based on a large model, and belongs to the technical field of intelligent report generation. Firstly, texts, images and sensor data related to a report theme are obtained and subjected to standardized preprocessing; analyzing the report generation instruction, and matching and querying a task modal mapping library matching modal configuration scheme according to a task demand; quantitatively evaluating the data quality of each modal, dynamically calculating the final decision weight of each modal in combination with the basic weight, and distributing the final decision weight to a corresponding processing path to form dominant, supplementary and reference data; inputting the dominant data and the supplementary data into a multi-modal model for analysis to obtain a preliminary conclusion with confidence score, performing consistency judgment, if no conflict exists, performing fusion to form a comprehensive conclusion, and if the conflict exists, combining a quality evaluation result and an arbitration rule to complete conflict judgment; and finally, inputting the comprehensive conclusion and the reference data into a large language model to generate a report text, and outputting a complete report after typesetting and proofreading.
Owner:NANJING ANCIENT NETWORK TECH CO LTD

Method and System for Optimizing Use of Retrieval Augmented Generation Pipelines in Generative Artificial Intelligence Applications

Systems and methods for dynamic knowledge integration in LLM systems including receiving multimodal input data comprising text, image, audio, video, and / or code data, extracting information by processing the multimodal input data through a document processor, storing the extracted information in a dynamic knowledge base, receiving a user query at a query processor, identifying knowledge domains related to the user query using domain-specific agents, retrieving real-time information from the dynamic knowledge base responsive to the identified knowledge domains, integrating the real-time information into the processing of an LLM by a dynamic knowledge integrator, and generating a response using the LLM with the real-time information.
Owner:MADISETTI VIJAY

Heterogeneous document set-oriented cross-modal semantic alignment and logic consistency verification system

The invention relates to document verification, in particular to a heterogeneous document set-oriented cross-modal semantic alignment and logic consistency verification system, which is used for heterogeneous document input, supports multi-format document input and comprises multi-modal elements including texts, pictures, tables and charts. The document analysis module is used for carrying out structured extraction on document contents; extracting multi-modal elements, identifying and classifying various elements in the document, and establishing position and type labels of a foundation; the knowledge graph construction module is used for uniformly modeling heterogeneous elements into a multi-modal knowledge graph; the graph neural network semantic alignment module is used for realizing accurate cross-modal semantic alignment by using a specially designed graph neural network based on the multi-modal knowledge graph; the hybrid consistency verification engine is used for performing logic consistency verification in combination with a symbol logic verification mechanism and a semantic consistency verification mechanism; according to the method, the defect that accurate cross-modal semantic alignment and logic consistency verification are difficult to carry out on professional documents with multi-modal elements can be effectively overcome.
Owner:ANHUI GAOSHAN TECH CO LTD

Intelligent interview scoring system based on large language model interpretable decision

The invention relates to an intelligent interview scoring system capable of explaining decisions based on a large language model, and the system comprises a multi-mode resume analysis and feature coding unit, a resume feature adaptive matching unit, an interactive scoring and knowledge enhancement unit, and an answer quality evaluation unit. Text, image and audio features are extracted through a cross-modal attention mechanism of a multi-modal large language model, resume features are encoded into dynamic word vectors, and entity-level feature vectors are extracted; the post description text is encoded into a demand feature vector by a resume feature adaptive matching unit; calculating semantic similarity between the resume entity feature vector and the demand feature vector; the interactive scoring and knowledge enhancement unit dynamically retrieves knowledge fragments to generate a preliminary evaluation report containing a scoring basis; and the answer quality evaluation unit fuses the information density, the fluency and the integrating degree to generate a final score. And the whole-process intelligence from demand analysis to final decision making is realized.
Owner:SHANGHAI JINYU INTELLIGENT TECH CO LTD

Knowledge question and answer library agent construction method and system

The invention provides a knowledge question and answer library agent construction method and system. Efficient knowledge management and question and answer are achieved through cooperation of multiple agents. According to the system, firstly, a multi-modal information extraction agent is constructed, and heterogeneous data such as texts, images and tables are converted into structured vectors and stored; meanwhile, the knowledge graph is dynamically constructed and continuously optimized by the self-adaptive knowledge graph construction agent, and a new relationship is derived through combination of symbolic logic and a graph neural network, so that an evolvable knowledge network is formed. In the question and answer stage, a query analysis agent deeply analyzes the intention of a user and generates sub-queries; retrieving the vector library and the knowledge graph in parallel by the retrieval enhancement generation agent; and the reasoning and synthesizing agent integrates multi-source information and generates an accurate answer with a complete source label through a large language model. Dynamic knowledge management, precise semantic analysis and system self-evolution are achieved, and the method is particularly suitable for professional field scenes needing high-reliability questions and answers.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Knowledge graph enhanced multi-modal file retrieval method

The invention relates to a knowledge graph enhanced multi-modal file retrieval method, and belongs to the field of artificial intelligence and multi-modal information retrieval. The invention aims to solve the problems of multi-modal information semantic segmentation, weak semantic reasoning ability and low semantic matching precision in the existing multi-modal archive resource retrieval process. Through four stages of archive multi-modal data preprocessing and feature extraction, knowledge graph construction and enhancement, semantic retrieval request analysis and intention modeling, and multi-modal semantic matching and sorting, a semantic relationship is enhanced by utilizing a knowledge graph, so that the semantic relevancy and context consistency of a retrieval result are remarkably improved; a user is allowed to input and inquire in various forms such as texts, images and voices, and semantic extension retrieval is supported. According to the method, more efficient and accurate archive resource retrieval can be realized, and the user retrieval experience and archive knowledge utilization are improved.
Owner:BEIJING INST OF COMP TECH & APPL

Semantic understanding system based on large language model

The invention belongs to the technical field of semantic understanding systems, and particularly relates to a semantic understanding system based on a large language model.The semantic understanding system is characterized in that firstly, a data preprocessing module is used for conducting cleaning, denoising and cross-modal conversion on input multi-modal data such as texts and images, and standardized data is generated; a semantic feature extraction module extracts general semantic features by using a pre-training model, adapts to field requirements through dynamic learning rate fine adjustment, and outputs scenarized semantic vectors; then, a dynamic semantic-knowledge bidirectional fusion module adjusts token weight according to a dynamic semantic weight algorithm, realizes real-time alignment of semantics and a knowledge graph by means of a knowledge entity association strength algorithm, and a knowledge enhancement fusion module further optimizes knowledge weight and dynamically updates association; then, the semantic reasoning module performs multi-round reasoning based on fusion information, and evaluates the result reliability in combination with a confidence coefficient algorithm; and finally, the output and optimization module generates a structured result, and iteratively optimizes parameters of each module according to feedback data to complete a semantic understanding processing flow.
Owner:SICHUAN JOYOU DIGITAL TECH CO LTD

Intelligent teaching-assistant question-answering system with enhanced multi-modal knowledge graph

The invention belongs to the technical field of artificial intelligence and educational informatization, and relates to a multi-mode knowledge graph enhanced intelligent teaching-assistant question-answering system. According to the method, the knowledge graph construction technology, the multi-modal content analysis technology and the large language model reasoning enhancement technology are comprehensively applied, and the semantic understanding, knowledge integration and reasoning generation capabilities of the intelligent teaching assisting system in an education and teaching scene are improved. The related technology comprises layout analysis of textbook documents, semantic description generation of image content, entity and relation extraction of text content, multi-modal knowledge graph construction and question and answer reasoning and natural language generation combined with the knowledge graph. Through cooperative application of the technologies, semantic interconnection can be carried out on various modal information such as texts, images and tables in the textbook, and a searchable and traceable textbook-level knowledge network is formed.
Owner:NORTHEASTERN UNIV CHINA

Zero-sample liquid crystal display screen defect detection method

The invention discloses a zero-sample liquid crystal display screen defect detection method. According to the scheme, the method comprises the following steps: 1) collecting and processing image data of a display screen; 2) constructing an image-text comparison pre-training model (CLIP) to carry out text-image similarity calculation; 3) designing a self-adaptive prompt network, combining static and dynamic prompts and designing an optimal fusion weight, realizing self-adaptive combination of prompt semantics, and improving the adaptability of the model; (4) semantic separation loss is added into the global features of the text, and the semantic separability of the normal text and the defect text is improved; and 5) before image classification, a feature improvement module for abnormal guidance is inserted into classified visual features to enrich the visual features, and the ability of alignment with the text is further improved. The method is suitable for a cold start stage of display screen defect detection, can solve the problems that in display screen zero sample detection, 'defect 'semantics are difficult to understand, and normal defect attributes are difficult to separate, and remarkably improves the classification and positioning capability of display screen defects.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

AI video output system and method combined with text image model

The invention discloses an AI video output system and method combined with a text image model, and relates to the technical field of AI video output, and the system comprises a visual angle conversion and multi-visual angle generation module which comprises a visual angle encoder and a multi-visual angle generator based on semantic-style hidden representation and visual angle vectors, the method is used for generating a multi-angle illustration sequence aiming at different camera parameters and keeping structure and style consistency among different visual angles through time domain consistency constraint in the generation process. According to the method, the artistic style is propagated and maintained among multiple frames and multiple visual angles according to a three-dimensional visual angle transformation rule through visual angle perception style propagation, and the generator outputs a consistent multi-angle illustration sequence under the constraint of a visual angle condition during sampling, so that the problem of inconsistent style drift and deformation under multiple visual angles or multiple lenses is solved; the incoherence of stroke, texture and main body structure caused by visual angle change is avoided, the later manual correction is obviously reduced, and the content consistency and the impression professional degree are improved.
Owner:BEIJING SHUYOU WENLV TECH CO LTD

Semantic segmentation model training method, electronic device and storage medium

A semantic segmentation model training method and apparatus, an electronic device and a storage medium are provided. The semantic segmentation model training method includes: acquiring a sample image, and extracting visual image features corresponding to the sample image by a semantic segmentation model to be trained; processing the sample image to obtain a text image feature corresponding to the sample image, the text image feature being an image feature generated from language description text for the sample image; fusing the visual image features with the text image feature to obtain multimodal features, and performing image segmentation prediction based on the multimodal features to obtain a target loss; and training the semantic segmentation model to be trained based on the target loss to obtain a target semantic segmentation model.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Method for constructing large mineralization potential evaluation model based on LLM

The invention discloses a method for constructing a large mineralization potential evaluation model based on LLM, and relates to the field of models.The method for constructing the large mineralization potential evaluation model based on LLM comprises the following steps that 1, a mineral geological map file is converted into a structured JSON format data set; step 2, constructing cue words, wherein the cue words at least comprise geological information extraction cue words, function task cue words and entity, relation and attribute extraction cue words; compared with the prior art, the method has the beneficial effects that the large mineralization potential evaluation model constructed by the method realizes text and map multi-mode mineral product geological map information identification and analysis, automatic mineralization geological feature analysis, mineralization feature information extraction and mineralization potential evaluation; the problems that in the prior art, man-made subjective influence exists, professional knowledge dependence is high, the geological information recognition accuracy rate in the recognition process is low, and information is lost during knowledge retrieval are solved.
Owner:JILIN UNIVERSITY

Method and system for constructing multi-modal knowledge graph in agricultural field

The invention provides a multi-modal knowledge graph construction method and system in the agricultural field, and relates to the technical field of data processing. The method comprises the following steps: performing recognition and structured expression on agricultural entities, relationships and events by fusing texts, images and time-space data to form an agricultural semantic representation model; calculating a confidence coefficient based on the evidence template and resolving conflicts, and generating a structured knowledge unit; an agricultural multi-modal knowledge graph is constructed on the basis of the event network, and dynamic updating is achieved through increment correction and self-correction; and finally, outputting an agricultural aid decision result with confidence by combining causal reasoning and confidence evaluation. The agricultural multi-modal knowledge graph construction method solves the problem that an agricultural multi-modal knowledge graph construction method in the prior art lacks a system modeling and closed-loop optimization mechanism for cross-modal alignment credibility, ontology constraint, evidence tracing, causal verification and dynamic updating.
Owner:XINJIANG UNIVERSITY

System and method for a large codeword model for deep learning

Modality agnostic Large Codeword Model (“LCM”) is an advanced deep learning architecture that processes discrete, compressed data representations called codewords across multiple modalities. Unlike traditional models using raw tokens and dense embeddings, LCMs efficiently handle diverse input types including text, images, audio, and video. The system employs a modality agnostic encoder, unified codebook, and multimodal machine learning core to capture inherent data structures and patterns. This approach enables more generalizable and interpretable feature learning, facilitating transfer learning across domains. The LCM's scalable and flexible architecture includes components for modality-specific processing, cross-modal attention, and joint representation learning. With its computational efficiency and versatility, the Modality Agnostic LCM offers significant potential for various AI applications, including natural language processing, computer vision, and multimodal reasoning.
Owner:ATOMBEAM TECH INC

Knowledge graph-based dynamic retrieval enhancement generation method and system, terminal and medium

The invention relates to the field of data retrieval, and particularly provides a dynamic retrieval enhancement generation method and system based on a knowledge graph, a terminal and a medium, and the method comprises the following steps: extracting a structured triple from multi-source heterogeneous data through a large language model, and constructing a global knowledge graph by means of an entity linking technology; integrating a real-time data stream interface, and dynamically updating graph nodes and attributes based on an event-driven mechanism; adopting a RotatE model to respectively encode the entity and the relationship to a complex number space, and fusing to generate a mixed vector to construct an efficient index; after user query is received, topic nodes are positioned through semantic analysis, related entities are retrieved through mixed indexes, and multi-hop reasoning is executed along a relation path to generate reasoning sub-graphs and extended contexts; and finally, generating structured text answers by using a large language model, and adaptively outputting multi-modal results such as texts, charts and the like according to user requirements. According to the method, the knowledge updating timeliness, the complex query reasoning capability and the retrieval precision are effectively improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Multi-modal rumor detection method based on anti-factual reasoning and causal intervention

The invention discloses a multi-modal rumor detection method fusing texts, images and social propagation structures, and belongs to the technical field of natural language processing, computer vision and causal reasoning. Specifically, the invention provides a unified causal inference framework, and hybrid deviation in multi-modal data is effectively stripped by integrating text anti-fact causal inference and an image dot product causal intervention mechanism. Under the framework, social propagation structure features are further fused, and a multi-head collaborative attention mechanism is adopted, so that deep alignment and semantic enhancement in cross-modal features are realized. Adversarial samples are generated through projection gradient descent for adversarial training, and model parameters are optimized in combination with anti-fact loss, so that the classification accuracy and generalization ability of the model are improved. According to the rumor detection method, a causal reasoning normal form is introduced into a rumor detection task, the effectiveness of an anti-fact and intervention mechanism in a complex information scene is verified, and a new theoretical support and method path are provided for constructing a credible multi-modal information system.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Visual positioning method based on semantic comprehension and attribute distinguishing enhancement

The invention belongs to the technical field of visual positioning, and relates to a visual positioning method based on semantic comprehension and attribute distinguishing enhancement. The framework mainly comprises a feature coding module, a semantic sensitive data enhancement module, a fine-grained attribute guiding module and a multi-stage cross-modal decoder module. Specifically, the feature coding module performs feature coding on an input graph. The semantic sensitive data enhancement module generates a plurality of queries which are consistent with long text semantics by keeping the consistency of spatial relation words in combination with a large language model, so that a training data set for the long text is expanded. The fine-grained attribute guiding module extracts attribute prior information from a text query in combination with a text graph model and an image encoder, constructs a visual feature representation with higher discrimination by using the information guiding model, and generates a target query with attribute difference at the same time.
Owner:DALIAN UNIV OF TECH

Complex formula intelligent calculation method for ship field

The invention provides a complex formula intelligent calculation method for the ship field, and the method comprises the steps: obtaining a complex formula calculation file in a ship design document, carrying out the multi-modal analysis of a text, an image, a formula and a table in the document, and converting the text, the image, the formula and the table into structural data; based on the structured data, calling a large language model to fill a standardized formula calculation document template, and generating a standardized formula calculation document including formula expression, parameter definition, calculation steps and unit expression; when a formula is incomplete, variables are missing or conditions are not explained, missing fragments and conditions are automatically complemented by utilizing a knowledge graph and a large model reasoning result; and performing cross-modal error correction and physical consistency verification on the document, and if logic conflicts are found, outputting correction suggestions and automatically updating formula expressions. According to the method, automatic analysis, complementation, inspection and calculation of complex formulas in the ship design field can be realized, and the standardization and reliability of engineering calculation are improved.
Owner:中国船舶集团海舟系统技术有限公司

Stylized visual text editing method, system and equipment and storage medium

The invention discloses a stylized visual text editing method, a stylized visual text editing system, stylized visual text editing equipment and a storage medium, which are corresponding schemes, and the related schemes aim to solve the problem of style consistency existing in image text editing of an existing diffusion model, and the stylized visual text editing efficiency is improved by combining visual features of a font image and an input text image. The method comprises the following steps of: extracting style embedded information from a text image, and inputting the style embedded information as an enhanced style condition into a diffusion model to realize fine control on a diffusion process, so that the diffusion model can generate a text image with high readability and style consistency, and can realize maintenance of an original text style or style migration based on a reference image.
Owner:UNIV OF SCI & TECH OF CHINA

Multi-modal data dynamic desensitization method and system based on machine learning

PendingCN121167765ADigital data protectionBiological modelsInformation security managementEngineering
The invention relates to the technical field of information security management, in particular to a multi-modal data dynamic desensitization method and system based on machine learning, multi-modal features such as texts, images and time sequences are jointly extracted through CNN-Transform-LSTM, 1-5 sensitive levels are evaluated by adopting a cross-modal attention model, a self-adaptive desensitization strategy is generated through an improved MOEA / D-DE algorithm, and the dynamic desensitization of the multi-modal data is realized. And realizing cross-modal collaborative desensitization by combining GNN. The system dynamically adjusts parameters through TD3 reinforcement learning, the response delay is less than or equal to 100ms, the Bi-LSTM-AE model evaluates privacy and utility in real time, and the block chain evidence storage whole process is realized. The method solves the problems of precision and utility imbalance, poor dynamic adaptability and insufficient multi-mode collaboration in the traditional technology, and is suitable for the fields of finance, medical treatment and the like.
Owner:北京睿航至臻科技有限公司

Multi-modal document analysis method, electronic equipment and storage medium

The invention provides a multi-modal document analysis method, electronic equipment and a storage medium, and the method comprises the steps: preprocessing an original document, and recognizing key elements (at least including formulas, tables, texts and images) by using a target detection model to obtain an element positioning labeling table; roughly identifying the task type based on the label table to obtain an identifier, and generating a field adaptation parameter in combination with an original document metadata identification field; determining an element range according to a labeling table, and performing layered detection and targeted repair on obstacles in combination with task identification to obtain a barrier-free element document; carrying out collaborative coding on barrier-free elements, splitting sub-tasks and carrying out parallel processing on the basis of task identifiers and coding results, and calling adaptive parameters to adjust precision; and checking a correction processing result, and integrating into a final structured document report according to an adaptive parameter format. According to the method and the device, the multi-modal document analysis efficiency and precision can be improved.
Owner:北京中科闻歌科技股份有限公司

Ultra-long-range medical record diagnosis and treatment question-answering system and method based on dynamic time sequence semantic map

The invention relates to an ultra-long-range medical record diagnosis and treatment question answering system and method based on a dynamic time sequence semantic map, belongs to the technical field of medical artificial intelligence and information processing, and solves the problems that in the prior art, the inference ability is insufficient when ultra-long-range medical records are processed. The system comprises a pre-processing module for pre-processing overlong-range medical record information of a patient to generate standardized data; the medical entity recognition module is used for recognizing event entities and descriptive entities in the standardized data; associating a time attribute for the event type entity to generate a time sequence entity tuple; outputting a medical entity recognition result; the dynamic semantic map construction module is used for constructing a specific dynamic semantic map of the corresponding patient; the text-map dual-mode fusion query engine is used for analyzing the medical problem proposed by the user and carrying out text-map dual-mode fusion query based on the specific dynamic semantic map; and the diagnosis and treatment reply module is used for generating a diagnosis and treatment reply result according to the text-map dual-mode fusion query result.
Owner:BEIJING YIYONG TECH CO LTD

Hierarchical decision-making violation scene identification method, device, equipment, medium and product

The invention discloses an illegal scene recognition method and device for hierarchical decision, equipment, a medium and a product. The method comprises the following steps: firstly, acquiring image and audio data of a scene to be recognized; processing by using a first-level decision model comprising an audio event detection and speech recognition model and a keyword grading and screening module to obtain a first-level result; if the first-level result is missed or the violation feature is not obvious, a second-level decision model containing a text and image feature extraction model is used for processing, and a second-level result is obtained; if the second-level result is still the same, inputting the image, the audio data, the first-level result and the second-level result into a third-level decision model containing a multi-mode large language model to obtain a third-level result; and finally, judging whether the scene is a target violation scene according to the three-level result. Through audio event detection, a keyword grading screening strategy and a grading strategy combining a double-tower model and a multi-mode large model, efficient recognition of violation scenes is realized.
Owner:YI REN HENG YE TECH DEV (BEIJING) CO LTD

Utilizing a diffusion prior neural network for text guided digital image editing

The present disclosure relates to systems, methods, and non-transitory computer readable media for utilizing a diffusion prior neural network for text guided digital image editing. For example, in one or more embodiments the disclosed systems utilize a text-image encoder to generate a base image embedding from the base digital image and an edit text embedding from edit text. Moreover, the disclosed systems utilize a diffusion prior neural network to generate a text-image embedding. In particular, the disclosed systems inject the base image embedding at a conceptual editing step of the diffusion prior neural network and condition a set of steps of the diffusion prior neural network after the conceptual editing step utilizing the edit text embedding. Furthermore, the disclosed systems utilize a diffusion neural network to create a modified digital image from the text-edited image embedding and the base image embedding.
Owner:ADOBE INC

Multi-modal input instruction and data analysis matching method and system

The invention relates to the technical field of multi-modal matching, in particular to a multi-modal input instruction and data analysis matching method and system, and the method comprises the following steps: collecting input multi-modal data and instructions, recognizing texts, images and audios, extracting a time sequence, spatial distribution and statistical parameters, and generating a feature distribution result. According to the method, the modal correlation degree is calculated, information entropy measurement is combined, the correlation features are subjected to differentiation processing, the accuracy and consistency of the high correlation features are ensured through space mapping, the potential deviation of the middle correlation features is reduced through alignment adjustment, and the weights of the low correlation features are re-distributed through normalization; according to the method, the overall feature distribution balance and semantic consistency are enhanced, accurate task matching and classification are realized based on the fitting degree score and the time sequence fitting degree, the association accuracy and the time sequence association recognition capability among complex multi-modal data are effectively improved, and the analysis requirement of diversified data input in a complex scene is met.
Owner:CRRC IND INST CO LTD

Bridge engineering knowledge base construction method and system

The invention discloses a bridge engineering knowledge base construction method and system, and the method comprises the steps: classifying the multi-modal data of each project in the bridge engineering field according to the data type, carrying out the structural conversion, forming a structural data document corresponding to each project, and constructing a modular parallel workflow; the data types comprise texts, images and voices; splitting each structured data document into document segments according to semantic features, and generating document segment vectors by using an embedded model; forming double indexes by using all document fragment vectors and an entity-relationship-attribute triple of the bridge knowledge graph, and constructing a structured data document library-document fragment vector library-knowledge graph three-layer knowledge base; and training a bridge large model based on implicit rule coding, explicit engineering constraint and a rule activation mechanism, and embedding the trained bridge large model into the three-layer knowledge base to complete construction of the bridge engineering knowledge base. Data are fully utilized, the coverage range of the knowledge base is expanded, a multi-dimensional evidence source is provided, and question and answer credibility is improved.
Owner:CHINA RAILWAY MAJOR BRIDGE RECONNAISSANCE & DESIGN INSTITUTE CO LTD