Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

97 results about "Visual evidence" patented technology

Appearance patent graph retrieval method based on multi-modal large model and infringement detection system

The invention discloses an appearance patent graph retrieval method and infringement detection system based on a multi-modal large model, and relates to the technical field of intelligent retrieval, and the method comprises the steps: carrying out the feature extraction of each patent picture in an appearance patent database through a multi-modal large language model, and storing the features in a vector database; generating an embedded vector of the target picture and a description text for describing the appearance characteristics of the target picture by utilizing a multi-modal large language model; calculating the similarity between a target picture embedding vector and each picture embedding vector in a vector database, performing retrieval to form a preliminary candidate set, and then calculating the semantic similarity between a target picture description text and a patent text corresponding to each picture in the preliminary candidate set; performing fusion to obtain a retrieval result of comprehensive sorting; correlating the corresponding patent law text, positioning and marking related text description, and forming a visual evidence chain combining vision and text evidence. According to the method, the vision and text bimodal information is fused, so that the whole-process intelligentization from image understanding to infringement judgment is realized.
Owner:GUANGDONG POLYTECHNIC NORMAL UNIV

Engineering vehicle safety simulation and prediction system based on digital twinning

The invention discloses an engineering vehicle safety simulation and prediction system based on digital twinning, and the system comprises a data collection layer which collects multi-source heterogeneous data in real time; in the knowledge graph layer, a streaming inference engine is constructed based on an Apache Jena graph database, and an entity-relationship-attribute triple dynamic graph structure is adopted; according to the AI model layer, a physical rule serves as a loss function constraint term to be embedded into a neural network through a physical information neural network, a digital organ model concept is combined to split a vehicle into key organs for heterogeneous modeling, a simplified physical model is adopted in the core physical process, and an LSTM-AI model is adopted in external behaviors; the explanatory analysis layer is used for integrating an SHAP / LIME explanatory tool to output a visual evidence chain during fault prediction, and deploying an online incremental learning framework to allow the model to learn from new data and dynamically adjust normal range definition; and the visualization and application layer is used for performing three-dimensional visualization rendering based on WebGL or Three.js, and ensuring data transmission security through block chain evidence storage and end-to-end encryption.
Owner:ZHONGXIN DIGITAL TECHNOLOGY (SICHUAN) CO LTD

Multi-agent collaborative visual navigation reasoning enhancement method and system

The invention discloses a multi-agent collaborative visual navigation reasoning enhancement method and system, and the method comprises the steps: an inference agent analyzes input data, and generates structured output containing a reasoning process and an initial answer; independently analyzing the input data by the evaluator agent, generating an evaluation answer and reviewing the output of the inference agent; comparing whether the initial answer is consistent with the evaluation answer, and if yes, outputting a final answer; if not, calling a bifurcation feedback agent, and generating a core bifurcation point abstract and a visual evidence verification plan; and calling a final decision maker agent to directionally retrieve the video evidence according to the visual evidence verification plan, and outputting a final answer. Four roles of an inference person, an evaluator, a divergence feedback person and a final decision maker are introduced, a complete'generation-evaluation-feedback-optimization 'decision link is constructed, and the cognition and inference accuracy in a complex scene is remarkably improved.
Owner:UESTC (SHENZHEN) ADVANCED RES INST

Medical analysis method and system based on multi-modal large language model and chain reasoning

The invention discloses a medical analysis method and system based on a multi-modal large language model and chain reasoning, and belongs to the crossing field of artificial intelligence and medical image analysis. The method comprises the following steps: firstly, carrying out preprocessing and modal alignment on medical multi-modal data to obtain image visual features and text embedding vectors of the same vector spatial dimension; capturing visual evidence in a structured pathological feature form through a self-excitation mechanism; on the premise of visual evidence, a thinking chain text containing causal logic is generated through evidence anchoring chain type reasoning; and finally, performing vision and text dual consistency verification on the thinking chain text, and outputting an analysis conclusion according to a result, or triggering a negative feedback correction mechanism to regenerate the thinking chain text. According to the method, medical reasoning full-process logic visualization is realized, medical illusion is eliminated, the diagnosis accuracy of complex cases is improved, and the method is suitable for medical scenes such as medical visual questions and answers and image diagnosis report generation.
Owner:CENT SOUTH UNIV

Electronic medical record intelligent coding method, device and system and storage medium

The invention discloses an electronic medical record intelligent coding method, device and system, and a storage medium. The method comprises the following steps: respectively carrying out entity extraction, feature coding and homomorphic encryption transmission on an electronic medical record text, a medical image and structured inspection data; according to the preprocessed data, performing multi-modal feature alignment, ICD semantic mapping and real-time medical insurance rule injection through a dynamic knowledge fusion coding engine; generating a three-dimensional visual evidence chain of the text evidence, the image evidence and the policy evidence through the interpretable traceability matrix; through a multi-role cooperative workflow engine, generating a code according to a confidence threshold routing AI, and auditing the code to a clinician or a professional coder; rare disease data are synthesized through an adversarial training enhancement module, and policy mutation is simulated for model adaptability training. By adopting the technical scheme of the invention, the problems of multi-modal data splitting, policy lag and non-traceability in the prior art are solved.
Owner:JINGWEI ZHIYUN (SHANGHAI) TECHNOLOGY CO LTD

Video understanding method and system based on multi-mode evidence chain

The invention relates to a video understanding method and system based on a multi-modal evidence chain. The method comprises the steps of obtaining a question text, an option set and a target video input by a user; using the question text and the option set to form a complementary analysis angle set; performing frame-by-frame matching on the target video based on a preset angle feature mapping library to obtain key timestamp sets corresponding to different analysis angles; extracting a corresponding key video frame set from the target video; scoring each key video frame in the key video frame set, and constructing an angle-frame mapping table between different analysis angles and corresponding visual evidence frames; obtaining a to-be-reasoned text needing visual evidence, and associating the to-be-reasoned text with the angle-frame mapping table to construct a multi-modal evidence chain; and obtaining a video reasoning result of the problem text based on the multi-modal evidence chain. High-precision and high-interpretability video understanding is realized, and the dependence of video understanding on large-scale annotation data is reduced.
Owner:CENT SOUTH UNIV

Graphical user interface forensic analysis method based on multi-modal large language model

The invention provides a graphical user interface forensic analysis method based on a multi-mode large language model, and belongs to the technical field of man-machine interaction, artificial intelligence safety and digital forensic. The method comprises the steps of evidence capture and distillation, semantic analysis and record generation, storage and indexing, and retrieval and inquiry. Through evidence capture and distillation, on the premise that evidence integrity is guaranteed, the visual data volume needing to be analyzed is greatly reduced, and long-term and efficient evidence obtaining analysis on a mobile terminal becomes possible; by introducing the analysis capability of a multi-modal large language model, original and structurality-free visual evidence is converted into a structuralization intelligent evidence which can be understood and deeply excavated; a designed two-stage natural language evidence obtaining query engine supports a complete analysis process from macroscopic semantic retrieval to microscopic detail inquiry; the problems that a traditional evidence obtaining method depends on manual troubleshooting and is low in efficiency are solved, and unprecedented high efficiency and convenience are provided for backtracking and auditing of mass interaction records.
Owner:PEKING UNIV +1

Multi-stream chain type perception enhanced multi-modal aspect level sentiment analysis method

The invention discloses a multi-stream chain type perception enhanced multi-modal aspect level sentiment analysis method, which is called MCPE model for short, and relates to the technical field of multi-modal sentiment analysis. According to the method, the problem of low accuracy of fine-grained sentiment analysis caused by information density difference in modals and information imbalance between modals in the prior art is solved. The method comprises the steps of obtaining a to-be-analyzed multi-modal text-image pair, inputting the text-image pair into a trained MCPE model, and obtaining aspect words and emotional polarities thereof; the MCPE model comprises a feature extraction module, a chain enhancement module, a multi-stream interaction module and a classifier; feature extraction adopts BART to extract text features and Faster R-CNN to extract image features; the chained enhancement module suppresses image and text noise and enhances fine-grained semantics through an IFE-TFE double-chain architecture; the multi-stream interaction module adopts bidirectional cross-modal attention to realize dynamic complementary fusion of text reasoning and visual evidence; and the classifier outputs an analysis result based on the fusion feature without an external tool. The method is suitable for sentiment analysis of social media comments.
Owner:HEILONGJIANG UNIV

Natural resource field proof task generation method and device

The invention discloses a natural resource on-site proof task generation method and device, and relates to the technical field of task generation, and the natural resource on-site proof task generation method is applied to a proof scene visualization building platform. A preset proof database and a preset template warehouse are stored on the proof scene visualization building platform. The method comprises the following steps: responding to a building request of a user for a natural resource field proof task; analyzing the building request, and determining to-be-generated proof task information; querying corresponding proof data from a preset proof database according to the to-be-generated proof task information; selecting a target task template from a preset template warehouse according to the proof data; and performing proof task configuration on the proof data on the target task template to generate a natural resource field proof task. Through the automatic process, the generation efficiency of the natural resource on-site proof task is improved, and the complexity and error rate of manual operation are reduced.
Owner:JIANGXI PROVINCIAL LAND & SPACE SURVEY & PLANNING RES INST

Traditional Chinese medicine cart capable of sorting and recording and sorting and recording method thereof

The invention relates to the technical field of traditional Chinese medicine treatment, and discloses a traditional Chinese medicine cart capable of sorting and recording and a sorting and recording method thereof. The traditional Chinese medicine cart comprises a movable cart body, a weighing table and a recording assembly integrating a control panel, a code scanner and a monitoring and shooting device. The method comprises the steps that prescription identification is obtained through a code scanner, an electronic prescription is called, a monitoring and photographing device is started to monitor traditional Chinese medicine weighing and storage areas, during sorting and weighing, the monitoring and photographing device synchronously records the operation process, weighing data of each medicinal material is bound with visual data of the corresponding operation process, and a traceable sorting record is generated. Through deep fusion of weighing data and visual evidences, traceability and error prevention of the whole process of the sorting process are achieved, and the sorting accuracy, safety and working efficiency of the traditional Chinese medicine formula are remarkably improved.
Owner:HUBEI YONGANXIN MEDICAL TECH CO LTD

Video content auditing method and device, electronic equipment and storage medium

The invention relates to the technical field of computers, and provides a video content auditing method and device, electronic equipment and a storage medium. For a to-be-audited video, partitioning the extracted key frame, extracting visual information of each image block, generating a visual token, and based on the association degree between each visual token and a set content query task and the attribute information of the to-be-audited video, performing hierarchical compression processing on each visual token, therefore, key information is prevented from being lost, calculation overhead is reduced, auditing efficiency and accuracy are effectively balanced, an auditing conclusion is given for the to-be-audited video, a visual basis supporting the auditing conclusion is given from the two dimensions of time and space, and the visual basis comprises a visual basis in the to-be-audited video and a visual basis in the to-be-audited video. According to the content description text of the video clip corresponding to the at least one timestamp interval associated with the content query task, the black box auditing process is well documented, and the credibility and reliability of an auditing conclusion are improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Multi-modal large model reasoning method, device and equipment based on generation understanding collaboration

The invention discloses a multi-modal large model reasoning method, device and equipment based on generation understanding collaboration. Comprising the steps of receiving a to-be-understood target image and a corresponding text question; generating an editing prompt based on the target image and the text question, wherein the editing prompt is used for indicating an operation strategy for processing the target image; calling a generation branch of a unified multi-modal model, and generating an auxiliary image based on the editing prompt; and constructing input information based on the auxiliary image, calling the understanding branch of the unified multi-modal model to execute reasoning, and obtaining an answer corresponding to the text question. According to the visual understanding method, retraining and external tools are not needed, the model can construct visual thinking through controllable generation, the self-generation capacity serves as an internal reasoning step to complement visual evidence, and reasoning accuracy and robustness are improved through the evidence.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Driver micro-sleep confirmation method and system for vehicle-mounted scene

The invention relates to the technical field of artificial intelligence, and provides a driver micro-sleep confirmation method and system for a vehicle-mounted scene. According to the method, an event-level micro-sleep confirmation mechanism is constructed, so that confirmation is only carried out in an initial window W defined by the rising side of a coherent peak of two natural behavior flows, resonance intensity R is taken as a necessary item, and visual evidence participation authority is controllably allowed or prohibited through an evidence availability mark E; the method is triggered only after the resonance intensity is met and the other evidence is co-occurred in the starting window, so that accurate judgment is realized, and the technical problems that a single evidence is easy to trigger and the reminding rhythm is incompatible with the driving task when the visibility of the visual channel is dynamically changed and the external disturbance is frequent in the vehicle-mounted driving scene are solved.
Owner:FAW CAR CO LTD

New energy vehicle accident time accurate identification method based on multi-modal data fusion

This invention discloses a method for accurate identification of the moment of an accident in a new energy vehicle based on multimodal data fusion. The method acquires vehicle motion sensor data, dynamic parameters, and visual image data, and then performs the following parallel processing: The physical layer processes the data to obtain motion features and compares them with preset trigger conditions; when a suspected collision is detected, it outputs the physical layer trigger time and acceleration amplitude; the dynamic layer processes the data based on the deviation between the vehicle's dynamic model and its actual motion state, outputting a vehicle trajectory status label; the visual layer processes the data to detect targets and analyze their motion state; when visual evidence of a collision is detected, it outputs the visual layer trigger time and the type of collision object; the fusion layer performs spatiotemporal alignment of the three layers' outputs and makes a comprehensive judgment, outputting the fused accurate collision trigger time and accident scene classification information. This invention effectively filters false alarms and improves the detection rate of low-speed accidents through multimodal fusion, achieving accurate identification of the moment of an accident and scene reconstruction.
Owner:BEIJING YUANSHU INTELLIGENT WHEEL TECHNOLOGY CO LTD

Systems and methods for making tamper-resistant containers

A container sealing system, comprising a container, wherein the container includes a closure mechanism configured to close an open end of the container; a strip of adhesive-bearing material separate from the container configured to cooperate with the closure mechanism of the container, wherein the adhesive is deposited on only one side of the strip of material and positioned to bind to itself; and a permanent seal created by the cooperation of the closure mechanism of the container and the strip of adhesive-bearing material, wherein the permanent seal cannot be broken without damaging the container, the trip of adhesive-bearing material, or both, thereby creating visual evidence of tampering with the container.
Owner:YAMBAO ALEXANDER +1

Drilling equipment for road detection

The utility model discloses drilling equipment for road detection, which comprises a mobile platform, a drilling assembly and a collecting assembly, the drilling assembly is movably mounted on one side of the interior of the mobile platform, the collecting assembly is fixedly mounted on one side of the drilling assembly, and the drilling equipment can be flexibly moved and used through the arrangement of the mobile platform. The drilling assembly is used for drilling machining, chippings generated after drilling are collected for use through the collecting assembly, and the drilling assembly comprises an upper support, a movable motor, a movable sleeve, a threaded rod, a telescopic air cylinder, a drilling motor and a drilling head. Before drilling and sampling, the movable motor drives the threaded rod to drive the movable sleeve to move back and forth to adjust the position needing to be drilled, the drilling equipment can accurately drill a core sample at a designated position, visual evidence is provided for the internal condition of a road structure, and therefore the detection accuracy is improved.
Owner:HEILONGJIANG LONGDU HIGHWAY ENG TESTING CO LTD

Multi-modal fine-grained instruction fine-tuning data construction method based on reverse verification

The application discloses a kind of based on reverse verification's multimodal fine-grained instruction fine-tuning data construction method, comprising: obtaining multimodal original data, the structured analysis of multimodal original data is obtained Verifiable visual evidence;Based on visual evidence generation fine-grained instruction, and then generate candidate answer set;For the candidate answer set of each fine-grained instruction, combined with verifiable visual evidence executes reverse verification, to obtain reverse verification score and failure reason label;To the candidate answer that passes through reverse verification, execute multi-criteria scoring and obtain multi-criteria score, and determine quality score based on reverse verification score and multi-criteria score;Filter based on quality score, output multimodal training data for instruction fine-tuning.The present application can construct traceable, verifiable, controllable difficulty and noise multimodal fine-grained instruction fine-tuning data without significantly increasing labor cost, improve the compliance ability and robustness of multimodal large model to fine-grained instruction.
Owner:MOLAR INTELLIGENCE INFORMATION TECHNOLOGY (HANGZHOU) CO LTD

Medical image segmentation method and system based on evidence-driven visual language model

The present invention discloses a medical image segmentation method and system based on an evidence-driven visual language model, and relates to the technical field of medical image segmentation. The method comprises the following steps: obtaining the original medical image to be segmented; constructing a multimodal segmentation model, calculating the similarity loss using a deviation-variance decomposition method based on a differential matrix, introducing uncertainty perspectives to represent the visual evidence embedding and the textual evidence embedding as visual perspectives and textual perspectives, respectively, and calculating the perspective loss, and calculating the segmentation loss using the segmentation difference between the visual-text fusion evidence and the true mask; optimizing the parameters of the multimodal segmentation model using the overall loss, and performing image segmentation on the original medical image to be segmented using the multimodal segmentation model. The present invention introduces evidence learning into the visual language model, and estimates the modal gap between the image and the text by aggregating the image-text perspectives converted from the evidence, thereby achieving multimodal fusion.
Owner:SHANDONG UNIV

Medical image report generation method, system and device and storage medium

PendingCN121439069AImage enhancementImage analysisData imbalanceClinical report
The invention discloses a medical image report generation method, system and device and a storage medium, and relates to the technical field of computer vision and natural language processing, and the method comprises the steps: constructing a causal graph model, taking an image as an input variable, taking a final report as an output variable, and taking a region-level pathological state in the image as an intermediary variable; the data imbalance factor is an unobservable hybrid factor; based on a causal graph model, intervening the intermediary variable by using a front door adjustment strategy, and establishing a causal path from visual evidence to report text; based on the intervened intermediary variable, executing a report generation process: a, identifying an abnormal region in the image by using a focus detection model and outputting a corresponding pathological discovery description; and b, inputting the intervened pathological discovery description and the original image into a visual language model to generate a complete clinical report by taking the intervened pathological discovery description and the original image as conditions. And the sensitivity of the model to pathological changes and the clinical reliability of report generation are improved.
Owner:ANHUI PROVINCIAL HOSPITAL

Causal decoupling based cross-media model error attribution and auxiliary correction method

The application discloses a cross-media model error attribution and auxiliary correction method based on causal decoupling, which comprises the following steps: image segmentation is performed on an input image, and the input image is divided into a plurality of sub-regions; a target function is solved according to the sub-regions, the minimum key region which supports the generation of a target word element to the greatest extent and removes the generated ability is sorted according to the contribution to the generation decision, and an attribution saliency map is generated; an influence score is calculated according to the change of the generation probability of the target word element in the process of inserting the sub-regions in the ordered subset; the ordered subset, the attribution saliency map and the influence score are evaluated from the dimensions of fidelity, positioning ability and error correction guiding ability, and an evaluation result is obtained; the evaluation result is obtained without relying on the internal gradient, attention weight or activation map of the model, the dependence on the internal structure of the model is reduced, the regional attribution of an arbitrary word element set is realized, and the relative dependence of the generation process on visual evidence and language prior is quantified.
Owner:SUN YAT SEN UNIVERSITY SHENZHEN +2

Weakly supervised video anomaly detection method based on visual language instance perception learning

The application discloses a kind of weak supervision video anomaly detection methods based on visual language instance perception learning, method includes the following steps: step A, focus and sweep multi-granularity video semantic extraction to video;Step B, video specificity visual and language contrast loss, and construct the loss function of video timing characteristics;Step C, construct parallel double-branch instance perception learning architecture, respectively simulate human perception and cognition channel;Step D, realize multi-stage differentiation training, and find abnormal clues in the remaining video segment.The application constructs parallel double-branch instance perception learning architecture, and video anomaly detection is promoted to the paradigm of "visual evidence + semantic priori" joint discrimination from the paradigm of simply relying on appearance / motion mutation, can simultaneously consider the coverage of significant anomaly and implicit anomaly under weak supervision condition, avoid that model is only dominated by a few high response segments, from mechanism, improve the integrity and stability of abnormal positioning.
Owner:GUANGZHOU RES INST OF XIAN UNIV OF ELECTRONIC SCI & TECH

Visual verification method and system

The invention discloses a visual evidence viewing method and system, and relates to the technical field of visualization, and the technical scheme is characterized by comprising the following steps: collecting a record text, and an associated text and a background text associated with the record text; processing the record text to obtain a record text feature set and an associated paragraph set, performing paragraph segmentation on the associated text according to the semantic features to obtain an associated text feature set, and performing paragraph segmentation on the record text according to the background text to obtain a background text feature set; performing comparative analysis on the associated text feature set and the associated paragraph set to obtain an intention association result set; judging whether the record text feature set and the background text feature set are in a subject-related condition or not, and performing processing and analysis to obtain a first intention factor set and a second intention factor set respectively; the method has the effect that potential evidence association can be found by comparing and analyzing different text feature sets and paragraph sets.
Owner:杭州威灿科技有限公司

Medical imaging report generation method and system based on large language model

The present application discloses a method and system for generating medical imaging reports based on a large language model, which relates to the field of imaging report generation. It first obtains the patient's original imaging data and clinical background text, and extracts the image embedding vector and background embedding vector respectively. Then, these embedding vectors are used to perform case retrieval based on prior knowledge, and K highly relevant historical case reports are screened out from massive historical data. Subsequently, the image embedding vectors and historical case reports are input into the observation large language model and output in a structured JSON format. Finally, a large language model is written to integrate visual evidence JSON, clinical background text and historical case reports to generate the final medical imaging report. This method effectively solves problems such as incoherent report logic and inaccurate information through a phased and multimodal fusion approach, significantly improving the quality and generation efficiency of the report.
Owner:ZHEJIANG FEITU IMAGING TECH CO LTD

Recycled article weighing data recording and associating method and system based on multiple cameras

The invention discloses a recycled article weighing data recording and associating method and system based on multiple cameras, the system comprises an intelligent recycling all-in-one machine, a background server and an applet client, the intelligent recycling all-in-one machine is internally provided with a main control module, a multi-camera module, a weight acquisition module and a man-machine interaction module; the background server is logically associated with the applet client, and associated data of the intelligent recycling all-in-one machine can be obtained by logging in the applet client. The core part of the method comprises the steps of system initialization and equipment binding, weighing instruction triggering and multi-camera synchronous snapshot, automatic association of image data and business documents, structured storage and unified display. According to the invention, through synchronously capturing two groups of images of the scale reading and the object state, a complete visual evidence chain including the object, the reading, the time and the environment is formed, and the problems of trust crisis and weight disputes caused by single or missing evidence in a traditional mode are fundamentally solved.
Owner:OUYE LIANJIN RENEWABLE RESOURCES CO LTD

Cosmetic product detail picture element extraction method and system based on multi-modal model

The invention relates to the technical field of deep learning, in particular to a makeup product detail picture element extraction method and system based on a multi-modal model, and the method comprises the steps: extracting a text region sequence and a visual region sequence in a detail picture, and obtaining the position coordinates of each text region and each visual region; calculating a visual weight according to the information entropy of each visual area, and carrying out weighted correction on the initial visual features of the visual areas to obtain enhanced visual features; calculating the semantic similarity between the text features and the enhanced visual features; constructing a logic correlation coefficient according to a vertical distance between the text region and the visual region, and correcting the semantic similarity to obtain a comprehensive similarity; and determining a visual area corresponding to the text area according to the comprehensive similarity, and outputting an element extraction result. Through the technical scheme of the invention, accurate extraction and alignment of the claim text, the experimental data, the component graph and other visual demonstration elements in the detail graph of the beauty makeup product are realized.
Owner:GUANGZHOU XINSHU INTELLIGENT TECH CO LTD

A hallucination mitigation method and device of a multi-modal large language model

ActiveCN121661442BSolve wasteAccurately preserve unique visual evidenceCharacter and pattern recognitionInference methodsLinguistic modelEngineering
The present application belongs to the technical field of artificial intelligence and multi-modal large language model, and particularly relates to a hallucination reduction method and device for multi-modal large language model. The method comprises the following steps: obtaining multi-expert visual tokens of an original image and text tokens of a text prompt, and inputting the connection of the two into a multi-modal large language model; calculating a similarity matrix of the multi-expert visual tokens and an adaptive clustering threshold, performing hierarchical clustering on the multi-expert visual tokens, merging redundant tokens and retaining complementary visual evidence; constructing a negative sample pool and distributing it to an auxiliary expert with the largest feature difference through a CLIP anchored complementary rejection fine-tuning strategy, and performing specialized training on the corresponding projectors of each expert; and inputting the clustered and fused visual tokens and the text tokens into the multi-modal large language model to generate target text output. The present application solves the problems of information redundancy, large computational overhead and insufficient exertion of expert advantages in the existing multi-expert multi-modal large language model hallucination reduction method.
Owner:SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN +1

Management method and system for data asset entry

The invention relates to the technical field of data management, and discloses a management method and system for data asset entry, and the method comprises the steps: collecting project approval data, evaluating the technical feasibility and market profit, verifying four essential components of capitalization, and forming a project approval voucher package and an approval record. Establishing a data resource catalog and classification, registering information elements and measuring integrity, checking consistency and issuing changes, and forming an information ledger and version; and enterprise and project double-view compliance diagnosis and closed loop rectification are carried out, ownership certification is collected and examined, and a compliance and ownership certificate package and risk measurement are output. According to the method, a standardized table entry system is established, then a standardized table entry process is realized, table entry steps are clarified, detailed guidance is provided for each table entry step through the system, clear, sufficient, complete and visual evidence chains are guided and collected, product content gathers knowledge and rich experience of domestic table entry experts, responsibilities of all departments are clarified, and internal cooperation efficiency is improved.
Owner:JILIN JIANXING INTELLIGENT TECH CO LTD

Blockchain-based strip material life cycle traceability and management method and system

The application discloses a kind of based on blockchain's strip material life cycle traceability and management method and system, for the technical field of ship material management.The present application generates unique digital identity when raw material large plate enters, packs and chains basic attribute, quality inspection hash, visual evidence, location information and operator digital signature;Automatic establish the parent-child relationship chain of raw material-strip-spare material when cutting processing;Material circulation and real-time acquisition multi-dimensional evidence and solidification are stored in the process of processing;Through smart contract, realize automatic check and abnormal alarm according to material balance relationship;Digitally identify and store management to available spare material, and complete whole life cycle verification through outsourcing return factory closed loop.The present application can realize strip material whole process tamper-proof traceability, data credible cross verification, loss accurate control and spare material efficient reuse, greatly improve material management efficiency and material utilization.
Owner:GUANGZHOU SHIPYARD INTERNATIONAL LTD

System and method for preventing copied IC (integrated circuit) card from passing after statistical analysis of passing record data

The invention discloses a system and a method for preventing a copied IC card from passing after statistical analysis of passing record data, the system comprises a card swiping device, a passing control and execution device, an upper computer management system, a card issuing device and a camera, the upper computer management system comprises an account opening and card issuing center and a front communication service; the card swiping device is connected with the passing control and execution device, the passing control and execution device is connected with the front communication service, the card issuing device is connected with the account opening and card issuing center, and the camera is connected with the passing control and execution device. Through traffic record statistical analysis and multi-source data fusion, a copy card use path is actively identified and visual evidences are pushed, so that the problems of passivity, isolation and tracing in the prior art are fundamentally solved.
Owner:XIAMEN DNAKE INTELLIGENT TRANSPORTATION TECH

Passport cross-modal retrieval method and system based on knowledge graph constraint

This invention belongs to the field of artificial intelligence technology and discloses a passport cross-modal retrieval method and system based on knowledge graph constraints. It aims to address the problems of domain semantic gap, lack of complex semantic parsing capabilities, and black-box nature of the retrieval decision-making process in existing technologies. This invention achieves a unified approach to high-precision retrieval and interpretability of the reasoning process under complex semantic understanding by constructing a passport multimodal knowledge graph, designing spatial inductive bias, building a multi-granularity feature deep fusion mechanism, and generating a visual evidence chain.
Owner:HUAZHONG UNIV OF SCI & TECH