Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

23 results about "Visual instruction" patented technology

Shape operation information display method and system based on multi-source information fusion

The invention discloses a weather modification operation information display method and system based on multi-source information fusion, and relates to the technical field of meteorological information intelligent decision making, and the method comprises the steps: carrying out the reasoning analysis of a time sequence knowledge graph through a graph neural network, recognizing the cloud system development change, reasoning the evolution of cloud physical characteristics, and predicting the space-time range of an operation potential region, generating an accurate forecast conclusion and a figure operation suggestion; and converting the accurate forecast conclusion and the weather modification operation suggestion into a visual instruction set, and generating a visual comprehensive situation map through highlighting core weather indexes, displaying evolution paths by dynamic arrows and marking operation potential areas and operation corridors by color coverage. According to the method, reasoning analysis is carried out on the time sequence knowledge graph by using the graph neural network, the operation potential area is automatically identified from a historical mode, the dynamic evolution of the operation potential area is predicted, and the perspectiveness and the accuracy of figure operation decision making are improved.
Owner:辽宁省人工影响天气办公室

Multi-modal large model illusion detection method based on reverse visual localization

The invention belongs to the technical field of artificial intelligence, and particularly relates to a multi-modal large model illusion detection method based on reverse visual positioning. The method comprises the following steps: constructing a visual instruction fine tuning data set rich in context; training a visual positioning large model with pixel-level positioning and rejection capability based on the data set; performing sentence-by-sentence verification on a response generated by a to-be-detected multi-modal large model by using the trained model, and judging whether illusion exists or not by judging whether text description can be reversely positioned back to image pixels or not; according to the method, illusion rich in details can be effectively detected, pixel-level masks and natural language interpretation are provided, and the accuracy and transparency of evaluation are remarkably improved.
Owner:FUDAN UNIV YIWU RES INST +1

Visual variant-based model training method and device, equipment and medium

The embodiment of the invention provides a model training method and device based on visual variants, equipment and a medium. The method relates to a sample processing technology, is applied to the financial field and the medical care field, and comprises the following steps: obtaining an original image, generating an original title based on the original image, extracting a corresponding object segmentation mask, and obtaining a corresponding variant title to generate a visual variant image; generating an image set according to the original image and the visual variant image, obtaining detection illusion question information and real answer information of the corresponding image in the image set, and generating a question-answer pair; and performing consistency verification on each question and answer pair according to a plurality of preset visual language models, generating a visual instruction data set, and finishing training of the visual language model to be trained by utilizing the visual instruction data set. A universal solution is provided for the fields of finance, medical treatment and the like with extremely high reliability requirements, and the performance limitation of a traditional text optimization strategy is broken through through essential improvement of the visual understanding ability.
Owner:PING AN TECH (SHENZHEN) CO LTD

Natural language question and answer-based operation and maintenance scene visual report generation method and system

PendingCN122451138AEngineeringSemantic feature
The application provides a kind of operation and maintenance scene visualization report generation method and system based on natural language question and answer, belongs to intelligent operation and maintenance technical field, method includes: receiving the natural language query input by user;Through the pre-training of large language model and the preset operation and maintenance terminology dictionary, the natural language query is parsed and entity is extracted, and the entity-field association table containing demand type is generated;According to demand type and semantic feature, the query type is judged to be data query or visual query;Based on the judgment result, generate structured query language sentence or visual instruction;Query is executed to the interface business database to obtain raw data;Raw data is processed and analyzed to obtain analysis result data;Call data feature adaptation algorithm to match chart type, generate and output visual report.The application realizes the full-link automation from natural language input to visual report output, reduces the operation and maintenance data interaction threshold, improves the operation and maintenance data processing efficiency and accuracy.
Owner:CHINESE PEOPLES LIBERATION ARMY INFORMATION SUPPORT CORPS ENGINEERING UNIVERSITY

Digital technology assisted physical creation method, device, product and medium

The invention discloses a method, equipment, a product and a medium for digital technology assisted physical creation, and relates to the field of digital technologies. The method comprises the steps of collecting a multi-view image of a physical creation entity and performing three-dimensional reconstruction to generate current state data; comparing the data with a preset digital target model to generate deviation data; generating a calibration visual instruction based on the deviation data, and displaying a calibration mark on a front-end display interface to guide a user; and collecting interaction data including confirmation or denial of the calibration instruction by the user. According to the method, the creation deviation can be dynamically calibrated, the conformity of the physical entity and the digital target model is improved, meanwhile, the final decision-making right of a creator is reserved through user interaction, interaction feedback of man-machine cooperation in the physical creation process is achieved, and then the physical creation efficiency of the creator is improved.
Owner:GUANGZHOU GUDONG INTELLIGENT TECHNOLOGY CO LTD

Spatial position instruction fine tuning method based on multi-modal large language model

The invention relates to a spatial position instruction fine tuning method based on a multi-modal large language model, and the method comprises the following steps: S1, converting a spatial position reasoning data set into a visual instruction format through employing a dialogue template, and obtaining a visual spatial position reasoning data set; s2, acquiring a large language model InternVL as a multi-modal large language model, performing pre-training on the general data set to obtain a pre-training model, reasoning the data set based on the visual spatial position, adjusting parameters of the pre-training model by adopting a low-rank adaptation method to obtain a trained large language model, and outputting a description corresponding to a spatial task by the large language model; and S3, introducing a text-based large language model, and optimizing the description corresponding to the space task based on the large language model. Compared with the prior art, the method has the advantages that the ability of the multi-modal large language model in understanding and generating context rich description is fully utilized, and the ability of the model in generating accurate and detailed description is enhanced.
Owner:SHANGHAI JIAOTONG UNIV

AI interaction system for oral science popularization

The invention, which relates to the technical field of medical information data, discloses an AI interaction system for oral science popularization, comprising a multi-modal sensing unit, an AI analysis engine, a knowledge graph matching unit and an AR interaction feedback unit. According to the method, the AI analysis engine is used for deeply analyzing the real oral image and the tooth brushing action of the user, and dynamic weighting of the knowledge graph is combined, so that accurate matching of science popularization content with the actual focus, the wrong action and the subjective requirement of the user is ensured, the knowledge transmission effectiveness is greatly improved, and the user experience is improved based on the graded medical knowledge base. The system can automatically adjust the weight according to the health influence degree of the oral cavity problem, preferentially present high-risk lesion early warning, and convert abstract medical knowledge into a visual real-time visual instruction by using an AR virtual-real overlapping technology, so that a user can obtain an action correction suggestion based on pixel-level alignment in the tooth brushing process, and the user experience is improved. Through fusion of voice interaction and AR interaction, the threshold of common people for understanding professional medical knowledge is reduced.
Owner:NANJING STOMATOLOGICAL HOSPITAL

system

We provide the system. [Solution] Means for acquiring user care data, medical information, and preference information, A means of analyzing acquired data to generate the optimal care plan, A means for transmitting the generated care plan to the caregiver's information device, A method for analyzing work-related memos and instructions entered by caregivers and updating the database, A method for caregivers to perform their duties while checking care plans in real time and receiving visual instructions using smart glasses, A system that includes this.
Owner:SOFTBANK GROUP CORP

Scene consistency image instruction generation method based on task decomposition

The present application relates to a task decomposition-based scene consistency image instruction generation method, and the generation method comprises the following steps: step S1, image analysis, object recognition, position and state recognition are performed; step S2, instruction understanding is performed, and set key information is extracted; step S3, feature extraction of a pre-scene image in a latent space is realized through a VAE encoder; step S4, a task decomposition generation step sequence is generated, and step-by-step description is performed; step S5, based on the step sequence, an instruction weight matrix with front and rear correlations is generated; the instruction weight matrix with front and rear correlations is designed based on the step sequence of step S4; step S6, based on the image reference of the pre-scene and the instruction weight, a visual instruction of the current scene is generated; the present application is reasonable in design, compact in structure and convenient to use.
Owner:QINGDAO HAIDA NOVA SOFTWARE CONSULTING CO LTD

Multi-source equipment state visual monitoring method of docking station and related device

The invention discloses a multi-source equipment state visual monitoring method of a docking station and a related device, and the method comprises the steps: obtaining original data traffic on an uplink data bus between a target docking station and host equipment; calculating an average flow value and a flow change rate sequence of the original data flow; calculating an I / O pressure index of the uplink data bus; determining whether the docking station is in a performance sensitive state based on the I / O pressure index; when the docking station is not in the performance sensitive state, obtaining first state data of the host equipment and second state data of the external equipment through a first preset sampling frequency; when it is determined that the docking station is in the performance sensitive state, first state data of the host device and second state data of the external device are obtained through a second preset sampling frequency; generating a visualization instruction of the running states of the host equipment and the external equipment; and controlling a display screen of the docking station to carry out visual presentation according to the visual instruction. According to the invention, the adaptability of visual monitoring of the docking station can be improved.
Owner:SHENZHEN SINOBRY ELECTRONICS LTD

A human shadow operation information display method and system based on multi-source information fusion

The application discloses a kind of based on multi-source information fusion's human shadow operation information display method and system, it is related to meteorological information intelligent decision-making technical field, including, utilize graph neural network to carry out inference analysis to time series knowledge graph, identify cloud system development change and infer the evolution of cloud physical characteristics, predict the spatiotemporal range of operation potential area, generate accurate forecast conclusion and human shadow operation suggestion;Accurate forecast conclusion and human shadow operation suggestion are converted into visual instruction set, and evolution path is shown through highlighting core weather index, dynamic arrow display, and color overlay marks operation potential area and operation corridor, generates visual comprehensive situation chart.The application is analyzed by utilizing graph neural network to time series knowledge graph, realizes from historical mode automatically identifying operation potential area and predicting its dynamic evolution, improves the foresight and accuracy of human shadow operation decision.
Owner:辽宁省人工影响天气办公室

Data-efficient visual instruction tuning for multimodal large language models

According to one aspect, instruction tuning may include generating a set of instructions for a reference set of images selected from a set of images based on one or more task specific instruction generation protocols, generating one or more task importance weights for the reference set of images based on the set of instructions and the reference set of images and a ratio of a first loss of a first loss function associated with a response and an image from reference set of images and a second loss of a second loss function associated with the response, a question, and the image from reference set of images, and generating a set of instructions for a remaining set of images from the set of images based on one or more of the task importance weights, k-means clustering, and neighbor centrality from a cluster of the k-means clustering.
Owner:HONDA MOTOR CO LTD

Intelligent data question and answer method and system fusing field large language model

The invention provides an intelligent data question and answer method and system fusing a field large language model, and relates to the technical field of large language models.The method comprises the steps that action data and view angle data of a user in a pre-constructed three-dimensional virtual scene corresponding to a target field are obtained, and question and answer data of the target field are collected; performing association analysis on the action data and the view angle data to determine a user query intention; performing associated coding processing on the query intention of the user and the scene position of the three-dimensional virtual scene to construct an intelligent question and answer library; obtaining a resource allocation scheme by using a software defined network technology; searching target question and answer data corresponding to the query request data from an intelligent question and answer library through a domain large language model, and generating a target question and answer result and a visual instruction; based on the resource allocation scheme and the visualization instruction, intelligent data question answering is completed, and high-immersion, low-delay and accurate-intention intelligent data question answering and visualization presentation oriented to the specific field are achieved.
Owner:FIVE DIMENSIONS INTELLIGENT TECHNOLOGY (SHANGHAI) CO LTD +1

Marketing originality automatic generation and optimization system based on deep learning

The invention discloses a marketing originality automatic generation and optimization system based on deep learning, and relates to the technical field of computers. Through a small sample stylization unit, the system can construct a style reference set only by relying on a small number of representative originality materials; the generated content distribution and the reference distribution are aligned in a unified depth feature space, so that the generated image and copywriting are closer to the target brand tonality in the dimensions of color, composition, tone and the like, and a differentiated vision and utterance system is quickly molded in the absence of large-scale brand data; in combination with a cross-modal semantic alignment unit, text information such as marketing targets, audience descriptions and the like and brand visual instructions are jointly coded into a unified semantic-style condition, so that the double constraints of what and what can be satisfied in the same generation process, and the risks of disjunction between copywriting and pictures and style deviation are reduced.
Owner:BEIJING HEJIN TECHNOLOGY CO LTD

A method, system, and storage medium for robotic handling in low gravity environments

ActiveCN120680516BSolve the problem of irregular rotation and difficulty in grabbingProgramme-controlled manipulatorComputer graphics (images)Angular velocity
The application discloses a robot carrying method and system in a low-gravity environment and a storage medium. A simple sketch drawn by an operator and a real scene image in a current cabin are collected, the simple sketch and the real scene image are input into a double-branch visual encoder model which has been trained, a basic action sequence instruction for controlling a robot to complete a material carrying operation is generated, and then a target rotation angular velocity is monitored in real time during target grabbing by executing the basic action sequence instruction. If the target rotation angular velocity is greater than a set threshold, the robot gripper is driven to rotate in a reverse direction for fine adjustment. Thus, pure visual instruction interaction is realized to adapt to the situation that voice interaction is completely unavailable in a low-gravity scene, and the problem that an object is easily subjected to irregular rotation and is difficult to be grabbed under low gravity is solved.
Owner:58 INTELLIGENT TECH (HANGZHOU) CO LTD

A robot material handling method, robot, and storage medium

The machine material carrying method, the robot and the storage medium disclosed by the application obtain a sketch depicting an appearance of a work target and carrying destination position information, extract a shape feature of the work target from the sketch, collect environment data in a current task scene, extract a visual feature from the environment data, align the visual feature with the shape feature by using a contrast loss function, extract a shape feature of the work target from a target region, calculate a target pose and confirm a material type from the environment data based on the target region, finally query a preset control information library based on the target object type to obtain corresponding action constraint information, combine the action constraint information, the target pose and the destination position information, generate a material carrying instruction, and control each actuator to move to complete a material carrying task. The material carrying can be realized in a noisy industrial scene by only using a pure visual instruction without inputting a complex text instruction.
Owner:58 INTELLIGENT TECH (HANGZHOU) CO LTD

Robot operation control method, device and equipment and readable storage medium

The invention relates to a robot operation control method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the steps of obtaining a task operation instruction and an operation image corresponding to a task operation environment; identifying the operation image to obtain a target identification result; the target recognition result comprises operation object information and execution component information of the robot; determining the task operation progress of the robot based on the operation object information and the execution component information; fusing the task operation instruction and the target identification result to obtain visual instruction joint information; and according to the execution component information, the task operation progress and the visual instruction joint information, determining the next action of the robot until a target operation task corresponding to the task operation instruction is completed. By adopting the method, the task completion accuracy can be improved.
Owner:SHENZHEN POWER SUPPLY BUREAU

Method and system for generating visual instruction chain based on fine-tuning large language model

The invention discloses a method and a system for generating a visual instruction chain based on a fine-tuning large language model. The method comprises the following steps: receiving a task target described in a natural language; analyzing the natural language description on the basis of understanding the power data characteristics and the power business context related to the task target; according to an analysis result, generating an instruction chain for visual presentation through a fine-tuned large language model; according to the method, it can be ensured that conversion from the user intention to the visualization result is accurate and efficient, and a visual and friendly display effect is provided for data analysis in the power field.
Owner:STATE GRID JIANGSU ELECTRIC POWER CO LTD +1

Document processing method and device, equipment, medium and program product

The invention provides a document processing method which can be applied to the technical field of large models. The method comprises the following steps: comparing a to-be-processed document with a cross-age feature map and a visual standard blueprint to generate a visual instruction; driving a large language model to adjust at least one of document terms, document logics, document styles and document sentiment values of the to-be-processed document based on the visualization instruction, and generating a replacement document; and executing parallel arbitration verification on the alternative document, and outputting the alternative document in response to verification passing. The invention further provides a document processing device and equipment, a storage medium and a program product.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Information pushing method and device, electronic equipment and storage medium

The application provides an information pushing method and device, electronic equipment and storage medium, relates to the field of financial technology, and includes obtaining text features of multiple market text information of a candidate pushing product and visual features obtained by fusing multiple visual data of the candidate pushing product, the visual features being obtained by a pre-trained investment information pushing model performing feature extraction on visual data based on corresponding visual instructions in multiple pre-trained visual instructions according to a task scene to which the visual data belongs in multiple preset task scenes, the multiple visual instructions corresponding to the multiple task scenes one by one; performing multi-modal feature fusion on the text features and the visual features, performing market trend prediction based on emotional analysis on the candidate pushing product according to the fused features, determining a pushing strategy for pushing the candidate pushing product to a target object, and performing investment information pushing according to the pushing strategy. The embodiments of the application can push more accurate investment information to the target object.
Owner:PING AN TECH (SHENZHEN) CO LTD

Interactive intelligent document question and answer method and system and storage medium

The invention provides an interactive intelligent document question-answering method and system and a storage medium, and the intelligent document question-answering method comprises the steps: S1, obtaining document data, and converting the document data into image data; s2, acquiring a text input by a user, and converting the input text into a visual instruction feature; and S3, performing visual observation processing on the image data according to the visual instruction features to obtain a target text answer. Due to the adoption of a pure vision scheme, the limitation of an OCR (Optical Character Recognition) technology is completely avoided, and unstructured contents which cannot be processed by a traditional Text-RAG method can be effectively processed. In addition, the reading behavior is self-adapted by simulating multiple steps of retrieval-search-focusing of human beings. And when needed, related pages are'watched 'or details are magnified, so that huge waste caused by the fact that all the pages are'watched at one time' in a large VLM is avoided, and the efficiency is higher.
Owner:HUBEI CHINA TOBACCO INDUSTRY CO LTD

Surgical training device with digital feedback and augmented reality guidance

PendingUS20260196140A1Digital feedbackSimulation
A modular, ergonomic training system designed to teach and evaluate laparoscopic surgical skills through a collapsible frame featuring adjustable trocar ports that simulate real-world surgical scenarios for both right- and left-handed users. It includes a pegboard for attaching various training modules to practice tasks like object manipulation, precision cutting, and suture tying. The system integrates augmented reality (AR) and artificial intelligence (AI) for real-time, step-by-step guidance and performance analysis. AR overlays on a live video feed from a tablet mounted on the frame offer visual instructions, while AI algorithms evaluate metrics such as precision and technique, providing personalized feedback and automated skill assessment. Additional features include a rechargeable light source that mimics clinical lighting and reflective surfaces for enhanced performance monitoring. The system enables efficient, scalable training with cloud-based data storage for tracking progress, supporting remote coaching and long-term skill development.
Owner:CILAG GMBH INTERNATIONAL