Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

43 results about "Visual instruction" patented technology

Air-ground cooperation system and air-ground cooperation method based on own intelligence

The invention discloses an air-ground cooperation system and an air-ground cooperation method based on intelligence, and relates to the field of intelligent unmanned system control. The system comprises an intelligent unit, a multi-element sensing unit, a communication unit and a task execution unit. Based on a Vision-Language-Action (VLA) framework, an intelligent unit with a body realizes end-to-end analysis of a natural language instruction and a visual instruction, fuses sensing data of an unmanned aerial vehicle and an unmanned vehicle, generates a joint feature code with time-space alignment, and dynamically allocates a cooperative task. The communication unit supports low-delay end-to-end communication and real-time data synchronization between clusters, and ensures efficient interaction between agents. The multi-element sensing unit is integrated with a visual camera, a laser radar and high-precision positioning equipment, and air-ground environment multi-mode sensing is achieved. And the task execution unit supports the unmanned aerial vehicle to complete coordination actions in a complex scene. According to the scheme, the problems of insufficient multi-mode perception, limited natural language interaction and low dynamic cooperation efficiency of the existing system are solved, and the cooperation capability and system robustness of the unmanned aerial vehicle and the unmanned aerial vehicle in a complex environment are remarkably improved.
Owner:NANJING UNIV

Fine-tuning a neural network model using a weight-decomposed low-rank adaptation

Embodiments of the present disclosure relate to fine-tuning a neural network model using a weight-decomposed low-rank adaptation (DoRA). DoRA reduces the number of parameters that are fine-tuned, thereby reducing memory and the time needed to fine-tune the parameters. the Pre-trained weights are decomposed into two components, magnitude and direction, which are separately fine-tuned. The magnitude components are fine-tuned while the direction components remain unchanged (frozen). Then low-rank adaptation (LoRA) is used to fine-tune the direction components, efficiently minimizing the number of trainable parameters. Compared with using LoRA to fine-tune the weights directly, using DoRA exhibits a closer resemblance to full fine-tuning's learning behavior and improves upon LoRA in commonsense reasoning and visual instruction tuning tasks. By employing DoRA, both the learning capacity and training stability of LoRA is enhanced. The fine-tuned decomposed magnitude and direction components may be merged into the pre-trained weights to avoid any additional inference overhead.
Owner:NVIDIA CORP

Shape operation information display method and system based on multi-source information fusion

The invention discloses a weather modification operation information display method and system based on multi-source information fusion, and relates to the technical field of meteorological information intelligent decision making, and the method comprises the steps: carrying out the reasoning analysis of a time sequence knowledge graph through a graph neural network, recognizing the cloud system development change, reasoning the evolution of cloud physical characteristics, and predicting the space-time range of an operation potential region, generating an accurate forecast conclusion and a figure operation suggestion; and converting the accurate forecast conclusion and the weather modification operation suggestion into a visual instruction set, and generating a visual comprehensive situation map through highlighting core weather indexes, displaying evolution paths by dynamic arrows and marking operation potential areas and operation corridors by color coverage. According to the method, reasoning analysis is carried out on the time sequence knowledge graph by using the graph neural network, the operation potential area is automatically identified from a historical mode, the dynamic evolution of the operation potential area is predicted, and the perspectiveness and the accuracy of figure operation decision making are improved.
Owner:辽宁省人工影响天气办公室

Multi-modal large model illusion detection method based on reverse visual localization

The invention belongs to the technical field of artificial intelligence, and particularly relates to a multi-modal large model illusion detection method based on reverse visual positioning. The method comprises the following steps: constructing a visual instruction fine tuning data set rich in context; training a visual positioning large model with pixel-level positioning and rejection capability based on the data set; performing sentence-by-sentence verification on a response generated by a to-be-detected multi-modal large model by using the trained model, and judging whether illusion exists or not by judging whether text description can be reversely positioned back to image pixels or not; according to the method, illusion rich in details can be effectively detected, pixel-level masks and natural language interpretation are provided, and the accuracy and transparency of evaluation are remarkably improved.
Owner:FUDAN UNIV YIWU RES INST +1

XR equipment-oriented remote expert guidance intention visualization method, equipment and device

The invention discloses an XR equipment-oriented remote expert guidance intention visualization method, equipment and device. The method comprises the steps of receiving a multi-mode instruction sent by a remote expert in combination with a real-time scene preview image; inputting the multi-modal instruction and the associated scene preview image into a visual language model, generating a structured intention data packet, and generating a dynamic 3D visual instruction matched with the operation action intention in the structured intention data packet based on the structured intention data packet; and performing space anchoring on the 3D visual instruction and a real-time picture of the XR equipment of the field user, and determining a rendering position of the 3D visual instruction according to a matching result so as to render the 3D visual instruction in the real-time picture of the XR equipment. By means of the method, experts can express complex operation intentions in a mode conforming to human intuition, and communication ambiguity is reduced.
Owner:HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD

Layperson audiovisual tourniquet

A tourniquet includes: a strap; a housing containing a winding mechanism, a first end of the strap being attached to a spool shaft of the winding mechanism; a buckle assembly attached to the housing; a buckle tongue attached to a second end of the strap for releasably coupling the strap to the buckle assembly; a first portion of the winding mechanism configured to reduce a length of the strap extending outside of the housing; and an instruction portion that is configured to emit one of audible and visual instructions to a user. The emitting of a particular one of the instructions results from a user step being performed with the tourniquet, the rotating handle is initially decoupled from the first portion of the winding mechanism, and rotation of the rotating handle causes the rotating handle to become mechanically coupled to the spool shaft.
Owner:INNOVITAL LLC +1

Case maintenance process automatic generation method and system based on image recognition

The invention discloses a method and system for automatically generating a case maintenance process based on image recognition, and belongs to the technical field of case maintenance, and the method comprises the following steps: S1, collecting image information in a case; s2, constructing a multi-task model based on a convolutional neural network, and performing image recognition on the image information to obtain a recognition result of each component; s3, performing integrated analysis on the identification result, and outputting structured case state information; s4, retrieving a maintenance process template matched with the structured case state information in a process library, and performing dynamic adjustment to obtain maintenance steps; and S5, outputting the maintenance step as a visual instruction for reference of maintenance personnel. According to the case maintenance process automatic generation method and system based on image recognition, personalized maintenance processes can be provided for different types of cases, so that the maintenance efficiency and accuracy are greatly improved.
Owner:BEIJING RUIDE KENUO ELECTRONICS EQUIP

Robot three-dimensional space constraint refining operation control system and method

The invention discloses a robot three-dimensional space constraint refining operation control system and method, and belongs to the field of robot control. The system comprises a task planning module for decomposing a natural language instruction into a sub-task instruction and a visual instruction, transmitting the visual instruction to a constraint extraction module, and transmitting the sub-task instruction to a sub-task execution module; the constraint extraction and refinement module is used for forming a component-level spatial constraint through multi-modal large-model semantic reasoning in combination with three-dimensional space coordinate mapping, judging whether the component-level spatial constraint needs to be refined or not, executing corresponding processing and returning a constraint result to the task planning module so as to be transmitted to the subtask execution module; and the sub-task execution module receives the sub-task instruction and the constraint result, firstly calls the execution condition construction module to obtain an execution condition required by the sub-task instruction in combination with the constraint result, obtains a robot key speed control instruction from the server according to the execution condition, and controls the robot to execute corresponding operation. According to the invention, the robot can be controlled to operate efficiently and accurately in a complex environment.
Owner:UNIV OF SCI & TECH OF CHINA

Visual variant-based model training method and device, equipment and medium

The embodiment of the invention provides a model training method and device based on visual variants, equipment and a medium. The method relates to a sample processing technology, is applied to the financial field and the medical care field, and comprises the following steps: obtaining an original image, generating an original title based on the original image, extracting a corresponding object segmentation mask, and obtaining a corresponding variant title to generate a visual variant image; generating an image set according to the original image and the visual variant image, obtaining detection illusion question information and real answer information of the corresponding image in the image set, and generating a question-answer pair; and performing consistency verification on each question and answer pair according to a plurality of preset visual language models, generating a visual instruction data set, and finishing training of the visual language model to be trained by utilizing the visual instruction data set. A universal solution is provided for the fields of finance, medical treatment and the like with extremely high reliability requirements, and the performance limitation of a traditional text optimization strategy is broken through through essential improvement of the visual understanding ability.
Owner:PING AN TECH (SHENZHEN) CO LTD

Natural language question and answer-based operation and maintenance scene visual report generation method and system

PendingCN122451138AEngineeringSemantic feature
The application provides a kind of operation and maintenance scene visualization report generation method and system based on natural language question and answer, belongs to intelligent operation and maintenance technical field, method includes: receiving the natural language query input by user;Through the pre-training of large language model and the preset operation and maintenance terminology dictionary, the natural language query is parsed and entity is extracted, and the entity-field association table containing demand type is generated;According to demand type and semantic feature, the query type is judged to be data query or visual query;Based on the judgment result, generate structured query language sentence or visual instruction;Query is executed to the interface business database to obtain raw data;Raw data is processed and analyzed to obtain analysis result data;Call data feature adaptation algorithm to match chart type, generate and output visual report.The application realizes the full-link automation from natural language input to visual report output, reduces the operation and maintenance data interaction threshold, improves the operation and maintenance data processing efficiency and accuracy.
Owner:CHINESE PEOPLES LIBERATION ARMY INFORMATION SUPPORT CORPS ENGINEERING UNIVERSITY

Auricular point model teaching aid

The utility model belongs to the technical field of teaching aids, and particularly relates to an auricular point model teaching aid, which comprises an ear model body for simulating the outline of an ear of a human body, and the outer wall of the ear model body is provided with a key switch capable of being pressed at the position corresponding to each auricular point, and the key switch is used for sending a signal when being pressed. Corresponding auricular point names are carved at the positions, corresponding to auricular points, of the inner wall of the ear model body, and LED lamps are embedded in the positions respectively and used for being turned on when receiving signals so that the corresponding auricular point names can be displayed through a shell, not completely transparent, of the ear model body. A microcontroller, a voice chip, a loudspeaker and a lithium battery are arranged in the ear model body, a display switch is arranged on one side of the ear model body and used for controlling the display state of the LED lamp, and when the teaching aid is used, visual teaching can be provided, interactivity is achieved, and the teaching quality can be effectively improved.
Owner:GUANGWAI HOSPITAL XICHENG DISTRICT BEIJING (GUANGWAI GERIATRIC HOSPITAL XICHENG DISTRICT BEIJING)

Digital technology assisted physical creation method, device, product and medium

The invention discloses a method, equipment, a product and a medium for digital technology assisted physical creation, and relates to the field of digital technologies. The method comprises the steps of collecting a multi-view image of a physical creation entity and performing three-dimensional reconstruction to generate current state data; comparing the data with a preset digital target model to generate deviation data; generating a calibration visual instruction based on the deviation data, and displaying a calibration mark on a front-end display interface to guide a user; and collecting interaction data including confirmation or denial of the calibration instruction by the user. According to the method, the creation deviation can be dynamically calibrated, the conformity of the physical entity and the digital target model is improved, meanwhile, the final decision-making right of a creator is reserved through user interaction, interaction feedback of man-machine cooperation in the physical creation process is achieved, and then the physical creation efficiency of the creator is improved.
Owner:GUANGZHOU GUDONG INTELLIGENT TECHNOLOGY CO LTD

Spatial position instruction fine tuning method based on multi-modal large language model

The invention relates to a spatial position instruction fine tuning method based on a multi-modal large language model, and the method comprises the following steps: S1, converting a spatial position reasoning data set into a visual instruction format through employing a dialogue template, and obtaining a visual spatial position reasoning data set; s2, acquiring a large language model InternVL as a multi-modal large language model, performing pre-training on the general data set to obtain a pre-training model, reasoning the data set based on the visual spatial position, adjusting parameters of the pre-training model by adopting a low-rank adaptation method to obtain a trained large language model, and outputting a description corresponding to a spatial task by the large language model; and S3, introducing a text-based large language model, and optimizing the description corresponding to the space task based on the large language model. Compared with the prior art, the method has the advantages that the ability of the multi-modal large language model in understanding and generating context rich description is fully utilized, and the ability of the model in generating accurate and detailed description is enhanced.
Owner:SHANGHAI JIAOTONG UNIV

AI interaction system for oral science popularization

The invention, which relates to the technical field of medical information data, discloses an AI interaction system for oral science popularization, comprising a multi-modal sensing unit, an AI analysis engine, a knowledge graph matching unit and an AR interaction feedback unit. According to the method, the AI analysis engine is used for deeply analyzing the real oral image and the tooth brushing action of the user, and dynamic weighting of the knowledge graph is combined, so that accurate matching of science popularization content with the actual focus, the wrong action and the subjective requirement of the user is ensured, the knowledge transmission effectiveness is greatly improved, and the user experience is improved based on the graded medical knowledge base. The system can automatically adjust the weight according to the health influence degree of the oral cavity problem, preferentially present high-risk lesion early warning, and convert abstract medical knowledge into a visual real-time visual instruction by using an AR virtual-real overlapping technology, so that a user can obtain an action correction suggestion based on pixel-level alignment in the tooth brushing process, and the user experience is improved. Through fusion of voice interaction and AR interaction, the threshold of common people for understanding professional medical knowledge is reduced.
Owner:NANJING STOMATOLOGICAL HOSPITAL

system

We provide the system. [Solution] Means for acquiring user care data, medical information, and preference information, A means of analyzing acquired data to generate the optimal care plan, A means for transmitting the generated care plan to the caregiver's information device, A method for analyzing work-related memos and instructions entered by caregivers and updating the database, A method for caregivers to perform their duties while checking care plans in real time and receiving visual instructions using smart glasses, A system that includes this.
Owner:SOFTBANK GROUP CORP

Game information processing method, program product and electronic equipment

The game information processing method comprises the steps that under the condition that a shielding object exists between a first virtual object and a second virtual object, an indication identifier is displayed on a graphical user interface of a second game account, and the indication identifier comprises a first graphical element representing the second virtual object and a second graphical element representing the shielding object, the first virtual object is controlled by a first game account, and the second virtual object is controlled by a second game account; and updating the overlapping display state of the first graphic element and the second graphic element in the indication identifier according to the shielding degree of the second virtual object by the shielding object. Through the method provided by the invention, the visual indication is provided for the players, the communication cost among the players is reduced, and the collaboration experience among the players is enhanced.
Owner:NETEASE (HANGZHOU) NETWORK CO LTD

Scene consistency image instruction generation method based on task decomposition

The present application relates to a task decomposition-based scene consistency image instruction generation method, and the generation method comprises the following steps: step S1, image analysis, object recognition, position and state recognition are performed; step S2, instruction understanding is performed, and set key information is extracted; step S3, feature extraction of a pre-scene image in a latent space is realized through a VAE encoder; step S4, a task decomposition generation step sequence is generated, and step-by-step description is performed; step S5, based on the step sequence, an instruction weight matrix with front and rear correlations is generated; the instruction weight matrix with front and rear correlations is designed based on the step sequence of step S4; step S6, based on the image reference of the pre-scene and the instruction weight, a visual instruction of the current scene is generated; the present application is reasonable in design, compact in structure and convenient to use.
Owner:QINGDAO HAIDA NOVA SOFTWARE CONSULTING CO LTD

Multi-source equipment state visual monitoring method of docking station and related device

The invention discloses a multi-source equipment state visual monitoring method of a docking station and a related device, and the method comprises the steps: obtaining original data traffic on an uplink data bus between a target docking station and host equipment; calculating an average flow value and a flow change rate sequence of the original data flow; calculating an I / O pressure index of the uplink data bus; determining whether the docking station is in a performance sensitive state based on the I / O pressure index; when the docking station is not in the performance sensitive state, obtaining first state data of the host equipment and second state data of the external equipment through a first preset sampling frequency; when it is determined that the docking station is in the performance sensitive state, first state data of the host device and second state data of the external device are obtained through a second preset sampling frequency; generating a visualization instruction of the running states of the host equipment and the external equipment; and controlling a display screen of the docking station to carry out visual presentation according to the visual instruction. According to the invention, the adaptability of visual monitoring of the docking station can be improved.
Owner:SHENZHEN SINOBRY ELECTRONICS LTD

Remote expert guidance intention visualization method, equipment and device for XR devices

This application discloses a method, device, and apparatus for visualizing the intention of remote expert guidance for XR devices. The method includes: receiving multimodal instructions issued by a remote expert in combination with a real-time scene preview image; inputting the multimodal instructions and the associated scene preview image into a visual language model to generate a structured intent data packet, and based on the structured intent data packet, generating dynamic 3D visual instructions that match the operation action intention in the structured intent data packet; spatially anchoring the 3D visual instructions to the real-time image of the XR device of the on-site user, and determining the rendering position of the 3D visual instructions based on the matching results, so as to render the 3D visual instructions in the real-time image of the XR device. Through the above method, experts can express complex operation intentions in a way that is consistent with human intuition, reducing communication ambiguity.
Owner:HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD

Robot three-dimensional space constraint refinement operation control system and method

The application discloses a robot three-dimensional space constraint refinement operation control system and method, and belongs to the field of robot control. The system comprises: a task planning module, which decomposes a natural language instruction into a subtask instruction and a visual instruction, transmits the visual instruction to a constraint extraction module, and transmits the subtask instruction to a subtask execution module; a constraint extraction refinement module, which forms a component-level space constraint through multi-modal large model semantic reasoning combined with three-dimensional space coordinate mapping, judges whether the component-level space constraint needs to be refined, performs corresponding processing, and returns the constraint result to the task planning module to be transmitted to the subtask execution module; and a subtask execution module, which receives the subtask instruction and the constraint result, first calls an execution condition construction module to obtain an execution condition required by the subtask instruction combined with the constraint result, acquires a robot key speed control instruction from a server according to the execution condition, and controls the robot to perform a corresponding operation. The application can control the robot to operate efficiently and accurately in a complex environment.
Owner:UNIV OF SCI & TECH OF CHINA

A human shadow operation information display method and system based on multi-source information fusion

The application discloses a kind of based on multi-source information fusion's human shadow operation information display method and system, it is related to meteorological information intelligent decision-making technical field, including, utilize graph neural network to carry out inference analysis to time series knowledge graph, identify cloud system development change and infer the evolution of cloud physical characteristics, predict the spatiotemporal range of operation potential area, generate accurate forecast conclusion and human shadow operation suggestion;Accurate forecast conclusion and human shadow operation suggestion are converted into visual instruction set, and evolution path is shown through highlighting core weather index, dynamic arrow display, and color overlay marks operation potential area and operation corridor, generates visual comprehensive situation chart.The application is analyzed by utilizing graph neural network to time series knowledge graph, realizes from historical mode automatically identifying operation potential area and predicting its dynamic evolution, improves the foresight and accuracy of human shadow operation decision.
Owner:辽宁省人工影响天气办公室

Data-efficient visual instruction tuning for multimodal large language models

According to one aspect, instruction tuning may include generating a set of instructions for a reference set of images selected from a set of images based on one or more task specific instruction generation protocols, generating one or more task importance weights for the reference set of images based on the set of instructions and the reference set of images and a ratio of a first loss of a first loss function associated with a response and an image from reference set of images and a second loss of a second loss function associated with the response, a question, and the image from reference set of images, and generating a set of instructions for a remaining set of images from the set of images based on one or more of the task importance weights, k-means clustering, and neighbor centrality from a cluster of the k-means clustering.
Owner:HONDA MOTOR CO LTD

Intelligent data question and answer method and system fusing field large language model

The invention provides an intelligent data question and answer method and system fusing a field large language model, and relates to the technical field of large language models.The method comprises the steps that action data and view angle data of a user in a pre-constructed three-dimensional virtual scene corresponding to a target field are obtained, and question and answer data of the target field are collected; performing association analysis on the action data and the view angle data to determine a user query intention; performing associated coding processing on the query intention of the user and the scene position of the three-dimensional virtual scene to construct an intelligent question and answer library; obtaining a resource allocation scheme by using a software defined network technology; searching target question and answer data corresponding to the query request data from an intelligent question and answer library through a domain large language model, and generating a target question and answer result and a visual instruction; based on the resource allocation scheme and the visualization instruction, intelligent data question answering is completed, and high-immersion, low-delay and accurate-intention intelligent data question answering and visualization presentation oriented to the specific field are achieved.
Owner:FIVE DIMENSIONS INTELLIGENT TECHNOLOGY (SHANGHAI) CO LTD +1

Visual management system based on electromechanical industry data

The invention discloses a visual management system based on electromechanical industry data, and relates to the technical field of train traction systems, and the system comprises a traction assembly modeling module, a working condition acquisition module, a heat effect identification module and a visual instruction module. The traction assembly modeling module identifies a train traction assembly, extracts the structure and operation parameters of the train traction assembly, and establishes a physical connection model; the working condition acquisition module acquires current, voltage, torque and temperature rise data under various typical working conditions, and constructs a working condition database. And the heat effect identification module constructs a heat accumulation curve based on the working condition data and the hot melting time constant, judges thermal inertia sudden change, cooling imbalance, critical overload and thermal collapse precursor risks, and generates corresponding multi-stage early warning instructions. And the visual instruction module dynamically displays the early warning level, marks the temperature rise trend, the heat rate change, the abnormal combination and the waste heat release time, and assists the train system in state evaluation and intelligent adjustment. Accurate monitoring and risk early warning of the thermal state of the electromechanical equipment are realized.
Owner:CHONGQING XIANGFU TECHNOLOGY CO LTD

Marketing originality automatic generation and optimization system based on deep learning

The invention discloses a marketing originality automatic generation and optimization system based on deep learning, and relates to the technical field of computers. Through a small sample stylization unit, the system can construct a style reference set only by relying on a small number of representative originality materials; the generated content distribution and the reference distribution are aligned in a unified depth feature space, so that the generated image and copywriting are closer to the target brand tonality in the dimensions of color, composition, tone and the like, and a differentiated vision and utterance system is quickly molded in the absence of large-scale brand data; in combination with a cross-modal semantic alignment unit, text information such as marketing targets, audience descriptions and the like and brand visual instructions are jointly coded into a unified semantic-style condition, so that the double constraints of what and what can be satisfied in the same generation process, and the risks of disjunction between copywriting and pictures and style deviation are reduced.
Owner:BEIJING HEJIN TECHNOLOGY CO LTD

E-commerce short video structured script generation method and system based on hierarchical text generation and RAG

The invention discloses an e-commerce short video structured script generation method and system based on hierarchical text generation and RAG, and relates to the technical field of artificial intelligence. Commodity information and a target user portrait are obtained firstly, commodity qualitative characteristics and user demand matching degree are determined through cognitive reasoning, and a user demand matching degree is determined; planning a sequence containing at least three sub-lens narrative functions such as curiosity triggering and demonstration functions according to the sub-lens narrative functions; aiming at each sub-mirror function, fusing commodities, user portraits and sub-mirror function vectors to construct a dynamic query vector, carrying out cross-modal retrieval on Top-K MMUs in a multi-modal memory network, and carrying out weighted aggregation through an attention mechanism positively correlated with MMU performance data to generate a multi-modal context vector; the method comprises the following steps: respectively inputting an e-commerce field fine-tuning LLM and a lightweight visual instruction generator to obtain a split copy and a structured visual description instruction, combining the split copy and the structured visual description instruction into a complete script according to a narrative sequence, newly adding an MMU in the system, and optimizing retrieval and fusing weight to realize self-evolution.
Owner:GUANGZHOU YUZHI CULTURE TECH CO LTD

A method, system, and storage medium for robotic handling in low gravity environments

ActiveCN120680516BSolve the problem of irregular rotation and difficulty in grabbingProgramme-controlled manipulatorComputer graphics (images)Angular velocity
The application discloses a robot carrying method and system in a low-gravity environment and a storage medium. A simple sketch drawn by an operator and a real scene image in a current cabin are collected, the simple sketch and the real scene image are input into a double-branch visual encoder model which has been trained, a basic action sequence instruction for controlling a robot to complete a material carrying operation is generated, and then a target rotation angular velocity is monitored in real time during target grabbing by executing the basic action sequence instruction. If the target rotation angular velocity is greater than a set threshold, the robot gripper is driven to rotate in a reverse direction for fine adjustment. Thus, pure visual instruction interaction is realized to adapt to the situation that voice interaction is completely unavailable in a low-gravity scene, and the problem that an object is easily subjected to irregular rotation and is difficult to be grabbed under low gravity is solved.
Owner:58 INTELLIGENT TECH (HANGZHOU) CO LTD

A robot material handling method, robot, and storage medium

The machine material carrying method, the robot and the storage medium disclosed by the application obtain a sketch depicting an appearance of a work target and carrying destination position information, extract a shape feature of the work target from the sketch, collect environment data in a current task scene, extract a visual feature from the environment data, align the visual feature with the shape feature by using a contrast loss function, extract a shape feature of the work target from a target region, calculate a target pose and confirm a material type from the environment data based on the target region, finally query a preset control information library based on the target object type to obtain corresponding action constraint information, combine the action constraint information, the target pose and the destination position information, generate a material carrying instruction, and control each actuator to move to complete a material carrying task. The material carrying can be realized in a noisy industrial scene by only using a pure visual instruction without inputting a complex text instruction.
Owner:58 INTELLIGENT TECH (HANGZHOU) CO LTD

Carry-scraper instrument parameter visualization system and method based on head-up display

The invention relates to the technical field of engineering machinery intelligence, and discloses a carry-scraper instrument parameter visualization system and method based on head-up display, and the system comprises a data collection module, a physical state simulation module, a health state evaluation and life prediction module, a control processing module, an alarm module and an HUD hardware module. The method comprises the following steps: generating dynamic parameters according to machine state data; generating dynamic parameters according to the machine state data; analyzing historical data records, and determining predictive health state parameters; fusing all parameters, executing a priority algorithm, and generating a visual instruction; monitoring parameters, giving an alarm according to conditions, and generating an instruction to change display attributes; and converting the visual instruction into optical projection, and presenting the optical projection in the visual field of a driver. The physical state simulation module is arranged, and according to measurable real-time parameters and in combination with a machine physical model matched with working conditions, the deep insight ability exceeding direct dimension measurement of a sensor is provided.
Owner:QINGDAO FAMBITION HEAVY MASCH CO LTD

A continuous visual command fine-tuning method based on separable mixture low-rank adaptation

The present invention discloses a continuous visual instruction fine-tuning method based on separable hybrid low-rank adaptation, which relates to the field of model deep learning and continuous learning technology; S1, creating a pre-training model; S2, inputting a continuous visual instruction fine-tuning data set sequence into the pre-training model; S3, selecting a suitable low-rank adaptation matrix according to the continuous visual instruction fine-tuning data set and performing fine-tuning; S4, using separable routing technology to obtain S3's visual understanding routing selection and instruction following routing selection; S5, performing weighted operation on the separable routing result of S4 and its corresponding adaptability matrix to obtain two corresponding outputs; S6, using adaptive fusion technology to obtain the final output of the low-rank adaptation module; the present invention adopts the above-mentioned continuous visual instruction fine-tuning method based on separable hybrid low-rank adaptation to ensure that the model does not interfere when processing visual and instruction tasks, effectively avoiding double catastrophic forgetting.
Owner:HEFEI UNIV OF TECH