Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

335 results about "Visual marking" patented technology

Controllable image-to-video generation

Techniques are generally described for controllable image-to-video generation. In various examples, a first image representing at least a first object may be received. First input data including a selection of the first object in the first image for animation may be received. Second input data including at least a first bounding box indicating a target location of the first object may be received. A latent diffusion text-to-image model and the first image may be used to generate a first plurality of visual tokens. One or more first grounding tokens may be generated representing a location of the first bounding box. The latent diffusion text-to-image model may be used to generate a video animating the first object based on the first plurality of visual tokens and the one or more first grounding tokens.
Owner:AMAZON TECH INC

Multi-modal knowledge graph completion model training method, completion method and device

The invention provides a multi-modal knowledge graph completion model training method, a completion method and equipment, and relates to the field of computer systems based on specific calculation models. The method comprises the following steps of: extracting an image block feature vector, a word feature vector, a structure feature vector, a corresponding visual mark, a text mark and a structure mark by adopting a multi-modal knowledge graph completion model; based on each visual mark, the text mark and the structure mark, acquiring multi-modal fusion feature data of each entity by a graph attention mechanism; and complementing the original multi-modal knowledge graph according to each piece of multi-modal fusion feature data, and calculating the target loss of the complemented multi-modal knowledge graph to optimize the multi-modal knowledge graph complementing model. According to the method, fine-grained multi-modal feature extraction and application can be realized in the completion model training process, the reliability and effectiveness of the completion model training process can be effectively improved, and then the accuracy and reliability of multi-modal knowledge graph completion can be improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Memory enhanced vision-language-motion submerged space dynamic fusion automatic driving method

The invention relates to a memory enhanced vision-language-action submerged space dynamic fusion automatic driving method. Comprising the following steps: generating a bird's-eye view feature map; extracting scene, agent and map marks from the aerial view feature map, and fusing the mark at the current moment and the previous i historical marks to generate memory enhanced visual marks; the memory enhanced visual mark and the vehicle state information are converted into a submerged space, text input of a driver is marked and unified into the submerged space, the submerged space is represented and fused, and fusion representation is generated; an automatic driving instruction data set is introduced for adjustment, and a large language model adapting to an automatic driving task is obtained; according to the fusion representation, track planning is carried out in an autoregression mode by using a large language model, and path point coordinates are obtained; and designing a transverse controller and a longitudinal controller based on PID (Proportion Integration Differentiation) to track and control the coordinates of the path points. According to the invention, the visual representation capability of end-to-end driving is improved, and the visual-language-action fusion effect is improved.
Owner:NANJING UNIV OF SCI & TECH

Data processing method and system, electronic equipment, storage medium and computer program product

The invention discloses a data processing method and system, electronic equipment, a storage medium and a computer program product, and relates to the technical field of large model technology and key value caching. The method comprises the following steps: acquiring a plurality of visual marks, a plurality of text marks and initial key value data corresponding to a visual language model; determining a key text mark in the plurality of text marks by utilizing the initial key value data; according to the initial key value data and the distribution positions of the key text marks in the multiple text marks, importance evaluation is conducted on the multiple visual marks, an evaluation result is obtained, and the evaluation result is used for representing cross-modal attention weight distribution between the multiple visual marks and the key text marks; and according to an evaluation result, performing cache compression processing on the initial key value data to obtain target key value data. According to the method and the device, the technical problem that the model reasoning efficiency is influenced due to high key value data caching overhead of a visual language model in related technologies is solved.
Owner:ALIBABA CLOUD FEITIAN (HANGZHOU) CLOUD COMPUTING TECH CO LTD

Light rail vehicle safety monitoring system based on multi-sensor fusion

The invention discloses a light rail vehicle safety monitoring system based on multi-sensor fusion, and relates to the technical field of light rail vehicle safety monitoring, the system is composed of a plurality of functional modules, and the system comprises a multi-mode sensing module for capturing images of the surrounding environment of a light rail in real time and delimiting a visual mark area; performing three-dimensional scanning on the visual marking area, generating environment point cloud data in real time, and acquiring a dynamic object distance by using a Doppler compensation algorithm; the data processing module is used for calculating object deformation between adjacent frames by adopting an optical flow method and a CNN deformation detection network; carrying out object relative motion calculation according to object shape analysis; carrying out instance segmentation based on MaskR-CNN, and outputting an object contour; meanwhile, a dynamic target clustering algorithm is used for conducting multi-target tracking on the object contour, and three-dimensional coordinates and motion vectors of the object are output in combination with environment point cloud data; and the fusion and analysis module performs time synchronization based on the PTP protocol.
Owner:HANGZHOU JIAHE INTELLIGENT TECH CO LTD

Image-text processing method and device for marking compression framework

The invention discloses an image-text processing method and device for marking a compression framework. The method comprises the following steps of: extracting visual features; a visual mark screening processing step; a text feature extraction step; and a multi-modal fusion and model processing step. The method has the beneficial effects that the inference efficiency of the MLLMs is remarkably improved under the condition that the visual mark compression frame does not need additional training; through global and local information fusion of the DVTS module and text guide supplement of the TGVC module, the number of visual marks is greatly reduced, meanwhile, key visual information is reserved, and visual-text alignment is enhanced; experiments show that in various image and video benchmark tests, compared with an existing method, the framework has the advantages that the calculation cost is greatly reduced, the model performance is maintained and even improved, and the framework has remarkable technical advantages and application potential.
Owner:ZHEJIANG YOULU ROBOT TECH CO LTD

Smart sensing for pallet loading and unloading

A pallet loading system may comprise a processor, and memory with instructions stored thereon that, when executed by the processor, cause the processor to receive package loading data, the package loading data including characteristics of a package to be placed on the pallet and characteristics of a loading operator. A placement location for the package on the pallet may be determined using an artificial intelligence (AI) or machine learning (ML) algorithm based on the characteristics of the package and the characteristics of the loading operator and a visual marking (such as a visual projection) may be displayed at the placement location. The system may output an instruction to the loading operator to place the package at the displayed placement location.
Owner:INTEL CORP

Secured transactions on wearable devices

Systems herein describe a secured transaction system for accessing a camera stream from a front-facing camera unit of a wearable computing device, identifying a visual marker in the camera feed associated with a payment application, in response to identifying the visual marker, transmitting a request to an operating system of the wearable computing device to isolate access of the camera stream to the payment application and in response to transmitting the request, initiating a transaction using the payment application.
Owner:SNAP INC

Historical building brick wall surface damage intelligent identification and diagnosis method

PendingCN120472220AImage analysisCharacter and pattern recognitionVisual markingSustainable management
The invention relates to the technical field of historical building protection, and discloses an intelligent identification and diagnosis method for historical building brick wall surface damage, which breaks through the limitations of low efficiency, strong subjectivity and easy detail omission of traditional manual exploration by fusing unmanned aerial vehicle high-precision image acquisition, deep learning semantic segmentation and multi-modal feature analysis technologies. And automatic rapid positioning and accurate classification of large-range wall surface damage are realized. Environmental distortion is eliminated by using a self-adaptive image registration algorithm, fine texture features are extracted in combination with a damage probability prediction model, and complex damage forms such as salting-out crystallization, plant root erosion and weathering spalling are effectively distinguished; through three-dimensional space clustering and visual marking technologies, the damage distribution density and the evolution trend are visually presented, and a reliable basis is provided for scientifically formulating a grading repair strategy. Besides, through establishment of a damage database and iterative optimization of an intelligent diagnosis model, the building health condition can be tracked for a long time, and technical support is provided for preventive protection and sustainable management of cultural heritage.
Owner:SHANGHAI MINGYUE ARCHITECTURAL DESIGN OFFICE CO LTD

Automatic driving space planning enhancement method based on visual marker

The invention discloses an automatic driving space planning enhancement method based on a visual marker. The method comprises the following steps: acquiring an original image and text input; processing the original image to obtain image features; processing the text input to obtain text features; generating a text output with a visual mark by using the image features and the text features; the text output with the visual marks is converted, and text output with coordinates is obtained; the accuracy and semantic consistency of spatial understanding in an automatic driving scene are remarkably improved, high synchronization of visual perception and semantic expression is achieved, and the problem of visual and language modal semantic segmentation in an existing method is effectively solved. The analysis precision of the object position, the motion state and the interaction relation in the automatic driving question and answer task is greatly improved, and the decision reliability and the planning naturalness in a complex driving scene can be remarkably enhanced.
Owner:SOUTH CHINA UNIV OF TECH

Image recognition method and system for stone slab crack detection

The invention discloses an image recognition method and system for stone slab crack detection, and the method comprises the steps: carrying out the preprocessing of a color image of a stone slab, and obtaining a gray image; calculating the gradient amplitude of the grayscale image, determining an optimal threshold value, and obtaining a preliminary edge image; calculating the first order gradient and direction of the preliminary edge image, constructing a Hessian matrix and generating a crack response image; performing edge detection on the crack response image to generate an initial edge image; performing connected domain analysis on the initial edge image to generate a final edge image; and extracting coordinates of crack points from the final edge image, mapping the crack points to the stone slab color image, and performing visual marking to generate a marked image. According to the method, high-precision extraction and fine distinguishing of the cracks under the background of complex stone textures are achieved, the limitation that the cracks and the patterns are difficult to separate in a traditional method is effectively overcome, meanwhile, crack marking is directly carried out on the stone slab original drawing, and the accuracy and practicability of stone slab quality detection are remarkably improved.
Owner:HUBEI PROVINCE HUAJIAN STONE CO LTD

Domain-Specific Shorthand for Generation of Data Visualizations based on Context Free Grammar

System, method and interface for generating data visualizations are provided. The system receives a user input to specify a natural language command directed to a data source. The system also generates a prompt for generating a data visualization based on relevant data fields and data values, rules that characterize the data visualization, and a context free grammar. The system also prompts a trained large language model using the prompt to generate a structured document following a domain-specific schema based on a shorthand notation. The system also uses a parser that uses the context free grammar to map the structured document to a visual specification. The visual specification specifies the data source, visual variables, and data fields from the data source. The system also generates and displaying a data visualization based on the visual specification, including displaying visual marks representing data, retrieved from the data source, for the data fields.
Owner:SALESFORCE INC

Customer service information generation method and system based on multi-modal intention recognition

The invention discloses a customer service information generation method and system based on multi-modal intention recognition, and the method comprises the steps: obtaining customer service session data, and carrying out the heterogeneous data classification operation of the session data; performing feature extraction operation on each piece of parting modal information; establishing a semantic bridging matrix among the modal features, performing hierarchical attention routing, and outputting an intention label; a knowledge graph is injected, the knowledge graph comprises a static knowledge base, a real-time service flow and a user historical portrait, and three-source dynamic knowledge is output; constructing a decision tree based on the intention label, the three-source dynamic knowledge and the space-time identifier; and executing a corresponding text generation operation, a visual mark generation operation or a voice synthesis operation based on a response action corresponding to the leaf node of the decision tree, and generating customer service information. According to the method and the system, the buyer intention recognition accuracy of the multi-modal session data can be greatly improved, and the response delay time of generating the customer service information is shortened.
Owner:深圳乐搏科技有限公司

Aircraft positioning method and device, aircraft, medium and program product

The embodiment of the invention provides an aircraft positioning method and device, an aircraft, a medium and a program product, and relates to the technical field of flight positioning. The method comprises the following steps: acquiring a to-be-identified image, and identifying a visual marking object; matching and associating the two-dimensional coordinate information corresponding to each visual marking object with the actual three-dimensional coordinate information; wherein each piece of actual three-dimensional coordinate information is obtained based on a preset coordinate transformation rule and is represented by a unified transformation matrix parameter; and constructing and solving an objective function about the two-dimensional coordinate information, the actual three-dimensional coordinate information and the to-be-solved aircraft pose, and obtaining the aircraft pose information according to a solving result. According to the embodiment of the invention, the position information of each mark is represented through the transformation matrix based on the unified coordinate system, so that different marks can be arbitrarily arranged in different planes, and the application flexibility of a visual mark positioning technology in a complex scene is effectively improved.
Owner:TIANJIN YUNSHENG INTELLIGENT TECH CO LTD

Visualization of external audio commands

This disclosure relates to methods, systems, and techniques for visualizing external audio captured by a microphone of the vehicle. Using the techniques described herein, external audio signals may be interpreted and converted into clear visual indicium that can be understood by individuals associated with the vehicle, such as passengers inside the vehicle or individuals awaiting pickup.
Owner:ZOOX INC

Video space-time understanding method and device based on multi-modal large model, and medium

The invention belongs to the field of computer vision, and particularly relates to a video space-time understanding method and device based on a multi-modal large model and a medium. A multi-modal large language model is connected with a mask segmentation model, and video features are encoded by using a multi-modal encoder; the time tasks and the space tasks are represented by adopting different numbers of sampling frames and visual mark forms. The visual mark is aligned to the text space and then input into the large language model together with the text mark, and a corresponding text answer is obtained through decoding. For a time task, a timestamp is directly extracted from a text answer, and spatial information passes through lt; sEGgt, SEGgt; and the mark code is embedded as a prompt input to the mask decoder, so that mask generation of the sampling frames and mask propagation of the whole video are realized. Compared with the prior art, the method has the advantages that joint training of fine-grained video time-space understanding is realized, and more accurate event time-space positioning can be realized besides overall understanding of the video.
Owner:FUDAN UNIVERSITY

Generating video descriptions using a machine learning model

The present disclosure describes techniques for generating video descriptions using a machine learning model. A plurality of sets of visual tokens corresponding to a plurality of frames of a video is generated. A first type of tokens is generated by implementing temporal pooling on the plurality of sets of visual tokens corresponding to the plurality of frames. A second type of tokens is generated by compressing each of the plurality of sets of visual tokens corresponding to each of the plurality of frames. A third type of tokens is generated by applying cross-attention between each of the plurality of sets of visual tokens and a fourth type of tokens including text tokens generated based on an input text query. A text description of the video is generated based on the first type of tokens, the second type of tokens, the third type of tokens, and the fourth type of tokens.
Owner:LEMON INC(GB)

Real-time image scan

Systems and methods herein describe a real-time scan system that identifies a visual marker in an image frame, identifies key points from the image frame, compares the image frame with a database of visual marker images using the identified key points, identifies a set of matched images in the database of visual marker images, filters the set of matched images using a neural network trained to analyze the image frame and set of matched images, detects the visual marker in an image of the filtered set of matched images, causes presentation of a notification on a display unit of the computing device, receives a selection of the notification, and in response to receiving the selection of the notification, causes presentation of multi-media content associated with the visual marker on the display unit of the computing device.
Owner:SNAP INC

Visual information fusion method, device, equipment, medium and computer program product

The invention provides a visual information fusion method, device and equipment, a medium and a computer program product, and the method comprises the steps: carrying out the coding of an input image and an input text, and obtaining a target mark sequence; the target marking sequence comprises a visual marking sequence and a language marking sequence; determining a fused visual context based on the attention of the language marker sequence to the visual marker sequence; based on the fused visual context, determining a modulation parameter of each target mark sequence; and based on the adjustment parameter, determining a semantic understanding result of the input text. According to the method, a dynamic feature modulation mechanism is introduced into each layer of the large language model, so that visual information can adaptively adjust text representation, and the ability of the large language model to understand multi-modal information is enhanced.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Calibration method for positioning and attitude determination of heading machine

The invention relates to the technical field of engineering construction measurement and control, and discloses a calibration method for positioning and attitude determination of a heading machine, which comprises the following steps: constructing a data acquisition module, a route making module, a real-time dynamic monitoring module, an execution module, a feedback module and an optimization module; the data acquisition module is used for completing comprehensive acquisition of multi-source data of the heading machine, a route, an environment and an obstacle, the route making module is used for fusing the multi-source data to carry out path planning and error compensation, and the real-time dynamic monitoring module is used for realizing heading machine pose tracking and emergency correction through visual marking points and a laser scanner. The execution module is used for converting an optimal tunneling path into a driving motor control instruction, the feedback module is used for collecting deviation data of an actual operation pose and a planned path in real time and returning the deviation data, the optimization module realizes long-term self-adaptive optimization of the system, and finally, precision, safety and high efficiency of the tunneling operation of the tunneling machine in a complex underground environment are realized.
Owner:TAIYUAN INST OF CHINA COAL TECH & ENG GROUP +1

Perception-Based Worksite Control System

A control system manages a plurality of mobile machines each equipped with a visual perception system to capture perception data. The control system via an onboard controller applies an object detection operation to the perception data to detect a detected marker position corresponding to a visual marker. The onboard controller assesses an assessed marker heath status with respect to the detected marker position and transmits that to a central worksite server. The central worksite server aggregates the assessed marker health statuses from a plurality of mobile machines to determine an aggregate marker health status for the visual marker.
Owner:CATERPILLAR INC

Method and apparatus for generating video, electronic device, and computer program product

The present disclosure relates to a method and apparatus for generating a video, an electronic device, and a computer program product. The method includes obtaining a visual token for generating an image frame in the video. The method further includes obtaining a control token for constraining position information of an object in the image frame. In addition, the method also includes generating the image frame in the video based on the visual token and the control token, where the object in the image frame satisfies the position information.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Automated captioning of augmented reality effects in videos

Described herein are techniques for generating captions for augmented reality (AR) effects in videos using an adapted multimodal large language model (MLLM). The technique involves sampling frames from base and AR-applied videos, combining them into concatenated frames, and processing them with a hybrid vision encoder. Visual tokens are projected into a language model token space, reshaped, downsampled, and interleaved with text tokens. A fine-tuned large language model processes this input sequence to generate AR effect captions. Optical character recognition extracts text from AR frames, which is combined with generated captions and metadata to produce merged captions and content tags. This approach enables accurate description of temporal AR effects and facilitates downstream applications like search and ranking.
Owner:SNAP INC

Method for controlling lighting effect of spliced lamp, device, apparatus, and medium

The present application relates to a method for controlling a lighting effect of a spliced lamp, as well as a device, an apparatus, and a medium. The method comprises: displaying a spatial topological graph of the spliced lamp in an effect editing region of a graphical user interface; displaying a visual marker of a motion base point of the spliced lamp in the effect editing region in an initialized manner; updating the motion base point according to the latest position of the visual marker relative to the spatial topological graph; driving the spliced lamp to play a lighting effect by using a lighting effect playing instruction that is generated on the basis of the motion base point, such that a mapping position of the motion base point in the physical space serves as a benchmark reference point of a motion process of the lighting effect.
Owner:SHENZHEN INTELLIROCKS TECH CO LTD

Agaricus bisporus autonomous picking method and system based on VLA

The invention provides a VLA-based agaricus bisporus autonomous picking method and system, and relates to the technical field of cultivation, and the method comprises the steps: obtaining picking task description information and visual observation data, and coding the visual observation data to obtain a visual feature sequence; constructing a multi-scale visual feature map based on the visual feature sequence, and generating an object-level visual mark sequence; encoding the picking task description information into a text marking sequence, mapping the text marking sequence and the object-level visual marking sequence to a unified semantic space, and splicing the text marking sequence and the object-level visual marking sequence to form a joint semantic sequence; and inputting the joint semantic sequence into a pre-trained VLA model, performing autoregression prediction to generate an action mark sequence, decoding according to the action mark sequence to obtain a picking control instruction, and driving a picking device to complete picking, so as to improve the key visual information extraction capability and enhance the accuracy and robustness of action generation in a complex unstructured cultivation scene.
Owner:SHANGHAI HENGZE FUHUI INTELLIGENT TECHNOLOGY CO LTD

Building construction risk intelligent identification and early warning method and system based on BIM model

The invention provides a building construction risk intelligent identification and early warning method and system based on a BIM model, and relates to the technical field of construction risk early warning, and the method comprises the steps: obtaining a building information model and historical meteorological data of a target scene; constructing a numerical wind tunnel model, carrying out transient wind field simulation, carrying out region identification by combining a transient wind field simulation result, and obtaining a prior attention region set; according to the construction progress information of the target scene and the prior attention area set, defining a real-time attention area set of the target scene; collecting real-time wind power data, and traversing the real-time attention region set to perform dynamic risk assessment in combination with the real-time wind power data and the numerical wind tunnel model to obtain a real-time risk coefficient set; and performing visual marking in a visual interface of the building information model, and performing risk early warning in combination with a visual marking result and the real-time risk coefficient set. The technical problem that in the prior art, a traditional building construction risk early warning method is insufficient in early warning real-time performance is solved.
Owner:HEBEI CONSTR GRP

AUV (Autonomous Underwater Vehicle) dynamic docking system for multiple docking stations

The AUV dynamic docking system based on the multi-port docking station comprises the multi-port docking station and an AUV, and the dynamic docking control method of the AUV and the multi-port docking station comprises the following specific steps: S1, extracting center coordinates of a guide light source of each dock port of the multi-port docking station; s2, in a long-distance stage, the AUV performs global pose calculation by combining an EPnP algorithm with all the guide light sources, and the AUV autonomously docks to a target dock entrance according to a calculation result; s3, when the AUV gradually gets close to the target dock entrance and is in the middle distance stage, the AUV estimates the azimuth included angle between the AUV and the target dock entrance based on the monocular double-light-source set relation, and the AUV adjusts the course according to the direction included angle and continues to be in butt joint with the target dock entrance; and S4, when the AUV is close to the target dock entrance and is about to enter the target dock entrance, the AUV is in a final docking stage at the moment, the docking dock is still in a moving state, the AUV identifies the Stag visual mark of the target dock entrance to carry out pose calculation of tail end docking, and the AUV is accurately docked with the target dock entrance according to the pose calculation.
Owner:HANGZHOU DIANZI UNIV +1

Anti-leak digital document marking system and method using distributed ledger

The system is disclosed for visual marking sensitive documents for leak prevention. Each time an action is taken with regard to a document (e.g., creation, viewing, downloading), that action is added to a distributed ledger, essentially creating a unique hash for a new instance of the document. This new hash is visually embedded in the document as a code comprising a plurality of differently shaded pixels, wherein some of the pixels directly encode information regarding the document (e.g., an account that generated the new instance of the document, a date, a time, a unique ID for the document, etc.) and some of the pixels do not encode information. The code is capable of being scanned either digitally or physically on a printed version of the document, such that the immediate source of the document, corresponding to who leaked the document, is able to be discerned.
Owner:EQUITY SHIFT INC

Intelligent management of approach boundaries in industrial domains

A method is disclosed, including determining a configuration of boundary(s) associated with one or more danger zones in the vicinity of industrial equipment, causing projection of visual markers on surfaces, the visual markers being configured based on the configuration of the respective boundary(s), processing images captured of the visual markers to detect a boundary event related to the visual markers, the boundary event including the presence of an intruder that is a person or an object proximate to, crossing over, moving to, or moving from the visual markers, determining whether the boundary event is allowed, processing the captured images and / or captured audio to determine or infer a location of the intruder before, during, or after the boundary event, selecting, contingent on a determination that the boundary event is not allowed, automated action(s) based on the location of the intruder, and causing performance of the automated action{s}.
Owner:SCHNEIDER ELECTRIC USA INC

Yoga mat surface defect automatic identification method based on intelligent sensor

The invention relates to the technical field of intelligent surface defect identification of an intelligent sensing detection and optical scattering method, and discloses a yoga mat surface defect automatic identification method based on an intelligent sensor. Establishing a two-dimensional coordinate and a discrete sampling array on the surface of a measured object, and unifying an installation posture; scanning the scattering intensity of each point according to a multi-angle sequence under fixed oblique incidence, and checking the integrity and the range; summarizing according to fixed angle weights to form an angular spectrum value matrix; carrying out point-by-point difference on the defect-free reference matrix to obtain a residual error, carrying out difference synthesis along rows and columns to obtain mutation intensity, and constructing a matrix; generating candidate anomalies according to a reference mutation statistics set threshold value; a connected domain is extracted on the grid in a four-adjacent mode, insufficient area is eliminated, and the center and the area are calculated; and outputting the number, the position and the area of the defects and performing visual marking.
Owner:ZHEJIANG SUNRISE HIGH TECH NEW MATERIAL CO LTD