Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

239 results about "Visual marking" patented technology

Domain-Specific Shorthand for Generation of Data Visualizations based on Context Free Grammar

System, method and interface for generating data visualizations are provided. The system receives a user input to specify a natural language command directed to a data source. The system also generates a prompt for generating a data visualization based on relevant data fields and data values, rules that characterize the data visualization, and a context free grammar. The system also prompts a trained large language model using the prompt to generate a structured document following a domain-specific schema based on a shorthand notation. The system also uses a parser that uses the context free grammar to map the structured document to a visual specification. The visual specification specifies the data source, visual variables, and data fields from the data source. The system also generates and displaying a data visualization based on the visual specification, including displaying visual marks representing data, retrieved from the data source, for the data fields.
Owner:SALESFORCE INC

Customer service information generation method and system based on multi-modal intention recognition

The invention discloses a customer service information generation method and system based on multi-modal intention recognition, and the method comprises the steps: obtaining customer service session data, and carrying out the heterogeneous data classification operation of the session data; performing feature extraction operation on each piece of parting modal information; establishing a semantic bridging matrix among the modal features, performing hierarchical attention routing, and outputting an intention label; a knowledge graph is injected, the knowledge graph comprises a static knowledge base, a real-time service flow and a user historical portrait, and three-source dynamic knowledge is output; constructing a decision tree based on the intention label, the three-source dynamic knowledge and the space-time identifier; and executing a corresponding text generation operation, a visual mark generation operation or a voice synthesis operation based on a response action corresponding to the leaf node of the decision tree, and generating customer service information. According to the method and the system, the buyer intention recognition accuracy of the multi-modal session data can be greatly improved, and the response delay time of generating the customer service information is shortened.
Owner:深圳乐搏科技有限公司

Aircraft positioning method and device, aircraft, medium and program product

The embodiment of the invention provides an aircraft positioning method and device, an aircraft, a medium and a program product, and relates to the technical field of flight positioning. The method comprises the following steps: acquiring a to-be-identified image, and identifying a visual marking object; matching and associating the two-dimensional coordinate information corresponding to each visual marking object with the actual three-dimensional coordinate information; wherein each piece of actual three-dimensional coordinate information is obtained based on a preset coordinate transformation rule and is represented by a unified transformation matrix parameter; and constructing and solving an objective function about the two-dimensional coordinate information, the actual three-dimensional coordinate information and the to-be-solved aircraft pose, and obtaining the aircraft pose information according to a solving result. According to the embodiment of the invention, the position information of each mark is represented through the transformation matrix based on the unified coordinate system, so that different marks can be arbitrarily arranged in different planes, and the application flexibility of a visual mark positioning technology in a complex scene is effectively improved.
Owner:TIANJIN YUNSHENG INTELLIGENT TECH CO LTD

Visualization of external audio commands

This disclosure relates to methods, systems, and techniques for visualizing external audio captured by a microphone of the vehicle. Using the techniques described herein, external audio signals may be interpreted and converted into clear visual indicium that can be understood by individuals associated with the vehicle, such as passengers inside the vehicle or individuals awaiting pickup.
Owner:ZOOX INC

Video space-time understanding method and device based on multi-modal large model, and medium

The invention belongs to the field of computer vision, and particularly relates to a video space-time understanding method and device based on a multi-modal large model and a medium. A multi-modal large language model is connected with a mask segmentation model, and video features are encoded by using a multi-modal encoder; the time tasks and the space tasks are represented by adopting different numbers of sampling frames and visual mark forms. The visual mark is aligned to the text space and then input into the large language model together with the text mark, and a corresponding text answer is obtained through decoding. For a time task, a timestamp is directly extracted from a text answer, and spatial information passes through lt; sEGgt, SEGgt; and the mark code is embedded as a prompt input to the mask decoder, so that mask generation of the sampling frames and mask propagation of the whole video are realized. Compared with the prior art, the method has the advantages that joint training of fine-grained video time-space understanding is realized, and more accurate event time-space positioning can be realized besides overall understanding of the video.
Owner:FUDAN UNIVERSITY

Generating video descriptions using a machine learning model

The present disclosure describes techniques for generating video descriptions using a machine learning model. A plurality of sets of visual tokens corresponding to a plurality of frames of a video is generated. A first type of tokens is generated by implementing temporal pooling on the plurality of sets of visual tokens corresponding to the plurality of frames. A second type of tokens is generated by compressing each of the plurality of sets of visual tokens corresponding to each of the plurality of frames. A third type of tokens is generated by applying cross-attention between each of the plurality of sets of visual tokens and a fourth type of tokens including text tokens generated based on an input text query. A text description of the video is generated based on the first type of tokens, the second type of tokens, the third type of tokens, and the fourth type of tokens.
Owner:LEMON INC(GB)

Calibration method for positioning and attitude determination of heading machine

The invention relates to the technical field of engineering construction measurement and control, and discloses a calibration method for positioning and attitude determination of a heading machine, which comprises the following steps: constructing a data acquisition module, a route making module, a real-time dynamic monitoring module, an execution module, a feedback module and an optimization module; the data acquisition module is used for completing comprehensive acquisition of multi-source data of the heading machine, a route, an environment and an obstacle, the route making module is used for fusing the multi-source data to carry out path planning and error compensation, and the real-time dynamic monitoring module is used for realizing heading machine pose tracking and emergency correction through visual marking points and a laser scanner. The execution module is used for converting an optimal tunneling path into a driving motor control instruction, the feedback module is used for collecting deviation data of an actual operation pose and a planned path in real time and returning the deviation data, the optimization module realizes long-term self-adaptive optimization of the system, and finally, precision, safety and high efficiency of the tunneling operation of the tunneling machine in a complex underground environment are realized.
Owner:TAIYUAN INST OF CHINA COAL TECH & ENG GROUP +1

Perception-Based Worksite Control System

A control system manages a plurality of mobile machines each equipped with a visual perception system to capture perception data. The control system via an onboard controller applies an object detection operation to the perception data to detect a detected marker position corresponding to a visual marker. The onboard controller assesses an assessed marker heath status with respect to the detected marker position and transmits that to a central worksite server. The central worksite server aggregates the assessed marker health statuses from a plurality of mobile machines to determine an aggregate marker health status for the visual marker.
Owner:CATERPILLAR INC

Automated captioning of augmented reality effects in videos

Described herein are techniques for generating captions for augmented reality (AR) effects in videos using an adapted multimodal large language model (MLLM). The technique involves sampling frames from base and AR-applied videos, combining them into concatenated frames, and processing them with a hybrid vision encoder. Visual tokens are projected into a language model token space, reshaped, downsampled, and interleaved with text tokens. A fine-tuned large language model processes this input sequence to generate AR effect captions. Optical character recognition extracts text from AR frames, which is combined with generated captions and metadata to produce merged captions and content tags. This approach enables accurate description of temporal AR effects and facilitates downstream applications like search and ranking.
Owner:SNAP INC

Method for controlling lighting effect of spliced lamp, device, apparatus, and medium

The present application relates to a method for controlling a lighting effect of a spliced lamp, as well as a device, an apparatus, and a medium. The method comprises: displaying a spatial topological graph of the spliced lamp in an effect editing region of a graphical user interface; displaying a visual marker of a motion base point of the spliced lamp in the effect editing region in an initialized manner; updating the motion base point according to the latest position of the visual marker relative to the spatial topological graph; driving the spliced lamp to play a lighting effect by using a lighting effect playing instruction that is generated on the basis of the motion base point, such that a mapping position of the motion base point in the physical space serves as a benchmark reference point of a motion process of the lighting effect.
Owner:SHENZHEN INTELLIROCKS TECH CO LTD

Agaricus bisporus autonomous picking method and system based on VLA

The invention provides a VLA-based agaricus bisporus autonomous picking method and system, and relates to the technical field of cultivation, and the method comprises the steps: obtaining picking task description information and visual observation data, and coding the visual observation data to obtain a visual feature sequence; constructing a multi-scale visual feature map based on the visual feature sequence, and generating an object-level visual mark sequence; encoding the picking task description information into a text marking sequence, mapping the text marking sequence and the object-level visual marking sequence to a unified semantic space, and splicing the text marking sequence and the object-level visual marking sequence to form a joint semantic sequence; and inputting the joint semantic sequence into a pre-trained VLA model, performing autoregression prediction to generate an action mark sequence, decoding according to the action mark sequence to obtain a picking control instruction, and driving a picking device to complete picking, so as to improve the key visual information extraction capability and enhance the accuracy and robustness of action generation in a complex unstructured cultivation scene.
Owner:SHANGHAI HENGZE FUHUI INTELLIGENT TECHNOLOGY CO LTD

Anti-leak digital document marking system and method using distributed ledger

The system is disclosed for visual marking sensitive documents for leak prevention. Each time an action is taken with regard to a document (e.g., creation, viewing, downloading), that action is added to a distributed ledger, essentially creating a unique hash for a new instance of the document. This new hash is visually embedded in the document as a code comprising a plurality of differently shaded pixels, wherein some of the pixels directly encode information regarding the document (e.g., an account that generated the new instance of the document, a date, a time, a unique ID for the document, etc.) and some of the pixels do not encode information. The code is capable of being scanned either digitally or physically on a printed version of the document, such that the immediate source of the document, corresponding to who leaked the document, is able to be discerned.
Owner:EQUITY SHIFT INC

Yoga mat surface defect automatic identification method based on intelligent sensor

The invention relates to the technical field of intelligent surface defect identification of an intelligent sensing detection and optical scattering method, and discloses a yoga mat surface defect automatic identification method based on an intelligent sensor. Establishing a two-dimensional coordinate and a discrete sampling array on the surface of a measured object, and unifying an installation posture; scanning the scattering intensity of each point according to a multi-angle sequence under fixed oblique incidence, and checking the integrity and the range; summarizing according to fixed angle weights to form an angular spectrum value matrix; carrying out point-by-point difference on the defect-free reference matrix to obtain a residual error, carrying out difference synthesis along rows and columns to obtain mutation intensity, and constructing a matrix; generating candidate anomalies according to a reference mutation statistics set threshold value; a connected domain is extracted on the grid in a four-adjacent mode, insufficient area is eliminated, and the center and the area are calculated; and outputting the number, the position and the area of the defects and performing visual marking.
Owner:ZHEJIANG SUNRISE HIGH TECH NEW MATERIAL CO LTD

Entity relationship diagram generation for databases

A database system includes at least one data storage device storing at least one database and one or more processors configured to identify, from Structured Query Language (SQL) commands received for the at least one database, entities of the SQL commands, attributes of the entities, and relationships between the entities. The identified entities, attributes, and relationships are translated into a visual markup language code using a Large Language Model (LLM). In some aspects, the LLM or another LLM may be provided with the SQL commands to identify the entities, attributes, and relationships. An Entity Relationship Diagram (ERD) is generated or updated for the at least one database based on the translated visual markup language code. In other aspects, at least two of the identified entities, attributes, or relationships are merged for representation in the ERD.
Owner:WESTERN DIGITAL TECHNOLOGIES INC

MOS transistor detection device and method thereof

The application provides a MOS tube detection device and method, and relates to the technical field of detection equipment. The MOS tube detection device comprises a fixing frame, a conveying device, a guide contact detection assembly, a lifting assembly, an electric detection module, a first camera, a second camera, a material separating device and a controller. The MOS tube detection method is used. After the steps of material conveying, preliminary visual detection, power-on detection, visual marking and material falling and separating, the MOS tube can be quickly and effectively detected. The detection efficiency of the MOS tube is greatly improved. The detection process does not need to be transported again, the detection time is saved, the MOS tube detection is convenient and fast, and the detection equipment is convenient for miniaturization. The MOS tube detection device and method can quickly and effectively perform power-on detection on the MOS tube, so as to improve the detection efficiency of the MOS tube.
Owner:SHENZHEN SHENWEI SEMICON CO LTD

Anti-leak digital document marking system and method using distributed ledger

ActiveUS12718311B1Visual markingPrint version
The system is disclosed for visual marking sensitive documents for leak prevention. Each time an action is taken with regard to a document (e.g., creation, viewing, downloading), that action is added to a distributed ledger, essentially creating a unique hash for a new instance of the document. This new hash is visually embedded in the document as a code comprising a plurality of differently shaded pixels, wherein some of the pixels directly encode information regarding the document (e.g., an account that generated the new instance of the document, a date, a time, a unique ID for the document, etc.) and some of the pixels do not encode information. The code is capable of being scanned either digitally or physically on a printed version of the document, such that the immediate source of the document, corresponding to who leaked the document, is able to be discerned.
Owner:EQUITY SHIFT INC

A question and answer method and system for traffic events based on natural language

The application discloses a kind of based on natural language's traffic event's question and answer method and system, belong to intelligent question and answer field, first acquisition roadside image and user query, respectively by visual encoder and natural language encoder extraction multiscale visual feature set and text feature;Then utilize pre-training cross-modal intent alignment and focus module, with text feature as guide to spatial attention anchor and cross-modal intermingling to visual feature, output joint representation feature;Again, the feature is input light-weight decoder obtained by structured logic decoupling and migration strategy distillation, so that it executes step-by-step logical inference and coordinate regression based on spatial response graph in parallel, synchronously output natural language answer text and target bounding box coordinates;Finally generate traffic event question and answer result containing semantic answer and visual mark.The application effectively solves the problem that prior art cannot accurately and efficiently based on natural language to traffic event intelligent question and answer.
Owner:HONG KONG UNIV OF SCI & TECH (GUANGZHOU)

Automated captioning of augmented reality effects in videos

Described herein are techniques for generating captions for augmented reality (AR) effects in videos using an adapted multimodal large language model (MLLM). The technique involves sampling frames from base and AR-applied videos, combining them into concatenated frames, and processing them with a hybrid vision encoder. Visual tokens are projected into a language model token space, reshaped, downsampled, and interleaved with text tokens. A fine-tuned large language model processes this input sequence to generate AR effect captions. Optical character recognition extracts text from AR frames, which is combined with generated captions and metadata to produce merged captions and content tags. This approach enables accurate description of temporal AR effects and facilitates downstream applications like search and ranking.
Owner:SNAP INC

Training-free gridding video feature compression method

The invention discloses a training-free gridding video feature compression method. The long video processing efficiency of a multi-modal large model is remarkably improved through space-time fusion and semantic compression technologies. Comprising the following steps: recombining a long video frame sequence into a grid image according to a time sequence, and extracting spatio-temporal joint features by using a visual Transform; calculating cosine similarity based on the global semantic center, and screening the first U key visual marks; and non-key mark information is dynamically fused to the key mark through the normalized weight, so that redundant compression is realized. And finally, aligning the compressed visual mark with a text space through linear projection, splicing the compressed visual mark with task prompt and user query, and inputting the spliced result into a large language model. According to the method, on the premise of being completely free of training, the number of visual marks is reduced, a GPU video memory is reduced, the reasoning speed is increased, meanwhile, the accuracy is improved, the problem of calculation bottleneck in long video understanding is effectively solved, and the method is suitable for plug-and-play deployment of mainstream multi-mode frames such as LLaVA and Video-LLaMA.
Owner:UESTC (SHENZHEN) ADVANCED RES INST

An AI interactive AR augmented reality travel culture wall system

This invention discloses an AI-interactive AR (Augmented Reality) cultural wall system for cultural tourism, belonging to the field of AR system technology. The surface of the cultural wall is equipped with a pre-embedded feature set, including visual markers, differentiated texture features, 3D depth variation areas, and specifically designed interactive hotspots. These features provide the system with stable visual benchmarks, semantic distinction criteria, depth cues, and precise positioning anchors. The system includes: a feature point acquisition module, used to collect the feature set on the cultural wall through a multimodal feature extraction algorithm to construct a 3D spatial map with semantic annotations; a user positioning and pose estimation module; an AI content generation engine; and a spatial adaptation and rendering module. Through the cultural wall feature set and the multi-module collaborative system, precise spatial integration of virtual content and the physical wall surface is achieved, along with intelligent dynamic content generation based on user behavior, effectively enhancing the immersion, interactivity, and personalization of the AR cultural experience.
Owner:ZUNCHUANG TECH GRP CO LTD

Vision mark based autonomous driving space planning enhancement method

The application discloses an automatic driving space planning enhancement method based on visual markers, comprising the following steps: acquiring an original image and text input; processing the original image to obtain image features; processing the text input to obtain text features; generating text output with visual markers by using the image features and the text features; converting the text output with visual markers to obtain text output with coordinates; significantly improving the accuracy and semantic consistency of space understanding in the automatic driving scene, realizing high synchronization of visual perception and semantic expression, and effectively solving the problem of semantic split between vision and language modal in the existing method. Not only greatly improves the analysis accuracy of object position, motion state and interaction relationship in the automatic driving question and answer task, but also can significantly enhance the decision reliability and planning naturalness in the complex driving scene.
Owner:SOUTH CHINA UNIV OF TECH

Intravascular endoscopy balloon

Provided herein is an intravascular endoscope with an improved balloon design. The endoscope comprises an anterior balloon with a camera, light source, and working channel port for visualization and therapeutic access. One or more posterior balloons are positioned inferior to, at the same level as, or superior to the anterior balloon. Inflation of the one or more posterior balloons pushes the anterior balloon against the blood vessel wall while maintaining a channel for blood flow past the anterior balloon. The endoscope enables real-time, direct visualization of the interior of blood vessels for diagnosis and treatment of vascular diseases while preserving critical blood flow. Additional imaging modalities, such as ultraviolet and infrared, can be incorporated to enhance tissue characterization. Visual markers on devices inserted through the working channel improve visibility and guidance under endoscopic imaging. The endoscope is compatible with arteries and veins throughout the body and provides advantages over traditional imaging techniques for managing vascular disease.
Owner:NORTHWESTERN UNIV

Method and apparatus for training multimodal large model, and method and apparatus for image question answering

Method and apparatus for training multimodal large model and method and apparatus for image question answering are disclosed, which relates to artificial intelligence technologies such as large models, deep learning, natural language processing, and computer vision. The method for training multimodal large model includes: obtaining an initial sample image, a sample object in the initial sample image, and a location information of the sample object; obtaining a target sample image including a sample visual marker based on the initial sample image and a target image region corresponding to the initial sample image; obtaining a sample question corresponding to the target sample image based on the sample visual marker, and obtaining a sample answer corresponding to the sample question; training an initial multimodal large model based on a target training sample constituted by the target sample image, the sample question and the sample answer to obtain a target multimodal large model. The method for image question answering includes: obtaining a target image including a target visual marker and a target question; inputting the target image and the target question into the target multimodal large model to obtain a target answer. The present disclosure enables the target multimodal large model to effectively understand the target visual marker in the target image, thereby improving the accuracy of the target answer.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Multi-camera cooperative state estimation method based on active infrared visual features

The invention discloses a multi-camera cooperative state estimation method based on active infrared visual features, and belongs to the technical field of image processing and computer vision. The existing visual marking technology depends on ambient light and is poor in performance under dark light or complex illumination, the observation visual field of a single camera or a fixed camera is limited, distortion exists, and consequently the pose resolving precision and robustness are insufficient. According to the method, active infrared visual feature marking is adopted, dependence on ambient light is avoided, and feature point information is stably provided. And efficient and accurate feature point identification and matching are realized through multi-camera collaborative observation in combination with binarization processing of dynamic threshold adjustment, geometric feature screening and a matching algorithm. And carrying out pose calculation by using a PnP algorithm, and checking the goodness of fit between the axis and the gravity direction. Meanwhile, multiple cameras are dynamically scheduled to rotate cooperatively according to the observation quality, and the observation view is expanded. The method is mainly used for accurate positioning of targets such as an unmanned aerial vehicle, significantly improves the environmental adaptability and observation stability of the system, reduces the deployment cost, and improves the pose calculation precision and robustness.
Owner:CHINA YANGTZE POWER

Video content intelligent mark display method and device based on AI pre-detection

The invention relates to the technical field of combination of front-end video processing and artificial intelligence application, in particular to a video content intelligent mark display method and device based on AI pre-detection, and the method comprises the steps: building a mapping relation between detection problems and a video time axis, and carrying out the intelligent grouping of adjacent detection problems through a time merging threshold value, the playing time can be monitored in real time in the video playing process, the visual mark and information display of the detection problem group accurately matched with the current playing time period can be dynamically controlled, interference caused by frequent mark switching is avoided, key problems are ensured to be presented in time, and the user experience is improved. Therefore, the perception efficiency and interaction experience of the user on the video content problem are improved.
Owner:BEIJING QINGSONG YIKANG INFORMATION TECHNOLOGY CO LTD

A method and system for autonomous picking of agaricus bisporus based on vla

The application provides an autonomous picking method and system for Agaricus bisporus based on VLA, and relates to the technical field of cultivation. The method comprises the following steps: acquiring picking task description information and visual observation data, encoding the visual observation data to obtain a visual feature sequence; constructing a multi-scale visual feature map based on the visual feature sequence, and generating an object-level visual label sequence; encoding the picking task description information into a text label sequence, mapping the object-level visual label sequence to a unified semantic space, and splicing to form a joint semantic sequence; inputting the joint semantic sequence into a pre-trained VLA model, and generating an action label sequence through autoregressive prediction; decoding the action label sequence to obtain picking control instructions, and driving a picking device to complete picking, so as to improve the key visual information extraction capability in a complex and unstructured cultivation scene, and enhance the accuracy and robustness of action generation.
Owner:SHANGHAI HENGZE FUHUI INTELLIGENT TECHNOLOGY CO LTD

Intelligent suspension conveying method and system based on vision measurement

The invention relates to an intelligent suspension conveying method and system based on vision measurement. The method comprises the following steps: acquiring image sequence data in real time, extracting the offset of a visual mark point, forming an offset time sequence, analyzing a dynamic damping coefficient, a wind vibration main frequency and a current phase, and predicting the predicted offset of a cable chain at the next moment; judging whether a preset safe swing threshold value is exceeded or not; if yes, the phase compensation amount is determined according to the symbol of the predicted offset and the current phase, and a sine wave signal is generated based on the wind vibration main frequency and the synthetic phase superposed with the phase compensation amount; multiplying the absolute value of the predicted offset by a preset proportionality coefficient to obtain a linear basic quantity; and performing speed compensation according to the linear basic quantity, the dynamic damping coefficient and the sine wave signal, generating a final execution speed instruction, and controlling the driving motor to operate according to the final execution speed instruction. Therefore, the stability and the safety of the suspension conveying system in the conveying process are improved.
Owner:SUZHOU UNIV

Protective garment having improved closing flap

A garment comprising protective apparel fabric, a fastener assembly for joining a first and a second area of the protective apparel fabric, and a closing flap for covering the fastener assembly, the closing flap attached to the outer garment surface of first area of protective apparel fabric, the closing flap having a size and shape such that the covering area of the closing flap fully covers the fastener assembly when the fastener assembly is closed, the closing flap having an average bending rigidity that is at least 7.5 percent greater than the average bending rigidity of the protective apparel fabric; the color and / or visual marking or pattern of the closing flap can be different from or distinct from other parts of the garment.
Owner:DUPONT SAFETY & CONSTRUCTION INC

Sample table automatic positioning analysis cavity based on visual feedback and sample transferring method

The invention discloses a sample table automatic positioning analysis cavity based on visual feedback and a sample transferring method. The mechanical structure comprises an upper cavity, a lower cavity and a first communicating cavity communicating the upper cavity with the lower cavity and provided with a gate valve. The upper end of the upper cavity is provided with a composite motion device integrated with X-axis, Y-axis and Z-axis displacement tables and a polar angle rotating device through a second communication cavity; the side wall of the lower cavity communicates with the analyzer. Aiming at the problem of depth information loss existing in monocular vision, the system adopts a binocular vision system to determine three-dimensional space coordinates (X, Y and Z) of visual marks on a sample table, so that the problem of depth blur is fundamentally solved, and a core guarantee is provided for realizing micron-sized positioning. The processing module pre-stores standard positioning coordinates and integrates image acquisition, information processing and position control functions. And micron-order positioning precision, intelligent safety guarantee, thermal deformation real-time compensation and full-process automation are realized.
Owner:UNIV OF SCI & TECH OF CHINA