Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

205 results about "Visual Objects" patented technology

Visual Objects is an object-oriented computer programming language that is used to create computer programs that operate primarily under Windows. Although it can be used as a general-purpose programming tool, it is almost exclusively used to create database programs.

Model-free six-dimensional object pose estimation

A composite pose-estimation algorithm includes a video-object segmentation sub-algorithm (311) configured to determine a mask of a visual object in an image, and an object-pose tracking sub-algorithm (312) configured to track a pose of a visual object over multiple depth-video frames, wherein the pose-estimation algorithm is configured to input a depth video, from which frames are extracted and fed to the video-object segmentation sub-algorithm, which determines respective object masks to be used by the object-pose tracking sub-algorithm alongside the depth video. A method of tracking a pose of a physical object comprises: obtaining a depth video depicting a physical object in a plurality of poses from an input interface (330); forming a storable data item representing the physical object by applying the pose-estimation algorithm to the depth video; and tracking the physical object or a copy thereof using an instance of the pose-estimation algorithm which has been initialized by means of the storable data item.
Owner:ABB (SCHWEIZ) AG

Scientific and technological intelligence deep analysis method and system based on cross-modal semantic enhancement

The invention provides a science and technology information deep analysis method and system based on cross-modal semantic enhancement, and relates to the technical field of science and technology information analys.The method comprises the steps that firstly, a cross-modal semantic anchor point set is constructed, and the cross-modal semantic anchor point set comprises text theme anchor points extracted from science and technology information texts, visual object anchor points extracted from images and the association mapping relation of the text theme anchor points and the visual object anchor points; constructing a semantic conduction path between anchor points based on the cross-modal semantic anchor point set, realizing bidirectional information transmission, generating a cross-modal semantic enhanced representation, performing hierarchical semantic analysis on the enhanced representation to obtain a topic association rule, a technical element dependency relationship and a concept evolution sequence, and integrating the topic association rule, the technical element dependency relationship and the concept evolution sequence into an analysis conclusion; the analysis conclusion is reversely mapped to adjust the association mapping relation strength, an updated set is obtained, finally, a structured science and technology information analysis report is generated based on the updated set, logic connection of all modules is achieved, and comprehensive and accurate science and technology information analysis is provided for users.
Owner:BEIJING SCI & TECH PATENT OFFICE

Intelligent text verification method based on hybrid model knowledge graph

The invention relates to the technical field of text verification, in particular to an intelligent text verification method based on a hybrid model knowledge graph, which comprises the following steps of: analyzing a document, separating a text from a visual object, and generating semantics and visual vectors by using a bidirectional encoder and a hybrid visual model; performing form normalization verification by constructing a self-adaptive template matrix; judging the semantic homology of the image-text content by using a cross-modal gating arbiter; the text is converted into a semantic fact triple mapped to a unified space-time coordinate system, and logic irregularity is detected in a domain knowledge graph based on ontology constraint; and finally, summarizing all results to generate a structured verification report. According to the method, cross-modal semantic understanding and knowledge graph reasoning are effectively fused, full-dimension intelligent verification of content forms, image-text semantics and deep space-time causal logic is achieved, and the depth and accuracy of large-scale digital content verification are remarkably improved.
Owner:NANJING DIGITAL TECHNOLOGY CO LTD

Passable area reasoning method and system based on visual language model

PendingCN121767911AAchieve collaborative understandingEnable high-level semantic reasoningCharacter and pattern recognitionBiological modelsSemantic alignmentVision based
The invention provides a passable area reasoning method and system based on a visual language model, and the method comprises the steps: obtaining the multi-modal data of a vehicle and the current position information of the vehicle; analyzing the multi-modal data, and determining visual features and traffic symbol features; performing spatial position coding on the visual object and the traffic symbol elements, and determining aerial view angle coordinate information; performing semantic alignment on the visual features and the traffic symbol features, and determining a shared embedding representation; constructing a traffic semantic map by fusing, sharing and embedding representation based on a graph neural network and bird's-eye view coordinate information of a visual object and a traffic symbol element; and according to the current position information of the vehicle, the traffic semantic map and a preset traffic rule, generating a bird's-eye view semantic map including a passable area, a no-pass area and a semantic association relationship. According to the method and the device, semantic alignment and consistency expression of visual perception and traffic symbol recognition are realized, and further feasible region reasoning of a complex traffic scene is realized.
Owner:SHANGHAI JIAOTONG UNIV

Technological framework for a dynamic virtual card deck system

An interactive computer-implemented system manages a virtual deck of digital cards whose attributes, such as value, status, score, or color, refresh continuously in response to live external data. A server-side ingestion pipeline normalizes event feeds and maps them to a programmable card object model executed on one or more processors. A rendering engine delivers sub-five-second visual updates, while a synchronization layer broadcasts state changes to all connected clients to maintain uniform gameplay. A rules-based constraint engine recalculates permissible user actions as card states evolve, preventing pre-event optimization and preserving competitive fairness. The architecture is device-agnostic, supports accessibility overlays, and extends beyond fantasy sports to any domain where real-time data drives interactive visual objects.
Owner:GIVANT PHILIP PAUL

System and method for detecting and identifying container number in real-time

Exemplary embodiments of the present disclosure are directed towards a method for detecting and identifying container number in real-time. Monitoring vehicle carrying containers and triggering first camera, second camera, third camera, fourth camera, fifth camera, and laser sensors to capture container views by pre-processing module. Transmitting containers image data to computing device by the pre-processing module. Detecting container number region in container image frames by visual object detection module. Cropping container number region by visual object detection module. Applying two-dimensional Fast Fourier Transform on cropped container number region. Segmenting each character situated in container number region by segmentation and character classification module. Classifying each character situated in container number region by segmentation and character classification module. Arranging characters in order based on relative positions of characters to obtain container number information by segmentation and character classification module. Aggregating container image frames and generating container number by post-processing module.
Owner:ATAI LABS PTE LTD

Visual target tracking method based on natural language and target state information

The invention discloses a visual target tracking method based on natural language and target state information. The method comprises the following steps: (1) constructing a training sample set; (2) constructing a visual target tracking model based on a natural language and target state information; step (3), adjusting parameters of an image-text encoder and loading a pre-training weight to obtain a feature after the text and the first template are fused, a feature of the second template and a feature of the search image; (4) fusing the position information of the target in the sample set and the bounding box information of the target into the features of the second template; step (5), obtaining features after joint modeling; step (6), acquiring a token containing target position information after query; (7) obtaining a predicted target bounding box regression result; and (8) obtaining a final tracking result. According to the invention, the tracking accuracy of the visual tracker based on the natural language is effectively improved.
Owner:XIDIAN UNIV

Electronic device, method, and computer readable storage medium for detection of vehicle appearance

According to various embodiments, an electronic device include a display, an input circuit, at least one memory and at least one processor configured to obtain a first image; display, in response to cropping an area comprising a visual object corresponding to a potential vehicle appearance from the first image, fields for inputting an attribute for the area, wherein, the fields include a first field for inputting a vehicle type as the attribute and a second field for inputting a positional relationship between a subject corresponding to the potential vehicle appearance and a camera obtained the first image as the attribute; obtain information about the attribute, by receiving a user input for each of the fields including the first field and the second field through the input circuit; store a second image configured of the area in a data set for training a computer vision model for vehicle detection.
Owner:THINKWARE

Wearable device, method, and non-transitory computer readable storage medium for eye calibration

A method executed by a wearable device including a display system including a first display and a second display facing eyes of a user wearing the wearable device, and a plurality of cameras configured to obtain an image including the eyes, includes: displaying objects at different time points on a screen of the display system, based on the image, identifying gazes looking at the objects, identifying, based on the identified gazes, errors associated with the gazes, wherein the errors indicate differences between display positions of the objects and focal positions of the gazes, and wherein the focal positions have a one-to-one correspondence with the objects, displaying a visual object on a background screen on the display system to move the visual object through partial display positions of the display positions, which are selected based on the errors, and based on another gaze looking at the visual object, correcting the errors.
Owner:SAMSUNG ELECTRONICS CO LTD

Method and system for identifying multi-modal named entities

The invention discloses a method and a system for identifying a multi-modal named entity, and belongs to the technical field of digital data processing. In order to solve the technical problems that in the prior art, when images and text information are processed, shared information and private information are sequentially connected in series, and the shared information and the private information of visual objects in the images are directly connected in series, so that feature information confusion is caused, fine-grained alignment in visual modes is influenced, and cross-modal understanding of a GMNER system is influenced. The shared visual features and the private visual features of the image are extracted respectively, the features of the visual objects in the image and the relation features between the visual objects are distinguished, and then the images are dynamically integrated and projected to the text embedding space, so that the corresponding relation between the visual object entities and the text entities is clearer, and the text embedding efficiency is improved. And the accuracy of fine granularity alignment is improved, so that the comprehensive cross-modal understanding capability of the GMNER system is improved. The method is mainly used for multi-modal named entity recognition.
Owner:HARBIN INST OF TECH +1

Foldable electronic device, method, and non-transitory computer-readable storage medium for adaptively displaying visual object

PCT designated stageWO2026141873A1Computer hardwareVisual Objects
This foldable electronic device comprises: at least one processor including processing circuitry; a housing including a first housing part and a second housing part; a flexible display including a first display area corresponding to the first housing part and a second display area corresponding to the second housing part; at least one sensor; an NFC circuit in the first housing part; and a memory which stores one or more programs configured to be individually or collectively executed by the at least one processor, and includes one or more storage media, wherein the one or more programs may include instructions for causing the foldable electronic device to: receive a signal from an external electronic device by using the NFC circuit; and display a visual object associated with the external electronic device and located in the first display area, on the basis of the signal received through the at least one sensor while identifying that the first housing part and the second housing part are partially folded.
Owner:SAMSUNG ELECTRONICS CO LTD

Visual target tracking dynamic calculation and distribution method based on scene complexity perception

The invention discloses a visual target tracking dynamic calculation distribution method based on scene complexity perception, and relates to the technical field of computer vision and artificial intelligence, and the method comprises the steps: firstly constructing a multi-layer visual Transform backbone network comprising a fixed layer and a dynamic layer; a scene complexity analyzer is activated behind the first dynamic layer, pooling, similarity calculation and enhancement processing are carried out on the template and search area features, and the exit score of each dynamic layer is predicted; and then, according to comparison between the score and a preset threshold value, dynamically determining whether to terminate reasoning in advance. And meanwhile, the teacher model knowledge is migrated to a plurality of dynamic layers of the student model by adopting a layered distillation method, so that the prediction precision of the early layer is improved. According to the method, adaptive perception of scene complexity and dynamic allocation of computing resources are realized, the balance of high precision and high real-time performance is achieved on resource-constrained equipment, and the reasoning efficiency of visual target tracking is remarkably improved.
Owner:JIANGNAN UNIV

Electronic device, method, and computer-readable medium for displaying visual object

An electronic device may include: a memory storing instructions, a processor(s), and a display. The instructions, when executed by the processor(s), may cause the electronic device to: display a first visual object at a first spot via the display, the first visual object being the topmost visual object among a plurality of visual objects stacked according to an arrangement sequence; display one or more visual objects other than the first visual object among the plurality of visual objects via the display while the first visual object is being displayed to move in response to a user input to the first visual object; and after the user input, display the first visual object at a second spot, dependent on the user input, via the display. The one or more visual objects may be displayed to sequentially follow the movement path of the first visual object according to the arrangement sequence while the first visual object is being displayed to move.
Owner:SAMSUNG ELECTRONICS CO LTD

Enhanced object detection with retrieval augmented generation and language model prompting system

Certain aspects of the disclosure provide a method for enhanced object detection. The method includes providing, to a machine learning (ML) model, a first prompt comprising a first image and a first instruction to output a first description; obtaining, from the ML model, the first description comprising an identification of an unidentified visual object; generating an embedding of the unidentified visual object; obtaining, from a retrieval augmented generation (RAG) database, an embedding associated with a known visual object and satisfying a similarity threshold; retrieving information associated with the known visual object; generating an enhanced context comprising the information associated with the known visual object; providing, to the ML model, a second prompt comprising the enhanced context and a second instruction to output a second description of the first image; and obtaining, from the ML model, the second description including an identification of a visual object associated with the unidentified visual object.
Owner:INTUIT INC

Electronic device and method for displaying modification of virtual object and method thereof

According to an embodiment, at least one processor of a wearable device may display, based on an input for entering a virtual space, on a display the virtual space. The at least processor may display within the virtual space a first avatar which is a current representation of a user and has a first appearance. The at least processor may display within the virtual space a first avatar together with a visual object for a second avatar which is a previous representation of the user and has a second appearance different from the first appearance of the first avatar. For example, the metaverse service is provided through a network based on 5G (fifth generation), and / or 6G (sixth generation).
Owner:SAMSUNG ELECTRONICS CO LTD

System for low-photon-count visual object detection and classification

A computing system can be configured for low-photon-count visual object classification. The computing system can include a photon-detection system that includes one or more cells. Each of the one or more cells can include one or more photon detectors. Each of the one or more photon detectors can be configured to output a photon signature in response to a photon incident on the one or more photon detectors. The computing system can include one or more processors and one or more storage devices storing computer-readable data. The data can include a low-photon-count classification model and one or more instructions that, when implemented, cause the one or more processors to perform operations for low-photon-count visual object recognition. The operations can include obtaining a photon signature from the photon-detection system. The operations can include providing the photon signature to the low-photon-count classification model. The operations can include determining, by the low-photon-count classification model, a classification of a visual object placed in a field of view of the photon-detection system based at least in part on the photon signature. The operations can include providing the classification as an output of the low-photon-count classification model.
Owner:GOOGLE LLC

Wearable device, method, and non-transitory computer-readable storage medium for displaying visual object corresponding to external object

The present invention may comprise: a memory for storing instructions and including one or more storage media; one or more cameras; a display assembly including a display; and at least one processor including processing circuitry, wherein the instructions, when individually or collectively executed by the processor, cause the wearable device to: display, on the display assembly, an avatar representing a user; obtain images by using the one or more cameras while displaying the avatar; detect, using at least a portion of the images, an external object gripped by the user's hand; identify a first size of the user's hand by using the at least a portion of the images on the basis of the detection; identify a second size of a visual object corresponding to the external object according to the first size; and display, on the display assembly, the avatar gripping the visual object having the second size.
Owner:SAMSUNG ELECTRONICS CO LTD

Electronic apparatus and method for identifying content

An electronic device includes a camera, a communication circuit, memory storing one or more computer programs, and at least one processor. The one or more computer programs include computer-executable instructions that, when executed by the one or more processors, cause the electronic device to obtain information on a second image comprising a plurality of visual objects and changed from a first image, transmit, to a server, the information on the second image, receive information on a three-dimensional (3D) model for a space comprising the plurality of external objects and information on a reference image, obtain a third image by removing, from the second image, at least one visual object among the plurality of visual objects, identify at least one feature point based on a comparison between the reference image and the third image, identify a pose of a virtual camera, and identify content superimposed on the first image.
Owner:SAMSUNG ELECTRONICS CO LTD

A spatio-temporal graph convolution-based video question answering method and system

The application discloses a video question answering method and system based on a space-time graph convolution, wherein the method comprises the following steps: preprocessing video data and question text data; performing visual feature extraction based on the video data to generate a target matrix; performing space information aggregation based on the target matrix by using a gated multi-layer perceptron layer to obtain visual feature representation; constructing a space-time graph based on the visual feature representation, performing space graph convolution processing on the space graph, and performing time dynamic mining on the time graph; generating word-level question embedding based on the question text data to construct a question graph; generating video multi-level feature representation based on the target space-time graph by using a hierarchical aggregation algorithm, and performing feature aggregation on the question graph; fusing visual aggregation features and question text aggregation features, and generating a predicted answer based on an answer decoder. The application can dynamically adjust the attention weight of a visual object, enhances multi-modal collaborative perception capability, and improves the reliability of video question answering.
Owner:SUN YAT SEN UNIV +1

Method for automatically generating responsive media

A method includes: accessing a static visual objects, and media formats; defining a multi-dimensional feature space representing possible arrangements of combinations of the set of static visual objects within the set of media formats; generating a primary feature container distributed within the multi-dimensional feature space; generating a primary responsive media, by inserting the primary subset of static visual objects into the primary media format according to a primary arrangement of the primary subset of static visual objects represented in the feature container; presenting the primary responsive media to an operator; in response to receiving a selection of the primary responsive media generating a secondary feature container distributed within the multi-dimensional feature space proximal the primary feature container; generating a secondary responsive media, and serving the secondary responsive media to a first device for playback to a first user responsive to inputs by the first user at the first device.
Owner:YIELDMO

Robot control device and method thereof

The present invention relates to a robot control device and a method thereof, the robot control device comprising: a laser radar; an imaging device; a memory in which a classifier group including a plurality of classifiers and a neural network model are stored; and a processor configured to acquire a virtual object represented in a two-dimensional form on the basis of acquiring a point cloud corresponding to an external object by the lidar, project the point cloud to a specified surface, and identify a visual object corresponding to the virtual object within an image acquired by the image pickup device, and generate a visual object represented in a two-dimensional form on the basis of identifying the visual object corresponding to the virtual object within the image acquired by the image pickup device. A portion of an image including a visual object is input into a neural network model, and a plurality of feature maps are input into a classifier group based on a specified number of feature maps associated with the portion of the image acquired from the neural network model, thereby identifying whether an external object corresponding to the visual object is a target object.
Owner:HYUNDAI MOTOR CO LTD +1

Combined zero sample recognition method based on shared learnable soft query vector and related equipment

The invention discloses a combined zero sample recognition method based on a shared learnable soft query vector and related equipment. The method comprises the following steps: acquiring an initial visual feature, an initial text attribute feature, an initial text object feature and an initial text combination feature; performing visual alignment and decoupling by sharing the soft query vector to obtain a target visual attribute feature, a target visual object feature and a target visual combination feature; performing text alignment and decoupling by sharing the soft query vector to obtain a target text attribute feature, a target text object feature and a target text combination feature; obtaining the similarity between the target visual attribute feature and the target text attribute feature, the similarity between the target visual object feature and the target text object feature, and the similarity between the target visual combination feature and the target text combination feature; and carrying out weighted fusion on the similarity to obtain a combined zero sample recognition result. The method can improve the recognition capability of unseen combinations, and can be widely applied to the technical field of computer vision.
Owner:GUANGZHOU UNIVERSITY

Wearable device for displaying visual object, and method thereof

A processor of an electronic device may detect a direction of the electronic device by using a sensor, while executing a first application for providing a virtual space. The processor may display, on the entire display area of a display, a first screen corresponding to a portion of the virtual space, the portion corresponding to the detected direction of the electronic device. The processor may determine, by using the sensor, whether a position of the electronic device is included in a first range, the first range ensuring maintenance of provision of a virtual reality service that is based on the virtual space. The processor may display, on the entire display area of the display, a second screen provided from a second application on the basis of detection of a position of the electronic device included in a second range that is distinguished from the first range. The processor may display, on the second screen, a visual object associated with the virtual space, the visual object having a reduced size on the second screen on the basis of a portion of the display area.
Owner:SAMSUNG ELECTRONICS CO LTD

Electronic device, method, and computer readable storage medium for detection of vehicle appearance

According to various embodiments, an electronic device include a display, an input circuit, at least one memory and at least one processor configured to obtain a first image; display, in response to cropping an area comprising a visual object corresponding to a potential vehicle appearance from the first image, fields for inputting an attribute for the area, wherein, the fields include a first field for inputting a vehicle type as the attribute and a second field for inputting a positional relationship between a subject corresponding to the potential vehicle appearance and a camera obtained the first image as the attribute; obtain information about the attribute, by receiving a user input for each of the fields including the first field and the second field through the input circuit; store a second image configured of the area in a data set for training a computer vision model for vehicle detection.
Owner:THINKWARE

Electronic device for applying function to image, operation method thereof, and recording medium

According to an embodiment, an electronic device may comprise a display, at least one processor, and a memory for storing instructions, wherein the instructions, when executed individually or collectively by the at least one processor, cause the electronic device to: on the basis of identifying a first input for a first image displayed through the display, identify a function corresponding to the first input through a software platform of the electronic device for executing at least one function related to an image; execute the first function through the software platform to identify at least one area of the first image related to the function; display, through the display, a visual object provided by the function on the at least one area of the first image; and on the basis of identifying a second input for selecting a specific area from among the at least one area on which the visual object is displayed, obtain, through the software platform, a second image obtained by applying the first function to the specific area of the first image.
Owner:SAMSUNG ELECTRONICS CO LTD

Wearable device, method, and non-transitory computer readable recording medium for eye calibration

This method is executed in a wearable device comprising: a display system including a first display and a second display which face eyes of a user when worn; and a plurality of cameras arranged to acquire an image including the eyes of the user when worn, and the method may comprise the operations of: displaying objects at least at different viewpoints on a screen displayed through the display system; identifying gazes directed to the objects; identifying errors related to the gazes on the basis of the identified gazes, wherein the errors indicate differences between display positions of the objects and focal positions of the gazes corresponding one-to-one to the objects; displaying, on a background screen of the display system, a visual object moving via some display positions selected on the basis of the errors from among the display positions; and correcting the errors.
Owner:SAMSUNG ELECTRONICS CO LTD

Electronic device and method for displaying image based on interaction

An electronic device receives a shooting input while displaying a preview image based on at least a portion of images obtained through a camera having a first field-of-view (FoV). The electronic device obtains a video of the first FoV through the camera, in response to the shooting input. The electronic device identifies a visual object included in the preview image, while obtaining the video. The electronic device displays the preview image, based on a second FoV that includes the visual object, and is included in the first FoV, in response to an input indicating selection of the visual object. The electronic device obtains meta data indicating reproduction of the video based on the second FoV corresponding to the input, among the first FoV and the second FoV, where the meta data is associated with the video obtained based on the first FoV.
Owner:SAMSUNG ELECTRONICS CO LTD

Electronic device, method, and non-transitory computer-readable storage medium for generating image filter

PCT designated stageWO2026141923A1Computer graphics (images)Radiology
An electronic device is disclosed. The electronic device may display, through a display, a screen including at least one visual object for inputting an image and reference data for generating an image filter. The electronic device may, while displaying the screen, receive a plurality of separate user inputs for the at least one visual object. The electronic device may generate the image filter on the basis of reference data selected on the basis of an initial user input among the plurality of user inputs, and display, through the display, a filtered image in which the generated image filter is applied to the image. The electronic device may update the image filter on the basis of a subsequent user input after the initial user input among the plurality of user inputs, and display, through the display, another filtered image obtained by applying the updated image filter to the image.
Owner:SAMSUNG ELECTRONICS CO LTD

Robot control device and control method thereof

To provide a robot control device for identifying a target object by using a camera and a lidar, and a control method thereof.SOLUTION: A device for controlling a robot according to the present invention includes a lidar, a camera, a memory in which a classifier group including a plurality of classifiers and a neural network model are stored, and a processor, wherein, when a point cloud corresponding to an external object is acquired through the lidar, the processor acquires a virtual object represented in two dimensions by projecting the point cloud onto a designated surface, after a visual object corresponding to a virtual object is identified in an image acquired through a camera, a part of the image including the visual object is input to a neural network model, a specified number of feature maps for the part of the image are acquired from the neural network model, and then the feature maps are input to a classifier group to identify whether an external object corresponding to the visual object is a target object.SELECTED DRAWING: Figure 1
Owner:HYUNDAI MOTOR CO LTD +1

Game debugging method, game debugging device, program product and electronic equipment

PendingCN120909907AError detection/correctionVideo gamesGraphicsProgram graph
The invention provides a game debugging method, a game debugging device, a program product and electronic equipment, and relates to the technical field of games. The method comprises the following steps: in response to a first editing instruction, adding a breakpoint identifier in a first programming graph; wherein the programming graph is a visual object corresponding to a code used for realizing game logic, and the first programming graph is a programming graph associated with a first game map; adding a breakpoint in a first code corresponding to the first programming graph according to the breakpoint identifier to generate a second code; in response to a game test instruction for the first game map, running the first game map in a debugging running mode, and executing the second code; in response to execution to a breakpoint in the second code, pausing execution of the second code; and displaying debugging information of the first programming graph according to the current execution state of the second code. The difficulty of checking and debugging the game editing content is reduced.
Owner:NETEASE (HANGZHOU) NETWORK CO LTD