Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

599 results about "Object based" patented technology

The term " object-based language " may be used in a technical sense to describe any programming language that uses the idea of encapsulating state and operations inside "objects". Object-based languages need not support inheritance or subtyping, but those that do are also said to be "object-oriented".

Hoisting attitude prediction method and device, electronic equipment and storage medium

The invention discloses a hoisting posture prediction method and device, electronic equipment and a storage medium, and relates to the technical field of hoisting postures and the like. The hoisting posture prediction method comprises the steps that a three-dimensional model of a target object is constructed based on region division of the target object; extracting a key mode based on the three-dimensional model and constructing an attitude reduced-order model; obtaining a gravity parameter based on the current hoisting path; and solving the attitude reduced-order model based on the gravity parameter and a first constraint condition to obtain a hoisting prediction attitude. According to the hoisting attitude prediction method disclosed by the invention, the three-dimensional model of the target object considering the precision and the calculated amount is adaptively established by dividing the region of the target object, and the attitude in the hoisting process is predicted in real time by solving the attitude reduced-order model, so that real-time risk early warning and active protection are realized.
Owner:聚变新能(安徽)有限公司

Training-free text-image generation method based on diffusion model

The invention provides a training-free text-image generation method based on a diffusion model, and relates to the technical field of computer graphic processing and artificial intelligence. The method comprises the following steps: extracting semantic phrases and layout information in an input text by utilizing a natural language model, inputting the input text, the semantic phrases and the layout information as additional conditions into a diffusion model, and extracting cross attention maps of different time steps; a positive and negative sample concept and a foreground and background concept based on an object are constructed, a new loss function is calculated on a cross attention map for semantic information and layout information, the loss function combines semantic loss, regional loss and original loss of a diffusion model and is used for updating a potential space image, and the image is generated through iterative denoising and a decoder. According to the method, additional training is not needed, image generation output based on the diffusion model better meets text requirements, and a better text and image alignment effect is achieved.
Owner:SHENYANG JIANZHU UNIVERSITY

Enhanced image and video object detection using multi-stage paradigm

This disclosure describes systems, methods, and devices related to object detection in images. A device may input an image, representing an object, to a manual labeling learner system; identify, using the system, first coordinates of an upper left corner of a bounding box representing the object based on a heatmap indicative of a probability of the first coordinates representing the upper left corner; identify, using the system, second coordinates of a bottom right corner of the bounding box based on the first coordinates and a first distance regression map indicative of coordinate differences between the second coordinates and ground truth coordinates input to the machine learning model as training data; generate, using the system, adjustments to the first coordinates and the second coordinates based on a second regression map; and generate, using the system, the adjusted first and second coordinates, the bounding box.
Owner:INTEL CORP

Neural network-based noise synthesis

Apparatuses, systems, and techniques are presented to introduce noise or variation into images. In at least one embodiment, one or more neural networks are used to determine noise to be added to one or more images of one or more objects based, at least in part, upon one or more features of the one or more objects.
Owner:NVIDIA CORP

Object three-dimensional reconstruction method, device and system based on deep learning

The invention discloses an object three-dimensional reconstruction method, device and system based on deep learning. The reconstruction method comprises the following steps: acquiring a multi-view color image of an object through a controllable image acquisition device; reconstructing a sparse three-dimensional point cloud by using a motion recovery structure method and obtaining a camera pose; initializing parameters of the three-dimensional Gaussian sputtering model based on the sparse point cloud and performing training optimization; a target object semantic segmentation data set is constructed, and a low-rank adaptive technology is adopted to finely segment all models; generating prompts through an open vocabulary detection model at each view angle, obtaining an accurate segmentation mask, and optimizing a three-dimensional segmentation weight by adopting a joint loss function fusing color consistency loss and edge perception loss; and finally outputting the color three-dimensional point cloud of the target object. According to the method, the original image is segmented, so that the influence of the quality of the rendered image is avoided; the segmentation precision of the model in a specific scene is improved through field adaptive fine tuning; and the accuracy of the segmentation boundary is ensured by adopting a double-loss joint optimization mechanism.
Owner:HUNAN AGRI UNIV

Semantic search for prompt builder system

Disclosed herein are system, method, and computer program product aspects for semantic search in a model-based prompt builder system. A system generates a search retriever object based on a search index comprising unstructured data. The search retriever object includes metadata specifying one or more details of a vector search operation to be performed on the search index. The system obtains search results by performing the vector search on the search index based on the one or more details of the vector search operation provided by the search retriever object and a search query. The system provides the search results to a prompt generator configured to use a model to generate a reply to a prompt request requiring the search results.
Owner:SALESFORCE INC

Method, device, and medium for training large scale object foundation model

Embodiments of the present disclosure provide a method, device, and medium for training a large scale object foundation model. The method comprises obtaining a training dataset comprising a plurality of subsets for a plurality of object perception tasks, wherein a sample in the training dataset comprises an image with an object, a prompt indicating the object, and labeled object perception information of the image. The method further comprises generating, by the image encoder, an image feature based on the image. The method further comprises generating, by the text encoder or the visual prompt encoder, a prompt embedding based on the prompt. The method further comprises generating, by the object decoder, object perception information of the object based on the image feature and the prompt embedding. In addition, the method further comprises training the object processing model based on the generated object perception information and the labeled object perception information.
Owner:LEMON INC(GB)

Automated data generation for bimanual dexterous manipulation in robotics systems and applications

In various examples, a technique for neural motion control includes generating a first trajectory associated with a first end effector of an articulated object based at least on a first reference trajectory and one or more transformations between one or more reference poses for one or more objects and one or more new poses for the one or more objects. The technique also includes generating a second trajectory associated with a second end effector of the articulated object based at least on a second reference trajectory and the one or more transformations. The technique further includes determining that a task is successfully completed via execution of the first trajectory and the second trajectory and updating one or more parameters of a machine learning model based at least on the first trajectory and the second trajectory to produce a trained machine learning model.
Owner:NVIDIA CORP

User-configurable object generation

A system may receive a request to generate an object. A system may generate, based on the request, a first portion of a prompt associated with a first layer of a hierarchy and a second portion of the prompt associated with a second layer of the hierarchy, where the second layer has a lower priority than the first layer. A system may generate the object based in part on providing the portions of the prompt as input to a machine learning model, where the object is formatted according to an object format and is internally consistent. To generate the object the machine learning model does not violate instructions associated with a layer of the hierarchy based on instructions associated with a layer of the hierarchy having a relatively lower priority.
Owner:AMAZON TECH INC

Mechanical arm grabbing method and equipment based on few-sample semantic segmentation model and medium

The invention discloses a mechanical arm grabbing method and equipment based on a few-sample semantic segmentation model and a medium, and relates to the technical field of computer systems based on specific calculation models. The method comprises the following steps: constructing a support set according to a scene image containing a target object and an annotation image containing multiple types of target objects; performing object segmentation on the input scene image to obtain a candidate mask image set corresponding to a plurality of segmented objects; in response to the grabbing instruction, determining a target mask corresponding to the specified segmentation object according to the candidate mask image set and the support set based on the few-sample segmentation device, and generating a mask thermodynamic diagram corresponding to the target mask; inputting the mask thermodynamic diagram and a depth map corresponding to the input scene image into a UNet network to determine a grabbing point and a grabbing attitude parameter corresponding to the specified segmented object; and according to the grabbing points and the grabbing posture parameters, the mechanical arm is controlled to execute the corresponding grabbing action, and grabbing of the designated segmented object is achieved.
Owner:INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD

Small sample incremental learning and training method and device based on multi-modal prompt learning

The embodiment of the invention discloses a small sample incremental learning and training method and device based on multi-modal prompt learning, and the learning method comprises the steps: obtaining a to-be-detected original image; determining a first region-of-interest (RoI) feature of a first target object in the original image; determining a matching degree between the plurality of preset multi-modal prototypes and the first RoI feature of the first target object; based on the matching degree between each preset multi-modal prototype and the first RoI feature, performing feature fusion processing on each preset multi-modal prototype to obtain a first fusion feature; determining a second RoI feature of the first target object based on the first fusion feature and the first RoI feature; and determining a target detection result of the first target object based on the second RoI feature of the first target object. According to the method and the device, the RoI features can be optimized based on the multi-modal prototype, so that the accuracy of target object detection is improved.
Owner:AEROSPACE INFORMATION RES INST CAS

Interactive labeling method for 3D dynamic object based on time series data, key frames, and interpolated frames

Disclosed is interactive labeling of a 4D dynamic object based on time series data, which aims at time series-related point cloud dynamic object data. Multi-frame local point clouds in the same time series are transformed into the same global coordinate system with corresponding poses to obtain global point clouds in the same time series are obtained, which clearly shows the moving trajectory of the dynamic object. Taggers can label key 3D boxes based on the moving trajectory of the dynamic object, and automatically generate 3D prediction boxes of other frames based on these key 3D boxes, which significantly reduces the number of frames that need manual operation and solves the problem that 3D prediction boxes generated based on deep learning model are inaccurate and efficiency can hardly be improved.
Owner:MOLAR INTELLIGENCE INFORMATION TECHNOLOGY (HANGZHOU) CO LTD

Abnormality detection method and device, electronic equipment, storage medium and computer program product

The embodiment of the invention provides an anomaly detection method and device, electronic equipment, a storage medium and a computer program product, and is at least applied to the field of artificial intelligence, and the method comprises the steps: carrying out the standardization processing of an index data sequence of a to-be-detected object in a preset time window, and obtaining a standardized index sequence; performing data reconstruction on the standardized index sequence to obtain reconstructed index data; determining a reconstruction error of the to-be-detected object in a preset time window based on the index data sequence and the reconstruction index data; and performing anomaly detection on the to-be-detected object based on the reconstruction error. According to the invention, the detection efficiency of the anomaly detection process can be improved, and the detection quality of the anomaly detection process is ensured.
Owner:SHENZHEN TENCENT COMP SYST CO LTD +1

Real-time process defect detection automation system and method using machine learning model

An artificial intelligence-based process defect detection system may include: a photographing module that collects image data by capturing a process that progresses on an object; a machine learning model that generates work data that is a result of recognizing and reading the object based on the image data; and a detection module that receives instruction data recorded regarding a process for an object optimized for product production, detects a defect or a non-defect by comparing the work data with the instruction data, and generates defect information when the process is defective.
Owner:CREFLE INC

Computer-readable media, information processing system, information processing apparatus, and information processing method

In an example of a game according to an exemplary embodiment, a movement of a movable dynamic object placed in a virtual space is controlled based on physical calculations, and based on an operation input, a player character is caused to perform an object operation action including a first operation of moving a dynamic object specified based on the object operation action and a second operation of forming an assembly object by linking the specified dynamic object to the other dynamic object, and if the assembly object including the dynamic object specified based on the object operation action includes the dynamic object that the player character is mounting, the dynamic object is prevented from moving based on the object operation action.
Owner:NINTENDO CO LTD

Generating virtual objects using autoregressive models and multi-scale tokenization

The disclosed method for generating virtual objects includes generating, based on object data, compressed object data, performing, based on the object data and scales, operations to train a first untrained machine learning model to generate a first trained machine learning model comprising a trained codebook and a trained decoder, wherein the first trained machine learning model is trained to generate a reconstruction of the compressed object data, generating, based on the compressed object data and the scales and using the first trained machine learning model, token maps data, performing, based on the token maps data and conditions, operations to train a second untrained machine learning model to generate a second trained machine learning model comprising a trained autoregressive model, wherein the second trained machine learning model is trained to generate predicted token maps, and generating, based on the scales, conditions, and using both trained models, a virtual object.
Owner:AUTODESK INC

Information processing system, information processing method and program

To use natural language to operate on the elements that make up an object placed in a representation space. [Solution] The information processing system has at least one processor capable of executing a program to perform each of the following steps: in the acquisition step, object information about at least one object placed in a representation space for visually representing the object is acquired, the object information including information about the elements that make up the object; in the assignment step, attributes different from identification codes are assigned to the elements that make up the object based on information about the object and predefined reference information; the attributes are described in natural language; and the reference information indicates the correspondence between each attribute and the state of the object's elements.
Owner:徳山佳央

Device and method for retrieving multimodal object based on composite embedding

Provided are a device and method for extracting a multimodal object on the basis of composite embedding. The device extracts training natural language text and training images from a training data storage, generates image composite embeddings including embeddings of the training images and key objects included in the training images, generates natural language composite embeddings on the basis of the training natural language text, and measure multimodal similarities between the image composite embeddings and the natural language composite embeddings.
Owner:ELECTRONICS & TELECOMM RES INST

Action instruction generation method and device, equipment and storage medium

The invention discloses an action instruction generation method and device, equipment and a storage medium. The accuracy of model reasoning and the accuracy of action instruction generation can be improved. The method comprises the steps of performing feature processing on a natural language instruction, determining an instruction semantic feature, performing feature processing on an image frame sequence, and determining a current visual context feature; predicting the position of the target object based on the instruction semantic features and the current visual context features through a large language model, and outputting current position information corresponding to the target object; the current position information represents the direction and the proximity degree of the target object relative to the origin of coordinates in a sector where the perceptual view coordinate system is located; reasoning based on the current position information, the instruction semantic feature and the current visual context feature through a large language model to generate an action instruction; the action instruction is used for driving an execution mechanism to execute a target task indicated by the natural language instruction.
Owner:BEIJING GALBOT AI CO LTD

Temporally sparse scale estimation for object tracking

PendingUS20260073529A1Image enhancementImage analysisScale estimationObject based
Examples in the present disclosure relate to temporally sparse scale estimation for object tracking. A computing device detects a current orientation of an object. The computing device determines that a difference between the current orientation and at least one of a plurality of previously detected orientations of the object is less than a threshold value. Each previously detected orientation has a respective scale estimate. In response to determining that the difference is less than the threshold value, the computing device generates an effective scale estimate for the object based on a combination of the respective scale estimates for the plurality of previously detected orientations. Each respective scale estimate contributes to the effective scale estimate according to a respective difference between the current orientation and the previously detected orientation for the respective scale estimate. The computing device tracks a pose of the object based on the effective scale estimate.
Owner:SNAP INC

Multi-modal combined image retrieval method fusing fine-grained semantic positioning and optimization generation features

The invention relates to a multi-modal combined image retrieval method fusing fine-grained semantic positioning and optimization generation features, which comprises the following steps: acquiring a reference image, predicting a bounding box of a target object based on the reference image and a positioning keyword, and dividing a background retaining region and a foreground editing region; semantic divergence prompt words of description information are obtained, text features in the semantic divergence prompt words are extracted to serve as guide targets to be used for establishing a composite optimization target function, a gradient descent algorithm is adopted to conduct iterative updating on learnable random noise vectors, and potential visual feature vectors are generated; the foreground editing area is filled with the image data, a background-foreground mixed feature map is generated, the fine-grained visual similarity between the background-foreground mixed feature map and the dense visual feature map is calculated, and the global text feature similarity between global text feature vectors in description information and global visual feature vectors is calculated; and carrying out weighted fusion on the fine-grained visual similarity and the global text feature similarity, and outputting a retrieval result.
Owner:BEIJING UNION UNIVERSITY

Flux sensing system

A flux sensing system includes a processor, a memory, and a sensor. The memory stores at least one semantic profile having a first preference setting, a second preference setting, and a third preference setting in which the first preference setting is associated with a first object, the second preference setting is associated with a second object and the third preference setting is associated with a semantic group including the first object and the second object. The processor infers a semantic associated with a third object, the inferred semantic being based on one or more inputs from the sensor. The processor further detects the presence of the first object and the second object based on inputs from the sensor, and applies or does not apply preference settings based on detected object presence and inferences of the semantic group.
Owner:LUCOMM TECHNOLOGIES INC

Machine learning and multi-stage prompting techniques for generating target classification signatures

Various embodiments of the present disclosure provide machine learning architectures and data processing techniques for improving computer-based text comprehension. The techniques include generating, using a trained classifier model, target classification probabilities for labelled text-based objects from a testing portion of a labelled training dataset and identifying predictive text-based objects from the labelled text-based objects based on the target classification probabilities. The techniques include applying a staged prompting mechanism with a generative extraction model to identify a target set of explanatory text segments from the predictive text-based objects that may be clustered into semantic segment clusters. The techniques include generating explanatory summary segments respectively corresponding to the semantic segment clusters and generating a target classification signature based on a plurality of terms from the one or more explanatory summary segments.
Owner:OPTUM INC

Road operation hidden danger recognition method and device based on artificial intelligence, terminal and medium

The invention discloses a highway operation hidden danger recognition method and device based on artificial intelligence, a terminal and a medium, and relates to the technical field of artificial intelligence and hidden danger recognition, and the method comprises the steps: obtaining a target construction site picture and a pre-constructed hidden danger checking knowledge base; searching a troubleshooting key point and a troubleshooting basis corresponding to the troubleshooting object from a hidden danger troubleshooting knowledge base; inputting the troubleshooting key points, the troubleshooting basis and the target construction site picture into a hidden danger identification large model; wherein the large hidden danger recognition model is configured to execute the following steps: judging whether each troubleshooting object in the target construction site picture conforms to a corresponding troubleshooting key point or not; if not, determining the troubleshooting object which does not conform to the troubleshooting key point as a hidden danger object, and generating a hidden danger description of the hidden danger object; based on the investigation basis, the hidden danger object and the hidden danger description thereof, generating a rectification suggestion of the hidden danger object; and outputting the hidden danger description and rectification suggestions of the hidden danger object according to a preset output format. And the rectification quality is improved.
Owner:CHINA MERCHANTS EXPRESSWAY NETWORK TECH HLDS CO LTD

Credit auditing method and device, equipment and medium

The embodiment of the invention provides a credit auditing method and device, equipment and a medium, relates to the technical field of artificial intelligence, and is used for improving the determination efficiency and quality of a credit auditing result and improving the user experience. The method comprises the following steps: in response to an audit request for a to-be-audited object, performing multiple rounds of dialogue interaction with the to-be-audited object to generate a dialogue file, the dialogue file comprising a risk control decision key data set of the to-be-audited object; determining a first decision result of the to-be-audited object based on the dialogue file and the risk decision original data of the to-be-audited object; when the first decision result is that the to-be-audited object needs to be supplemented, generating and sending a supplementing notification to the to-be-audited object according to the object type of the to-be-audited object; and when a supplementary file uploaded by the to-be-audited object is received, analyzing the supplementary file, extracting risk control decision auxiliary data, and determining a second decision result of the to-be-audited object based on the risk control decision auxiliary data set, the dialogue file and the risk decision original data.
Owner:DUXIAOMAN TECH (BEIJING) CO LTD

Point cloud completion method and device based on text prompt, equipment and storage medium

The invention relates to the technical field of vision and artificial intelligence, and discloses a point cloud completion method and device based on text prompt, equipment and a storage medium, and the method comprises the steps: generating the diffusion loss of a diffusion model according to a first loss model, and generating the completion loss of a completion model according to a second loss model; according to the diffusion loss of the diffusion model, the completion loss of the completion model and the total loss model, generating the total loss of the overall model, updating the model parameters of the overall model, and storing the updated overall model; determining text semantic features of a preset object based on text prompt information corresponding to the target object, determining fusion features of the target object based on a plurality of non-overlapping local areas in the current incomplete point cloud and the text semantic features of the preset object, and updating the overall model based on the fusion features of the target object. And complementing the current incomplete point cloud to generate a complete point cloud of the target object. According to the invention, the complementation efficiency of the complete point cloud of the target object can be improved.
Owner:湖南工商大学

Semantic scene reconstruction and interaction method and system based on visual language model

The invention provides a semantic scene reconstruction and interaction method and system based on a visual language model, and relates to the technical field of artificial intelligence and robots. The method comprises the following steps: acquiring image data shot by a camera; estimating a world pose of the camera; inputting the image data to a Visual Language Model (VLM) to generate a structured description containing an object semantic tag and quantitative spatial information thereof relative to the camera; calculating the world pose of the object according to the camera pose and the quantitative space information; a node representative of the object is created or updated in a hierarchical scene graph. The invention further provides a corresponding interaction system. According to the method, a rich and visual three-dimensional semantic map facing natural language interaction is constructed by deeply fusing the semantic understanding ability of the VLM and the geometric mapping ability of the SLAM, and the intelligent level of robot environment perception and man-machine interaction is greatly improved.
Owner:BEIHANG UNIV

Neural network-based point cloud generation

Apparatuses, systems, and techniques to identify a three-dimensional (3D) point cloud. In at least one embodiment, one or more three-dimensional (3D) point clouds of one or more objects based is generated using one or more neural networks based, at least in part, on one or more two-dimensional (2D) images and the one or more 3D point clouds.
Owner:NVIDIA CORP

Model optimization method and related device

The invention discloses a model optimization method and a related device, which can be applied to automatic driving, auxiliary driving, intelligent traffic, traffic simulation, vehicle-mounted scenes and the like. Perception data of the target object is obtained, an interpretation information prompt text is obtained, and the interpretation information prompt text is used for indicating the content understanding model to output the text content meeting the expectation. And outputting an object behavior control signal of the target object and decision explanation information matched with the explanation information prompt text through a content understanding model according to the perception data and the explanation information prompt text. If the action executed by the target object based on the object behavior control signal does not meet the preset target, that is, the object behavior control signal cannot control the target object to execute the correct action to cause avoidance failure, which link has a problem can be determined based on the decision interpretation information, so that the avoidance failure is avoided. Therefore, the content understanding model is adjusted more accurately and conveniently based on the decision interpretation information, and the performance of the content understanding model is improved.
Owner:LINKTECH NAVI TECH