Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

115 results about "Image vector" patented technology

Building plane element identification method based on graph neural network and convolutional neural network

The invention discloses a building plane element recognition method based on a graph neural network and a convolutional neural network, and belongs to the technical field of building plane element recognition. The identification method comprises the following steps: inputting a building plane pixel image, and carrying out preprocessing through image layer screening and information extraction to obtain a binary pixel image; vectorizing the binarized pixel image to obtain a vectorized image; performing image segmentation on the vectorized image; using the segmented vectorized image to construct a regional adjacency graph; performing regional adjacency graph optimization on the regional adjacency graph through graph analysis and regional merging by using a graph neural network, and identifying a room edge; calculating the area of the room; using the convolutional neural network image recognition model to recognize articles in the room; and performing room function prediction according to an identification result. According to the method, the problem of shape blurring caused by convolution in a traditional method is avoided. The graph neural network model training effect is improved, and the classification accuracy is improved. The method is suitable for various planar graph styles and is high in universality.
Owner:BEIJING UNIV OF CIVIL ENG & ARCHITECTURE

Identification method and device based on semantic space consistency, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses an identification method, device and equipment based on semantic space consistency, and a medium, and the method comprises the steps: obtaining and preprocessing a target image, a positioning text region and an object region, and generating region positioning information; identifying and correcting the text region to generate a target text vector, extracting an object region image feature and normalizing to generate a target image vector, and mapping the two to the same semantic space for consistency analysis to generate a fusion confidence coefficient and a preliminary identification result; and retrieving, comparing and outputting final identification information and the fusion confidence coefficient in a preset knowledge base based on the fusion confidence coefficient and the preliminary identification result. According to the method, complementary fusion of the text and the image is realized through semantic space mapping, and standardized verification is completed by combining with knowledge base comparison, so that the problem that a single recognition mode is easy to make mistakes is effectively solved, and the accuracy and reliability of a result are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Image vectorization method and device based on neural diffusion curve

PendingCN121708128A2D-image generationBiological modelsGraphicsColor structure
The invention discloses an image vectorization method and device based on a neural diffusion curve. According to the method, geometric control point parameters of multiple sections of Bezier curves are predicted from end to end in an input grating image through a deep neural network, and the main contour and color structure of the image are accurately represented in a Bezier curve form; color control points on the two sides of the curve are automatically extracted through a micro color sampling module, and physically consistent diffusion curve representation is constructed; and performing neural approximation on the resolving process of the diffusion equation by adopting a rendering network based on a Fourier neural operator to realize end-to-end micro-reconstruction of the image generated by the prediction curve, thereby reversely optimizing the geometric prediction of the curve by using a reconstruction error. According to the method, manual intervention is not needed, image vector representation which is high in fidelity, infinite in scaling, lossless in editing and capable of supporting natural layering can be automatically generated while physical rendering consistency is kept, precision and efficiency are improved, and the method is suitable for the fields of image vectorization, digital art creation, micro rendering, graphic design and the like.
Owner:ZHEJIANG SCI-TECH UNIV +1

Propagation trend deduction method and system based on multi-agent cooperation

The invention relates to the technical field of artificial intelligence, and discloses a propagation trend deduction method and system based on multi-agent cooperation. The method comprises the following steps: constructing a social network environment with a plurality of agents, and defining portrait vectors for the agents; generating a social adjacency matrix; the internal cognitive module of each agent is responsible for processing the internal state of the agent, and the external behavior module defines parameterized actions of interaction between the agent and the social network environment and between the agent and other agents; taking an embedded vector obtained based on the internal state and the external information of the agent as the input of a decision model, and fusing a social context bias matrix into an attention mechanism of a decision encoder of the decision model to reflect a social local atmosphere; and generating an internal vertical field based on an embedded vector, and outputting an external action in combination with a social environment to drive propagation simulation and trend deduction. The method has remarkable advantages in the aspects of processing multi-modal information, depicting individual heterogeneity and realizing large-scale prospective prediction.
Owner:UNIV OF SCI & TECH OF CHINA +1

Modulating fidelity and detail in image vectorization

A method, apparatus, non-transitory computer readable medium, and system for modulating the level of fidelity to an input image include obtaining an input image and a fidelity parameter. The input image depicts an entity, and the fidelity parameter indicates a level of fidelity, i.e., faithfulness, to the input image. Embodiments then add noise to the input image based on the fidelity parameter to obtain an intermediate noise image. Subsequently, embodiments generate a synthetic image based on the intermediate noise image using an image generation model. The synthetic image includes a vectorizable depiction of the entity and has the level of fidelity to the input image indicated by the fidelity parameter. The vectorizable depiction is more suitable for conversion to vector format, as the resulting vector image will have a reduced number of paths and shapes.
Owner:ADOBE INC

Unmanned aerial vehicle inspection image vector retrieval method based on space-time fragmentation

The invention relates to an unmanned aerial vehicle inspection image vector retrieval method based on space-time fragmentation, and the method comprises the steps: obtaining an unmanned aerial vehicle inspection image and image metadata, and carrying out the space-time fragmentation division of the unmanned aerial vehicle inspection image according to the image metadata, and obtaining the space-time fragmentation of the unmanned aerial vehicle inspection image; performing feature vector extraction on the unmanned aerial vehicle inspection image to obtain an image feature vector; mapping the space-time fragment of the unmanned aerial vehicle inspection image to a physical fragment of a vector database, and storing the image feature vector and the image metadata to the vector database according to a mapping relationship; and according to a retrieval condition, obtaining the physical fragment of the vector database, and carrying out collaborative retrieval in the target physical fragment according to the mixed index of the content and the spatio-temporal information to obtain a target image set. According to the method, the retrieval efficiency and accuracy of large-scale unmanned aerial vehicle inspection images can be remarkably improved.
Owner:KUNMING UNIV OF SCI & TECH

Multi-modal data processing and retrieval method

The invention particularly relates to a multi-modal data processing and retrieving method. The multi-modal data processing and retrieval method comprises the following steps: respectively carrying out depth feature extraction on image data and text data to generate an image vector and a text vector; carrying out interaction on the image vector and the text vector, and capturing semantic association information among multiple modes; mapping the vectors after interaction to a unified semantic space to realize semantic alignment among different modes; dividing the unified semantic space vector into fragments, storing the fragments in distributed nodes, and establishing a vector index; and generating a query vector by using an image or text queried by a user, carrying out parallel calculation on the similarity with a storage vector, and returning a retrieval result according to the similarity. According to the multi-modal data processing and retrieval method, the problems that the semantic difference between different modal data such as images and texts is large, the retrieval efficiency is low and storage is difficult to expand are solved, the accuracy and efficiency of multi-modal data retrieval are remarkably improved, and the method has remarkable technical advantages and wide application scenes.
Owner:INSPUR QILU SOFTWARE IND

Three-dimensional geospatial model

A system uses models to relocalize a mobile device. The system accesses an input image of a scene in a real-world environment, where the image was captured by a mobile device. The system applies a two-dimensional (2D) foundation model to the input image. The 2D foundation model is trained to determine an image vector representing characteristics of the input image. The system accesses a map representation of the real-world environment, where the map representation includes visual data that describes the real-world environment. The system applies a three-dimensional (3D) geospatial model to the map representation and the image vector. The 3D geospatial model is configured to output 3D splats representing the real-world environment. The system determines a pose of a camera that captured the input image using the 3D splats.
Owner:NIANTIC SPATIAL INC

Text generation method and system based on AI

The invention provides an AI-based text generation method and system, and the method comprises the steps: extracting a global semantic text vector sequence from text data, and extracting a spatial correlation image vector sequence from image data; obtaining a text and image attention weight matrix and an image and text attention weight matrix based on the two; further acquiring a global text and graph vector sequence and a spatial text and graph vector sequence, and acquiring multi-modal fusion features based on the global text and graph vector sequence and the spatial text and graph vector sequence; and taking the multi-modal fusion features as condition information of a language model, outputting a plurality of query texts through the language model, and outputting confidence coefficients corresponding to the query texts. Through fine-grained alignment, the relevance between image-text features is ensured, the problem of modal difference is solved, and the semantic relevance between different modals is improved; the dependency degree on texts or images in the image-text fusion process is dynamically calculated based on the gating weight, and deep semantic association between different modes is further mined.
Owner:ENERGY RES INST OF JIANGXI ACAD OF SCI

A spectral reconstruction method based on double-branch heterogeneous feature fusion

This invention provides a spectral reconstruction method based on dual-branch heterogeneous feature fusion, relating to the fields of computational imaging and artificial intelligence. The method includes acquiring two-dimensional compressed measurement image vectors, constructing a local-global feature fusion deep neural network model, which includes a local feature extraction module, a global context extraction module, a feature fusion module, and a decoder; designing a composite loss function, including reconstruction loss, physical prior loss, and feature regularization loss; performing end-to-end training on the local-global feature fusion deep neural network model; and deploying the trained model to achieve rapid spectral reconstruction. This invention combines the fast inference capability of deep neural networks with a novel parallel feature extraction and fusion architecture. Through a carefully designed network model and a composite loss function, it effectively integrates and constrains local and global information in the feature space, achieving rapid and high-precision reconstruction of hyperspectral images.
Owner:BAY AREA LOW ALTITUDE RESEARCH INSTITUTE (GUANGDONG) CO LTD

Material image recommendation method and device, storage medium and program product

The embodiment of the invention provides a material image recommendation method and device, a storage medium and a program product, and is applied to the technical field of artificial intelligence. The method comprises the following steps: acquiring a plurality of first material images; for each first material image, determining a text image weight matrix according to an image vector corresponding to the first material image and a character vector corresponding to the description character included in the first material image; multiplying the image vector, the character vector and the text image weight matrix to generate a target fusion vector; determining the score of each to-be-recommended material image according to the target fusion vector of each first material image and the exposure of each to-be-recommended material image; and pushing the to-be-recommended material image with the score meeting a preset condition to the first user. The method is used for achieving the effect of improving the personalized recommendation accuracy.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD +1

Question and answer processing method and electronic equipment

The invention provides a question and answer processing method and electronic equipment. The method comprises the following steps: receiving a question text and a question image input by a user; intention recognition is conducted on the question text, a task description text is generated, and the task description text comprises slot position information; vectorizing the task description text and the problem image to obtain a text vector and an image vector; based on a multi-modal database and a multi-modal index, recalling target text knowledge and associated image knowledge corresponding to the text vector, recalling target image knowledge and associated text knowledge corresponding to the image vector, and recalling key field information corresponding to slot information; optimizing the target text knowledge, the associated image knowledge, the target image knowledge, the associated text knowledge and the field information, and then performing aggregation to obtain an aggregation result; and based on the aggregation result, generating a question-answer result by using a large language model. According to the method, the accuracy and credibility of the question and answer result can be improved.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

ADVANCED IMAGE TRACKING FOR VECTORIZATION OF RASTER IMAGES

The present disclosure relates to systems, non-transitory computer-readable media, and methods for detecting and tracking edges in raster images using an advanced edge detection algorithm. For example, the disclosed systems generate a local histogram ranking of pixels within a sliding pixel window in a raster image, based on pixel values ​​located within the sliding pixel window. In some embodiments, the disclosed systems determine an edge thickness for an edge depicted in a region of the raster image enclosed by the sliding pixel window by comparing pixel ranks specified by the local histogram ranking of the pixels.In certain embodiments, the disclosed systems also provide the edge for display based on determining the edge thickness for the area of ​​the raster image enclosed by the sliding pixel window.
Owner:ADOBE INC

Scalable search and ranking of biometric signatures

Systems and methods are directed to detecting fraudulent activity based on biometric signatures. A capture component captures data associated with user interactions with a user interface displayed by a computing device. The captured data includes mouse movement. The captured data is converted into an image of the mouse movement. An image encoder encodes the image into an image vector. Subsequently, an analysis system searches a vector database to identify one or more historical image vectors that are similar to the image vector. The vector database comprises a plurality of image vectors corresponding to previously encoded images, whereby at least some of the plurality of image vectors are flagged as indicating fraudulent activity. A result based on the identified historical image vectors is outputted. The result can comprise a risk score for each of the identified historical image vectors that is based in part on a similarity difference measurement.
Owner:WELLS FARGO BANK NA

Verifiable cross-modal retrieval method

The invention provides a verifiable cross-modal retrieval method, and belongs to the technical field of Internet of Things security. The method comprises the steps that a data owner constructs a data set, a 1-norm image vector, a 1-norm text vector, an index table, a prime number table, a ciphertext vector and a Merkel prime number tree, sends the ciphertext vector, the index table and the Merkel prime number tree to a cloud server, and sends a root abstract of the Merkel prime number tree, the prime number table, the 1-norm image vector and the 1-norm text vector to a query client; the query client sends the cross-modal query request to the cloud server; the cloud server receives the cross-modal query request and sends the inner product vector, the group number interval and the verification object set to the query client; and the query client performs integrity verification, correctness verification and Hamming distance calculation according to the inner product vector, the group number interval and the verification object set to find the most similar data. According to the method, the prime number is fused into the traditional Merkel tree to construct the Merkel prime number tree to ensure the integrity and correctness of a cross-modal retrieval result.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Station lost article retrieval method based on multiple modes

The invention discloses a station lost article retrieval method based on multiple modes, which comprises the following steps of: S1, acquiring an image of a lost article, and encoding the image into an image vector by using an image encoder; s2, obtaining text description of the lost goods, and extracting text information of the lost goods through a large language model; s3, inputting the lost article text information into a text encoder of the CLIP model, and generating an initial text vector; s4, inputting the initial text vector into a lost article encoder, and outputting an enhanced text vector; s5, performing residual connection on the initial text vector and the enhanced text vector, inputting the initial text vector and the enhanced text vector to a lost article adapter, and outputting a final text vector; s6, calculating the cosine similarity between the final text vector and the corresponding image vector in the vector database; the key challenges such as non-uniform text description formats of the lost goods and diversity of image features are effectively solved, the retrieval efficiency of the lost goods is greatly improved, and a large amount of manpower and material resource cost is reduced.
Owner:INFORMATION TECH INST OF CHINA RAILWAY ZHENGZHOU BUREAU GRP CO LTD

Photovoltaic power station construction safety intelligent monitoring method and system based on AI video identification

The invention relates to the technical field of intelligent monitoring, in particular to a photovoltaic power station construction safety intelligent monitoring method and system based on AI video recognition, and the method comprises the steps: employing an input image frame sequence, carrying out the feature extraction of the image frame sequence through a convolutional network, generating an image vector sequence, inputting the image vector sequence into a bidirectional gating circulation unit, and carrying out the recognition of the image vector sequence; outputting a stage representation vector, and outputting a construction stage number corresponding to the stage representation vector through a stage classifier; on the basis of the construction stage number and the segment representation vector, identifying the specific behavior type of each constructor in the current operation window and the station number of the constructor; constructing a job graph structure, and identifying risk scores of nodes in the job graph structure through a node risk scoring function; and on the basis of the constructed operation graph structure, final judgment on the violation behavior is completed.
Owner:GUANGDONG WANYE CONSTR ENG CO LTD

An unsupervised pre-training method for a neural network for image processing

The present application relates to neural network unsupervised learning technical field, specifically to a kind of neural network unsupervised pre-training method for image processing, comprising the following steps: first, image is divided into image block, then mask operation is carried out, then perception loss is calculated, contrast loss and reconstruction loss are calculated, finally, training is carried out using loss.After training, the input image is processed using the trained model, and the category feature vector and the reconstructed image vector are obtained.The present application can measure the influence of mask operation on neural network by using perception loss, and make the feature more obvious by using contrast loss, and finally learn how to abstract the image into feature by reconstruction loss, while reducing the loss of information in the abstraction process, improve the feature extraction capability of neural network for image.
Owner:GUILIN UNIV OF ELECTRONIC TECH +1

Model training method, image retrieval method, device and electronic equipment

Embodiments of the present application disclose a model training method, an image retrieval method, a device, an electronic device and a storage medium. The method comprises: obtaining a training data set, the training data set comprising a plurality of training data, each training data comprising text data and image data; inputting the training data set into a to-be-trained model, obtaining a text vector and an image vector corresponding to each training data output by the to-be-trained model; inputting the training data set into a reference model, obtaining a reference text vector and a reference image vector corresponding to each training data output by the reference model; adjusting parameters of the to-be-trained model based on the text vector, the image vector, the reference text vector and the reference image vector until a training end condition is met, and obtaining a text-image retrieval model. Through the above method, the prediction accuracy of a text-image retrieval model with a small scale is improved.
Owner:GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

Synchronization of lip movement images to audio voice signal

Systems and methods for synchronization of lip movement images to an audio voice signal are provided. A method includes acquiring a source video; dividing the source video into a set of image frames and a set of audio frames; generating, by the computing device, a vector database based on the set of image frames and the set of audio frames, wherein a vector of the vector database includes a face vector and an audio vector; receiving a target image frame and a target audio frame; determining a target image vector based on the target image frame and a target audio vector based on the target audio frame; searching the vector database to select a pre-determined number of vectors corresponding to the target image vector and the target audio frame; and generating, based on the pre-determined number of vectors, an output image frame of an output video.
Owner:PHEON INC

A sam-based intelligent image search method

The application provides a kind of intelligent image search method based on SAM, it is related to image processing technical field.The method comprises: inputting the image to be searched into SAM segmentation model, obtains the segmentation mask of all objects in the image to be searched;According to the object, select the object to be searched, obtain the segmentation mask corresponding to the search object image;According to the image to be searched and the segmentation mask corresponding to the search object image, determine the target object image;Obtain the feature vector of the target object image;Create image vector warehouse, obtain the feature vector of each image of the image vector warehouse;Determine the cosine distance between each image of the image vector warehouse and the target object image;According to the cosine distance, determine the most similar image of the target object image in the image vector warehouse.According to the application, the detail features of the search object can be highlighted, and part of the features of the surrounding object can be retained, and the search accuracy can be improved.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Three-Dimensional Geospatial Model

A system uses models to relocalize a mobile device. The system accesses an input image of a scene in a real-world environment, where the image was captured by a mobile device. The system applies a two-dimensional (2D) foundation model to the input image. The 2D foundation model is trained to determine an image vector representing characteristics of the input image. The system accesses a map representation of the real-world environment, where the map representation includes visual data that describes the real-world environment. The system applies a three-dimensional (3D) geospatial model to the map representation and the image vector. The 3D geospatial model is configured to output 3D splats representing the real-world environment. The system determines a pose of a camera that captured the input image using the 3D splats.
Owner:NIANTIC SPATIAL INC

An infrared weak and small target detection method of line-by-line detection

The application discloses an infrared weak small target detection method based on line-by-line detection, and solves the problems of high detection delay and large resource consumption caused by global image caching in the prior art. The method skips the traditional "read-out-caching-global detection" process and performs real-time detection while reading out the infrared image data line by line. The method comprises the following steps: performing first-order and second-order differential calculation on a single-line image vector, extracting and fusing the intra-line spatial features; inputting the continuous multi-line features into an inter-line fusion module based on a self-attention mechanism in parallel, and restoring the high-dimensional features of the target; then, completing target detection through a U-Net network, and adapting the dimension through a line expansion and line compression module; and in the training, adaptively selecting a loss function according to whether the current line block contains a target, so as to improve the training efficiency. The application realizes parallel processing of detection and data reading, significantly improves the timeliness of detection, and reduces the demand for cache and computing resources of edge devices.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Systems and methods for water damage claims triage portal with computer vision

PendingUS20260141322A1FinanceUser deviceUser identifier
A method including receiving text image data corresponding to a claim instance from a user device. Inputting the text and image data into one or more trained machine-learning models to determine a text vector and an image vector corresponding to the text and image data. Inputting the text the image vectors into the one or more machine-learning models to determine one or more claim ratings corresponding to the claim instance. Receiving claims representative user data corresponding to one or more claims representative users from one or more data stores, wherein the claims representative user data includes a claims representative user identifier, a claims representative user availability, and a claims representative user skill level. Analyzing the claims representative user data and the one or more claim ratings to determine the best fit claims representative user. Transmitting the claim instance to a user device corresponding to the best fit claims representative user.
Owner:STATE FARM MUTAL AUTOMOBILE INSURANCE COMPANY

Method, device and equipment for auditing multi-document association and image-text calculation fusion

The invention discloses an auditing method, device and equipment for multi-document association and image-text calculation fusion, and relates to the technical field of multi-document auditing. The method comprises the following steps: carrying out structured analysis on a candidate document to obtain a document image-text vector set; determining a dynamic auditing rule set according to the types of the candidate documents and the professional field knowledge base; according to the dynamic auditing rule set and the document image-text vector set, candidate text segments and associated content key points are determined, compliance judgment and consistency verification are carried out, and to-be-audited text segments and to-be-audited associated key points are obtained; calling a preset agent to obtain an algorithm verification result; and according to the to-be-audited text segments, the to-be-audited associated key points and the algorithm verification result, constructing an auditing context so as to determine a document image-text calculation comprehensive auditing result. According to the scheme, multi-document structured analysis, single-document content key point and multi-document associated content key point extraction analysis and intelligent agent calculation verification are carried out, so that the multi-associated document auditing efficiency and depth are improved.
Owner:BEIJING YUNKE CENTURY TECHNOLOGY CO LTD

An image retrieval method, system, device and medium for enhancing fine-grained object retrieval performance

The present disclosure relates to the technical field of image retrieval, and provides an image retrieval method, system, device and medium for enhancing fine-grained object retrieval performance. The method processes original images through a multi-scale image interception technology to generate target images and a slice set. These images are input into an encoder to construct an image vector library. Meanwhile, a multi-modal large model performs semantic analysis on the images and slices to generate text descriptions and construct a text vector library. In the retrieval stage, visual similarity matching is performed based on the image vectors, and text similarity matching is performed based on the text vector library. The results are integrated to obtain a candidate set. Finally, through weighted fusion and sorting, the final retrieval result is obtained. The present disclosure significantly improves the accuracy of fine-grained object retrieval. By combining visual and text information, not only is the comprehensiveness of the retrieval enhanced, but also the relevance of the results is improved, enabling users to more quickly and accurately find the required information.
Owner:CHINA TOWER CO LTD

Multi-modal entity linking method fusing vision and text feature enhancement

The method comprises the following steps: taking an image and a text as input, combining a low-level high-resolution image and a high-level strong-semantic image to enhance image features by using a feature pyramid bidirectional fusion strategy, extracting phrase-level features of the text through a convolutional neural network, and extracting a multi-modal entity link of the visual and text feature enhancement; the method comprises the following steps: firstly, extracting a visual vector, fusing the visual vector with text character-level features obtained by an encoder to form text input, filtering noise embedded in the vector by utilizing a bottleneck fusion network, improving the fusion efficiency of a model, and fusing the text vector and the visual vector to form a multi-modal fusion vector; and finally, the entity mentions are linked to the candidate entity with the highest score in the knowledge base based on the matching scores. According to the method, noise of text vectors and image vectors of an existing model can be effectively reduced, key information such as text features and visual features is reserved, feature dimensions are aligned, a semantic gap between the text vectors and the image vectors is made up, subsequent multi-modal fusion is facilitated, and the overall performance of entity linking is improved.
Owner:LIAONING UNIVERSITY

A single-stage gating multi-modal fusion method with two-stage cross-attention

ActiveCN120654812BAlgorithmComputer vision
The present application relates to the technical field of artificial intelligence, in particular to a two-stage cross-attention single-stage gating multimodal fusion method, comprising: inputting text through an embedding layer and a text encoder to obtain a text vector; inputting an image through an image encoder to obtain an image vector; inputting the text vector and the image vector into a modal feature fusion module; the modal feature fusion module adopts a two-stage cross-attention single-stage gating modal fusion mechanism to output a fusion vector; and inputting the fusion vector through a decoder to obtain a predicted text. The present application improves the interaction effect between modes, reduces the hallucinations generated by the model, effectively reduces the calculation parameters of the model, and has good generality and practicability.
Owner:JIANGNAN UNIV

Navigation execution control method and device, computer equipment and storage medium

The invention relates to a navigation execution control method and device, computer equipment and a storage medium, and the method comprises the following steps: carrying out instruction analysis and text description on user language instruction information to obtain a target description text set; inputting the target description text set into a preset text encoder to generate a text vector set; encoding the view image set based on a preset image encoder to obtain an image vector set; calculating a similarity value of the text vector set and the image vector set; judging whether the similarity value is smaller than a preset similarity threshold value or not; if yes, stopping the current path exploration; and performing navigation path optimization according to the map layout information, a preset scene knowledge base and the target description text set to obtain an optimized navigation path to control the to-be-navigated target to move. The method can be applied to application scenes of a financial service system and a digital medical system, the to-be-navigated robot can be effectively navigated, and the flexibility and efficiency of navigation task execution are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD