Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

193 results about "Image vector" patented technology

Image-text question and answer method, system and device based on multi-mode RAG and storage medium

The invention belongs to the technical field of artificial intelligence, and relates to an image-text question answering method, system and device based on multi-modal RAG and a storage medium, and the method comprises the following steps: 1) extracting multi-modal information from a PDF document, representing the multi-modal information as dense vectors, and storing the dense vectors in a text vector database and an image vector database; 2) obtaining a semantic embedding vector of the problem text and a multi-modal embedding vector of the problem text; obtaining a semantic embedding vector of the description text of the problem image and a multi-mode embedding vector of the problem image; 3) performing coarse screening in the text vector database and the image vector database by using the semantic embedding vector and the multi-modal embedding vector, finding out coarse screening text data and coarse screening image data, performing multi-modal fine ranking on the coarse screening text data and the coarse screening image data, and obtaining text data and image data which are retrieved and recalled; and 4) generating a final answer by the multi-modal large language model. According to the method, multi-modal data in a long document can be effectively analyzed, and information most relevant to a problem can be accurately retrieved.
Owner:BEIJING ZHIPU PILOT TECHNOLOGY CO LTD

Method, device, and product for retrieval

The present disclosure provides a method, a device, and a product for retrieval. The method includes acquiring context information related to an image and determining a representation of the image based on image data and the context information of the image, where the context information includes at least one of environment parameters, user behavior data, time elements, or field metadata. The method further includes encoding the representation as an image vector in a high-dimensional vector space and storing it into an image vector database. When retrieval is performed, a query that includes text information and that is for the image vector database is received, and an image associated with the text information is determined from the image vector database. The method according to the present disclosure can improve accuracy and efficiency for image retrieval.
Owner:DELL PROD LP

Medical image report generation method based on multi-modal large model retrieval enhancement

The invention discloses a medical image report generation method based on multi-modal large model retrieval enhancement, which relates to the technical field of medical image processing and comprises an image vector model module, a vector database and similarity detection module, a multi-modal large model module, a dynamic retrieval module and a retrieval enhancement generation prompt project module. The multi-modal large model module comprises a multi-modal input splicing unit, a reasoning information processing unit and a diagnosis report set generation unit, the dynamic retrieval module comprises a similarity threshold value analysis unit and a retrieval quantity analysis unit, and the similarity threshold value analysis unit is used for analyzing a maximum similarity value between a new image and a database. According to the invention, an image vectorization processing technology based on a CLIP-ViT model is adopted, so that the system can accurately capture fine features of a focus, and the risk of key information omission is reduced; a dynamic threshold adjustment algorithm is introduced, and the rigidity defect of a traditional fixed retrieval strategy is overcome.
Owner:砺进(杭州)科技有限公司

Utilizing implicit neural representations to parse visual components of subjects depicted within visual content

This disclosure describes one or more implementations of systems, non-transitory computer-readable media, and methods that utilize a local implicit image function neural network to perform image segmentation with a continuous class label probability distribution. For example, the disclosed systems utilize a local-implicit-image-function (LIIF) network to learn a mapping from an image to its semantic label space. In some instances, the disclosed systems utilize an image encoder to generate an image vector representation from an image. Subsequently, in one or more implementations, the disclosed systems utilize the image vector representation with a LIIF network decoder that generates a continuous probability distribution in a label space for the image to create a semantic segmentation mask for the image. Moreover, in some embodiments, the disclosed systems utilize the LIIF-based segmentation network to generate segmentation masks at different resolutions without changes in an input resolution of the segmentation network.
Owner:ADOBE INC

Building plane element identification method based on graph neural network and convolutional neural network

The invention discloses a building plane element recognition method based on a graph neural network and a convolutional neural network, and belongs to the technical field of building plane element recognition. The identification method comprises the following steps: inputting a building plane pixel image, and carrying out preprocessing through image layer screening and information extraction to obtain a binary pixel image; vectorizing the binarized pixel image to obtain a vectorized image; performing image segmentation on the vectorized image; using the segmented vectorized image to construct a regional adjacency graph; performing regional adjacency graph optimization on the regional adjacency graph through graph analysis and regional merging by using a graph neural network, and identifying a room edge; calculating the area of the room; using the convolutional neural network image recognition model to recognize articles in the room; and performing room function prediction according to an identification result. According to the method, the problem of shape blurring caused by convolution in a traditional method is avoided. The graph neural network model training effect is improved, and the classification accuracy is improved. The method is suitable for various planar graph styles and is high in universality.
Owner:BEIJING UNIV OF CIVIL ENG & ARCHITECTURE

Identification method and device based on semantic space consistency, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses an identification method, device and equipment based on semantic space consistency, and a medium, and the method comprises the steps: obtaining and preprocessing a target image, a positioning text region and an object region, and generating region positioning information; identifying and correcting the text region to generate a target text vector, extracting an object region image feature and normalizing to generate a target image vector, and mapping the two to the same semantic space for consistency analysis to generate a fusion confidence coefficient and a preliminary identification result; and retrieving, comparing and outputting final identification information and the fusion confidence coefficient in a preset knowledge base based on the fusion confidence coefficient and the preliminary identification result. According to the method, complementary fusion of the text and the image is realized through semantic space mapping, and standardized verification is completed by combining with knowledge base comparison, so that the problem that a single recognition mode is easy to make mistakes is effectively solved, and the accuracy and reliability of a result are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Image vectorization method and system based on deep learning and edge detection

The invention relates to the technical field of images, in particular to an image vectorization method and system based on deep learning and edge detection, and the method comprises the steps: obtaining image data of a household appliance module, and carrying out the preprocessing of the obtained image data; taking the preprocessed image data as input to construct an edge detection model, performing optimization processing on an edge probability graph output by the edge detection model, performing edge detection on the preprocessed image data in sequence to generate multiple groups of edge graphs, and performing average processing on the multiple groups of edge graphs to obtain a comprehensive edge image; the optimized marginal probability graph and the comprehensive marginal image serve as input, marginal information is converted into a vector graph, a DXF file is obtained, the optimized marginal probability graph and the comprehensive marginal image serve as input, vectorization conversion is conducted through the Potrace algorithm, optimization processing on the marginal detection result in the early stage is combined, and the edge detection efficiency is improved. Vector graphs with higher precision and less noise can be generated.
Owner:OCEAN UNIV OF CHINA

Multi-source data health intervention method

The invention relates to the technical field of medical information, in particular to a multi-source data health intervention method, which comprises the following steps: firstly, acquiring texts, vital signs, image vectors and wearable signals, generating a unified vector through multi-modal coding, and extracting a topological life vector; constructing a spiking neuro-causal diagram based on a unified vector, and injecting a life vector as an external field into a Sheng differential equation to generate a dynamic pathological manifold; compressing the manifold into a tensor network state, constructing a quantum optimization model in combination with a causal diagram, minimizing energy expectation and topological risk to obtain an optimal intervention sequence, and reinjecting a target gradient to adjust the manifold in real time; after the sequence is subjected to clinical logic verification, a commitment value and a zero-knowledge proof are generated and written into a Byzantine account book, and meanwhile a control instruction is issued to an execution terminal. According to the method, millisecond-level decision making is achieved, the false alarm rate is reduced, and chain traceability is provided.
Owner:GENERAL GLOBAL JADE BIRD HEALTH TECHNOLOGY CO LTD

Chinese character pattern generation method and device, equipment and medium

The invention relates to the technical field of image detection, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a Chinese character font generation method, device, equipment and medium, the method comprises the following steps: obtaining a Chinese character image to be processed, and carrying out image coding on the Chinese character image to be processed to obtain a Chinese character image sequence; performing style feature extraction and semantic feature extraction on each Chinese character image vector in the Chinese character image sequence to obtain corresponding Chinese character style features and Chinese character semantic features; constructing a corresponding style loss function and a semantic loss function according to the Chinese character style features and the Chinese character semantic features; performing loss weighting calculation on the style loss function and the semantic loss function to obtain a target loss function; and performing font style generation on the to-be-processed Chinese character image according to the target loss function to obtain a target Chinese character font of the to-be-processed Chinese character image. The Chinese character font generation efficiency and font generation diversity can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Training and using a vector encoder to determine vectors for sub-images of text in an image subject to optical character recognition

Provided are a computer program product, system, and method for training and using a vector encoder to determine vectors for sub-images of text in an image to subject to optical character recognition. A vector encoder is trained to encode images representing text into vectors in a vector space. Vectors of images representing similar text have a high degree of cohesion in the vector space. Vectors of images representing dissimilar text have a low degree of cohesion in the vector space. An input image is processed to determine sub-images of the input image that bound text represented in the input image. The sub-images are inputted to the vector encoder to output sub-image vectors. The vector encoder generates a search vector for search text. Optical character recognition is applied to at least one region of the input image including the sub-images having sub-image vectors matching the search vector.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Image vectorization method and device based on neural diffusion curve

PendingCN121708128A2D-image generationBiological modelsGraphicsColor structure
The invention discloses an image vectorization method and device based on a neural diffusion curve. According to the method, geometric control point parameters of multiple sections of Bezier curves are predicted from end to end in an input grating image through a deep neural network, and the main contour and color structure of the image are accurately represented in a Bezier curve form; color control points on the two sides of the curve are automatically extracted through a micro color sampling module, and physically consistent diffusion curve representation is constructed; and performing neural approximation on the resolving process of the diffusion equation by adopting a rendering network based on a Fourier neural operator to realize end-to-end micro-reconstruction of the image generated by the prediction curve, thereby reversely optimizing the geometric prediction of the curve by using a reconstruction error. According to the method, manual intervention is not needed, image vector representation which is high in fidelity, infinite in scaling, lossless in editing and capable of supporting natural layering can be automatically generated while physical rendering consistency is kept, precision and efficiency are improved, and the method is suitable for the fields of image vectorization, digital art creation, micro rendering, graphic design and the like.
Owner:ZHEJIANG SCI-TECH UNIV +1

Learning video recommendation methods, media, and devices based on modal semantic space alignment

The present invention proposes a learning video recommendation method, medium, and device based on modal semantic space alignment, which relates to the field of multimodal semantic space alignment technology. The method includes: extracting user behavior vectors and vectors of different modalities of learning videos, including text, image, audio, and structure vectors; using a multi-layer perceptron to project the extracted vectors of different modalities into a common semantic space, and within the common semantic space, vector splicing the structure vector with the text vector, image vector, and audio vector respectively; performing modal comparison and modal matching on the spliced ​​vectors, and obtaining a multimodal feature vector of the learning video by optimizing the semantic alignment between different modalities; fusing the multimodal feature vectors, calculating the similarity between the user behavior vector and the fused multimodal feature vector using cosine similarity, and ranking and recommending all videos. The present invention can fully understand the modal semantic structure of learning videos, achieving accurate matching and recommendation.
Owner:HUBEI UNIV

Cross-modal image classification method and device, equipment, medium and program product

The invention discloses a cross-modal image classification method, device and equipment, a medium and a program product, which can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like. The method comprises the following steps: respectively carrying out dimension transformation on a trained image prompt vector corresponding to an image coding branch and a trained text prompt vector corresponding to a text coding branch in a pre-training model; fusing the trained image prompt vector with the transformed text prompt vector, cascading with the image vector of the to-be-classified image, and inputting to an image coding branch; fusing the trained text prompt vector with the transformed image prompt vector, cascading with the text vectors of the plurality of category description texts, and inputting to a text coding branch; and determining target probability distribution of the to-be-classified image on the plurality of category description texts according to the output of the two branches so as to classify the to-be-classified image. According to the scheme, the robustness and generalization of the model on a cross-modal image classification task can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Intelligent inspection method and device based on image and point cloud rapid matching

The invention discloses a municipal facility intelligent inspection method and equipment based on image and point cloud matching. The method comprises the following steps: constructing a municipal facility database containing a multi-view image template and a point cloud grid model; a scene image and a point cloud are synchronously acquired through an unmanned vehicle carrying an optical camera, a laser radar and a GPS; respectively carrying out image target detection and point cloud target detection; a cross-modal target matching stage: calculating a 2D / 3D nearest neighbor vector of an image mask centroid and a point cloud cluster centroid, performing direction included angle matching on the point cloud vector and an image vector after a y coordinate of the point cloud vector is returned to zero, and verifying matching correctness through vector modulus length ratio mode distribution; and when a defect is detected in any mode, a maintenance work order containing a defect position and a GPS coordinate is automatically generated. According to the invention, bimodal independent detection and result fusion are realized, the sensor calibration requirement is avoided, and the inspection efficiency and accuracy are improved.
Owner:CHANGCHUN UNIV OF SCI & TECH +1

Multiangular dynamic imaging system and method for measuring and evaluating ovds

A Multi-Angle Dynamic Image system provides multi-angle observation capabilities comprising lamps strategically positioned to obtain high quality images of the security documents or items under analysis, to obtain images with variations in the observation angle of the security document or item under study. The system includes a box platform for placing the security document or item thereupon, the box platform changes its orientation, allowing the system to capture the variations in color and structure as the orientation and the angles of the observation and illumination change. The system also includes a high-resolution linear camera to capture the intricate details while avoiding distortion and parallax issues. The system also includes a computer configured to run algorithms to process image acquisition data for a complete set of observation angles, to pre-process image acquisition data including crop and image standardization, to calculate differences between adjacent angle images, eliminate noise, to calculate a total optical flow using the images at different angles and their differences to generate the Dynamic Images vectorial field, one per orientation, and the Dynamic Map, a heat map generated using the magnitude from the Dynamic Image vectorial field, and calculate from the Dynamic Image the different parameters associated with dynamic properties of the OVD.
Owner:GONZALEZ-CANDELA ERNESTO

Multi-modal fine-grained representation enhancement method and system based on LLAVA model

The invention discloses a multi-modal fine-grained representation enhancement method and system based on an LLAVA model, and relates to the field of artificial intelligence. Inputting the original image and the prompt text into an LLAVA model to generate a detailed text, and extracting entity nouns from the detailed text to generate an enhanced text; inputting the detailed text and the prompt text into an LLAVA model to generate a negative sample; obtaining a fusion text vector according to the enhanced text and the detailed text; obtaining a negative sample text vector according to the negative sample; encoding the original image to obtain a global image vector; segmenting the original image to obtain a plurality of segmented images, and encoding the segmented images to obtain an image fine-grained vector; fusing the global image vector and the image fine-grained vector to obtain a fused image vector; and an enhanced CLIP model is obtained through comparative learning and metric learning. According to the method, the fine-grained understanding capability of the CLIP model is greatly improved, so that the performance on a computer vision task is improved.
Owner:HAINAN UNIV

Propagation trend deduction method and system based on multi-agent cooperation

The invention relates to the technical field of artificial intelligence, and discloses a propagation trend deduction method and system based on multi-agent cooperation. The method comprises the following steps: constructing a social network environment with a plurality of agents, and defining portrait vectors for the agents; generating a social adjacency matrix; the internal cognitive module of each agent is responsible for processing the internal state of the agent, and the external behavior module defines parameterized actions of interaction between the agent and the social network environment and between the agent and other agents; taking an embedded vector obtained based on the internal state and the external information of the agent as the input of a decision model, and fusing a social context bias matrix into an attention mechanism of a decision encoder of the decision model to reflect a social local atmosphere; and generating an internal vertical field based on an embedded vector, and outputting an external action in combination with a social environment to drive propagation simulation and trend deduction. The method has remarkable advantages in the aspects of processing multi-modal information, depicting individual heterogeneity and realizing large-scale prospective prediction.
Owner:UNIV OF SCI & TECH OF CHINA +1

Package design method and device based on target GNN model, equipment and medium

The invention provides a package design method and device based on a target GNN model, equipment and a medium, is used for the technical field of package design, and can solve the problem that existing package design is poor in efficiency and quality. Comprising the steps of performing feature recognition on a to-be-designed image, and determining a target design vector according to each image vector and a corresponding target text vector; calculating a material similarity value and a process similarity value according to the target GNN model, and determining a target design scheme according to the material similarity value and the process similarity value; according to the user score value and the predicted score value corresponding to each target design scheme, calculating a score loss value, and according to a preset material, a preset process, the plurality of historical design schemes and the target design schemes, calculating a divergence value; performing weighted calculation on the score loss value and the divergence value to obtain a target loss value, and adjusting parameters of the target GNN model according to the target loss value and a preset loss threshold; therefore, the design quality and efficiency are improved.
Owner:CHENGDU AJIAXI INTELLIGENT TECH CO LTD

Image retrieval method, system and equipment for enhancing fine-grained object retrieval performance and medium

The invention relates to the technical field of image retrieval, and provides an image retrieval method, system, equipment and medium for enhancing fine-grained object retrieval performance, the method comprises the following steps: processing an original image through a multi-scale image interception technology, generating a target image and a slice set, inputting the images into an encoder, constructing an image vector library, and meanwhile, inputting the target image and the slice set into a database; the multi-modal large model carries out semantic analysis on images and slices, generates text description and constructs a text vector library, in the retrieval stage, visual similarity matching is carried out based on image vectors, text similarity matching is carried out based on the text vector library, results are integrated to obtain a candidate set, and finally, a final retrieval result is obtained through weighted fusion and sorting. According to the method, the accuracy of fine-grained object retrieval is remarkably improved, the comprehensiveness of retrieval is enhanced by combining visual and text information, and the correlation of results is improved, so that a user can find required information more quickly and more accurately.
Owner:CHINA TOWER CO LTD

Modulating fidelity and detail in image vectorization

A method, apparatus, non-transitory computer readable medium, and system for modulating the level of fidelity to an input image include obtaining an input image and a fidelity parameter. The input image depicts an entity, and the fidelity parameter indicates a level of fidelity, i.e., faithfulness, to the input image. Embodiments then add noise to the input image based on the fidelity parameter to obtain an intermediate noise image. Subsequently, embodiments generate a synthetic image based on the intermediate noise image using an image generation model. The synthetic image includes a vectorizable depiction of the entity and has the level of fidelity to the input image indicated by the fidelity parameter. The vectorizable depiction is more suitable for conversion to vector format, as the resulting vector image will have a reduced number of paths and shapes.
Owner:ADOBE INC

Unmanned aerial vehicle inspection image vector retrieval method based on space-time fragmentation

The invention relates to an unmanned aerial vehicle inspection image vector retrieval method based on space-time fragmentation, and the method comprises the steps: obtaining an unmanned aerial vehicle inspection image and image metadata, and carrying out the space-time fragmentation division of the unmanned aerial vehicle inspection image according to the image metadata, and obtaining the space-time fragmentation of the unmanned aerial vehicle inspection image; performing feature vector extraction on the unmanned aerial vehicle inspection image to obtain an image feature vector; mapping the space-time fragment of the unmanned aerial vehicle inspection image to a physical fragment of a vector database, and storing the image feature vector and the image metadata to the vector database according to a mapping relationship; and according to a retrieval condition, obtaining the physical fragment of the vector database, and carrying out collaborative retrieval in the target physical fragment according to the mixed index of the content and the spatio-temporal information to obtain a target image set. According to the method, the retrieval efficiency and accuracy of large-scale unmanned aerial vehicle inspection images can be remarkably improved.
Owner:KUNMING UNIV OF SCI & TECH

Spectrum reconstruction method based on dual-branch heterogeneous feature fusion

The invention provides a spectrum reconstruction method based on double-branch heterogeneous feature fusion, and relates to the technical field of computational imaging and artificial intelligence, and the method comprises the steps: obtaining a two-dimensional compression measurement image vector, constructing a local-global feature fusion deep neural network model, and constructing a local-global feature fusion deep neural network model; the method comprises a local feature extraction module, a global context extraction module, a feature fusion module and a decoder, a composite loss function is designed, the composite loss function comprises reconstruction loss, physical prior loss and feature regularization loss, end-to-end training is performed on a local-global feature fusion deep neural network model, and the trained model is deployed. And rapid spectrum reconstruction is realized. According to the method, the rapid reasoning capability of the deep neural network is combined with a novel parallel feature extraction and fusion architecture, and local and global information is effectively integrated and constrained in a feature space through a well-designed network model and a composite loss function, so that rapid and high-precision reconstruction of a hyperspectral image is realized.
Owner:BAY AREA LOW ALTITUDE RESEARCH INSTITUTE (GUANGDONG) CO LTD

System and method for increasing a resolution of a three-dimension (3D) image

A system for increasing a resolution of a three-dimension (3D) image determines contours of a 3D image. The system determines a mesh image vector, where the mesh image vector indicates location coordinates of feature points on a surface of an object in the 3D image. The system compares the mesh image vector with each contour. The system determines an intersecting feature point where the mesh image vector meets a contour. The system determines that a baseline dataset includes a first feature point that corresponds to the intersecting feature point. In response, the system generates a structural vector by populating the structural vector with the intersecting feature point. The system determines a color code associated with the intersecting feature point and generates a textural vector by populating the textural vector with the color code. The system generates an image vector by combining the structural vector and the textural vector.
Owner:BANK OF AMERICA CORP

Multi-modal data processing and retrieval method

The invention particularly relates to a multi-modal data processing and retrieving method. The multi-modal data processing and retrieval method comprises the following steps: respectively carrying out depth feature extraction on image data and text data to generate an image vector and a text vector; carrying out interaction on the image vector and the text vector, and capturing semantic association information among multiple modes; mapping the vectors after interaction to a unified semantic space to realize semantic alignment among different modes; dividing the unified semantic space vector into fragments, storing the fragments in distributed nodes, and establishing a vector index; and generating a query vector by using an image or text queried by a user, carrying out parallel calculation on the similarity with a storage vector, and returning a retrieval result according to the similarity. According to the multi-modal data processing and retrieval method, the problems that the semantic difference between different modal data such as images and texts is large, the retrieval efficiency is low and storage is difficult to expand are solved, the accuracy and efficiency of multi-modal data retrieval are remarkably improved, and the method has remarkable technical advantages and wide application scenes.
Owner:INSPUR QILU SOFTWARE IND

Three-dimensional geospatial model

A system uses models to relocalize a mobile device. The system accesses an input image of a scene in a real-world environment, where the image was captured by a mobile device. The system applies a two-dimensional (2D) foundation model to the input image. The 2D foundation model is trained to determine an image vector representing characteristics of the input image. The system accesses a map representation of the real-world environment, where the map representation includes visual data that describes the real-world environment. The system applies a three-dimensional (3D) geospatial model to the map representation and the image vector. The 3D geospatial model is configured to output 3D splats representing the real-world environment. The system determines a pose of a camera that captured the input image using the 3D splats.
Owner:NIANTIC SPATIAL INC

Mine video semantic communication transmission method based on feature compression

The invention discloses a feature compression-based mine video semantic communication transmission method, which comprises the following steps of: sampling videos collected in a mine into pictures according to frames, packaging and sending the pictures into a semantic encoder; mapping the image vector to a feature vector of a feature space, and partitioning; the semantic encoder carries out feature sensing compression processing on the feature vectors, meanwhile, a tracking matrix is adopted to record a compression path, and the compressed features and the tracking matrix are sent to a channel through the transmitting end; the semantic decoder maps the feature vectors to a low-dimensional space through an embedded layer and then recovers the spatial position information of the features according to a tracking matrix; semantic decoding is carried out on the feature vectors of the recovered spatial position information; and taking minimization of a loss function of a target task as a target, jointly training a proposed semantic encoder and a semantic decoder, and updating parameters of a language model by adopting a self-adaptive estimation method. According to the method, the transmission data volume is greatly reduced under the condition that the semantic task execution effect is hardly damaged.
Owner:SOUTHEAST UNIV +1

Image vector quantization encoding, text-image model training and using method and device

The application discloses an image vector quantization encoding method and device, a text-image model training method and device, and a text-image model using method and device. The method comprises the following steps: inputting an image into an encoder to obtain intermediate feature vectors corresponding to each image block contained in the image; searching for indexes of image representations closest to the intermediate feature vectors corresponding to each image block in the image in a first codebook; the first codebook contains multiple rows of image representations and corresponding indexes, and the positions of the indexes corresponding to the similar image representations in the first codebook are also adjacent; and replacing the intermediate feature vectors of each image block in the image with the indexes searched to obtain vector quantization encoding corresponding to each image block in the image. The application can greatly save the calculation amount, improve the speed and efficiency of vector quantization encoding, improve the training efficiency of the model, and reduce the consumption of computing resources.
Owner:ALIBABA (CHINA) CO LTD

Text generation method and system based on AI

The invention provides an AI-based text generation method and system, and the method comprises the steps: extracting a global semantic text vector sequence from text data, and extracting a spatial correlation image vector sequence from image data; obtaining a text and image attention weight matrix and an image and text attention weight matrix based on the two; further acquiring a global text and graph vector sequence and a spatial text and graph vector sequence, and acquiring multi-modal fusion features based on the global text and graph vector sequence and the spatial text and graph vector sequence; and taking the multi-modal fusion features as condition information of a language model, outputting a plurality of query texts through the language model, and outputting confidence coefficients corresponding to the query texts. Through fine-grained alignment, the relevance between image-text features is ensured, the problem of modal difference is solved, and the semantic relevance between different modals is improved; the dependency degree on texts or images in the image-text fusion process is dynamically calculated based on the gating weight, and deep semantic association between different modes is further mined.
Owner:ENERGY RES INST OF JIANGXI ACAD OF SCI

Structured data generation method for engineering field drawings

The invention discloses a method for generating structured data of engineering field drawings, which comprises the following steps of: extracting data in different formats and drawing field types from the engineering field drawings in different file formats, and normalizing the data in different formats and the engineering field drawings, the method comprises the following steps: performing corresponding field matching according to a drawing field type and a multi-field sample image-text corpus, selecting a sample with high similarity in a matched field, constructing a cue word context for the sample, generating structured data based on large model knowledge distillation, and filtering to obtain fine-tuning corpora of different engineering fields; image vector embedding and text vector embedding are obtained based on fine-tuning corpora in different engineering fields, fine-tuning training is carried out in combination with the base model, an engineering field drawing analysis model is obtained, and structural data are obtained by using the engineering field drawing analysis model. The method solves the problems that the structured data can be generated only when multiple engineering field plug-ins are developed to generate full-design process data, and the structured data need to be manually annotated in paper drawings.
Owner:CHINA HAISUM ENG

A spectral reconstruction method based on double-branch heterogeneous feature fusion

This invention provides a spectral reconstruction method based on dual-branch heterogeneous feature fusion, relating to the fields of computational imaging and artificial intelligence. The method includes acquiring two-dimensional compressed measurement image vectors, constructing a local-global feature fusion deep neural network model, which includes a local feature extraction module, a global context extraction module, a feature fusion module, and a decoder; designing a composite loss function, including reconstruction loss, physical prior loss, and feature regularization loss; performing end-to-end training on the local-global feature fusion deep neural network model; and deploying the trained model to achieve rapid spectral reconstruction. This invention combines the fast inference capability of deep neural networks with a novel parallel feature extraction and fusion architecture. Through a carefully designed network model and a composite loss function, it effectively integrates and constrains local and global information in the feature space, achieving rapid and high-precision reconstruction of hyperspectral images.
Owner:BAY AREA LOW ALTITUDE RESEARCH INSTITUTE (GUANGDONG) CO LTD