Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

505 results about "Image description" patented technology

An image description is a textual, audio or graphical content portraying the image in a representation intelligible by the addressees. The description should be comprehensive as well as perceptible by the target audience.

Intelligent coagulant adding control method and system based on image recognition and multi-parameter modeling

The invention relates to an image processing and data processing technology, in particular to an intelligent coagulant dosing control method and system based on image recognition and multi-parameter modeling, the floc state is accurately quantified through image recognition and deep learning modeling, and a dosing prediction model with self-adaptive capacity is established in combination with raw water feed-forward information. And accurate control of coagulant addition is realized. The method comprises the following steps: collecting a floc image, carrying out image processing and floc feature extraction, and constructing a floc image description vector; time sequence input structure data fusing the floc image and the water quality data is constructed, and floc image sampling at each moment is defined as a time frame; performing enhancement and reconstruction processing on the training data of the dosing amount prediction model by adopting a data enhancement and sample equalization strategy to obtain continuously distributed synthetic samples; and constructing a hierarchical feature fusion enhanced dosing amount prediction model, fusing the previous water quality parameters, the current water quality parameters and the floc image joint feature vectors, and optimizing the dosing amount prediction precision layer by layer.
Owner:GUANGDONG LONGQUAN TECH CO LTD

Picture editing method and system based on multi-modal condition adaptation

The invention provides a picture editing method and system based on multi-modal condition adaptation, and the method comprises the steps: obtaining a first text vector: obtaining a picture description corresponding to a picture through processing, processing the picture description through an editor, and obtaining a first text vector; obtaining a first text vector based on the picture information; a second text vector acquisition step: processing the used editing instruction to obtain a second text vector based on the editing instruction; a fusion feature acquisition step: fusing the first text vector and the second text vector through weight to obtain a fusion feature; an image potential feature code acquisition step; in the image editing step, the injected condition information is received, meanwhile, the received potential features and potential noise of the image are denoised, and the image desired by the user is gradually generated in the iterative denoising process under the guidance of the received condition information; and an image restoration step. According to the invention, the stability, controllability and accuracy of image editing can be improved.
Owner:SHENZHEN EMDOOR DIGITAL TECH

Intelligent hardware grinding system and method based on machine vision

The invention discloses an intelligent hardware grinding system based on machine vision, which comprises a machine vision module used for carrying out image acquisition on grinding of hardware and generating surface feature information and grinding state image feature information of the hardware; the robot grinding module is used for grinding the hardware according to the surface feature information and the grinding state image feature information; the control module is connected with the machine vision module and the robot grinding module and used for adjusting grinding equipment control instruction features and grinding parameters in real time according to the surface feature information and the grinding state image information; according to the method, the target hardware grinding information is subjected to image recognition, the image description set is mined, and accurate path planning is generated and covers shape, size, defect position and other information. According to the method, deviation judgment is carried out on the multiple grinding task events, so that the deviation correction execution strategy is determined, the hardware grinding precision is effectively improved, the hardware rejection rate is reduced, and the hardware production efficiency is improved.
Owner:JINHONGXING (HUIZHOU) TECH CO LTD

Target detection method and device, model training method and device, electronic equipment and medium

The invention relates to the technical field of data processing, and provides a target detection method and device, a model training method and device, electronic equipment and a medium. The target detection method comprises the steps that a to-be-recognized image and a query text are acquired, and the query text is used for querying a target object corresponding to the query text in the to-be-recognized image; performing image recognition on the to-be-recognized image to obtain image description features and region detection visual features; performing regional multi-modal fusion processing on the image description features and the regional detection visual features to obtain regional multi-modal fusion features; performing feature fusion processing on text features obtained based on the query text and the regional multi-modal fusion features to obtain text regional fusion features corresponding to the query text; and a target detection result is obtained based on the text features and the text region fusion features, so that the fusion degree of text semantics and image region features is improved, and the accuracy and robustness of target detection in a complex scene are improved.
Owner:BEIJING JIZHI DIGITAL TECH CO LTD

Intellectual visual question and answer method and device based on big and small model collaboration and medium

The invention discloses a knowledge-based visual question-answering method and device based on big and small model collaboration and a medium, belongs to the technical field of visual question-answering, and solves the problem of how to improve the accuracy of complex questions faced by a visual question-answering technology. In the image description extraction step, an image description and key object label set which is highly associated with a natural language problem is generated, and in the entity enhancement processing step, key entities are extracted and a clarity problem and an entity example which are used for resolving semantic ambiguity are generated. In the candidate answer generation step, a plurality of semantically complementary answers are generated under different reasoning dimensions on the basis of a multi-round generation strategy, optimized candidate answers are screened on the basis of a model built-in scoring mechanism, in the context example retrieval step, examples most related to current input are retrieved, and finally, the candidate answers are obtained. And the large language model receives a unified input prompt formed by splicing output results in the steps, and inference is performed in an autoregression mode to generate answers, so that the accuracy of the visual question-answering technology facing complex questions is effectively improved.
Owner:STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST

Multi-language object illusion relieving method based on cross-language attention mode

The invention discloses a multi-language object illusion relieving method based on a cross-language attention mode, solves the problem of how to relieve multi-language object illusion during detection of a large visual language model under non-English questions, and belongs to the technical field of multi-mode questions and answers. The method comprises the following steps: identifying a cross-modal attention head set of which a visual language model shows obviously different behaviors for English and a target language when same semantic questions in different languages are processed; constructing image description queries of English and target languages for the same image, respectively inputting the image description queries into respective visual language models for reasoning, obtaining attention output under the English and target languages, and taking an average difference between the attention output as a language migration vector of the target languages; and in the reasoning process of the target language question, intervening the attention heads in the attention head set by using the language migration vector, so that the visual understanding ability of the visual language model under the non-English question is closer to the English question.
Owner:HARBIN INST OF TECH

Image processing model

A method comprising generating image descriptions of images in an original training set of images; determining, using at least one LLM, at least one domain and / or class which is under-represented in the original training set; generating, using a second LLM and based on the determination of the at least one domain and / or class, at least one instruction for a third LLM to generate at least one text prompt; generating, using the third LLM and based on the at least one instruction, the at least one text prompt for a text-to-image model; generating, using the text-to-image model and based on the at least one text prompt, at least one synthetic image; and generating an enhanced training set of images for use in training an image processing machine learning, ML, model, the enhanced training set of images comprising the original training set of images and the at least one synthetic image.
Owner:FUJITSU LTD +1

Method and device for multi-stage generation of medical image question and answer thinking chain data

The invention discloses a method and device for generating medical image question and answer thinking chain data in multiple stages. According to the method, by introducing a multi-stage model collaborative reasoning mechanism, extraction of key image content, enhanced generation of image description and construction of a thinking chain are completed in sequence, so that the accuracy, pertinence and interpretability in a medical image question and answer task are improved; the problems that an existing model is inaccurate in image detail recognition, opaque in reasoning process, generalized in question and answer content and the like are solved.
Owner:HANGZHOU DIANZI UNIV

Multi-modal fault diagnosis method based on cross-modal data enhancement

The invention relates to a multi-modal fault diagnosis method based on cross-modal data enhancement, and the method comprises the following steps: obtaining an industrial multi-modal data set, inputting a multi-modal named entity recognition model, and obtaining a fault diagnosis result; in the multi-modal named entity recognition model, a cross-modal data enhancement module is established, and image description is generated according to image data to improve the diversity of training data, so that the robustness of the model in a low-resource scene is enhanced; constructing a global alignment optimizer based on an information theory principle, aligning global semantics among modals by maximizing consistency information among the modals, and filtering irrelevant visual noise by minimizing irrelevant information of image modals; a fine-grained cross-perception module is constructed based on a cross attention mechanism, accurate interaction between text and image modalities is captured, and multi-modal representation with rich information is generated, so that the accuracy and reliability of fault diagnosis are improved. Compared with the prior art, the robustness and accuracy of fault diagnosis can be further improved.
Owner:TONGJI UNIV

Image description generation method and system based on enhanced fine-grained information

The invention discloses an image description generation method and system based on enhanced fine-grained information. The method comprises the following steps: 1) extracting regional features and global features of an image, and carrying out weighted fusion to generate multi-view feature information; 2) combining text feature extraction and cosine similarity screening, dynamically fusing text-image features, and optimizing cross-modal features through a dynamic multi-view enhancement mechanism; (3) a hierarchical memory enhancement encoder is adopted, multi-head attention and Transform structure hierarchical coding features are combined, and detail information is reserved through a memory unit; and 4) aggregating image and text weights in two stages through a double cross-modal attention decoder to generate high-quality description. The system comprises a feature extraction module, a text processing module, an encoding module and a decoding module. Through fine-grained feature enhancement and cross-modal interaction, the problems of detail loss and semantic deviation in a traditional method are effectively solved, and the method is suitable for image semantic analysis of complex scenes.
Owner:NANJING UNIV OF SCI & TECH

Dynamic split mirror generation system and method based on controllable diffusion model

The invention discloses a dynamic split mirror generation system and method based on a controllable diffusion model, and belongs to the technical field of film and television production. The implementation method comprises the following steps of: 1, setting a keyword text by a user, and inputting the keyword text into ChatGPT to generate a script; 2, converting the script into a scene category, a camera motion mode, a character role position and action and image description in a shot language by utilizing ChatGPT; 3, using a CLIP model to carry out contrast training on the image encoder and the text encoder; 4, performing an OpenPose model on the action reference image to obtain skeleton key points of the character role, and converting the skeleton key points into image character actions; 5, performing action fine-grained control on the character action of the image by using a Stable Diffusion model and a ControlNet model, and generating a split image; 6, utilizing a Pika Labs model to generate a dynamic video from the split image; compared with the prior art, the accuracy of user role action matching under the scene based on multi-text and high-complexity actions is improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM +2

Processing method for medical report structured information extraction and privacy protection

The invention discloses a processing method for structural information extraction and privacy protection of a medical report, which relates to the field of medical report processing and can efficiently and accurately extract text information in a table by integrating advanced image preprocessing, dynamic layout analysis, OCR (optical character recognition) technologies and a dual-channel privacy detection model. Establishing a coordinate mapping relation between the image description paragraph and the corresponding image; the generated JSON structure containing the text, table and image-text mapping relation ensures the integrity and accuracy of the data; the problems of data dislocation, context splitting, privacy disclosure and the like when the medical document with complex typesetting is processed in the prior art are solved; besides, an improved Transform framework is constructed on the basis of a Qwen model, a logical reasoning chain is injected, meanwhile, a case reading large model is obtained by integrating clinical diagnosis rules and multi-modal contrast learning strategy training, the model is utilized to generate a high-quality structured abstract, and the abstract generation accuracy is greatly improved.
Owner:SHANGHAI HENGFANG HEALTH TECHNOLOGY CO LTD

Community data management system for intelligent old-age care community

The invention relates to the field of data management, in particular to a community data management system for an intelligent old-age care community, and the system comprises a video monitoring unit which comprises a plurality of community monitoring devices and is used for collecting monitoring videos of all monitoring areas; the video analysis unit is used for periodically determining the video frame category of each basic video frame according to the personnel change reference value and the personnel influence reference value of the basic video frame corresponding to the community monitoring equipment; the video optimization unit is used for determining the state of the monitoring equipment according to the first-class video frame proportion and the first-class video frame density corresponding to the community monitoring equipment, and determining a corresponding video frame optimization mode as compensation optimization or activity feature optimization according to the state of the monitoring equipment; the key frame optimization unit is used for determining a processing mode to supplement or generate image description for the associated frame according to the effective state; and the monitoring data processing efficiency is improved.
Owner:NINGBO DAHONGYING UNIV

Character image generation method and device, electronic equipment and storage medium

The embodiment of the invention provides a character image generation method and device, electronic equipment and a storage medium, belongs to the technical field of artificial intelligence, and is suitable for the fields of financial science and technology and medical science and technology. The method comprises the following steps: carrying out semantic feature coding on a current image description text to obtain description text semantic features; and carrying out image element identification on the semantic features of the description text to obtain image components. And updating the semantic features of the description text according to the historically generated image and the image composition elements to obtain the current semantic features. Performing image screening according to the current semantic feature to obtain a reference character image, and performing image feature coding on the reference character image to obtain a reference image feature; and performing feature fusion according to the description text semantic features and the reference image features to obtain guidance fusion features. And performing image generation based on the guidance fusion features. According to the embodiment of the invention, the accuracy of generating the character image can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-modal data generation and fine adjustment method for domain image

The invention provides a multi-modal data generation and fine tuning method for a domain image, and belongs to the technical field of artificial intelligence, and the method comprises the steps: carrying out the primary processing of the domain image through a visual multi-modal model, obtaining a general description, carrying out the natural language conversion of a set label, obtaining a conversion text, and carrying out the fine tuning of the conversion text; carrying out semantic integration and expansion on the image description and the converted text based on a large language model to generate an initial description; performing multi-layer extraction on the domain image according to description dimensions to obtain layer description of each layer, and generating comprehensive description in combination with the initial description; matching the domain image with the corresponding comprehensive description to obtain multi-modal data; and performing fine tuning processing on the visual multi-modal model by using the multi-modal data to obtain an optimized multi-modal model. And the identification and question-answering capabilities in the domain image are improved.
Owner:BEIJING ZHONGKE STRON CLOUD INTELLIGENT TECH CO LTD

Multi-modal large model object illusion relieving method for removing training prejudice based on efficient machine forgetting

The invention discloses a multi-modal large model object illusion relieving method for removing training prejudice based on efficient machine forgetting. The method comprises the following steps: S1, collecting a deviation reasoning result: firstly, selecting an instruction related to an image description task from LVLM training data as a source data set; thirdly, the image and text instructions in the source data set are subjected to reasoning again through LVLM, and a reasoning result set is generated; by comparing the object existence conditions in the real output and the reasoning result, screening out the reasoning result containing the illusion phenomenon; and S2, efficient depolarization training: based on the collected deviation reasoning result, carrying out depolarization training on the LM head by adopting a machine forgetting method based on gradient rise. The invention provides an efficient depolarization method by specifically eliminating the illusion phenomenon in a large visual language model (LVLM). According to the method, a deviation reasoning result of a model on training data is collected, and a language model head (LM head) is finely adjusted, so that efficient utilization of data and parameters is realized.
Owner:RENMIN UNIVERSITY OF CHINA

Image description generation system, training method, generation method and electronic equipment

The invention discloses an image description generation system, a training method, a generation method and electronic equipment, and belongs to the technical field of image description. According to the method, visual features of an image are mapped to a visual and language comparable space, after a semantic information sequence is obtained, cross-modal semantic calculation of the semantic information sequence and the visual feature sequence is achieved through a Transform decoder, the middle hidden state of each candidate vocabulary is obtained, and then a corresponding directed acyclic graph is constructed; and after an optimal path is selected from the directed acyclic graph, the optimal path is directly mapped into image text description by a linear classifier. According to the method, visual information of the image and semantic information contained in the image are fully utilized, the sequential relation between words is learned by introducing the directed acyclic graph, the smoothness of description generation is improved, the method has non-autoregressive decoding attributes, and high-quality image text description can be generated at a high speed.
Owner:HUAZHONG UNIV OF SCI & TECH

Image text description generation method, electronic equipment and readable storage medium

The invention provides an image text description generation method, electronic equipment and a readable storage medium. According to the method, the object perception prototype learning module and the global context feature extraction module are introduced, so that fine-grained information and global semantic understanding in the image are effectively balanced. The visual backbone network module can extract multi-scale and multi-level image features and perform fusion, thereby enhancing the expression ability of the image features. The object perception prototype learning module further extracts an object prototype from the fusion features to ensure that the model can accurately capture key objects and attributes thereof in the image, and the global context feature extraction module ensures that the overall context of the image is fully understood. On the basis, the encoding and decoding module combines the global context and the object prototype to generate the text description, so that the semantic splitting phenomenon in the traditional method is avoided, and the detail information in the image is effectively reserved, thereby improving the accuracy and integrity of the image description.
Owner:WUHAN UNIV

Unified streaming processing method, system and device for multi-mode AI interactive content, medium and program product

The invention discloses a unified streaming processing method, system and device for multi-modal AI interactive content, a medium and a program product, and the method comprises the steps: receiving a request which is sent by a user and comprises an image and a text, and constructing a request context; extracting image features; carrying out image analysis and outputting an image description text; performing intention recognition and outputting an intention; multi-stage reasoning is carried out, and thinking content fragments are generated in a streaming mode; buffering and releasing thinking content fragments; generating text content fragments based on multi-stage reasoning; buffering and releasing the text content fragments; performing recommendation triggering based on the image description text and the intention; recommending content fragments; inserting a recommended position, buffering and issuing; and process event management: issuing event notifications when the beginning, any step fails and the end. The method is a universal method capable of transmitting different types of outputs in a single ordered stream, the analysis cost can be reduced, and the interaction experience and expansibility are improved.
Owner:BEIJING DIANFU TECHNOLOGY CO LTD

Automatic software testing method and system based on multi-mode AI cooperation

The invention relates to the technical field of automatic testing, in particular to an automatic software testing method and system based on multi-mode AI cooperation. According to the scheme, optical character recognition, voice recognition and natural language processing analysis tools are packaged into independent containers, and analysis results are input into a multi-mode semantic association model in real time; calculating the semantic similarity between the image description text and the corresponding voice transcription text; constructing a test demand complexity evaluation model; defining use case quality indexes, and establishing an error sample library to store manually labeled problem use cases and correction labels thereof; the intention demand is finally determined according to the screening result of the three-stage filter; and setting a test case template, optimizing a path according to the fusion quality, improving the case generation efficiency, and synchronously updating the optimized path to a knowledge base associated with the error sample library. According to the scheme, automatic software testing is achieved through multi-modal AI cooperation in combination with containerized deployment, model dynamic optimization and the like, and efficiency and recognition accuracy are improved.
Owner:RUIJIAN TECHNOLOGY (BEIJING) CO LTD

Document structure extraction and model training method and device, equipment and medium

The invention discloses a document structure extraction and model training method and device, equipment and a medium, and relates to the technical field of artificial intelligence and computer vision. The method comprises the following steps: constructing a special training data set containing data of at least two document understanding tasks (including optical character recognition, layout analysis, text positioning, regional text extraction, image description and chart title generation), and a fine tuning data set for converting a document image into a machine-readable structured text format; constructing a multi-modal large model comprising a shape adaptive cutting module, a visual encoder, a visual token compression module, a modal connector and a language decoder; pre-training the model by using the special training data set to jointly learn various document understanding tasks; and performing fine tuning on the pre-trained model by using the fine tuning data set, and adapting to a document structure extraction task to obtain a document structure extraction model. By means of the technical scheme, efficient and accurate document structure extraction can be achieved.
Owner:CETC CYBERSPACE SECURITY TECH CO LTD

Text-based image retrieval

A method, apparatus, non-transitory computer readable medium, and system for media processing include obtaining a text prompt describing content, generating, using a multi-modal encoder, a text embedding based on the text prompt, and obtaining an image depicting the content based on the text embedding. The multi-modal encoder is trained to encode image descriptions based on a similarity between a caption of a training image and a paraphrase of the caption.
Owner:ADOBE INC

Multi-modal task fine tuning method based on singular value decomposition enhanced routing function

The invention discloses a singular value decomposition-based multi-modal task fine tuning method for enhancing a routing function, which comprises the following steps of: mapping input language and visual features from a high-dimensional space to a low-rank space by using a PEFT method, performing singular value decomposition on language features in the low-rank space, performing routing function alignment through a tensor after efficient reconstruction, and performing multi-modal task fine tuning on a multi-modal task based on a singular value decomposition-enhanced routing function. And finally, after the low-rank space is recovered to the original dimension again, performing residual connection with the original language features, and outputting the features. According to the method, singular value decomposition is applied to language features before a routing function, a low-rank dominant mode of the language features is extracted, the alignment precision of vision and language features is enhanced, interference of high-dimensional noise is eliminated, and meanwhile calculation efficiency and model stability are kept. Routing calculation is carried out through the reconstructed tensor, key information in the features can be better extracted and aligned, and therefore the precision and effect of feature alignment are improved. The method is suitable for VL tasks such as visual questioning and answering and image description generation, and model performance can be obviously improved.
Owner:GUANGDONG POLYTECHNIC NORMAL UNIV

Knowledge enhancement and emotion inconsistency-based multi-mode siphonage detection method

The invention discloses a knowledge enhancement and sentiment inconsistency-based multi-modal chaffy detection method, which comprises the following steps of: obtaining a to-be-detected sample containing a text and an image, and extracting image and text features and text embedding features by using a feature encoder; and obtaining image description and texts in the image based on the image, splicing the image description and the texts to form knowledge texts, and sequentially inputting the knowledge texts into the emotion dictionary and the feature encoder to obtain emotion features and knowledge embedding features. And carrying out feature interaction among the image, the text and the emotion features by adopting cross attention, and carrying out adaptive weighting on the image and text features through a gating mechanism to obtain image-text comprehensive features and emotion features processed by the cross attention. And inputting the text embedding features and the knowledge embedding features into an emotion inconsistency module, and calculating an emotion inconsistency expression. And based on the image-text comprehensive features, emotion features subjected to cross attention processing and emotion inconsistent representation, performing chiffon prediction, and outputting a detection result. According to the invention, the method achieves better prediction performance on a multi-mode anti-tech detection public data set, and can more accurately recognize an image-text sample with irony emotion.
Owner:SHANTOU UNIV

Image context based text generation

Methods, systems, and storage media for generating contextually relevant text from image descriptions and user intent are disclosed. Exemplary implementations may: receive an image and a user-defined intent for text output; analyze the received image to generate a contextual description of the image; generate a query based on the contextual description of the image and the user-defined intent; and generate the text output based on the query.
Owner:SHUTTERSTOCK

Panoramic vision relation detection method based on thinking chain reasoning

According to the thinking chain reasoning-based panoramic vision relation detection method provided by the invention, an image description technology is introduced in a thinking chain reasoning process, so that a model can perform target detection and mutual verification according to generated description information and image information at the same time, and compared with the prior art, a posterior link is added, and the detection efficiency is improved. The target detection effect and reliability are improved, and the generated relation can better conform to the actual description; meanwhile, relation detection is carried out on the basis of panoramic segmentation, target recognition and visual relation detection can be achieved at the same time, and compared with the prior art, the mining degree of the relation between the target in the picture and the existence is higher; finally, visual relation detection based on panoramic segmentation can be realized through zero samples or small samples, and a good relation detection result can be obtained without training through a large amount of data.
Owner:BEIJING INST OF TECH

Method for generating image description text based on large model

The invention discloses a method for generating an image description text based on a large model, which relates to the technical field of image processing, and comprises the following steps: an image preprocessing step: dividing a target level through semantic segmentation and extracting key visual information by adopting hybrid denoising and self-adaptive normalization; a feature extraction step: fusing the multi-scale visual features and the semantic features, and generating a high-dimensional fusion feature vector through cross-modal alignment; a large model initialization and adaptation step: loading the pre-training model and performing incremental fine tuning, and dynamically adjusting the Prompt template; a text generation step: generating candidate texts through logic constraint and beam search; and a text optimization adjustment step of outputting a final description text based on multi-dimensional evaluation and user preference iterative correction. According to the method, the semantic matching degree, logic coherence and common sense accuracy of image description are improved, multi-element scenes are adapted through dynamic adaptation and iterative optimization, and high-quality text description support is provided for high-precision and diversified scenes.
Owner:BEIJING LINGMANG TECH CULTURE CO LTD

Multi-modal cross-domain question and answer data construction method, device and equipment

The invention discloses a multi-modal cross-domain question and answer data construction method, device and equipment, and the method comprises the steps: carrying out the syntactic analysis of a question text in obtained general domain image-text question and answer data, generating a question template, and constructing a general question and answer template library based on the question template; performing feature extraction on the to-be-processed target domain image data to obtain image description information; generating a target domain question text by combining the general question and answer template library and the image description information; and inputting the target domain question text and the target domain image data into a multi-modal question and answer model to generate an answer text, and taking the answer text and the target domain question text as target domain image question and answer pair data. According to the method, efficient, flexible and accurate generation of question and answer data can be realized for different fields.
Owner:XIAMEN YUANTING INFORMATION TECH CO LTD +1

Wafer defect retrieval method and device and storage medium

The invention provides a wafer defect retrieval method and device and a storage medium, and the method comprises the steps: obtaining a to-be-retrieved wafer image and related text information, the to-be-retrieved wafer image comprises defect features, and the related text information of the to-be-retrieved wafer image comprises an image description text; inputting the to-be-retrieved wafer image and the related text information into a pre-trained multi-modal feature extraction model to obtain multi-modal features of the to-be-retrieved wafer image; and retrieving a preset wafer defect database by using the multi-modal features of the to-be-retrieved wafer image to obtain similar cases of the to-be-retrieved wafer image. By means of the scheme, the accuracy and distinction degree of the detection result can be improved, and the wafer defect retrieval efficiency is improved.
Owner:SKYVERSE TECH CO LTD