Material identification method and device

Through multiple vision encoders to extract visual features and integrate weights, the hallucination problem in large visual language models is solved, and the accuracy of recognition and the reliability of the model is improved.

CN120198725APending Publication Date: 2025-06-24SHANGHAI HODE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510266946.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Large visual language models often face hallucinations in practical applications, generating erroneous text descriptions that are inconsistent with visual inputs, affecting the reliability and wide application of the model.

Method used

Multiple visual features are extracted from the material to be identified through multiple visual encoders, and the material features are obtained based on these features and preset weight arrays, and finally the recognition result is obtained. This method uses weights to combine multiple visual features, integrate knowledge of different visual encoders, and reduce object hallucinations.

Benefits of technology

It improves the ability to understand visual content, improves the accuracy of material recognition, reduces hallucinations, and enhances the reliability and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198725A_ABST
    Figure CN120198725A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a material identification method and device, computer equipment, a computer readable storage medium and a computer program product, and relates to the technical field of computers. The material identification method comprises the following steps: acquiring a to-be-identified material; extracting a plurality of visual features from the to-be-identified material through a plurality of visual encoders, wherein each visual encoder corresponds to one visual feature; according to the plurality of visual features and a preset weight array, obtaining material features of the to-be-identified material; wherein the preset weight array comprises a plurality of preset weights, and each preset weight corresponds to one visual encoder; and obtaining an identification result of the to-be-identified material according to the material features. According to the technical scheme provided by the embodiment of the invention, a plurality of visual encoders can be integrated through weights, so that the materials are identified by using different visual emphasis of different visual encoders, the understanding ability of visual contents is improved, the accuracy of material identification is improved, and the hallucination phenomenon is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technology, and in particular, to a method and apparatus for material recognition, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] In the modern field of artificial intelligence, large vision-language models (LVLMs) have become important tools for understanding and generating visual content. However, these models often face the problem of hallucination in practical applications, that is, generating incorrect text descriptions that are inconsistent with the visual input. This hallucination phenomenon not only affects the reliability of the model but also limits its wide application in practical scenarios.

[0003] Although it can be alleviated by increasing the amount of data and decoding optimization, the computational cost is often high, and the hallucination problem caused by insufficient visual understanding ability has not been fundamentally solved.

[0004] It should be noted that the above content is not necessarily prior art and is not used to limit the patent protection scope of the present application. Summary of the Invention

[0005] Embodiments of the present application provide a method and apparatus for material recognition, a computer device, a computer-readable storage medium, and a computer program product to solve or alleviate one or more of the above technical problems.

[0006] One aspect of embodiments of the present application provides a method for material recognition, the method including: Obtaining the material to be recognized; Extracting multiple visual features from the material to be recognized through multiple visual encoders, each visual encoder corresponding to one of the visual features; Obtaining the material feature of the material to be recognized according to the multiple visual features and a preset weight array; wherein, the preset weight array includes multiple preset weights, each preset weight corresponding to one of the visual encoders; Obtaining the recognition result of the material to be recognized according to the material feature.

[0007] Optionally, the multiple visual encoders include a target encoder and multiple expert encoders for different visual extraction tasks, and the multiple visual features include global features, local features, and multiple expert features; extracting multiple visual features from the material to be recognized through multiple visual encoders includes: Extracting the global feature and the local feature from the material to be recognized through the target encoder; and Extracting multiple of the expert features from the material to be recognized through the multiple expert encoders respectively.

[0008] Optionally, multiple said expert encoders are used to separately extract multiple said expert features from the material to be recognized, including: Obtain the feature dimension of the target encoder; Use multiple said expert encoders and linear adapters to extract multiple said expert features from the material to be recognized; wherein, the linear adapter is used to adjust the dimension of the expert features so that the dimension of the expert features is the same as the feature dimension.

[0009] Optionally, multiple said expert encoders are determined through the following operations: Obtain multiple candidate expert encoders and a preset test set; wherein, the preset test set includes multiple preset question classifications, each of the preset question classifications is associated with multiple test samples, and each of the preset question classifications corresponds to one of the visual extraction tasks; According to the target preset question classification and its associated multiple test samples, select one or more of the candidate expert encoders from the multiple candidate expert encoders as one or more target expert encoders corresponding to the target preset question classification; wherein, the target preset question classification is any one of the multiple preset question classifications; Determine multiple said expert encoders according to one or more said target expert encoders respectively corresponding to the multiple preset question classifications.

[0010] Optionally, according to the multiple visual features and a preset weight array, obtain the material feature of the material to be recognized, including: Determine the weighted feature of the material to be recognized according to the multiple global features, multiple said expert features and the preset weight array; Obtain the material feature according to the weighted feature and the local feature.

[0011] Optionally, the preset weight array is obtained through the following operations: Obtain a preset training set, and the preset training set includes multiple training samples; Perform multiple rounds of training according to the multiple training samples, and adjust starting from the default weight array until the preset weight array is obtained; In each round of training: Extract global training features from the current training sample through the target encoder; Extract expert training features from the current training sample through multiple said expert encoders; Obtain the target sample feature of the current training sample according to the global training feature, multiple expert training features, and the current weight array; wherein, the current weight array is the weight array obtained after the previous round of training. Obtain the target recognition result of the current training sample according to the target sample feature. Adjust the current weight array according to the target recognition result to obtain the weight array after this round of training.

[0012] Optionally, obtaining the recognition result of the material to be recognized according to the material feature includes: Convert the material feature into a target material feature recognizable by the target language model through a visual projector. Input the target material feature into the target language model, and output the recognition result through the target language model.

[0013] Another aspect of the embodiments of the present application provides a material recognition device, and the device includes: A first acquisition module, configured to acquire the material to be recognized. An extraction module, configured to extract multiple visual features from the material to be recognized through multiple visual encoders, and each visual encoder corresponds to one of the visual features. A second acquisition module, configured to acquire the material feature of the material to be recognized according to the multiple visual features and a preset weight array; wherein, the preset weight array includes multiple preset weights, and each preset weight corresponds to one of the visual encoders. A third acquisition module, configured to acquire the recognition result of the material to be recognized according to the material feature.

[0014] Another aspect of the embodiments of the present application provides a computer device, including: At least one processor; and A memory communicatively connected to the at least one processor; Wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0015] Another aspect of the embodiments of the present application provides a computer-readable storage medium, and computer instructions are stored in the computer-readable storage medium, and when the computer instructions are executed by a processor, the method as described above is implemented.

[0016] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.

[0017] The embodiments of the present application adopting the above technical solutions may include the following advantages: By using weights to combine multiple visual features to obtain material features, multiple visual encoders can be integrated through weights, so as to identify materials using different visual focuses of different visual encoders, improve the ability to understand visual content, improve the accuracy of material recognition, and reduce hallucination phenomena. BRIEF DESCRIPTION OF THE DRAWINGS The drawings exemplarily show embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary implementation manners of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0018] Figure 1 Schematically shows an operating environment diagram of the material recognition method according to Embodiment 1 of the present application; Figure 2 Schematically shows a flowchart of the material recognition method according to Embodiment 1 of the present application; Figure 3 Schematically shows Figure 2 a sub-step flowchart of step S202 in Figure 4 Schematically shows Figure 3 a sub-step flowchart of step S302 in Figure 5 Schematically shows an additional flowchart of the material recognition method according to Embodiment 1 of the present application; Figure 6 Schematically shows Figure 2 a sub-step flowchart of step S204 in Figure 7 Schematically shows another additional flowchart of the material recognition method according to Embodiment 1 of the present application; Figure 8 Schematically shows Figure 2 a sub-step flowchart of step S206 in Figure 9 Schematically shows an application example flowchart of the material recognition method according to Embodiment 1 of the present application; Figure 10 Schematically shows a block diagram of the material recognition device according to Embodiment 2 of the present application; and Figure 11 Schematically shows a hardware architecture diagram of a computer device according to Embodiment 3 of the present application; and Figure 12 Schematically shows an application diagram of the material recognition method according to Embodiment 1 of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] In order to make the purpose, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts belong to the scope of protection of the present application.

[0020] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.

[0021] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and to distinguish each step, and thus cannot be understood as a limitation of the present application.

[0022] First, the following provides the term explanations involved in the present application: LVLMs: A class of artificial intelligence systems that combine visual perception and language generation capabilities. By combining image or other visual inputs with natural language processing, they can understand and generate text descriptions related to visual content.

[0023] Hallucination: Refers to the phenomenon that visual language models (LVLMs) generate incorrect content that is inconsistent with the actual visual input when generating text descriptions. This phenomenon may manifest as incorrect object recognition, color misdirection, position errors, or fabrication of text content, etc.

[0024] Visual encoder: A key component in visual language models (LVLMs), whose main function is to convert the input image data into semantic feature vectors for fusion and interaction with the language model. It is usually pre-trained based on deep learning architectures (such as convolutional neural networks or Transformers) and can capture visual information in images, such as the shape, color, position of objects, etc.

[0025] CLIP: Contrastive Language–Image Pre-training, is a multi-modal deep learning model proposed by OpenAI. It maps images and texts to a shared embedding space through contrastive learning to achieve semantic alignment between images and texts. By jointly training an image encoder (usually based on Vision Transformer) and a text encoder (a Transformer-based language model), it learns the similarity relationship between image and text pairs.

[0026] CLS token: Classification Token, is a special token used to process sequence data in the Transformer architecture (such as BERT, CLIP, etc.). At the beginning of the input sequence, the CLS token is added to the sequence to aggregate the information of the entire sequence.

[0027] Patch tokens: is a technical means for processing image data. The input image is divided into several non-overlapping small blocks (patches), and each patch is flattened into a one-dimensional vector, and then embedded into the feature space of the model through a linear transformation.

[0028] Expert Encoder: is a neural network model specifically designed to process specific visual tasks, usually based on deep learning architectures. By pre-training on specific visual datasets or tasks, it can learn high-level visual features and semantic information related to that task.

[0029] Multi-modal: refers to the technology of simultaneously processing and integrating multiple different types of information or data in artificial intelligence and machine learning. These modalities usually include vision (such as images, videos), language (such as texts, speeches), and audition (such as audio), etc.

[0030] Convolutional Neural Network (CNN): is a deep learning architecture widely used in image processing and computer vision fields. It extracts local features of images through convolutional layers, reduces the feature dimension using pooling layers, and finally performs classification or regression through fully connected layers.

[0031] Transformer Architecture: A neural network architecture based on the self-attention mechanism, used for processing natural language processing (NLP) tasks. It abandons the sequential processing method of traditional recurrent neural networks (RNNs) and significantly improves the training efficiency through parallel computing.

[0032] ConvNeXt: An efficient convolutional neural network architecture designed to combine the powerful receptive fields of traditional convolutional networks with the flexibility of the Transformer architecture. By introducing hierarchical feature extraction, depthwise separable convolutions, and normalization and activation mechanisms similar to those of the Transformer, it significantly improves the performance of convolutional networks in tasks such as image classification, object detection, and segmentation, while maintaining a relatively low computational complexity.

[0033] Dimensionality Reduction Technique: A data processing method aimed at mapping high-dimensional data into a low-dimensional space while preserving as much of the key information of the original data as possible. It simplifies complex datasets by eliminating redundant features or extracting the main structure of the data, thereby reducing computational costs, improving model efficiency, and reducing the risk of overfitting.

[0034] Linear Adapter: A neural network module used for feature transformation and dimensionality adjustment. It consists of a linear transformation layer that learns the linear mapping relationship between input features and target features to achieve feature re-encoding or dimensionality adjustment.

[0035] LLaVA-ReCap-118K: A dataset for training vision-language models, containing approximately 118,000 images and their corresponding detailed descriptive captions. It is used to provide high-quality image-text pairs for vision-language models to enhance the models' understanding and generation capabilities of visual content.

[0036] PCA (Principal Component Analysis): A statistical method used for data dimensionality reduction and feature extraction. It projects the original data onto a new set of orthogonal coordinate axes (principal components) through a linear transformation. These principal components are sorted by variance size and can retain the original information of the data to the greatest extent.

[0037] Visual Projector: A module in a vision-language model whose main function is to map the visual features (such as semantic information and spatial information of images) extracted by the visual encoder from the visual feature space to an embedding space compatible with the language model.

[0038] ReLU (Rectified Linear Unit): An activation function applied to neural networks that sets the part of the input value less than zero to zero and keeps the part greater than zero unchanged.

[0039] Stochastic Gradient Descent (SGD): It is an algorithm used to optimize the parameters of machine learning models. It is a variant of Gradient Descent. The core idea is to randomly select one or a small batch of training samples to calculate the gradient of the loss function, thereby updating the model parameters.

[0040] Secondly, to facilitate the understanding of the technical solutions provided by the embodiments of the present application by those skilled in the art, the related technologies are described below: Large Visual-Language Models (LVLMs), such as GPT-4V and LLaVA, are powerful visual understanding tools and play an important role in content generation. However, they suffer from the problem of generating object hallucinations, which limits their use in practical applications. Many studies have proposed various methods to mitigate hallucinations in LVLMs, such as solving this problem by increasing the model size and dataset scale, but these methods have high computational costs and require efficient deployment during the inference stage.

[0041] Some studies have explored the use of visual encoders to reduce training costs. The CLIP visual encoder is commonly used as the base encoder for the latest leading LVLMs. However, the cross-modal pre-training paradigm of CLIP may have limitations in fine-grained visual representation capabilities, resulting in different visual encoders showing different hallucination characteristics in downstream tasks. To address this limitation, an auxiliary visual encoder can be used to optimize the CLIP representation. However, their different architectural designs and domain-specific training data may induce heterogeneous hallucination patterns, where different encoder architectures show different hallucination characteristics in downstream tasks.

[0042] For this reason, the embodiments of the present application provide a technical solution for material recognition. In this technical solution: (1) By setting a more refined test set, the types of hallucination problems are carefully divided, the deficiencies in visual perception capabilities of existing multimodal architectures are accurately identified, and expert encoders with better processing capabilities in different visual tasks are selected using the test set; (2) Extract the feature representations of the expert encoders, and then based on the global features of the materials, learn to assign weights to each expert encoder; (3) Integrate the features extracted by the visual encoder through a linear adapter, thereby fusing the knowledge of different expert encoders, comprehensively encoding visual inputs from multiple perspectives, and helping to reduce object hallucinations. See the following for details.

[0043] Finally, for ease of understanding, an exemplary operating environment is provided below.

[0044] As Figure 1 shown, the operating environment diagram includes: a server 2, a network 4, and a client 6, where: Server 2 can be composed of a single or multiple computing devices. The multiple computing devices can include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing devices can load virtual machines based on virtual images and / or other data that define specific software (e.g., operating systems, dedicated applications, servers) for emulation. As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more computing devices. A hypervisor can be implemented to manage the use of different virtual machines on the same computing device.

[0045] Server 2 can be configured to communicate with a client 6, etc. via a network 4. The network 4 includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network 4 can include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links, combinations thereof, etc., or wireless links, such as cellular links, satellite links, Wi-Fi links, etc.

[0046] Server 2 can provide services such as storage, reading, writing, querying, deleting, etc., such as providing a material identification service for the client.

[0047] The client 6 can be an electronic device running an operating system such as Windows, Android™, or iOS, such as a smartphone, tablet device, laptop computer, virtual reality device, gaming device, set-top box, in-vehicle terminal, smart TV. Based on the above operating systems, various application programs can be run, such as a program for uploading materials to be identified.

[0048] The client 6 can provide / configure a user access page for manipulating the server 2 or uploading objects, etc.

[0049] It should be noted that the above devices are exemplary, and in different scenarios or according to different requirements, the number and types of devices can be adjusted.

[0050] The following takes the server 2 or the client 6 as the execution entity and introduces the technical solutions of the present application through multiple embodiments. It should be noted that these embodiments can be implemented in a variety of different forms and should not be construed as being limited only to the embodiments described herein.

[0051] Embodiment 1 Figure 2 A flowchart of a material identification method according to Embodiment 1 of the present application is schematically shown.

[0052] As Figure 2 shown, the material identification method can include steps S200 to S206, where: Step S200, obtain the material to be recognized; Step S202, extract multiple visual features from the material to be recognized through multiple visual encoders, where each visual encoder corresponds to one of the visual features; Step S204, obtain the material feature of the material to be recognized according to the multiple visual features and a preset weight array; Step S206, obtain the recognition result of the material to be recognized according to the material feature.

[0053] The material recognition method provided in this embodiment uses weights to combine multiple visual features to obtain a material feature. Thus, multiple visual encoders can be integrated through weights, so as to recognize the material using different visual focuses of different visual encoders, improve the ability to understand visual content, improve the accuracy of material recognition, and reduce the hallucination phenomenon.

[0054] The following combines Figure 2 , and elaborates in detail each step in steps S200 to S206 and other optional steps.

[0055] Step S200 , obtain the material to be recognized.

[0056] The material to be recognized can be a static picture, a dynamic picture, a video, or a combination of multi-modal data, etc. When the obtained material is a video, key frames can also be extracted from it as the material to be recognized. After obtaining the material to be recognized, it can be preprocessed, such as format conversion, size adjustment, noise removal, enhancement processing, etc. When the obtained material to be recognized has too low quality, such as too low picture resolution, damaged picture, or damaged video, etc., the material to be recognized can be repaired or directly excluded.

[0057] Step S202 , extract multiple visual features from the material to be recognized through multiple visual encoders, where each visual encoder corresponds to one of the visual features. One of the visual features can be represented by a feature vector.

[0058] In some embodiments, visual encoders with multiple architectures can be used, such as convolutional neural networks (CNNs), Transformer architectures (such as ViT), hybrid architectures (such as ConvNeXt), etc. In other embodiments, visual encoders pre-trained on different datasets or tasks can be used, such as models pre-trained on ImageNet, models pre-trained in specific fields (such as medical images, remote sensing images), or models pre-trained on large-scale unsupervised data (such as DINOv2). The visual encoders used can also include multi-modal visual encoders, such as the visual encoder of CLIP, to process image and text information simultaneously.

[0059] Multiple visual encoders can be dynamically selected based on the type of visual extraction tasks to be performed, or selected based on the content of the material to be recognized. For example, for tasks that require precise target localization, an encoder with stronger localization capabilities can be selected; for images of complex scenes, an encoder that can handle complex backgrounds can be selected.

[0060] After extracting features, the features can be enhanced, for example, highlighting the features of important regions through an attention mechanism. Dimensionality reduction techniques (such as PCA, t-SNE) or compression techniques (such as autoencoders) can also be used to optimize the features to reduce computational costs and improve the expressive power of the features.

[0061] In this embodiment, by using multiple visual encoders to jointly extract features from the material to be recognized, multi-dimensional visual information of the material can be more comprehensively captured, so as to more accurately describe different visual features, and improve the understanding ability of complex scenes as well as the accuracy and robustness of recognition.

[0062] As mentioned above, there can be various types of visual encoders, and correspondingly, there can be various implementation methods for extracting multiple visual features of the material to be recognized through multiple visual encoders. The following provides an exemplary method.

[0063] In an alternative embodiment, the multiple visual encoders include a target encoder and multiple expert encoders for different visual extraction tasks, and the multiple visual features include global features, local features, and multiple expert features. As Figure 3 shown, step S202 includes: S300, extracting the global feature and the local feature from the material to be recognized through the target encoder.

[0064] S302, extracting the multiple expert features from the material to be recognized respectively through the multiple expert encoders.

[0065] The target encoder can be a CLIP encoder. Different target encoders can also be selected according to factors such as the type and feature complexity of the material to be recognized. The global feature can be a feature representation carrying the overall semantic information and macroscopic structure of the material to be recognized, such as the CLS token extracted by the CLIP encoder; the local feature can be a feature representation extracted from different regions of the material to be recognized and used to describe the detailed information and local patterns of different regions, such as the patch tokens extracted by the CLIP encoder. The global feature can be divided into different levels such as coarse-grained, medium-grained, and fine-grained, and the local feature can be divided into different ranges such as small-scale, medium-scale, and large-scale.

[0066] Multiple expert encoders can include ConvNext, EVA-02, SAM, DINOv2, and Vary, etc. These expert encoders can adopt different architectures and be pre-trained on different downstream tasks to have different visual capabilities. They can also be fine-tuned through task-specific datasets to further improve their performance on specific tasks.

[0067] In some embodiments, multiple encoders that have the best effects on all types of visual tasks can be selected as expert encoders. In other embodiments, a task recognition module can also be used to determine the main visual task of the material to be recognized and select the most relevant expert encoder for feature extraction.

[0068] In this embodiment, through the global features and local features extracted by the target encoder, the global semantic information and local detail information of the material to be recognized are grasped, improving the accuracy and precision of recognition. At the same time, through the multiple expert features extracted by multiple expert encoders, it can better adapt to multiple visual tasks, improving the versatility and flexibility of the material recognition method.

[0069] When using expert encoders to extract expert features, various adaptive settings can also be made. The following provides an exemplary setting.

[0070] In an alternative embodiment, as Figure 4 shown, S302 includes: S400, obtaining the feature dimension of the target encoder; S402, extracting multiple expert features from the material to be recognized through multiple expert encoders and a linear adapter; wherein, the linear adapter is used to adjust the dimension of the expert features so that the dimension of the expert features is the same as the feature dimension.

[0071] For example, when the feature dimension output by the CLIP encoder as the target encoder is 1024, the output dimensions of multiple expert encodings can be fixed at 1024 through a linear adapter.

[0072] In some embodiments, the linear adapter can also dynamically adjust the feature dimension of the target encoder and the dimension of the expert features simultaneously according to the complexity of the material to be recognized or the visual task requirements of the material to be recognized to achieve more efficient feature integration.

[0073] In this embodiment, the expert encoder outputs expert features with the same dimension as the target encoder through a linear adapter. Thus, it is convenient to align the expert features with the features extracted by the target encoder, facilitating the integration of the visual features extracted by all visual encoders and improving the adaptability and computational efficiency of the material recognition method.

[0074] As mentioned above, there can be various types of expert encoders, and different expert encoders can be selected through multiple methods. The following provides an exemplary method for selecting an expert encoder.

[0075] In an alternative embodiment, as Figure 5 shown, multiple of the expert encoders are determined through the following operations: S500, obtain multiple candidate expert encoders and a preset test set; wherein, the preset test set includes multiple preset question classifications, each of the preset question classifications is associated with multiple test samples, and each of the preset question classifications corresponds to one of the visual extraction tasks.

[0076] S502, according to the target preset question classification and the multiple test samples associated therewith, select one or more of the candidate expert encoders from the multiple candidate expert encoders as one or more target expert encoders corresponding to the target preset question classification; wherein, the target preset question classification is any one of the multiple preset question classifications.

[0077] S504, determine multiple of the expert encoders according to one or more of the target expert encoders respectively corresponding to the multiple preset question classifications.

[0078] In some embodiments, the preset test set may include hallucination classifications corresponding to four basic capabilities of detection, segmentation, localization, and classification. Among them, detection hallucinations can include three subtypes: category hallucinations (misidentification of object existence), counting hallucinations (misestimation of object quantity), and occlusion hallucinations (partial observation fallacy). Segmentation hallucinations can include two subtypes: text hallucinations and shape hallucinations. Localization hallucinations can include two subtypes: absolute localization hallucinations and relative localization hallucinations. Classification hallucinations can include three subtypes: color hallucinations, action hallucinations, and relative interaction hallucinations. Each subtype represents a preset question classification. In other embodiments, other available test sets can also be selected as the preset test set according to actual needs.

[0079] For example, a category hallucination can be that the image only depicts a beach and the sea, but after the model identifies the image, it wrongly outputs a picture description of "a man surfing"; a text hallucination can be misidentifying "Cloud" as "Clown" due to font artifacts.

[0080] Each test sample can include a triple, which includes a visual material, a true text description of the visual material, and a hallucinated text description generated by GPT-4 according to the given true text description and hallucination subtype. The visual material in the test sample can be selected from existing open-source datasets, such as LLaVA-ReCap-118K, or can be taken or made by the user himself / herself.

[0081] When selecting a target expert encoder, it can be automatically selected based on the performance metrics of the candidate expert encoders (such as accuracy, recall, F1 value, etc.), or a comprehensive selection can be made by combining the complexity of the candidate expert encoders (such as the number of parameters and computational cost). A multi-round selection mechanism can also be used to gradually screen out the most suitable target expert encoder.

[0082] For example, selection can be made based on the accuracy of the candidate expert encoder (i.e., the proportion of test results without hallucinations). In a test, 1800 test samples of category hallucination classification in a preset test set are used to test 4 candidate expert encoders, obtain the test results of the 4 candidate expert encoders for the 1800 test samples, and compare the test results with the real text. The accuracies of the 4 candidate expert encoders are found to be 60%, 75%, 40%, and 82% respectively. Then, the candidate expert encoder with the highest accuracy (82%) can be selected as the target expert encoder for category hallucination classification. Of course, candidate expert encoders with an accuracy above the accuracy requirement threshold (75%) can also be used as the target expert encoders for category hallucination classification.

[0083] Previous studies on object hallucination patterns in LVLMs (Large Visual Language Models) mainly focused on simply classifying hallucination types into several major categories without delving into the differences in the visual sub-tasks involved. For example, all non-factual responses were classified as one type of hallucination, regardless of whether the error originated from an object counting or color recognition task. This simple approach failed to capture the nuances and specificities of hallucinations in different visual tasks.

[0084] In view of this, in this embodiment, designing a test set classification for each type of visual extraction task can improve the fineness of the expert encoder test, thereby enhancing the proficiency of the selected expert encoder for the corresponding visual extraction task, matching the most suitable expert encoder for each visual task, enhancing the adaptability and robustness of the material recognition method, and being able to better meet the diverse visual task requirements.

[0085] Step S204 , obtaining the material features of the material to be recognized according to the multiple visual features and a preset weight array; wherein, the preset weight array includes multiple preset weights, and each preset weight corresponds to one of the visual encoders.

[0086] In some embodiments, the preset weight array can be a universal weight array set by the user that has good effects for various visual tasks, or a specific weight array that has excellent effects for a certain specific visual task.

[0087] In some other embodiments, it is also possible to adaptively adjust the preset weight array according to the content, scene type, or task requirements of the material through a machine learning model (such as a neural network), etc. The preset weight array can also be adjusted according to the user's feedback and preferences. The optimization objectives of the preset weight array can be the balance of multiple objectives such as recognition accuracy, computational efficiency, and robustness. It is also possible to dynamically adjust the preset weight array according to the performance feedback of the material recognition method (such as misrecognition rate, response time, etc.).

[0088] Before weighting, each visual feature can be normalized to eliminate the dimensional differences between different features. It is also possible to perform dimensionality reduction on the visual features (such as PCA, t-SNE, etc.) to reduce the computational complexity and improve the interpretability of the features. After weighted fusion, a non-linear transformation (such as ReLU, Sigmoid, etc.) can also be performed on the material features to enhance the expression ability of the features.

[0089] Integrating multiple visual features using the preset weight array can make the material recognition method more adaptable in different visual tasks or application scenarios, providing a more comprehensive feature input for subsequent recognition tasks, thereby helping to reduce object hallucinations and improve the performance and recognition effect of the material recognition method.

[0090] As previously mentioned, there can be various visual encoders, and different combinations of visual encoders will naturally correspond to different specific methods for obtaining material features. The following provides an exemplary situation.

[0091] In an alternative embodiment, as Figure 6 shown, step S204 includes:[[]] S600, determining the weighted features of the material to be recognized according to the multiple global features, multiple expert features, and the preset weight array.

[0092] S602, obtaining the material features according to the weighted features and the local features.

[0093] In some embodiments, non-linear transformations (such as activation functions like ReLU, Tanh, etc.) can be used to enhance the expression ability of the features, or feature noise reduction techniques can be applied to remove redundant information in the weighted features. Feature importance analysis techniques such as SHAP values or LIME can also be introduced to interpret the obtained material features. Visualization techniques (such as feature map visualization) can also be used to show the roles of different features in the process of obtaining the material features.

[0094] In some other embodiments, a knowledge graph can be used to enhance the features, and the information in the knowledge graph can be incorporated into the process of obtaining the material features. Rules or constraints defined by domain experts can also be introduced to guide this process.

[0095] In some embodiments, the weighted features and local features can be integrated by a visual projector to obtain material features, and training or fine-tuning can be performed through datasets such as the LLaVA pre-training dataset or the LLaVA fine-tuning dataset to adjust the parameters of the visual projector, thereby improving the quality of the obtained material features and the ability of the material recognition method to resist hallucinations.

[0096] In this embodiment, multiple expert features and global features of the material to be recognized are integrated according to weights and matched with local features to obtain material features. Thus, visual features in multiple aspects and dimensions of the material to be recognized can be integrated, enabling the material recognition method to handle complex visual tasks more flexibly and comprehensively, improving the robustness and accuracy of the material recognition method, as well as the ability of the material recognition method to resist hallucinations.

[0097] As mentioned above, there are various methods for obtaining or setting the preset weight array. The following provides an exemplary method.

[0098] In an alternative embodiment, as Figure 7 shown, the preset weight array is obtained through the following operations: S700, obtain a preset training set, where the preset training set includes multiple training samples; S702, perform multiple rounds of training based on the multiple training samples, starting from the default weight array and adjusting until the preset weight array is obtained; S703, in each round of training: S703A, extract global training features from the current training sample through the target encoder; S703B, extract expert training features from the current training sample through the multiple expert encoders; S703C, according to the global training features, the multiple expert training features, and the current weight array, obtain the target sample features of the current training sample; wherein, the current weight array is the weight array obtained after the previous round of training; S703D, according to the target sample features, obtain the target recognition result of the current training sample; S703E, according to the target recognition result, adjust the current weight array to obtain the weight array after this round of training.

[0099] For example, applying the material recognition method of the present application in a material recognition system, the material recognition system includes a context-aware routing module for assigning weights to each expert encoder. The preset weight array is A. After the first round of training, the probability of hallucination when using the default weight array to recognize the preset training set is 20%, and most of them are color hallucinations. So the context-aware routing module increases the weight of the expert encoder E1 that is good at color recognition, and obtains the weight array A1 after this round of training; In the second round of training, use the weight array A1 (i.e., the current weight array) for training. After the second round of training, the probability of hallucination when recognizing the preset training set is 19%, and most of them are text hallucinations. So the context-aware routing module increases the weight of the expert encoder E2 that is good at text recognition, and obtains the weight array A2 after this round of training; Follow the above method and conduct multiple rounds of training subsequently until the probability of hallucination when recognizing the preset training set is 15% and no longer decreases, then the weight array after the last round of training can be determined as the preset weight array.

[0100] In some embodiments, the preset training set can be image or video materials covering different scenarios, different lighting conditions, and different resolutions. Data augmentation techniques such as random cropping, rotation, flipping, color adjustment, etc. can also be introduced to expand the quantity and diversity of training samples and improve the generalization ability of the model.

[0101] Before training, the quality of training samples can be evaluated and screened to remove samples that are blurred, have too much noise, or have inaccurate annotations. The training samples can also be preprocessed, such as removing duplicate samples, correcting incorrect annotations, etc.

[0102] The default weight array can be that the weights are evenly distributed among each visual encoder, or it can be weights set according to prior knowledge or random weights. During the training process, different optimization algorithms such as Stochastic Gradient Descent (SGD), Adam, RMSprop, etc. can be adopted to find the optimization strategy that is most suitable for the current task. An early stopping mechanism can also be introduced to monitor the performance of the validation set during the training process. When the performance of the validation set no longer improves, stop training in advance to avoid overfitting.

[0103] In some embodiments, the monitoring metrics during the training process, in addition to the common loss value and accuracy, can also include the change trend of weights, the distribution of gradients, etc., in order to better understand the training status of the model. The training process can also be visually displayed through visualization tools.

[0104] In this embodiment, by dynamically adjusting the weight array through multiple rounds of training to obtain a preset weight array, the adaptability of the preset weight array to various visual tasks can be improved, thereby improving the performance of the material recognition method in different visual recognition tasks and enhancing the ability of the material recognition method to resist hallucinations.

[0105] Step S206 , according to the material features, obtain the recognition result of the material to be recognized.

[0106] The recognition result may include the object category, location information, behavior description, text content, etc. in the material. In some embodiments, a classifier, a regressor, a generative model, or other machine learning algorithms may be used to generate the recognition result according to the material features.

[0107] After obtaining the recognition result, the result may also be filtered according to a confidence threshold to improve the accuracy of recognition. The recognition result may also be fused with text information or other modal data to further verify the reliability of the recognition result.

[0108] In this embodiment, determining the recognition result of the material to be recognized according to the material features weighted by multiple visual features can improve the accuracy and credibility of the recognition result and enhance the ability of the material recognition method to resist hallucinations.

[0109] As described above, there are multiple methods for obtaining the recognition result according to the material features. An exemplary method is provided below.

[0110] In an alternative embodiment, as Figure 8 shown, step S206 includes: S800, through a visual projector, transform the material features into target material features that can be recognized by the target language model.

[0111] S802, input the target material features into the target language model, and output the recognition result through the target language model.

[0112] In some embodiments, the visual projector may use linear transformation, non-linear transformation (such as multi-layer perceptron MLP), attention mechanism, or other neural network structures to implement feature mapping. Normalization operations (such as LayerNorm, BatchNorm) may also be used to optimize the distribution of features, and an adaptive feature scaling mechanism may be used to dynamically adjust the dimension or range of features according to the requirements or architecture of the target language model.

[0113] The target language model can be a pre-trained large language model (such as GPT series, Qwen, LLaMA, etc.), or a customized model fine-tuned for specific tasks. In some embodiments, multiple language models can also be integrated to improve the recognition accuracy through voting or fusion of multiple language models. In other embodiments, the most suitable language model can also be dynamically selected according to the characteristics of the input material, or the parameters or architecture of the model can be dynamically adjusted according to the complexity or feature distribution of the input material.

[0114] In this embodiment, the target language model is used to understand the transformed target material features to output the recognition result of the material to be recognized. Thus, the semantic understanding ability of the language model can be fully utilized, so that the material recognition method can obtain high-quality and readable recognition results, and the recognition accuracy and generalization ability of the material recognition method can be improved.

[0115] To make this application easier to understand, the following is combined with Figure 9 An exemplary application is provided. Wherein: S11, according to a preset test set, select 6 expert encoders E1~E6 that are the most effective for ten types of questions in the test set; S12, use the expert encoders E1~E6 and the CLIP encoder (i.e., the target encoder), and use the preset training set to train for N rounds to obtain a preset weight array G; S13, obtain an image P (i.e., the material to be recognized); S14, use the CLIP encoder to extract the CLS token (i.e., the global feature) and Patch tokens (i.e., the local features) from the image P; S15, determine that the feature dimension of the CLIP encoder is D, use the expert encoders E1~E6 to extract 6 expert features EF1~EF6 from the image P, and make the dimensions of the expert features EF1~EF6 also D through a linear adapter; S16, according to the preset weight array G, weight the CLS token and the 6 expert features EF1~EF6 to obtain a weighted feature R; S17, use the weighted feature R and the Patch tokens to combine to obtain the image feature F (i.e., the material feature) of the image P; S18, through a visual projector, transform the image feature F into an image feature FS (i.e., the target material feature) that can be recognized by a large language model L (i.e., the target language model); S19, input the image feature FS into the large language model L, and the large language model L outputs the recognition result of the image P.

[0116] In addition to the above exemplary applications, the material recognition method of the embodiments of the present application can also be applied in a variety of scenarios.

[0117] For example, when receiving a video uploaded by a user (such as a foreign language movie clip video), the neural network model can be used to extract key frames from the video, and then the picture features of the key frames can be extracted through the material recognition method of the present application, and a Chinese description of the key frames can be output, so that reviewers or machines can quickly understand the video content, thereby improving the review efficiency and accuracy. A content summary can also be generated according to the Chinese description of the video that has passed the review for the audience to quickly understand the video content.

[0118] For another example, after the generative artificial intelligence model generates video content, the material recognition method of the present application can be used to extract features from the generated video content and output a text description of the video content. By comparing the output text description with the original text used to generate the video content, the effect of video generation can be evaluated.

[0119] For yet another example, the material recognition method of the present application can be applied to an intelligent driving system. By capturing road images in real time and combining multiple expert encoders with good road scene recognition capabilities, road scenes can be recognized and described in real time. The generated road scene description can be used by the intelligent driving system of the vehicle to make corresponding driving decisions, such as avoiding pedestrians and obeying traffic signs, thereby improving driving safety and reliability.

[0120] For yet another example, in medical image diagnosis, multiple expert encoders with good medical image recognition capabilities can be used, and the weight array can be adjusted to improve the understanding and description capabilities of the material recognition method of the present application for medical images. When a new medical image is uploaded, the features of the medical image can be extracted and a description can be generated for doctors to refer to and quickly understand the key information in the image, thereby assisting in diagnosis and improving the accuracy and efficiency of diagnosis.

[0121] It should be noted that only a few application scenarios are listed above. According to actual needs, it can also be used in various other scenarios.

[0122] Embodiment 2 Figure 10 The block diagram of the material recognition device according to Embodiment 2 of the present application is schematically shown. The device can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 10As shown, the device 1000 may include: a first acquisition module 1100, an extraction module 1200, a second acquisition module 1300, and a third acquisition module 1400, where: The first acquisition module 1100 is configured to acquire the material to be recognized; The extraction module 1200 is configured to extract multiple visual features from the material to be recognized through multiple visual encoders, and each visual encoder corresponds to one of the visual features; The second acquisition module 1300 is configured to acquire the material feature of the material to be recognized according to the multiple visual features and a preset weight array; where the preset weight array includes multiple preset weights, and each preset weight corresponds to one of the visual encoders; The third acquisition module 1400 is configured to acquire the recognition result of the material to be recognized according to the material feature.

[0123] As an optional embodiment, the multiple visual encoders include a target encoder and multiple expert encoders for different visual extraction tasks, the multiple visual features include global features, local features, and multiple expert features, and the extraction module 1200 is further configured to: Extract the global feature and the local feature from the material to be recognized through the target encoder; and Extract multiple expert features from the material to be recognized through the multiple expert encoders respectively.

[0124] As an optional embodiment, the extraction module 1200 is further configured to: Acquire the feature dimension of the target encoder; Extract multiple expert features from the material to be recognized through the multiple expert encoders and linear adapters; where the linear adapter is used to adjust the dimension of the expert features so that the dimension of the expert features is the same as the feature dimension.

[0125] As an optional embodiment, the device 1000 further includes an expert encoder determination module, which is configured to: Acquire multiple candidate expert encoders and a preset test set; where the preset test set includes multiple preset question classifications, each preset question classification is associated with multiple test samples, and each preset question classification corresponds to one of the visual extraction tasks; Select one or more of the candidate expert encoders from the multiple candidate expert encoders as one or more target expert encoders corresponding to the target preset question classification according to the target preset question classification and its associated multiple test samples; where the target preset question classification is any one of the multiple preset question classifications; Determine a plurality of the expert encoders according to one or more of the target expert encoders respectively corresponding to the plurality of the preset question classifications.

[0126] As an optional embodiment, the second acquisition module 1300 is further configured to: Determine the weighted features of the material to be recognized according to the plurality of the global features, the plurality of the expert features, and the preset weight array; Obtain the material features according to the weighted features and the local features.

[0127] As an optional embodiment, the apparatus 1000 further includes a preset weight array determination module, configured to: Obtain a preset training set, where the preset training set includes a plurality of training samples; Perform multiple rounds of training according to the plurality of training samples, and adjust starting from the default weight value until the preset weight value is obtained; In each round of training: Extract global training features from the current training sample through the target encoder; Extract expert training features from the current training sample through the plurality of the expert encoders; Obtain the target sample features of the current training sample according to the global training features, the plurality of the expert training features, and the current weight array; wherein, the current weight array is the weight array obtained after the previous round of training; Obtain the target recognition result of the current training sample according to the target sample features; Adjust the current weight array according to the target recognition result to obtain the weight array after this round of training.

[0128] As an optional embodiment, the third acquisition module 1400 is further configured to: Convert the material features into target material features recognizable by the target language model through a visual projector; Input the target material features into the target language model, and output the recognition result through the target language model.

[0129] Embodiment III Figure 11Schematically shown is a hardware architecture diagram of a computer device 10000 suitable for implementing a material recognition method according to Embodiment 3 of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers), etc. As Figure 11 shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can be communicatively linked to each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 10000. Of course, the memory 10010 may also include both an internal storage module and an external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the material recognition method. In addition, the memory 10010 may also be used to temporarily store various types of data that have been output or will be output.

[0130] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chip. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0131] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.

[0132] It should be noted that Figure 11 Only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0133] In this embodiment, the material recognition method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of this application.

[0134] Embodiment 4 The embodiments of this application also provide a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the material recognition method in the embodiments are implemented.

[0135] In this embodiment, the computer-readable storage medium includes flash memory, hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), random access memories (RAM), static random access memories (SRAM), read-only memories (ROM), electrically erasable programmable read-only memories (EEPROM), programmable read-only memories (PROM), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the material recognition method in the embodiment. In addition, the computer-readable storage medium can also be used to temporarily store various data that have been output or will be output.

[0136] Embodiment 5 The embodiment of the present application further provides a computer program product, including a computer program, which implements the method in the above embodiment when executed by a processor.

[0137] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general-purpose computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device, so that they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to be implemented. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0138] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present application by the same token.

Claims

1. A material identification method, characterized in that: The method comprises: Obtain the material to be identified; Extracting multiple visual features from the to-be-recognized material by using multiple visual encoders, each visual encoder corresponding to one of the visual features; According to the plurality of visual features and a preset weight array, the material features of the material to be identified are obtained; wherein the preset weight array includes a plurality of preset weights, each preset weight corresponding to one of the visual encoders; According to the material feature, a recognition result of the material to be recognized is obtained.

2. The method according to claim 1, characterized in that The multiple visual encoders include a target encoder and multiple expert encoders for different visual extraction tasks, and the multiple visual features include a global feature, a local feature and multiple expert features; Extracting multiple visual features from the to-be-recognized material through multiple visual encoders, including: Extracting the global features and the local features from the to-be-recognized material by means of the target encoder; and A plurality of the expert features are respectively extracted from the material to be identified through the plurality of the expert encoders.

3. The method according to claim 2, characterized in that The method further comprises: extracting a plurality of expert features from the material to be identified by using a plurality of expert encoders, including: Obtaining a feature dimension of the target encoder; A plurality of the expert features are extracted from the to-be-recognized material through a plurality of the expert encoders and the linear adapters; wherein the linear adapter is used to adjust the dimension of the expert feature so that the dimension of the expert feature is the same as the feature dimension.

4. The method according to claim 2, characterized in that: The plurality of expert encoders are determined by the following operations: Acquire multiple candidate expert encoders and preset test sets; wherein the preset test set includes multiple preset question categories, each of the preset question categories is associated with multiple test samples, and each of the preset question categories corresponds to one of the visual extraction tasks; According to the target preset problem classification and the multiple test samples associated therewith, one or more of the candidate expert encoders are selected from the multiple candidate expert encoders as one or more target expert encoders corresponding to the target preset problem classification; wherein the target preset problem classification is any one of the multiple preset problem classifications; A plurality of the expert encoders are determined according to one or more of the target expert encoders corresponding to the plurality of the preset question classifications respectively.

5. The method according to claim 2, characterized in that: According to the plurality of visual features and the preset weight array, the material features of the material to be identified are obtained, including: Determining weighted features of the material to be identified according to the plurality of global features, the plurality of expert features and the preset weight array; The material feature is acquired according to the weighted feature and the local feature.

6. The method according to claim 5, characterized in that The preset weight array is obtained by the following operations: Acquire a preset training set, where the preset training set includes multiple training samples; Perform multiple rounds of training according to multiple training samples, and adjust the default weight array as a starting point until the preset weight array is obtained; In each training round: Extracting global training features from current training samples through the target encoder; Extracting expert training features from the current training sample by using a plurality of the expert encoders; According to the global training feature, the plurality of expert training features and the current weight array, the target sample feature of the current training sample is obtained; wherein the current weight array is the weight array obtained after the previous round of training; According to the target sample characteristics, obtaining the target recognition result of the current training sample; According to the target recognition result, the current weight array is adjusted to obtain the weight array after this round of training.

7. The method according to claim 1, characterized in that Obtaining the recognition result of the to-be-recognized material according to the material feature, including: The material features are converted into target material features that can be recognized by a target language model through a visual projector; The target material features are input into the target language model, and the recognition result is output through the target language model.

8. A material recognition device, characterized in that: The device comprises: A first acquisition module, used to acquire the material to be identified; An extraction module, used for extracting a plurality of visual features from the to-be-recognized material through a plurality of visual encoders, each visual encoder corresponding to one of the visual features; A second acquisition module, configured to acquire the material features of the material to be identified according to the plurality of visual features and a preset weight array; wherein the preset weight array includes a plurality of preset weights, each preset weight corresponding to one of the visual encoders; The third acquisition module is used to obtain the recognition result of the material to be recognized according to the material feature.

9. A computer device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.