A model performance optimization method, an inference method, a device, a computer program product and a storage medium

CN122840209APending Publication Date: 2026-09-29ALIBABA CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510361969.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]目前,视觉语言模型在进行推理的过程中,会产生大量的视觉词元(VisualToken),这导致推理过程中的计算复杂度不断攀升,进而导致视觉语言模型的推理性能不佳

Benefits of technology

[0020]在本申请实施例中,提出了一种模型性能优化方案,在目标模型进行推理的过程中,可对目标模型内的目标中间层进行监测,在目标中间层输出视觉表示的情况下,可对这些视觉表示所对应的视觉词元进行筛选处理,以过滤掉部分视觉词元。这样,在筛选处理环节中,可有效缩减视觉词元的数量。在此基础上,可针对筛选出的目标视觉词元继续进行视觉表示的合并处理,以得到合并后视觉表示。这样,在合并处理环节中,可进一步缩减视觉词元的数量,而且可充分保留视觉表示中对推理有用的信息。据此,通过筛选处理环节和合并处理环节的相互结合,可在目标模型中有效缩减视觉词元的数量,且可充分保留对推理有用的信息,减少信息浪费,从而可实现对推理效率和推理准确性的双重保障,进而有效提升目标模型的推理性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840209A_ABST
    Figure CN122840209A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a model performance optimization method, an inference method, a device, a computer program product and a storage medium. In the process of inference of a target model, in the case that visual representations are output in a target intermediate layer, the visual tokens corresponding to the visual representations are screened to filter out part of the visual tokens. In this way, in the screening process, the number of visual tokens can be effectively reduced. In addition, according to a preset retention rate, the screened target visual tokens are further subjected to a merging process of visual representations to obtain merged visual representations. In this way, in the merging process, the number of visual tokens can be further reduced, and useful information in the visual representations for inference can be fully retained, and information waste can be reduced. Accordingly, the inference efficiency and inference accuracy can be guaranteed, and the inference performance of the target model can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a model performance optimization method, inference method, device, computer program product, and storage medium. Background Technology

[0002] Visual Language Models (VLMs) have demonstrated superior capabilities across numerous fields. They integrate and process information from both visual and linguistic modalities, enabling models not only to handle text input, perform advanced reasoning, and generate natural language responses, but also to understand and process image input provided in prompts.

[0003] Currently, visual language models generate a large number of visual tokens during the reasoning process, which leads to a continuous increase in computational complexity and consequently poor reasoning performance. Summary of the Invention

[0004] This application provides a model performance optimization method, inference method, device, computer program product, and storage medium to improve the inference performance of an inference model.

[0005] This application provides a model performance optimization method, including:

[0006] During the reasoning process of the target model, when the target intermediate layer in the target model outputs a visual representation, target visual words that meet the preset filtering conditions are selected from the visual words corresponding to the visual representation output by the target intermediate layer.

[0007] The visual representations output by the target intermediate layer for the target visual lexical units are merged to obtain the merged visual representation;

[0008] The merged visual representation is then input into the next intermediate layer corresponding to the target intermediate layer.

[0009] This application also provides a reasoning method, the method comprising:

[0010] The inference task is received through a visual language model, and the inference task contains a visual file to be processed.

[0011] During the reasoning process of the visual language model for the reasoning task, the corresponding visual lexical units are determined for the visual representation output by the target intermediate layer in the visual language model.

[0012] From the identified visual words, select the target visual words that meet the preset selection criteria;

[0013] The visual representations output by the target intermediate layer for the target visual lexical units are merged to obtain the merged visual representation.

[0014] The merged visual representation is input into the next intermediate layer corresponding to the target intermediate layer in the visual language model, serving as the basis for the next intermediate layer to continue reasoning for the reasoning task.

[0015] This application also provides a computing device, including a memory, a processor, and a communication component;

[0016] The memory is used to store one or more computer instructions;

[0017] The processor is coupled to the memory and the communication component to execute one or more computer instructions for performing the aforementioned model performance optimization method.

[0018] This application also provides a computer-readable storage medium for storing a computer program, which, when executed by one or more processors, causes the one or more processors to perform the aforementioned model performance optimization method.

[0019] This application provides a computer program product, including a computer program that, when executed by one or more processors, causes the one or more processors to perform the aforementioned model performance optimization method.

[0020] This application proposes a model performance optimization scheme. During the inference process of the target model, the target intermediate layer within the target model can be monitored. When the target intermediate layer outputs visual representations, the visual lexical units corresponding to these visual representations can be filtered out to remove some visual lexical units. This filtering process effectively reduces the number of visual lexical units. Based on this, the filtered target visual lexical units can be further processed by merging visual representations to obtain merged visual representations. This merging process further reduces the number of visual lexical units while fully preserving information useful for inference in the visual representations. Therefore, by combining the filtering and merging processes, the number of visual lexical units in the target model can be effectively reduced while preserving information useful for inference, reducing information waste. This achieves a dual guarantee of inference efficiency and accuracy, thereby effectively improving the inference performance of the target model. Attached Figure Description

[0021] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0022] Figure 1 A flowchart illustrating a model performance optimization method provided for an exemplary embodiment of this application;

[0023] Figure 2 A logical schematic diagram of a model performance optimization method provided for an exemplary embodiment of this application;

[0024] Figure 3 A logical diagram illustrating an optional implementation of a merging process step provided for an exemplary embodiment of this application;

[0025] Figure 4 A flowchart illustrating a reasoning method provided for an exemplary embodiment of this application;

[0026] Figure 5 This is a schematic diagram of the structure of a computing device provided for another exemplary embodiment of this application. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0028] It should be noted that, in the cases involving user information in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) comply with the relevant laws and standards.

[0029] Before proceeding with a detailed description of the technical solutions provided in the various embodiments of this application, the following is a brief explanation of several technical concepts involved in this application.

[0030] Inference models can be understood as models used to handle complex problems that require multiple steps to solve. There are many types of inference models, and some inference models can support the processing of multimodal data and are often called multimodal models. A typical multimodal model is the Visual Language Model (VLM).

[0031] Visual language models are models that can learn from both images and text simultaneously to handle various reasoning tasks. Large visual language models have good zero-shot capability, good generalization ability, and can handle multimodal data, including images, videos, documents, and web pages.

[0032] Visual tokens can be understood as discretized, meaningful basic units obtained after processing visual data (such as images or videos). They transform visual data into a representation that is easy for models to process and understand, similar to converting words in text into tokens. Visual tokens are a type of input unit in inference models.

[0033] Intermediate layers in an inference model can be understood as the hierarchical structure between the input and output layers within the inference model. For example, the transformation layer in a Transformer model is an intermediate layer, and there are usually multiple intermediate layers in an inference model.

[0034] During their research, the inventors discovered that inference models with image processing capabilities, such as visual language models, generate a large number of visual words in response to visual data during the inference process. They also found that the computational complexity of the inference model is directly proportional to the number of visual words. Therefore, an excessive number of visual words leads to a continuous increase in the computational complexity of the inference model, and the memory space occupied by the model to process these visual words also increases, resulting in poor inference performance.

[0035] To address this, this application proposes a model performance optimization scheme. The basic idea is to reduce the number of visual lexical units that need to be processed in the inference model. By reducing the number of visual lexical units that need to be processed during the inference process, the computational complexity and memory space occupied in the inference model can be effectively improved, thereby effectively improving the inference performance of the inference model.

[0036] Based on this basic idea, how to reasonably reduce the number of visual lexical units that need to be processed in the reasoning model becomes the problem of concern in the embodiments of this application.

[0037] To address this problem, embodiments of this application provide a solution that aims to improve the reasoning performance of the reasoning model by reasonably reducing the number of visual lexical units that need to be processed in the reasoning model.

[0038] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0039] Figure 1 A flowchart illustrating a model performance optimization method provided as an exemplary embodiment of this application is shown below. Figure 2 This is a logical schematic diagram of a model performance optimization method provided for an exemplary embodiment of this application. The method can be executed by a model performance optimization device, which can be implemented as software, hardware, or a combination of software and hardware. (Reference) Figure 1 The method may include:

[0040] Step 100: During the reasoning process of the target model, when the target intermediate layer outputs a visual representation within the target model, select target visual words that meet the preset selection conditions from the visual words corresponding to the visual representations output by the target intermediate layer.

[0041] Step 101: Merge the visual representations output by the target intermediate layer for the target visual lexical units to obtain the merged visual representation;

[0042] Step 102: Input the merged visual representation into the next intermediate layer corresponding to the target intermediate layer.

[0043] In this embodiment, the deployment method of the model performance optimization device is not limited. For example, the model performance optimization device can be deployed inside the target model in the form of a plug-in. Alternatively, the model performance optimization device can also be used as an external component of the target model. Further examples of deployment methods are not provided here.

[0044] In this embodiment, the target model can be any reasoning model with visual processing capabilities, such as the aforementioned visual language model. There are no restrictions on the model type or model name of the target model.

[0045] refer to Figure 1 In step 100, the target intermediate layer within the target model can be monitored during the inference process of the target model. The target intermediate layer can be any intermediate layer within the target model. In this embodiment, one or more intermediate layers can be specified within the target model as the target intermediate layer as needed.

[0046] In this embodiment, during the execution of the inference task by the target model, the target intermediate layer within the target model is monitored, mainly to monitor whether the target intermediate layer has completed the inference calculation.

[0047] Considering that this embodiment mainly focuses on visual processing, step 100 proposes that the filtering process proposed in this embodiment can be initiated when the visual representation is output by the target intermediate layer within the target model.

[0048] Each intermediate layer generates inference data after performing inference computation, which can be understood as an intermediate result within the inference model. The intermediate results generated by the intermediate layers are passed to the next intermediate layer as the basis for inference computation. Here, visual representation can be understood as the visual feature representation contained in the intermediate results generated by the intermediate layers.

[0049] It is worth emphasizing that visual lexical units and visual representations are different concepts in this embodiment. Visual lexical units can be understood as an input unit of the inference model, which is the basis for the inference model to perform visual processing. Visual features, on the other hand, can be understood as an intermediate result generated by the inference model in the process of processing visual lexical units, or as the product of processing visual lexical units, representing the inference model's understanding and memory of visual lexical units at a certain moment or level.

[0050] In this embodiment, different types of inference models may use different names for the visual features generated in the intermediate layers. Therefore, the specific name for visual features in the inference model is not limited here. For example, in the aforementioned Transformer model, the hidden state output by the intermediate layer is a typical visual feature. Here, the concept of hidden state is explained to help understand the definition of visual features in this embodiment. A hidden state is a state vector generated by the inference model during the inference process. It is mainly used to store and pass information between different time steps or levels of the inference model. When processing sequential data or data with contextual relationships, the hidden state can capture the temporal dependencies and contextual information in the data, allowing the inference model to consider previous information when processing the current input. For example, in natural language processing, it helps the model understand the semantic relationships between words in a sentence. Similarly, in visual processing, it can help understand the associations between visual words.

[0051] Based on this, it can be understood that after receiving the inference task, the target model will enter the inference process. During the inference process, the data to be processed (including visual data) first enters the input layer of the target model. Next, the intermediate layers of the target model will successively perform complex inference calculations. Finally, the output layer of the target model can output the inference result.

[0052] The following uses the Transformer model as an example to briefly introduce the inference process:

[0053] Input stage: Taking the need to perform visual processing on images as an example, the model first performs operations such as segmenting the image, and then converts these image blocks into visual words through embedding and other methods. These visual words are provided as input to the input layer of the model, providing the model with the original image feature information.

[0054] Intermediate Layer Processing Stage: The intermediate layers of the model mainly consist of multiple Transformer layers. Within each Transformer layer, self-attention mechanisms and feedforward neural networks are used to process the visual features output from the previous layer, further extracting features and fusing information. For example, in a multi-intermediate-layer Transformer model, the first Transformer layer receives visual words as input and calculates the corresponding visual representations. From the second Transformer layer onwards, the input is the visual representation output by the first Transformer layer, and subsequent Transformer layers follow the same pattern. Each Transformer layer processes the visual representation output by the previous one, continuously updating and enriching the visual representations, and uncovering more complex contextual relationships between visual words.

[0055] Information Transfer and Update: As the intermediate layers within the model continue processing, the visual representation becomes increasingly rich in information. It goes beyond the initial simple representation of visual lexical units, incorporating the correlation information between different visual lexical units and global image information. The visual representation output by each Transform layer is constantly updated and evolved, passing higher-level feature information to the next layer, until the final output layer performs the final task prediction based on the final visual representation, producing the final inference result.

[0056] Based on this, refer to Figure 2 In step 100, by monitoring the target intermediate layer within the target model during the inference process, the visual lexical units corresponding to the visual representation output by the target intermediate layer can be determined. That is, it can be determined which visual lexical units the target intermediate layer performed inference calculations on. It is worth noting that in this embodiment, if one target intermediate layer is set in the target model, the visual lexical units determined by this target intermediate layer can be all the visual lexical units converted in the aforementioned input stage; if multiple target intermediate layers are set in the target model, the visual lexical units determined by this target intermediate layer may be some of the visual lexical units converted in the aforementioned input stage, that is, they may be the remaining visual lexical units after being reduced by the previous target intermediate layer according to the concept of this embodiment.

[0057] As mentioned earlier, in this embodiment, the filtering process proposed in this embodiment can be initiated when the target intermediate layer outputs a visual representation within the target model. The filtering process proposed in this embodiment will be described below.

[0058] In the filtering process: target visual words that meet the preset filtering conditions can be filtered from the visual words corresponding to the visual representation output by the target intermediate layer.

[0059] It should be understood that the basic idea in this processing step is to directly filter out some visual words processed by the target intermediate layer. In this embodiment, the above-mentioned preset filtering conditions can be flexibly designed as needed. If it is desired to further reduce the number of visual words, more visual words can be filtered out in this step by designing preset filtering conditions; if it is desired to retain more information, a small number of visual words can be filtered out in this step by designing preset filtering conditions. No limitation is made to the preset filtering conditions here. Furthermore, if multiple target intermediate layers are set in the target model, the preset filtering conditions corresponding to different target intermediate layers can be the same or different; this embodiment does not limit this either.

[0060] An exemplary preset screening condition could be: the attention score falls within a retention range determined according to the screening ratio. Based on this exemplary preset screening condition, an optional implementation of the screening process in this embodiment could be:

[0061] Based on the required filtering ratio of the target intermediate layer, determine the number M of target visual words to be filtered out, where M is an integer greater than 1; from the visual words corresponding to the visual representation output by the target intermediate layer, select the M visual words with the highest attention scores in the target intermediate layer as target visual words.

[0062] The attention score corresponding to a visual word can be understood as the relative importance of that visual word to other visual words, as evaluated in the intermediate layer of the target model. The attention score is a type of output information that can be generated in the intermediate layer of the inference model. Typically, in the intermediate results produced by the intermediate layer, the attention score of a visual word and its corresponding visual representation are correlated. Here, the specific calculation method for the attention score in the intermediate layer of the target model is not limited or elaborated upon.

[0063] Based on this, in this optional implementation, the attention scores generated for each visual word processed by the target intermediate layer can be queried and used as the filtering basis in the filtering process.

[0064] In this optional implementation, the visual words processed by the target intermediate layer can be sorted in descending order according to the attention scores generated in the target intermediate layer; from the sorted visual words, the top M visual words are selected as the target visual words in the filtering process.

[0065] Of course, other implementation methods can also be used in the filtering process proposed in this embodiment. For example, the visual position corresponding to the visual word can be used as the filtering basis. Specifically, visual words that are close to the visual center can be retained first, while visual words that are far from the visual center (such as those located at the visual edge) can be filtered out. The implementation method of the filtering process is not limited here, nor are further examples provided.

[0066] Thus, the filtering process proposed in this embodiment can directly filter out some visual words in the target model, thereby effectively reducing the number of visual words that need to be processed by subsequent intermediate layers in the target model, and thus effectively reducing the computational complexity of subsequent intermediate layers. Furthermore, it can prioritize the retention of visual words useful for subsequent inference calculations, thereby ensuring that sufficient information is provided for subsequent inference calculations.

[0067] refer to Figure 1 and Figure 2 Furthermore, the technical concept of this embodiment proposes that a merging process can be performed after the above-mentioned screening process. The merging process in this embodiment will be described below.

[0068] In the merging process, as described in step 101, the visual representations output by the target intermediate layer for the target visual lexical units can be merged to obtain the merged visual representation.

[0069] Here, the merging process can be performed according to the required retention rate. The retention rate can be understood as the proportion of the number of visual words that need to be retained in the target visual words selected in the filtering process. In this embodiment, it is possible to set the required retention rate for the target intermediate layer as needed. Moreover, if there are multiple target intermediate layers in the target model, the retention rates set for different target intermediate layers can be the same or different, which is not limited here.

[0070] For example, if 100 target visual words are selected in the screening process, 50 visual words can be retained if the retention rate is set at 50% in the merging process; and 70 visual words can be retained if the retention rate is 70%.

[0071] As can be seen from the above examples, in the merging process proposed in this embodiment, the number of visual lexical units in the target model can be further reduced.

[0072] In this embodiment, unlike the aforementioned screening process, the merging process does not directly filter out some target visual words. Instead, it merges the visual representations corresponding to the target visual words to reduce the number of visual words.

[0073] It is understood that in this embodiment, the merging process is equivalent to fusing the information carried in the visual representations corresponding to some target visual words into the visual representations corresponding to the retained visual words. Thus, the visual representations corresponding to the retained visual words are substantially enriched with information extracted from other target visual representations through the merging process.

[0074] Therefore, in this embodiment, the number of visual lexical units that need to be processed by subsequent intermediate layers in the target model can be further reduced during the merging process, thereby further reducing the computational complexity of the subsequent intermediate layers. More importantly, the information carried in the visual representations corresponding to the discarded target visual lexical units is fully preserved during the merging process, which can effectively improve the inference accuracy of subsequent intermediate layers.

[0075] In this embodiment, the visual representation generated by the merging process in the merging process is described as the merged visual representation.

[0076] Based on this, refer to Figure 1 In step 102, the merged visual representation generated in the merging process can be input into the next intermediate layer corresponding to the target intermediate layer. This eliminates the need to pass all the visual representations output by the target intermediate layer backward.

[0077] It is worth emphasizing here that in this embodiment, during the merging process, some target visual units may not have been merged. To address this, step 102 of this embodiment further proposes:

[0078] If there are unmerged words in the target visual vocabulary that have not undergone merging processing, then the visual representation of the unmerged words output by the target intermediate layer is input into the next intermediate layer corresponding to the target intermediate layer. That is, unmerged words can be directly retained in the merging process.

[0079] In this embodiment, unmerged lexical units are not merged with other target visual lexical units. Therefore, the visual representations corresponding to unmerged lexical units in the target intermediate layer have not yet been integrated into the visual representations corresponding to other target visual lexical units. By directly passing the visual representations corresponding to unmerged lexical units in the target intermediate layer to the next intermediate layer, the information carried in the visual representations corresponding to unmerged lexical units can be effectively preserved, thereby reducing information waste and ensuring the accuracy of inference in subsequent intermediate layers. Of course, this is optional; in step 102, only the merged visual representation can be passed to the next intermediate layer. This embodiment does not limit this. Regardless of whether unmerged lexical units are retained, in step 102, the number of visual representations (including merged visual representations and, possibly, the visual representations corresponding to unmerged lexical units in the target intermediate layer) passed to the next intermediate layer must meet the required retention rate.

[0080] Clearly, in this embodiment, the number of visual words that the next intermediate layer needs to process in step 102 is effectively reduced, and the number of visual words that the next intermediate layer needs to process is lower than the number of the target visual words. Therefore, the reasoning complexity in the next intermediate layer can be effectively reduced. Moreover, the visual representation received in the next intermediate layer fully retains the information of the visual representation of the target visual words in the target intermediate layer, thus effectively ensuring the reasoning accuracy in the next intermediate layer.

[0081] In summary, this embodiment proposes a model performance optimization scheme. During the inference process of the target model, the target intermediate layer within the target model can be monitored. When the target intermediate layer outputs visual representations, the visual lexical units corresponding to these visual representations can be filtered out to remove some visual lexical units. This filtering process effectively reduces the number of visual lexical units. Based on this, visual representations can be merged for the filtered target visual lexical units to obtain merged visual representations. This merging process further reduces the number of visual lexical units while fully preserving information useful for inference from the visual representations. Therefore, by combining the filtering and merging processes, the number of visual lexical units in the target model can be effectively reduced while preserving information useful for inference, reducing information waste. This achieves a dual guarantee of inference efficiency and accuracy, thereby effectively improving the inference performance of the target model.

[0082] In the above or below embodiments, the merging process proposed in this embodiment can be implemented in various ways. The following provides an optional implementation method. Figure 3 This is a logical diagram illustrating an optional implementation of a merging process step provided for an exemplary embodiment of this application.

[0083] refer to Figure 3 In this optional implementation, it is proposed that: the target visual lexical can be divided into at least two visual lexical groups; visual lexical pairing is performed between at least two visual lexical groups to obtain multiple pairs of lexical to be merged; based on the visual representation output by the target intermediate layer for the target visual lexical, the visual representations of the multiple pairs of lexical to be merged are merged respectively to obtain the merged visual representations corresponding to each of the multiple pairs of lexical to be merged.

[0084] In this optional implementation, a pairing mechanism is proposed. The target visual lexical units (usually multiple) retained in the filtering process are divided into at least two visual lexical unit groups. That is, the multiple target visual lexical units retained in the filtering process can be distributed within two visual lexical unit groups. In this optional implementation, visual lexical unit pairing can also be performed between the at least two visual lexical unit groups. This visual lexical unit pairing can be understood as selecting one visual lexical unit from each of the at least two visual lexical unit groups in a pairing process; the selected at least two visual lexical units can then form a pair of lexical units to be merged.

[0085] Thus, in this optional implementation, multiple pairs of tokens to be merged can be generated through a pairing mechanism. Based on this, visual representation merging can be performed on each pair. From the perspective of a single pair of tokens to be merged, the visual representations corresponding to at least two visual tokens contained in that pair in the target intermediate layer can be merged to obtain the merged visual representation corresponding to that pair. Therefore, in this optional implementation, a single pair of tokens to be merged, after merging, can produce a merged visual representation.

[0086] Preferably, in this optional implementation, it is proposed that at least two visual lexical groups contain the same number of visual lexical elements through an equal division operation.

[0087] By performing an equal distribution operation, the number of visual lemmas in at least two visual lemmas can be more balanced, allowing more target visual lemmas to be paired, thereby supporting a lower retention rate and ensuring that the merging process can flexibly adapt to various retention rates.

[0088] Based on this, an exemplary grouping scheme could be: sorting the target visual words according to the attention scores corresponding to the target visual words in the target intermediate layer; and performing an equal division operation on the sorted target visual words to obtain at least two groups of visual words.

[0089] For ease of explanation, the following example uses two visual word groups obtained after grouping, and these two visual word groups are described as the first word group and the second word group, respectively. In this exemplary grouping scheme, attention scores can be used as the grouping criterion. After sorting the target visual words according to their corresponding attention scores in the target intermediate layer, an equal distribution operation is performed. In this way, the number of visual words in the resulting first word group and second word group is more balanced. Moreover, target visual words with lower attention scores can be concentrated in a certain divided visual word group, such as the second word group. In other words, the attention scores of the visual words contained in the first word group in the intermediate layer can all be higher than the attention scores of the visual words contained in the second word group in the intermediate layer.

[0090] By dividing the target visual units equally according to their attention scores, they can be differentiated into two categories: target visual units that are more useful for subsequent inference calculations and target visual units that have less impact on subsequent inference calculations. This classification effect allows for prioritizing the retention of information carried in the visual representations of target visual units that are more useful for subsequent inference calculations in the target intermediate layer during the merging process, thereby minimizing the loss of this information and better ensuring the accuracy of subsequent inference calculations.

[0091] Building upon this, the implementation further proposes that the second word group can be considered as the visual word group to be reduced. If any visual word in the second word group is paired with a visual word in the first word group, then that visual word in the second word group can be identified as the word to be discarded. In the merging process, after merging the word pair to be merged containing the word to be discarded, the word to be discarded can be discarded, while retaining the visual word paired with it.

[0092] Under this concept, in the merging process, target visual words with lower attention scores can be merged into target visual words with higher attention scores, thereby prioritizing the retention of target visual words with higher attention scores and fully preserving the information carried by the visual representation of the target visual words with lower attention scores in the target intermediate layer.

[0093] Furthermore, this optional implementation does not limit the specific pairing dimension used to construct multiple pairs of tokens to be merged. An exemplary pairing scheme is proposed below:

[0094] Similar visual lemmas can be paired between two visual lemmas to obtain multiple lemma pairs to be merged.

[0095] This exemplary pairing scheme proposes to pair visual morphemes between two visual morpheme groups based on similarity as a pairing dimension. According to this pairing dimension, the two target visual morphemes in the resulting morpheme pairs to be merged exhibit good similarity, which allows for efficient fusion of similar information in the visual representation, thereby effectively reducing the amount of information in the merged visual representation.

[0096] For ease of explanation, we will continue to use the first and second word groups mentioned above to illustrate a solution for pairing similar visual word groups:

[0097] From the visual lexical units contained in the first lexical unit, similar visual lexical units can be found for the first visual lexical unit in the second lexical unit, where the first visual lexical unit is any visual lexical unit contained in the second lexical unit.

[0098] Combine the first visual word and the found visual word into a tuple;

[0099] According to the required retention rate, the tuples formed under the second word tuple are filtered to obtain multiple word pairs to be merged.

[0100] In this solution, similar visual words in the first lexical group can be found for each visual word in the second lexical group, thus forming binary pairs for each visual word in the second lexical group. Furthermore, the binary pairs formed under the second lexical group can be filtered according to the desired retention rate to obtain multiple pairs of word words to be merged. One proposed binary pair filtering scheme is as follows:

[0101] The number of pairs N to be retained can be determined according to the required retention rate, where N is an integer greater than 1; from the pairs formed under the second word pair, the N pairs with the highest word similarity are selected as word pairs to be merged.

[0102] In this screening scheme, the similarity between visual words within a tuple is used as the screening criterion. This allows for the selection of tuples with higher similarity between their internal visual words as the word pairs to be merged. This enables the merging of visual words with higher similarity to be prioritized during the merging process, thereby better reducing the amount of information in the merged visual representation and effectively reducing the reasoning complexity in subsequent intermediate layers.

[0103] It's worth noting that when determining the number N of tuples to be retained, one can consider whether to retain the unmerged lexical units mentioned earlier. Regardless of whether the unmerged lexical units mentioned earlier need to be retained, the tuple filtering ensures that the number of visual representations passed to the next intermediate layer in step 102 (which can be N, or N + the number of unmerged lexical units) meets the required retention rate. Here, examples of the number of tuples to be retained under different circumstances are not provided.

[0104] It is understandable here that, based on the set retention rate, in this solution, not all tuples formed under the second tuple will necessarily be selected as tuple pairs to be merged; some tuples may be chosen. For example, when grouping in an even distribution manner, if the retention rate is 50%, all tuples formed under the second tuple can be selected as tuple pairs to be merged. As another example, when grouping in an even distribution manner, if the retention rate is higher than 50%, some tuples will not be selected. In this case, the visual tuples contained in the unselected tuples are the unmerged tuples mentioned earlier. In the preferred embodiment, the visual representations corresponding to these unmerged tuples in the target intermediate layer can be directly transferred to the next intermediate layer corresponding to the target intermediate layer.

[0105] In addition, when similarity is used as the pairing dimension, an exemplary calculation scheme for calculating the similarity between visual lexical units is proposed: the visual representation similarity between the visual representations corresponding to visual lexical units can be calculated as the similarity between visual lexical units.

[0106] Following the above example, the solution for matching similar visual lexical units can be implemented in the following exemplary calculation scheme: based on the visual representation output by the target intermediate layer for the target visual lexical unit, the visual lexical unit with the highest visual representation similarity to the first visual lexical unit can be found from the visual lexical units contained in the first lexical unit group, and used as the visual lexical unit that matches the first visual lexical unit.

[0107] In this exemplary calculation scheme, visual representation is used as the basis for evaluating the similarity between visual words. This effectively connects to the aforementioned screening process. By using the visual representations corresponding to the target visual words already retained in the screening process, the similarity can be calculated without referencing other parameters. This effectively improves the convenience of the merging process in this embodiment, enabling more efficient visual word pairing and thus improving the merging efficiency in the merging process.

[0108] Of course, besides the similarity pairing dimension mentioned above, other pairing dimensions can also be used in this exemplary pairing scheme. For example, visual distance between visual words. Under this pairing dimension, target visual words that are visually closer can be prioritized for pairing. Regarding the grouping scheme, the exemplary grouping scheme mentioned above can be used. For specific pairing schemes under the visual distance pairing dimension, the relevant technical details can be adapted to the above description and will not be elaborated here. No further examples of pairing dimensions will be provided here, and the pairing dimensions that can be selected in this embodiment are not limited to these.

[0109] In this optional implementation, various schemes can be used to visually merge the multiple word units to be merged. An exemplary visual representation merging scheme could be:

[0110] For any pair of lexical units to be merged, the attention scores of the two visual lexical units in the target intermediate layer are used as weights to perform weighted fusion of the visual representations of the two visual lexical units in the intermediate layer, so as to obtain the merged visual representation of the pair of lexical units to be merged.

[0111] This exemplary visual representation merging scheme employs a weighted fusion mechanism. The attention scores of the two visual words in the merging pair are used as weights for weighted fusion, ensuring that the merged visual representation retains sufficient information from the visual representations of the two visual words in the merging pair. This weighted fusion mechanism guarantees that the merged visual representation retains enough information, effectively ensuring the accuracy of inference in subsequent intermediate layers.

[0112] In summary, this embodiment provides an optional implementation method for the merging process. By pairing grouped visual words, multiple pairs of words to be merged can be formed according to the required retention rate. Moreover, using the similarity between visual words as the pairing dimension allows similar visual words to be preferentially combined. Therefore, in the merging process, the visual representations corresponding to similar visual words in the target intermediate layer can be merged first, effectively reducing the amount of information in the merged visual representation, thereby reducing the inference complexity in subsequent intermediate layers and improving the inference efficiency of the target model. Furthermore, a weighted fusion mechanism is proposed during the merging of pairs of words to be merged. This mechanism fully considers the attention scores of visual words in the target intermediate layer, thus more rationally planning the proportion of information retained by the two visual words in the weighted fusion process. This prioritizes retaining information more useful for subsequent inference calculations and minimizes information waste, effectively improving the inference accuracy of the target model.

[0113] Of course, besides the pairing method after grouping described above, other implementation methods can be used in the merging process of this embodiment. For example, pairing can be done randomly without grouping, based on the pairing dimensions such as similarity or visual distance. Another example is keeping the visual words in the first word group unchanged, while combining visual words in the second word group. In this way, the visual words in the first word group can be retained as the aforementioned unmerged words, while the visual representations of the visual words in the second word group corresponding to the target intermediate layer can be merged. Regarding the grouping criteria, in addition to the attention score described above, random grouping can also be used. Here, the implementation methods that can be used in the merging process of this embodiment are not limited, nor are further examples provided.

[0114] The following example illustrates an application of the model performance optimization method proposed in this embodiment. In this application, the Transformer model is used as the target model in this embodiment.

[0115] I. In this application scheme, the following parameters can be defined first:

[0116] K: Set the Kth intermediate layer (Transformer layer) of the Transformer model as the target intermediate layer.

[0117] attn_score: Represents the attention score obtained for each visual token processed in the Kth Transformer layer.

[0118] total: The total number of visual lexical units before compression.

[0119] n: The proportion of visual lexical units retained after filtering to the number of visual lexical units before compression.

[0120] t: The proportion of visual words retained after merging to the number of visual words before compression. This can be required to... (In cases where the number of unmerged tokens is greater than a certain threshold, some unmerged tokens may be directly retained.)

[0121] Cosine similarity: A method for calculating the similarity between two vectors. In this application, the cosine similarity between the visual representations (which can be hidden states) corresponding to two visual words can be calculated using the following formula:

[0122]

[0123] Here, low_token and high_token represent the hidden states corresponding to the Kth intermediate layer of the two visual tokens, respectively.

[0124] II. Calculation Process

[0125] This application scheme can occur during each round of inference. After the inference computation is completed in the Kth Transformer layer, the visual lexical units corresponding to the hidden state output by that Transformer layer are compressed as follows:

[0126] Screening and processing steps:

[0127] Based on attn_score, the visual lexical units of the Transformer layer are sorted, and the top n visual lexical units with the highest scores are retained. A total of n×total visual lexical units are retained.

[0128] Merging process:

[0129] The remaining visual tokens are sorted according to their attn_score and divided into two sets: low_attn_tokens and high_attn_tokens. The low_attn_tokens set contains multiple sets of visual tokens with low attention scores (low tokens); the high_attn_tokens set contains multiple sets of visual tokens with high attention scores (high tokens).

[0130] For any low token in the low_attn_tokens set, calculate the cosine similarity between it and each high token in the high_attn_tokens set, and find the high token that is most similar to it. Based on this, each low token and its most similar high token can form a pair, and such pairs can be obtained in total. Of course, multiple low tokens can be paired with the same high token.

[0131] Among these pairs, select the t×total pairs with the highest similarity scores. Then, merge the two visual words in each of these pairs.

[0132] In this application scheme, the hidden states corresponding to two visual lexical units in the Kth Transformer layer can be weighted and fused according to attn_score. The specific formula for calculating the merged visual features can be:

[0133]

[0134] Where low_attn and high_attn are the attention scores obtained by the two visual words in the Kth intermediate layer, respectively; low_token and high_token represent the hidden states corresponding to the two visual words in the Kth layer.

[0135] Ultimately, the number of visual lexical units retained after compression will be t×total.

[0136] As demonstrated by the above application scheme, the model performance optimization method provided in this embodiment can compress visual words in the target model through the coordinated use of a filtering process and a merging process. First, based on attention scores, visual words are filtered at a specific layer, discarding unimportant ones. Then, a merging process is used, with attention scores as weights for weighted merging. This reduces the space occupied by inference while ensuring stable inference performance, thus effectively improving inference efficiency. Furthermore, it better preserves the information required for inference computation, reducing information waste and effectively ensuring inference accuracy.

[0137] Figure 4 This is a flowchart illustrating a reasoning method provided as another exemplary embodiment of this application. In this reasoning method, a visual language model may be responsible for handling the reasoning task. (Reference) Figure 4 The reasoning method may include:

[0138] Step 400: Receive a reasoning task through a visual language model, wherein the reasoning task contains a visual file to be processed;

[0139] Step 401: During the reasoning process of the visual language model for the reasoning task, determine the corresponding visual lexical units for the visual representation output by the target intermediate layer in the visual language model.

[0140] Step 402: Select target visual words that meet the preset filtering conditions from the identified visual words.

[0141] Step 403: Merge the visual representations output by the target intermediate layer for the target visual lexical units to obtain the merged visual representation;

[0142] Step 404: Input the merged visual representation into the next intermediate layer corresponding to the target intermediate layer in the visual language model, as the basis for the next intermediate layer to continue reasoning for the reasoning task.

[0143] The reasoning principles of each intermediate layer in the visual language model will not be elaborated here. In this embodiment, existing reasoning principles can be used or new reasoning principles that may emerge in the future can be adopted adaptively.

[0144] In this embodiment, in step 400, the visual language model can receive an inference task, which includes a visual file to be processed. This embodiment supports various types of visual files, such as video files and image files, etc., and is not limited here. In addition, this embodiment is not limited to the specific content of the inference task. For example, it can be object detection, change detection, etc., and no further examples of the task content of the inference task will be given here.

[0145] As mentioned in the conceptual introduction section, in a visual language model, the intermediate layer is responsible for performing inference and can output visual representations of the visual lexical units corresponding to the inference task. From the perspective of the target intermediate layer, its input side may receive visual representations of all or part of the visual lexical units corresponding to the aforementioned video file, depending on the visual representation output by the intermediate layer preceding the target intermediate layer. If the visual lexical unit compression step in this embodiment has already been performed in an intermediate layer preceding the target intermediate layer, then the input side of the target intermediate layer may receive visual representations of part of the visual lexical units corresponding to the visual file.

[0146] Continue to refer to Figure 4 In step 401, visual lexical units can be determined for each visual representation output by the target intermediate layer. It is understood that the visual lexical units determined here are generated by the visual language model for the visual file in step 400. Their source may be the input layer of the visual language model (divided for the original visual file), or they may originate from new visual lexical units generated for the visual file during the inference process of the visual language model (for example, if autoregressive inference is used, new visual lexical units may be generated during the inference process). This embodiment does not limit the timing of visual lexical unit generation.

[0147] refer to Figure 4 Steps 402-403 constitute the visual lexical compression step proposed in this embodiment. For details regarding the visual lexical compression step, please refer to the relevant descriptions in the model optimization methods provided in the preceding embodiments; they will not be repeated here.

[0148] In summary, in this embodiment, after receiving the inference task through the visual language model, a visual lexical compression step can be applied during the inference process. Based on this visual lexical compression step, the number of visual representations input to the next intermediate layer corresponding to the target intermediate layer can be effectively reduced. Simultaneously, through a design concept that assists in filtering and merging, the information retention rate can be fully guaranteed, ensuring that the retained visual representations contain sufficient information to support the accuracy of the next intermediate layer's inference. Therefore, improvements can be achieved simultaneously in both inference efficiency and inference accuracy, effectively enhancing the inference performance of the visual language model.

[0149] It should be noted that some processes described in the above embodiments and accompanying drawings include multiple operations appearing in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear in this document, or they may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should also be noted that the descriptions such as "first" and "second" in this document are used to distinguish different tuples, and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0150] Figure 5 This is a schematic diagram of a computing device provided as another exemplary embodiment of this application. The aforementioned model performance optimization apparatus for performing a model performance optimization method can be integrated into this computing device. Figure 5 As shown, the computing device may include: memory 50 and processor 51.

[0151] The memory 50 is used to store the computer program corresponding to the above-mentioned model performance optimization device.

[0152] Processor 51, coupled to memory 50, is used to execute computer programs in memory 50 for:

[0153] During the reasoning process of the target model, when the target intermediate layer in the target model outputs a visual representation, target visual words that meet the preset filtering conditions are selected from the visual words corresponding to the visual representation output by the target intermediate layer.

[0154] The visual representations output by the target intermediate layer for the target visual lexical units are merged to obtain the merged visual representation;

[0155] The merged visual representation is then input into the next intermediate layer corresponding to the target intermediate layer.

[0156] In an optional embodiment, when the processor 51 performs merging processing on the visual representations output by the target intermediate layer for the target visual lexical units to obtain the merged visual representation, it may specifically be used to:

[0157] The target visual word unit is divided into at least two visual word unit groups;

[0158] Visual lexical pairing is performed between the at least two visual lexical groups to obtain multiple lexical pairs to be merged;

[0159] Based on the visual representation output by the target intermediate layer for the target visual lexical unit, the visual representations of the multiple lexical units to be merged are merged respectively to obtain the merged visual representations corresponding to the multiple lexical units to be merged.

[0160] In an optional embodiment, when the processor 51 performs visual lexical pairing between the at least two visual lexical groups to obtain the plurality of lexical pairs to be merged, it may specifically be used to:

[0161] Similar visual lemmas are paired between the two visual lemmas to obtain multiple lemma pairs to be merged.

[0162] In an optional embodiment, when the processor 51 performs visual representation merging on the plurality of word pairs to be merged to obtain the merged visual representations corresponding to each of the plurality of word pairs to be merged, it may specifically be used to:

[0163] For any pair of lexical units to be merged, the attention scores of the at least two visual lexical units in the target intermediate layer are used as weights to perform weighted fusion on the visual representations of the two visual lexical units in the target intermediate layer, so as to obtain the merged visual representation of the pair of lexical units to be merged.

[0164] In an optional embodiment, when dividing the target visual lexical into at least two visual lexical groups, the processor 51 may specifically be used for:

[0165] The target visual words are sorted according to the attention scores they correspond to in the target intermediate layer.

[0166] Perform an equal-division operation on the sorted target visual lexical units to obtain the at least two visual lexical units.

[0167] In an optional embodiment, the at least two visual lexical groups include a first lexical group and a second lexical group; when the processor 51 performs similar visual lexical pairing between the at least two visual lexical groups to obtain multiple lexical pairs to be merged, it can be specifically used for:

[0168] From the visual lexical units contained in the first lexical unit, find similar visual lexical units for the first visual lexical unit in the second lexical unit, wherein the first visual lexical unit is any visual lexical unit contained in the second lexical unit;

[0169] The first visual word and the found visual word are combined into a tuple;

[0170] According to the required retention rate, the tuples formed under the second word group are filtered to obtain the multiple word pairs to be merged.

[0171] In an optional embodiment, the attention scores of the visual lexical units contained in the first lexical unit group in the target intermediate layer are all higher than the attention scores of the visual lexical units contained in the second lexical unit group in the target intermediate layer; the retention rate is higher than or equal to 50%.

[0172] In an optional embodiment, when the processor 51 searches for similar visual lexical units in the second lexical unit from the visual lexical units contained in the first lexical unit, it may specifically be used to:

[0173] Based on the visual representation output by the target intermediate layer for the target visual lexical unit, the visual lexical unit with the highest visual representation similarity to the first visual lexical unit is found from the visual lexical units contained in the first lexical unit group, and is used as the visual lexical unit that matches the first visual lexical unit.

[0174] In an optional embodiment, when the processor 51 filters the tuples formed under the second word group according to the required retention rate to obtain the plurality of word pairs to be merged, it may specifically be used to:

[0175] Determine the number of tuples N to be retained based on the required retention rate and the number of target visual lexical units, where N is an integer greater than 1;

[0176] From the pairs of words formed under the second word group, select the N pairs with the highest visual representation similarity to obtain the multiple word pairs to be merged.

[0177] In an alternative embodiment, processor 51 may also be used for:

[0178] If there are unmerged words in the target visual words that have not been merged, then the visual representation of the unmerged words output by the target intermediate layer is input into the next intermediate layer corresponding to the target intermediate layer.

[0179] In an optional embodiment, when the processor 51 filters out target visual words that meet preset filtering conditions from the visual words corresponding to the visual representation output from the target intermediate layer, it may specifically be used to:

[0180] According to the required filtering ratio of the target intermediate layer, determine the number M of target visual words to be filtered out, where M is an integer greater than 1;

[0181] From the visual lexical units corresponding to the visual representation output by the target intermediate layer, select the M visual lexical units with the highest attention scores in the target intermediate layer as the target visual lexical units.

[0182] Figure 5 The computing device shown can also be used to execute the aforementioned reasoning method. In this case, the processor 51 in the computing device can be used to:

[0183] The inference task is received through a visual language model, and the inference task contains a visual file to be processed.

[0184] During the reasoning process of the visual language model for the reasoning task, the corresponding visual lexical units are determined for the visual representation output by the target intermediate layer in the visual language model.

[0185] From the identified visual words, select the target visual words that meet the preset selection criteria;

[0186] The visual representations output by the target intermediate layer for the target visual lexical units are merged to obtain the merged visual representation.

[0187] The merged visual representation is input into the next intermediate layer corresponding to the target intermediate layer in the visual language model, serving as the basis for the next intermediate layer to continue reasoning for the reasoning task.

[0188] Furthermore, such as Figure 5 As shown, the computing device also includes other components such as a communication component 52 and a power supply component 53. Figure 5 The diagram only shows some components and does not mean that the computing device includes only these components. Figure 5 The components shown.

[0189] It is worth noting that the technical details of the above embodiments of the computer device can be referred to the relevant descriptions in the foregoing method embodiments. To save space, they will not be repeated here, but this should not cause any loss to the scope of protection of this application.

[0190] The aforementioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0191] The aforementioned communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.

[0192] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.

[0193] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium may be volatile, non-volatile, or a combination thereof, and may be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, digital video disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium.

[0194] Accordingly, this application also provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the processor is able to implement the steps in the above method embodiments. It should be understood that each step or combination of steps in the above method flow can be implemented by the computer program or instructions. In addition, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, so that the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device can be implemented as a means to implement the corresponding functions in the above method embodiments.

[0195] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0196] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for optimizing model performance, characterized in that, include: During the reasoning process of the target model, when the target intermediate layer in the target model outputs a visual representation, target visual words that meet the preset filtering conditions are selected from the visual words corresponding to the visual representation output by the target intermediate layer. The visual representations output by the target intermediate layer for the target visual lexical units are merged to obtain the merged visual representation; The merged visual representation is then input into the next intermediate layer corresponding to the target intermediate layer.

2. The method according to claim 1, characterized in that, The visual representations output by the target intermediate layer for the target visual lexical units are merged to obtain a merged visual representation, including: The target visual word unit is divided into at least two visual word unit groups; Visual lexical pairing is performed between the at least two visual lexical groups to obtain multiple lexical pairs to be merged; Based on the visual representation output by the target intermediate layer for the target visual lexical unit, the visual representations of the multiple lexical units to be merged are merged respectively to obtain the merged visual representations corresponding to the multiple lexical units to be merged.

3. The method according to claim 2, characterized in that, Visual lexical pairing is performed between the at least two visual lexical groups to obtain the plurality of lexical pairs to be merged, including: Similar visual lemma pairing is performed between the at least two visual lemma groups to obtain multiple lemma pairs to be merged.

4. The method according to claim 2, characterized in that, Visual representation merging is performed on the multiple pairs of lexical units to be merged, respectively, to obtain the merged visual representations corresponding to each of the multiple pairs of lexical units to be merged, including: For any pair of lexical units to be merged, the attention scores of the at least two visual lexical units contained in the pair of lexical units in the target intermediate layer are used as weights to perform weighted fusion on the visual representations of the at least two visual lexical units contained in the pair of lexical units in the target intermediate layer, so as to obtain the merged visual representation corresponding to the pair of lexical units to be merged.

5. The method according to claim 3, characterized in that, The at least two visual lexical groups include a first lexical group and a second lexical group; according to the desired retention rate, similar visual lexical pairs are performed between the at least two visual lexical groups to obtain multiple lexical pairs to be merged, including: From the visual lexical units contained in the first lexical unit, find similar visual lexical units for the first visual lexical unit in the second lexical unit, wherein the first visual lexical unit is any visual lexical unit contained in the second lexical unit; The first visual word and the found visual word are combined into a tuple; According to the required retention rate, the tuples formed under the second word group are filtered to obtain the multiple word pairs to be merged.

6. The method according to claim 5, characterized in that, The target visual lexical unit is divided into at least two visual lexical units, including: The target visual words are sorted according to the attention scores they correspond to in the target intermediate layer. Perform an equal-division operation on the sorted target visual lexical units to obtain the at least two visual lexical units.

7. The method according to claim 5, characterized in that, The attention scores of the visual lexical units contained in the first lexical unit in the target intermediate layer are all higher than the attention scores of the visual lexical units contained in the second lexical unit in the target intermediate layer.

8. The method according to claim 5, characterized in that, From the visual lexical units contained in the first lexical unit, find similar visual lexical units for the first visual lexical unit in the second lexical unit, including: Based on the visual representation output by the target intermediate layer for the target visual lexical unit, the visual lexical unit with the highest visual representation similarity to the first visual lexical unit is found from the visual lexical units contained in the first lexical unit group, and is used as the visual lexical unit similar to the first visual lexical unit.

9. The method according to claim 8, characterized in that, According to the required retention rate, the tuples formed under the second word group are filtered to obtain the plurality of word pairs to be merged, including: Based on the required retention rate and the number of target visual lexical units, determine the number N of tuples to be retained, where N is an integer greater than 1; From the pairs of words formed under the second word group, select the N pairs with the highest visual representation similarity to obtain the multiple word pairs to be merged.

10. The method according to claim 1, characterized in that, Also includes: If there are unmerged words in the target visual words that have not been merged, then the visual representation of the unmerged words output by the target intermediate layer is input into the next intermediate layer corresponding to the target intermediate layer.

11. The method according to claim 1, characterized in that, From the visual lexical units corresponding to the visual representation output by the target intermediate layer, target visual lexical units that meet preset filtering conditions are selected, including: According to the required filtering ratio of the target intermediate layer, determine the number M of target visual words to be filtered out, where M is an integer greater than 1; From the visual lexical units corresponding to the visual representation output by the target intermediate layer, select the M visual lexical units with the highest attention scores in the target intermediate layer as the target visual lexical units.

12. A reasoning method, characterized in that, The method includes: The inference task is received through a visual language model, and the inference task contains a visual file to be processed. During the reasoning process of the visual language model for the reasoning task, the corresponding visual lexical units are determined for the visual representation output by the target intermediate layer in the visual language model. From the identified visual words, select the target visual words that meet the preset selection criteria; The visual representations output by the target intermediate layer for the target visual lexical units are merged to obtain the merged visual representation. The merged visual representation is input into the next intermediate layer corresponding to the target intermediate layer in the visual language model, serving as the basis for the next intermediate layer to continue reasoning for the reasoning task.

13. A computing device, characterized in that, Includes memory, processor, and communication components; The memory is used to store one or more computer instructions; The processor is coupled to the memory and the communication component and is used to execute one or more computer instructions to perform the model performance optimization method of any one of claims 1-11 or the inference method of claim 12.

14. A computer-readable storage medium for storing a computer program, characterized in that, When the computer program is executed by one or more processors, it causes the one or more processors to perform the model performance optimization method according to any one of claims 1-11 or the inference method according to claim 12.

15. A computer program product, characterized in that, The computer program, when executed by one or more processors, causes the one or more processors to perform the model performance optimization method of any one of claims 1-11 or the inference method of claim 12.