Feature processing method and device, readable medium, electronic equipment and program product

By performing feature cropping processing on the input content of the multimodal model, the problem of inferential efficiency of multimodal model is solved, and the effect of improving inference accuracy and efficiency is achieved.

CN120011783APending Publication Date: 2025-05-16BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510096564.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the process of inference based on multimodal models, a large number of features will be generated based on the input content, resulting in inference efficiency.

Method used

By determining the embedded features of the target modal content and sub-content in the input content of the multimodal model and cropping these features based on different cropping ratios, the features used for multimodal model inference are obtained.

Benefits of technology

This can not only retain the global features in the target modal content, but also take into account the local features in the target subcontent, reducing the number of features used for multimodal model inference, thereby improving inference accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011783A_ABST
    Figure CN120011783A_ABST
Patent Text Reader

Abstract

The invention discloses a feature processing method and device, a readable medium, electronic equipment and a program product, and the method comprises the steps: determining target modal content except a text modal according to the input content of a multi-modal model; segmenting the target modal content to obtain target sub-content; determining a first embedding feature of the target modal content and a second embedding feature of the target sub-content; and cutting the first embedded feature and the second embedded feature based on different cutting proportions to obtain features for reasoning of the multi-modal model. The first embedded feature and the second embedded feature can be clipped based on different clipping proportions, so that global features in the target modal content can be reserved, local features in the target sub-content can be considered, the number of features used for reasoning of the multi-modal model can be reduced, and the reasoning efficiency of the multi-modal model is improved. Therefore, when the multi-modal model carries out reasoning based on the cut features, the reasoning precision and reasoning efficiency of the multi-modal model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to a feature processing method, device, readable medium, electronic device and program product. Background Art

[0002] With the rapid development of intelligent technology, more and more intelligent models have emerged, such as multimodal models. However, in the related art, in the process of reasoning based on multimodal models, a large number of features are generated based on the input content, so there are problems such as low reasoning efficiency. Summary of the invention

[0003] This summary is provided to introduce concepts in a brief form that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] In a first aspect, the present disclosure provides a feature processing method, the feature processing method comprising: Determine the target modality content other than the text modality according to the input content of the multimodal model; Segmenting the target modal content to obtain target sub-content; Determining a first embedding feature of the target modal content and a second embedding feature of the target sub-content; The first embedded feature is cropped based on a first cropping ratio to obtain a first target feature, and the second embedded feature is cropped based on a second cropping ratio to obtain a second target feature, wherein the first cropping ratio and the second cropping ratio are different, and the first target feature and the second target feature are used for reasoning by the multimodal model.

[0005] In a second aspect, the present disclosure provides a feature processing device, the feature processing device comprising: A first determination module, used to determine target modality content other than text modality according to input content of the multimodal model; A first processing module, configured to segment the target modal content to obtain target sub-content; A second determination module, configured to determine a first embedding feature of the target modal content and a second embedding feature of the target sub-content; A second processing module is used to crop the first embedded feature based on a first cropping ratio to obtain a first target feature, and to crop the second embedded feature based on a second cropping ratio to obtain a second target feature, wherein the first cropping ratio and the second cropping ratio are different, and the first target feature and the second target feature are used for reasoning by the multimodal model.

[0006] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which implements the steps of the method described in the first aspect when the computer program is executed by a processing device.

[0007] In a fourth aspect, the present disclosure provides an electronic device, including: a storage device having a computer program stored thereon; A processing device is used to execute the computer program in the storage device to implement the steps of the method in the first aspect.

[0008] In a fifth aspect, the present disclosure provides a computer program product, comprising a computer program, which implements the steps of the method described in the first aspect when executed by a processor.

[0009] Through the above technical solution, after determining the first embedded feature of the target modal content and the second embedded feature of the target sub-content, the first embedded feature and the second embedded feature can be cropped based on different cropping ratios to obtain features for reasoning with the multimodal model. Since the first embedded feature and the second embedded feature can be cropped based on different cropping ratios, it is possible to retain the global features in the target modal content and take into account the local features in the target sub-content, and it is also possible to reduce the number of features used for reasoning with the multimodal model, so that when the multimodal model is reasoned based on the cropped features, the reasoning accuracy and reasoning efficiency of the multimodal model can be improved. In addition, since feature compression is achieved by cropping the first embedded feature and the second embedded feature after obtaining the first embedded feature and the second embedded feature, when feature compression is performed by the feature processing method disclosed in the present invention, it is not necessary to train the multimodal model, nor is it necessary to develop additional reasoning operators, so that the operational complexity of feature compression can be simplified and the feature compression efficiency can be improved.

[0010] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 is a flow chart of a feature processing method according to an exemplary embodiment of the present disclosure; Figure 2 is a flow chart of another feature processing method according to an exemplary embodiment of the present disclosure; Figure 3 is a schematic diagram of an experimental picture according to an exemplary embodiment of the present disclosure; Figure 4 is a schematic diagram showing a characteristic graph according to an exemplary embodiment of the present disclosure; Figure 5 is a schematic diagram showing a binarization result according to an exemplary embodiment of the present disclosure; Figure 6 is a schematic diagram showing a mask result according to an exemplary embodiment of the present disclosure; Figure 7 is a structural block diagram of a feature processing device according to an exemplary embodiment of the present disclosure; Figure 8 It is a schematic structural diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0012] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0013] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0014] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0015] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0016] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0017] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0018] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0019] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0020] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0021] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0022] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0023] As mentioned in the background technology, in the related technology, in the process of reasoning based on the multimodal model, a large number of tokens or features are generated based on the input content, so there are problems such as low reasoning efficiency.

[0024] For example, in a scenario where images and texts are used as inputs to a multimodal model, in order to more accurately capture the details of the image, the image is often segmented, that is, an image is cut into several image sub-images. Then, the several image sub-images and the original image are input into the multimodal model together, and processed by the visual encoder of the multimodal model to obtain the embedded features of the image and the embedded features of each image sub-image. Finally, the embedded features of the image and the embedded features of each image sub-image are transmitted to the reasoning module in the multimodal model to obtain the reasoning results.

[0025] However, since the number of tokens or embedded features generated by an image may be more than 7,000, this places a heavy burden on the reasoning efficiency of the multimodal model, resulting in poor multimodal reasoning performance.

[0026] To overcome the above technical problems, feature compression or token compression can be used to reduce the number of tokens, thereby increasing model reasoning efficiency. However, in related technologies, token compression is generally performed in the following two ways, but there are problems such as long model development cycle, high model deployment cost and / or the need to develop additional reasoning operators.

[0027] Method 1: Design some token compression modules and embed them in the visual encoder. Therefore, in order to reduce the loss of inference accuracy caused by token compression, it is necessary to build a training set to train the multimodal model, build a validation set to verify the multimodal model, and build a test set to test the multimodal model. This method has problems such as complex operation, long model development cycle and high deployment cost.

[0028] Method 2: Design some customized inference operators to identify and reduce unnecessary calculations in the visual encoder. For example, if there is a 10×10 feature matrix in the visual encoder, and the inference operator in the visual encoder is a multiplication operator for calculating the 10×10 matrix, but only the 10×5 part of the 10×10 feature matrix affects the output of the inference, then a new inference operator can be designed to only calculate the 10×5 part, thereby reducing the amount of calculation. However, this method also has problems such as complex operation and high model deployment cost.

[0029] In view of this, the present disclosure provides a feature processing method, device, readable medium, electronic device and program product to solve the above technical problems.

[0030] The embodiments of the present disclosure are further explained below with reference to the accompanying drawings.

[0031] Figure 1 is a flowchart of a feature processing method according to an exemplary embodiment of the present disclosure, referring to Figure 1 , the method may include the following steps: S101: Determine target modal content other than text modality according to input content of the multimodal model.

[0032] It should be understood that the multimodal model can process and understand data from multiple different modalities, such as text data, image data, audio data and / or video data, etc. Therefore, determining the target modality content other than the text modality may be determining the image modality content, audio modality content and / or video modality content in the input content, etc. In other words, the target modality content may include at least one of images, videos and audios.

[0033] In this embodiment, the multimodal model may be a multimodal large language model or an audio multimodal model. Of course, it may also refer to other models, and the embodiments of the present disclosure do not impose any limitations on this.

[0034] S102: Segment the target modal content to obtain target sub-content.

[0035] As mentioned above, the target modal content may include at least one of an image, a video, and an audio. Thus, when the target modal content is an image, the target sub-content may be an image sub-image obtained by segmenting the image; when the target modal content is a video, the target sub-content may be a sub-video frame obtained by segmenting each video frame; when the target modal content is audio, the target sub-content may be an audio segment obtained by segmenting the audio.

[0036] In this embodiment, the target modal content may be segmented evenly or unevenly, and the embodiment of the present disclosure does not impose any limitation on this.

[0037] For example, when the target modal content is an image, the image can be evenly divided into P parts, thereby obtaining P image sub-images. At the same time, in order to better capture the details in the image sub-image, in a possible manner, after obtaining the image sub-image, the image resolution of each image sub-image can be adjusted so that the resolution of each image sub-image is consistent with the original image, wherein the original image refers to the image that has not been divided.

[0038] S103: Determine a first embedding feature of the target modal content and a second embedding feature of the target sub-content.

[0039] For example, continuing to refer to the above example, after obtaining the image sub-image, the image sub-image and the image can be input into the multimodal model together, so that the multimodal model can pre-process the image to obtain the image tensor of the image, and input the image tensor of the image into the visual encoder in the multimodal model to obtain the first embedding feature of the image. And the multimodal model can pre-process each image sub-image to obtain the image tensor of each image sub-image, and input the image tensor of each image sub-image into the visual encoder in the multimodal model to obtain the second embedding feature of each image sub-image.

[0040] S104: Crop the first embedded feature based on a first cropping ratio to obtain a first target feature, and crop the second embedded feature based on a second cropping ratio to obtain a second target feature, wherein the first cropping ratio and the second cropping ratio are different, and the first target feature and the second target feature are used for reasoning in a multimodal model.

[0041] In this embodiment, the first cropping ratio and the second cropping ratio can be determined according to actual conditions, and the disclosed embodiment does not impose any restrictions on this. In view of the fact that the complete target modal content can contain more useful information than a single target sub-content, when the same number of embedded features are generated, the second embedded features containing useless information or background information will be more than the first embedded features containing useless information or background information. Therefore, in order to avoid losing useful information in the target modal content and reduce the number of embedded features containing useless information, in a possible manner, the first cropping ratio can be smaller than the second cropping ratio. For example, the first cropping ratio can be set to 37.5%, and the second cropping ratio can be set to 50%. This makes it possible to retain the global features in the target modal content, take into account the local features in the target sub-content, and reduce the number of features used for reasoning with the multimodal model, so that the multimodal model can improve the reasoning accuracy and reasoning efficiency of the multimodal model when reasoning based on the cropped features.

[0042] Through the above technical solution, after determining the first embedded feature of the target modal content and the second embedded feature of the target sub-content, the first embedded feature and the second embedded feature can be cropped based on different cropping ratios to obtain features for reasoning with the multimodal model. Since the first embedded feature and the second embedded feature can be cropped based on different cropping ratios, it is possible to retain the global features in the target modal content and take into account the local features in the target sub-content, and it is also possible to reduce the number of features used for reasoning with the multimodal model, so that when the multimodal model is reasoned based on the cropped features, the reasoning accuracy and reasoning efficiency of the multimodal model can be improved. In addition, since feature compression is achieved by cropping the first embedded feature and the second embedded feature after obtaining the first embedded feature and the second embedded feature, when feature compression is performed by the feature processing method disclosed in the present invention, it is not necessary to train the multimodal model, nor is it necessary to develop additional reasoning operators, so that the operational complexity of feature compression can be simplified and the feature compression efficiency can be improved. In other words, through the above method, token pruning can be achieved without training the multimodal model or developing additional inference operators, thereby reducing the number of tokens, simplifying the operational complexity of token compression, and improving token compression efficiency.

[0043] To facilitate understanding of the feature processing method provided by the present disclosure, possible implementation methods in the present disclosure are described below.

[0044] In a possible manner, cropping the first embedded feature based on the first cropping ratio to obtain the first target feature may include: For each first embedded feature, a first index value of the first embedded feature is determined, where the first index value is used to characterize the importance of the first embedded feature to the multimodal model reasoning process; based on the first cropping ratio and the first index value, the first embedded feature is cropped to obtain a first target feature.

[0045] It should be understood that the number of first embedded features is generally large, but in the model reasoning process, not every first embedded feature contributes to the reasoning of the multimodal model. Therefore, this example determines the importance of the first embedded features to the multimodal model reasoning process, and can perform trimming based on the first trimming ratio and the importance of each first embedded feature to the multimodal model reasoning process, thereby reducing the number of first embedded features and improving the model reasoning efficiency without affecting the model reasoning accuracy.

[0046] In a possible manner, based on the first cropping ratio and the first index value, cropping the first embedded feature to obtain the first target feature may include: The first embedded features are sorted according to the first index value to obtain a first feature sequence; the first number is determined according to the first cropping ratio and the number of the first embedded features; the first number of embedded features in the first feature sequence are cropped to obtain a first target feature, wherein the first index values ​​corresponding to the first number of embedded features are all smaller than the first index values ​​corresponding to the remaining embedded features in the first feature sequence except the first number of embedded features.

[0047] For example, if the first cropping ratio is set to 37.5% and the number of first embedded features is set to 1000, the first number is 375. Therefore, after determining the first index value of each first embedded feature, the 1000 first embedded features can be sorted in descending order to obtain a first feature sequence; then the first embedded features after the 625th first embedded feature in the first feature sequence are cropped, and the remaining 625 first embedded features are used as first target features.

[0048] It should be understood that during the reasoning process, the multimodal model generally combines the position order between features for reasoning. Therefore, in order to improve the accuracy of reasoning, after sorting and cropping the first embedded feature, it is necessary to restore the original position information of the first embedded feature.

[0049] Through the above method, the first embedded features that are less important to the multimodal model reasoning process can be cropped based on the first cropping ratio, and the first embedded features that are more important to the multimodal model reasoning process can be retained, so that when the multimodal model is reasoned based on the first embedded features, the reasoning efficiency can be improved without affecting the reasoning accuracy.

[0050] In a possible manner, the first embedding feature is obtained by a visual encoder in a multimodal model, and accordingly, determining a first indicator value of the first embedding feature may include: Determine at least one of a maximum eigenvalue in the first embedded feature, an L1 norm and an L2 norm of a hidden state corresponding to the first embedded feature in the visual encoder.

[0051] It should be understood that the first embedded feature is a multidimensional feature, and its dimension is the same as the dimension of the hidden layer, that is, the first embedded feature can be represented by the same number of eigenvalues ​​as the hidden layer. For example, if the dimension of the hidden layer is 256, then the first embedded feature can also be represented by 256 eigenvalues. At the same time, the inventors found in the experimental process that the larger the eigenvalue of the first embedded feature, the greater its role in the multimodal model reasoning process, so the first index value can be determined based on the maximum eigenvalue in the first embedded feature.

[0052] It should also be understood that the hidden state refers to the output of the hidden layer in the visual encoder; the L1 norm, i.e., L1norm, refers to the sum of the absolute values ​​of the elements of the vector; the L2 norm, i.e., L2norm, refers to the square root of the sum of the squares of the elements of the vector. Thus, the L1 norm of the hidden state corresponding to the first embedded feature in the visual encoder may refer to: the sum of the absolute values ​​of the elements in the feature vector output by the hidden layer before the visual encoder obtains the first embedded feature. The L2 norm of the hidden state corresponding to the first embedded feature in the visual encoder may refer to: the square root of the sum of the squares of the elements in the feature vector output by the hidden layer before the visual encoder obtains the first embedded feature.

[0053] In this embodiment, any one of the maximum eigenvalue, L1 norm and L2 norm can be used as the first index value, the value obtained by combining any two of them can be used as the first index value, or the value obtained by combining the maximum eigenvalue, L1 norm and L2 norm can be used as the first index value. The present disclosed embodiment does not impose any restrictions on this.

[0054] For example, the first indicator value can be obtained by the following formula:

[0055] in, represents the first index value, , and represents the coefficient, represents the maximum eigenvalue.

[0056] In a possible manner, cropping the second embedded feature based on the second cropping ratio to obtain the second target feature may include: For each second embedded feature, a second index value of the second embedded feature is determined, where the second index value is used to characterize the importance of the second embedded feature to the multimodal model reasoning process; based on the second cropping ratio and the second index value, the second embedded feature is cropped to obtain a second target feature.

[0057] It should be understood that the number of second embedded features is generally large, but in the model reasoning process, not every second embedded feature contributes to the reasoning of the multimodal model. Therefore, this example determines the importance of the second embedded features to the multimodal model reasoning process, and can perform trimming based on the second trimming ratio and the importance of each second embedded feature to the multimodal model reasoning process, thereby reducing the number of second embedded features and improving the model reasoning efficiency without affecting the model reasoning accuracy.

[0058] In a possible manner, based on the second cropping ratio and the second index value, cropping the second embedded feature to obtain the second target feature may include: The second embedded features are sorted according to the second index value to obtain a second feature sequence; the second number is determined according to the second cropping ratio and the number of the second embedded features; the second number of embedded features in the second feature sequence are cropped to obtain second target features, wherein the second index values ​​corresponding to the second number of embedded features are all smaller than the second index values ​​corresponding to the remaining embedded features in the second feature sequence except the second number of embedded features.

[0059] For example, if the second cropping ratio is set to 50% and the number of second embedded features is set to 1000, the second number is 500. Therefore, after determining the second index value of each second embedded feature, the 1000 second embedded features can be sorted in descending order to obtain a second feature sequence; then the second embedded features after the 500th second embedded feature in the second feature sequence are cropped, and the remaining 500 second embedded features are used as second target features.

[0060] It should be understood that during the reasoning process, the multimodal model generally combines the position order between features for reasoning. Therefore, in order to improve the accuracy of reasoning, after sorting and cropping the second embedded features, it is necessary to restore the original position information of the second embedded features.

[0061] Through the above method, the second embedded features that are less important to the multimodal model reasoning process can be cropped based on the second cropping ratio, and the second embedded features that are more important to the multimodal model reasoning process can be retained, so that when the multimodal model is reasoned based on the second embedded features, the reasoning efficiency can be improved without affecting the reasoning accuracy.

[0062] In a possible manner, the second embedding feature is obtained through a visual encoder in the multimodal model, and accordingly, determining the second indicator value of the second embedding feature may include: Determine at least one of a maximum eigenvalue in the second embedded feature, an L1 norm and an L2 norm of a hidden state corresponding to the second embedded feature in the visual encoder.

[0063] It should be understood that the second embedded feature is a multidimensional feature, and its dimension is the same as the dimension of the hidden layer, that is, the second embedded feature can be represented by the same number of eigenvalues ​​as the hidden layer. For example, if the dimension of the hidden layer is 256, then the second embedded feature can also be represented by 256 eigenvalues. At the same time, the inventors found in the experimental process that the larger the eigenvalue of the second embedded feature, the greater its role in the multimodal model reasoning process, so the second index value can be determined based on the maximum eigenvalue in the second embedded feature.

[0064] It should also be understood that the hidden state refers to the output of the hidden layer in the visual encoder; the L1 norm, i.e., L1norm, refers to the sum of the absolute values ​​of the elements of the vector; the L2 norm, i.e., L2norm, refers to the square root of the sum of the squares of the elements of the vector. Thus, the L1 norm of the hidden state corresponding to the second embedded feature in the visual encoder may refer to: the sum of the absolute values ​​of the elements in the feature vector output by the hidden layer before the visual encoder obtains the second embedded feature. The L2 norm of the hidden state corresponding to the second embedded feature in the visual encoder may refer to: the square root of the sum of the squares of the elements in the feature vector output by the hidden layer before the visual encoder obtains the second embedded feature.

[0065] In this embodiment, any one of the maximum eigenvalue, L1 norm and L2 norm can be used as the second index value, or the value obtained by combining any two of them can be used as the second index value, or the value obtained by combining the maximum eigenvalue, L1 norm and L2 norm can be used as the second index value. The present disclosed embodiment does not impose any restrictions on this.

[0066] To facilitate further understanding of the feature processing method provided by the present disclosure, a possible implementation of the feature processing method is described below in combination with various steps: For example, Figure 2As shown, the multimodal model can be a multimodal large language model, and the input content can include text and images. For images, the image can be split first to obtain image sub-images. After that, the image and the image sub-image are pre-processed by the multimodal model respectively to obtain the image tensor of the image and the image tensor of the image sub-image, and the image tensor of the image and the image tensor of the image sub-image are respectively input into the visual encoder in the multimodal model to obtain the image embedding vector image embedding of the image and the image embedding vector image embedding of each image sub-image, wherein the image embedding of the image includes multiple first embedding features, and the image embedding of each image sub-image includes multiple second embedding features. For each first embedding feature, the maximum eigenvalue of the first embedding feature can be used as the first index value, and the first embedding features can be sorted in order from small to large according to the size of the first index value to obtain a first feature sequence. Then, the first number can be determined based on the first cropping ratio and the number of first embedded features. After obtaining the first number, the first number of first embedded features in the first feature sequence are cropped, and the original position information of the retained first embedded features is restored, thereby obtaining the compressed image embedding vector of the image compressed image embedding, wherein the compressed image embedding vector of the image includes the first embedded features that have not been cropped, that is, the first target features. The second embedded features are processed based on the same processing method to obtain the compressed image embedding vector compressed image embedding of each image sub-image, wherein the compressed image embedding vector of the image sub-image includes the second embedded features that have not been cropped, that is, the second target features. After obtaining the compressed image embedding vector of the image and the compressed image embedding vector of the image sub-image, the compressed image embedding vector of the image and the compressed image embedding vector of the image sub-image are aggregated to obtain the final compressed image embedding vector.

[0067] For text, we can first tokenize the text to get multiple token feature tokens. For each token feature, we can assign a unique identifier to each token feature to get an input identifier sequence. Then, we can map the text embedding vector based on the mapping relationship between the identifier sequence and the embedding vector and the input identifier sequence.

[0068] After obtaining the compressed image embedding vector and the text embedding vector, the compressed image embedding vector and the text embedding vector can be combined into a multimodal embedding vector. After that, the multimodal embedding vector is input into a multimodal large language model for reasoning to obtain an output identifier sequence, and then the output result is obtained based on the mapping relationship between the identifier sequence and the embedding vector and the output identifier sequence.

[0069] To facilitate understanding of the feasibility of the feature processing disclosed in the present invention, the following is an explanation with reference to the accompanying drawings: For example, the input content may include: Figure 3 For the image shown in FIG. 1 , after converting the image into an embedded feature, the maximum eigenvalue of the embedded feature can be used as the first index value, and the corresponding embedded feature can be represented by blocks of different colors according to the size of the first index value, as shown in FIG. Figure 4 The feature graph shown in the figure shows that the whiter the block is, the greater the role of the corresponding embedded feature in the multimodal reasoning process. After that, the first number can be determined based on the set cropping ratio and the number of embedded features, and after obtaining the first number, whether the embedded feature in the feature graph is important can be binarized to obtain a binarization result of whether the embedded feature is important, such as Figure 5 As shown in the figure, the embedded features corresponding to the black squares are the embedded features that play a smaller role in the multimodal reasoning process, that is, the embedded features that need to be cut off, and the embedded features corresponding to the white areas are the embedded features that play a larger role in the multimodal reasoning process, that is, the embedded features that need to be retained. Figure 3 and Figure 5 After merging, we can get the mask result shown in Figure 6. Figure 6 It can be seen that when the feature processing method provided by the present disclosure is used for processing, more foreground information in the image will not be lost, that is, when the feature processing is performed in this way, the inference accuracy of the model can be reduced while reducing the amount of calculation. At the same time, since the feature processing method provided by the present disclosure can also be combined with the embedded features of the target sub-content, when the target modal content is an image, it can not only better retain the global features of the original image, but also take into account the local features of the image sub-image, thereby improving the inference accuracy of the cropped model.

[0070] Based on the same concept, the present disclosure also provides a feature processing device, such as Figure 7 As shown, the feature processing device 700 may include: A first determination module 701 is used to determine target modality content other than text modality according to input content of the multimodal model; A first processing module 702 is used to segment the target modal content to obtain target sub-content; A second determination module 703, configured to determine a first embedding feature of the target modal content and a second embedding feature of the target sub-content; The second processing module 704 is used to crop the first embedded feature based on a first cropping ratio to obtain a first target feature, and to crop the second embedded feature based on a second cropping ratio to obtain a second target feature, wherein the first cropping ratio and the second cropping ratio are different, and the first target feature and the second target feature are used for reasoning in a multimodal model.

[0071] Through the above-mentioned feature processing device 700, after determining the first embedded feature of the target modal content and the second embedded feature of the target sub-content, the first embedded feature and the second embedded feature can be cropped based on different cropping ratios to obtain features for reasoning with the multimodal model. Since the first embedded feature and the second embedded feature can be cropped based on different cropping ratios, it is possible to retain the global features in the target modal content and take into account the local features in the target sub-content, and it is also possible to reduce the number of features used for reasoning with the multimodal model, so that when the multimodal model is reasoned based on the cropped features, the reasoning accuracy and reasoning efficiency of the multimodal model can be improved. In addition, since feature compression is achieved by cropping the first embedded feature and the second embedded feature after obtaining the first embedded feature and the second embedded feature, when feature compression is performed by the feature processing method disclosed in the present invention, it is not necessary to train the multimodal model, nor is it necessary to develop additional reasoning operators, so that the operational complexity of feature compression can be simplified and the feature compression efficiency can be improved. In other words, through the above method, token pruning can be achieved without training the multimodal model or developing additional inference operators, thereby reducing the number of tokens, simplifying the operational complexity of token compression, and improving token compression efficiency.

[0072] In a possible manner, the second processing module 704 may include: A first determining unit, configured to determine, for each first embedded feature, a first index value of the first embedded feature, wherein the first index value is used to characterize the importance of the first embedded feature to the multimodal model reasoning process; The first cropping unit is used to crop the first embedded feature based on a first cropping ratio and a first index value to obtain a first target feature.

[0073] In a possible manner, the first cropping unit may include: A first sorting subunit, used to sort the first embedded features according to the first index value to obtain a first feature sequence; A first determining subunit, configured to determine a first quantity according to a first cropping ratio and a quantity of first embedded features; The first cropping subunit is used to crop a first number of embedded features in the first feature sequence to obtain a first target feature, wherein the first index values ​​corresponding to the first number of embedded features are all smaller than the first index values ​​corresponding to the remaining embedded features in the first feature sequence except the first number of embedded features.

[0074] In a possible manner, the first embedded feature is obtained through a visual encoder in a multimodal model, and accordingly, the first determination unit is used to determine the maximum eigenvalue in the first embedded feature, and at least one of the L1 norm and L2 norm of the hidden state corresponding to the first embedded feature in the visual encoder.

[0075] In a possible manner, the second processing module 704 may include: A second determining unit is used to determine, for each second embedded feature, a second index value of the second embedded feature, where the second index value is used to characterize the importance of the second embedded feature to the multimodal model reasoning process; The second cropping unit is used to crop the second embedded feature based on a second cropping ratio and a second index value to obtain a second target feature.

[0076] In a possible manner, the second cropping unit may include: A second sorting subunit, used to sort the second embedded features according to the second index value to obtain a second feature sequence; A second determining subunit is used to determine a second number according to a second cropping ratio and a number of second embedded features; The second cropping subunit is used to crop a second number of embedded features in the second feature sequence to obtain a second target feature, wherein the second index values ​​corresponding to the second number of embedded features are all smaller than the second index values ​​corresponding to the remaining embedded features in the second feature sequence except the second number of embedded features.

[0077] In a possible manner, the second embedded feature is obtained through a visual encoder in a multimodal model, and accordingly, the second determination unit is used to determine the maximum eigenvalue in the second embedded feature, and at least one of the L1 norm and L2 norm of the hidden state corresponding to the second embedded feature in the visual encoder.

[0078] In a possible manner, the first cropping ratio is smaller than the second cropping ratio.

[0079] In a possible manner, the target modality content includes at least one of image, video and audio.

[0080] Based on the same concept, an embodiment of the present disclosure further provides a computer-readable medium on which a computer program is stored, and when the program is executed by a processing device, the steps of any of the above-mentioned feature processing are implemented.

[0081] Based on the same concept, an embodiment of the present disclosure further provides an electronic device, which may include: a storage device having a computer program stored thereon; A processing device is used to execute the computer program in the storage device to implement the steps of any of the above-mentioned feature processing methods.

[0082] Based on the same concept, an embodiment of the present disclosure also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned feature processing methods when executed by a processor.

[0083] Reference below Figure 8 , which shows a schematic diagram of the structure of an electronic device 800 suitable for implementing the embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0084] like Figure 8 As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 to a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0085] Typically, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 8 The electronic device 800 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0086] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0087] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0088] In some embodiments, any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol) can be used for communication, and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0089] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0090] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: determines the target modal content other than the text modality based on the input content of the multimodal model; segments the target modal content to obtain target sub-content; determines the first embedding feature of the target modal content and the second embedding feature of the target sub-content; crops the first embedding feature based on a first cropping ratio to obtain a first target feature, and crops the second embedding feature based on a second cropping ratio to obtain a second target feature, wherein the first cropping ratio and the second cropping ratio are different, and the first target feature and the second target feature are used for reasoning with the multimodal model.

[0091] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0092] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0093] The modules involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a module does not, in some cases, constitute a limitation on the module itself.

[0094] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0095] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0096] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.

[0097] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0098] Although the subject matter has been described in language specific to structural features and / or method logic actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims. Regarding the device in the above embodiment, the specific manner in which each module performs the operation has been described in detail in the embodiment related to the method, and will not be elaborated here.

Claims

1. A feature processing method, characterized in that: The feature processing method comprises: Determine the target modality content other than the text modality according to the input content of the multimodal model; Segmenting the target modal content to obtain target sub-content; Determining a first embedding feature of the target modal content and a second embedding feature of the target sub-content; The first embedded feature is cropped based on a first cropping ratio to obtain a first target feature, and the second embedded feature is cropped based on a second cropping ratio to obtain a second target feature, wherein the first cropping ratio and the second cropping ratio are different, and the first target feature and the second target feature are used for reasoning by the multimodal model.

2. The feature processing method according to claim 1, characterized in that: The step of clipping the first embedded feature based on a first clipping ratio to obtain a first target feature includes: For each first embedded feature, determine a first index value of the first embedded feature, where the first index value is used to represent the importance of the first embedded feature to the multimodal model reasoning process; Based on the first cropping ratio and the first index value, the first embedded feature is cropped to obtain a first target feature.

3. The feature processing method according to claim 2, characterized in that: The step of clipping the first embedded feature based on the first clipping ratio and the first index value to obtain a first target feature includes: Sort the first embedded features according to the first index value to obtain a first feature sequence; Determining a first number according to the first cropping ratio and the number of the first embedded features; The first number of embedded features in the first feature sequence are trimmed to obtain first target features, wherein first index values ​​corresponding to the first number of embedded features are all smaller than first index values ​​corresponding to remaining embedded features in the first feature sequence except the first number of embedded features.

4. The feature processing method according to claim 2, characterized in that: The first embedded feature is obtained by a visual encoder in the multimodal model, and determining a first indicator value of the first embedded feature includes: Determine at least one of a maximum eigenvalue in the first embedded feature, an L1 norm and an L2 norm of a hidden state corresponding to the first embedded feature in the visual encoder.

5. The feature processing method according to any one of claims 1 to 4, characterized in that: The step of clipping the second embedded feature based on a second clipping ratio to obtain a second target feature includes: For each second embedded feature, determining a second index value of the second embedded feature, where the second index value is used to characterize the importance of the second embedded feature to the multimodal model reasoning process; Based on the second cropping ratio and the second index value, the second embedded feature is cropped to obtain a second target feature.

6. The feature processing method according to claim 5, characterized in that: The step of clipping the second embedded feature based on the second clipping ratio and the second index value to obtain a second target feature includes: Sort the second embedded features according to the second index value to obtain a second feature sequence; Determining a second number according to the second cropping ratio and the number of the second embedded features; The second number of embedded features in the second feature sequence are trimmed to obtain second target features, wherein second index values ​​corresponding to the second number of embedded features are all smaller than second index values ​​corresponding to remaining embedded features in the second feature sequence except the second number of embedded features.

7. The feature processing method according to claim 5, characterized in that: The second embedded feature is obtained by a visual encoder in the multimodal model, and determining a second indicator value of the second embedded feature includes: Determine at least one of a maximum eigenvalue in the second embedded feature, an L1 norm and an L2 norm of a hidden state corresponding to the second embedded feature in the visual encoder.

8. The feature processing method according to any one of claims 1 to 4, characterized in that: The first cropping ratio is smaller than the second cropping ratio.

9. The feature processing method according to any one of claims 1 to 4, characterized in that: The target modality content includes at least one of image, video and audio.

10. A feature processing device, characterized in that: The feature processing device comprises: A first determination module, used to determine target modality content other than text modality according to input content of the multimodal model; A first processing module, configured to segment the target modal content to obtain target sub-content; A second determination module, configured to determine a first embedding feature of the target modal content and a second embedding feature of the target sub-content; A second processing module is used to crop the first embedded feature based on a first cropping ratio to obtain a first target feature, and to crop the second embedded feature based on a second cropping ratio to obtain a second target feature, wherein the first cropping ratio and the second cropping ratio are different, and the first target feature and the second target feature are used for reasoning by the multimodal model.

11. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the steps of the method according to any one of claims 1 to 9 are implemented.

12. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, used to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 9.

13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 9 are implemented.