Object processing method, device and electronic equipment
The object processing method enhances cross-modal retrieval accuracy by fusing local and global features, addressing the limitations of current multimodal models in capturing detailed information.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- LENOVO (BEIJING) LTD
- Filing Date
- 2025-12-16
- Publication Date
- 2026-07-02
AI Technical Summary
Current multimodal models for cross-modal retrieval, such as text-to-image or image-to-text, primarily focus on global features, leading to the overlook of important details and low accuracy in cross-modal retrieval tasks.
An object processing method that includes obtaining essential and initial feature information, performing fusion processing on local and global feature information to obtain target feature information, and determining an object that satisfies a similarity condition, enhancing fine-grained feature extraction.
Improves the perception of fine-grained image and text information during cross-modal retrieval, increasing the representational power and user experience by retaining original global features and adding fine-grained details.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
REFERENCE TO RELATED REGISTRATION This application claims priority over Chinese patent application No. 202411998148.9, filed on December 31, 2024, the entire contents of which are hereby incorporated by reference. TECHNICAL AREA The present disclosure relates to the field of cross-modal retrieval technology and in particular to an object processing method, a device and an electronic device. TECHNICAL BACKGROUND The technique of retrieving information from different modalities (for example, text-to-image or image-to-text) using multimodal models is commonly referred to as cross-modal retrieval. Multimodal deep learning models map information from different modalities into a common representational space to enable effective comparison and mapping of the two different types of information within the same space. However, the features extracted from text and image information by current multimodal models are limited to global features, and information from different modalities is processed using the same method. Consequently, important details are easily overlooked, resulting in low accuracy in cross-modal retrieval, which is often insufficient for user needs. OVERVIEW OF THE INVENTION One aspect of the present disclosure provides an object processing procedure that includes: processing an object to be processed to obtain essential feature information and initial feature information of the object to be processed, wherein the initial feature information includes local feature information and global feature information; performing fusion processing on the essential feature information, the local feature information, and the global feature information to obtain target feature information of the object to be processed, wherein the essential feature information is target frame information or text feature information of the object to be processed of different modalities; and determining an object that satisfies a similarity condition with respect to the target feature information of the object to be processed as a result object corresponding to the object to be processed.where a modality of the object to be processed differs from a modality of the result object, and the modality includes text, image, video and / or audio. Another aspect of the present disclosure provides an object processing device comprising an information output module, a target feature information reference module, and a result generation module. The information output module is configured to process a target object in order to obtain essential feature information and initial feature information of the target object. The initial feature information includes local feature information and global feature information. The target feature information reference module is configured to perform fusion processing on the essential feature information, the local feature information, and the global feature information to obtain target feature information of the target object.The essential feature information consists of target framework information or text feature information of the object to be processed, representing different modalities. The result generation module is configured to identify an object that fulfills a similarity condition regarding the target feature information of the object to be processed as a result object corresponding to the object to be processed. A modality of the object to be processed differs from a modality of the result object. The modality includes text, image, video, and / or audio. Another aspect of the present disclosure provides an electronic device comprising one or more processors and one or more memories for storing one or more computer programs, which, when executed by the one or more processors, cause the one or more processors to process an object to be processed in order to obtain essential feature information of the object to be processed and initial feature information of the object to be processed, to perform fusion processing on the essential feature information, the local feature information and the global feature information in order to obtain target feature information of the object to be processed, and to determine an object that satisfies a similarity condition with respect to the target feature information of the object to be processed as a result object corresponding to the object to be processed.The initial feature information includes local and global feature information. The essential feature information consists of target frame information or text feature information of the object to be processed, representing different modalities. A modality of the object to be processed differs from a modality of the output object, and the modality includes text, image, video, and / or audio. BRIEF DESCRIPTION OF THE DRAWINGS In conjunction with the accompanying drawings and with reference to the following description of exemplary embodiments, the foregoing and further features, advantages, and aspects of the embodiments of this disclosure will become clearer. In all drawings, identical or similar reference numerals represent identical or similar elements. It is understood that the drawings are schematic and elements are not necessarily drawn to scale. Fig. 1 is a schematic diagram of an application scenario of an object processing method, a device, and an electronic device according to some embodiments of this disclosure. Fig. 2 is a schematic flowchart of an object processing method according to some embodiments of this disclosure.Figure 3A is a schematic diagram showing a determination process for fine-grained image fusion feature information according to some embodiments of the present disclosure. Figure 3B is a schematic diagram showing a determination process for coarse- to fine-grained text fusion feature information according to some embodiments of the present disclosure. Figure 3C is a schematic diagram showing a training process for a deep learning model according to some embodiments of the present disclosure. Figure 4 is a schematic block diagram of an object processing device according to embodiments of the present disclosure. Figure 5 is a schematic block diagram of an electronic device for implementing an object processing method according to some embodiments of the present disclosure. DETAILED DESCRIPTION Embodiments of the present disclosure are described with reference to the accompanying drawings. However, these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. Furthermore, descriptions of known structures and techniques are omitted in the following description in order to avoid obscuring the concepts of the present disclosure. The terms used herein serve only to describe certain embodiments and are not intended to limit the present disclosure. The terms "have" and "comprise" denote the presence of the mentioned features, steps, processes and / or components, but do not exclude the presence or addition of one or more other features, steps, processes or components. All terms used herein (including technical and scientific terms) have, unless otherwise defined, the meanings generally understood by a person skilled in the art. It should be noted that the terms used herein are to be interpreted in such a way as to have meanings consistent with the context of this description and should not be interpreted in an idealized or overly rigid manner. The accompanying drawings depict block diagrams and / or flowcharts. Some blocks in the block diagrams and / or flowcharts, or a combination thereof, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a specialized computer, or other programmable data processing device, such that when executed by the processor, the instructions cause the processor to create a device for implementing the functions / operations illustrated in the block diagrams and / or flowcharts. Therefore, the technology of the present disclosure can be implemented in the form of hardware and / or software (including firmware, microcode, etc.). Furthermore, the technology of the present disclosure can be in the form of a computer program product stored on a computer-readable medium containing instructions that can be used or combined with an instruction execution system. In the context of the present disclosure, a computer-readable medium can be any medium capable of containing, storing, transmitting, disseminating, or transferring instructions. For example, the computer-readable medium can include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, equipment, or transmission media.Specific examples of computer-readable media include: magnetic storage devices such as magnetic tapes or hard disk drives (HDDs); optical storage devices such as optical discs (CD-ROMs); memory such as random access memory (RAM) or flash memory; and / or wired / wireless communication links. Fig. 1 is a schematic diagram of an application scenario 100 of an object processing method, a device and an electronic device according to some embodiments of the present disclosure. Fig. 1 is merely an example of scenarios of embodiments of the present disclosure to facilitate the understanding of the technical content of the present disclosure by the person skilled in the art, and does not imply that embodiments of the present disclosure cannot be used in other facilities, systems, environments or scenarios. As shown in Fig. 1, the application scenario 100 according to embodiments of the present disclosure comprises a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is a medium configured to provide a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 can have various connection types, such as wired communication links, wireless communication links, or fiber optic cables. Users can interact with server 105 via network 104 using the first terminal 101, the second terminal 102, and the third terminal 103 to receive or send messages. Various communication client applications can be installed on the first terminal 101, the second terminal 102, and the third terminal 103. The first terminal 101, the second terminal 102 and the third terminal 103 can be various electronic devices that have a screen and support browsing the web, including, but not limited to, smartphones, tablets, laptops, desktop computers and the like. Server 105 can be a server that provides various services. For example, an object to be processed can be processed on Server 105 to obtain essential feature information and initial feature information. In one example, the essential feature information, local feature information, and global feature information can be fused on Server 105 to obtain target feature information. Then, a processing result corresponding to the object being processed can be determined, and the processing result (outcome object) can be returned to the end device. The object processing method according to embodiments of this disclosure can generally be executed by the server 105. Accordingly, an object processing device according to embodiments of this disclosure can generally be arranged in the server 105. The object processing method according to embodiments of this disclosure can also be executed by a server or server cluster other than the server 105, which are capable of communicating with the first terminal 101, the second terminal 102, the third terminal 103, and / or the server 105. Accordingly, the object processing device according to embodiments of this disclosure can also be arranged in the server or server cluster other than the server 105, which are capable of communicating with the first terminal 101, the second terminal 102, the third terminal 103, and / or the server 105. The number of end devices, networks, and servers in Fig. 1 is for illustrative purposes only. Any number of end devices, networks, and servers can be provided as needed. In some embodiments, a multimodal model (e.g., the Contrastive Language-Image Pretraining Model – CLIP model) can be applied to cross-modality retrieval tasks and can be configured to compare image features and their corresponding text features for learning purposes. The image feature and the text feature can each be represented by a one-dimensional vector. Paired image-text vectors can be spatially approximated to perform tasks in image classification, text-to-image search, and image-to-text search. However, the above-mentioned method can simply involve extracting a global feature for the image and text for calculation. In the feature extraction process, all image blocks and text words are treated equally. Therefore, the extracted features may be more suitable as a global feature, and the accuracy of the retrieval result may be low. In some other embodiments, an attention mechanism can be added to the architecture of the multimodal model to attempt to assign weights to each image block and each word in the multimodal model by introducing the attention mechanism, in order to determine which image parts and which text parts are most important for the task. For tasks resembling zero learning, however, the model may not receive task-specific data. Therefore, the model cannot be strengthened by a large number of task-related examples to aid understanding. Simply learning implicit weights is not necessarily sufficient to provide adequate knowledge, and image details may be neglected in the extracted features. Consequently, retrieval efficiency and accuracy in text-to-image retrieval tasks may be impaired, and the user experience may be worse. Based on the problems mentioned above, the present disclosure provides an object processing procedure that includes: processing the object to be processed to obtain the essential feature information and initial feature information of the object to be processed, wherein the initial feature information includes local feature information and global feature information; performing fusion processing on the essential feature information, the local feature information, and the global feature information to obtain the target feature information of the object to be processed, wherein the essential feature information is target frame information or text feature information based on the object to be processed of different modalities; and determining an object that satisfies a similarity condition with respect to the target feature information of the object to be processed as the result object.that corresponds to the object to be processed, wherein the modality of the object to be processed differs from the modality of the result object and the modality includes text, image, video and / or audio. According to embodiments of the present disclosure, fusion processing is performed on the essential feature information, the local feature information, and the global feature information of the object to be processed in order to obtain the target feature information, which comprises coarse- and fine-grained fusion. The resulting object, corresponding to the object to be processed, can then be determined. Since the target feature information is determined based on the essential feature information (target frame information or text feature information), fine-grained features can be added to the extracted target frame information or text feature information by retaining the original global features, thereby improving the perception of the fine-grained image and text information during cross-modal retrieval.This allows the representational power of the overall features to be increased in order to further improve the user experience. Fig. 2 is a schematic flowchart of an object processing method according to some embodiments of the present disclosure. As shown in Fig. 2, the method includes processes S201 to S203. In S201, the object to be processed is processed to obtain the essential feature information and initial feature information of the object to be processed, where the initial feature information includes the local feature information and the global feature information. In embodiments of the present disclosure, the object to be processed can be objects of various modalities, including, but not limited to, images, text, time-frequency representations, and audio, which can be used in scenarios such as image search engines, image annotation, visual question answering, and multimodal translation. The essential feature information can comprise the target frame information and text feature information according to the different modalities of the object to be processed. Thus, the essential feature information can be obtained through an information detection layer in the multimodal model. The local feature information and global feature information can be obtained through an information recognition layer in the multimodal model. For example, the image can be detected using the information detection layer in the multimodal model to obtain the target frame information of the image. The image can be recognized using the information recognition layer in the multimodal model to extract the local feature information and the global feature information. In one example, word segmentation and retrieval processing can be performed on an input text to obtain word segmentation features and a set of image candidates corresponding to the text. The image candidates can then be detected using the information detection layer in the multimodal model to obtain text information. The features from the word segmentation that match the text information can then be determined as the text feature information. The text can be recognized using the information recognition layer of the multimodal model to extract the local and global feature information corresponding to the text. In S202, the essential feature information, the local feature information, and the global feature information are fused to obtain the target feature information of the object to be processed, where the essential feature information is the target frame information or text feature information based on the object to be processed of different modalities. In embodiments of the present disclosure, the local feature information based on the object to be processed, which may be of different modalities, can include local text feature information or local image feature information. The global feature information based on the object to be processed, which may be of different modalities, can include global text feature information or global image feature information. The target feature information can represent the coarse- and fine-grained feature information obtained after fusing several types of feature information, including a coarse-fine-grained text fusion feature or a coarse-fine-grained image fusion feature. In one example, the local text feature information can be weighted according to the text feature information to obtain the fine-grained text feature information. Then, the fine-grained text feature information and the global text feature information (i.e., the coarse-fine-grained text feature information) can be merged and weighted to obtain a coarse-fine-grained text fusion feature. In one example, the local image feature information can be weighted based on the image feature information to obtain the fine-grained image feature information. Then, the fine-grained image feature information and the global image feature information (i.e., the coarse-fine-grained image feature information) can be fused and weighted to obtain a coarse-fine-grained image fusion feature. In S203, an object that satisfies a similarity condition regarding the target feature information of the object to be processed is determined to be the result object that corresponds to the object to be processed, where the modality of the object to be processed is different from the modality of the result object and the modality includes text, image, video and / or audio. In embodiments of the present disclosure, the similarity condition can be a condition threshold for determining the similarity between the object to be processed and an initial result object. Based on the result of a comparison with the condition threshold, one object that satisfies the condition threshold can be selected as the result object from among several initial result objects. The specific condition threshold can be determined according to the actual retrieval or recognition accuracy requirements and is not limited here. The modality of the object to be processed and the modality of the processing result can be different. For example, an image that meets the condition threshold for the coarse-to-fine text fusion feature can be determined as the result of a text-to-image query. In some other embodiments, text that meets the condition threshold for the coarse-to-fine image fusion feature can be determined as the result of an image-to-text query. In some other embodiments, the modality of the object to be processed and the modality of the processing result can also be the same. For example, in an application scenario of retrieving a second image using a first image, the second image that satisfies the condition threshold regarding the coarse-fine fusion features of the first image can be determined as the retrieval result of the first image. In embodiments of the present disclosure, the target feature information, including the coarse-fine fusion feature, can be obtained by fusing the essential feature information, the local feature information, and the global feature information of the object to be processed in order to determine the result object corresponding to the object to be processed. Since the target feature information is determined based on the essential feature information (i.e., the target frame information or text feature information), the fine-grained features can be added to the extracted target frame information or the extracted text feature information while retaining the original global features. Thus, the perception of the fine-grained image and text information can be improved during cross-modal retrieval.The representational power of the entire feature can be improved, and the user experience can be further enhanced. How to determine the result object is described above, and how to determine the target characteristic information is described below. In embodiments of the present disclosure, obtaining the target feature information based on the essential feature information, the local feature information, and the global feature information may include processing the local feature information based on the essential feature information to obtain the fine-grained feature information, and performing fusion processing on the fine-grained feature information and the global feature information to obtain the target feature information. In embodiments of the present disclosure, the fine-grained feature information can comprise the fine-grained text feature or the fine-grained image feature, depending on the different modalities of the object to be processed. The fine-grained feature information can be obtained by weighting the local feature information using the essential feature information. Depending on the modality of the object to be processed, the target feature information can comprise the target text feature (i.e., the coarse-to-fine-grained text fusion feature) or the target image feature (i.e., the coarse-to-fine-grained image fusion feature). The target feature information can be obtained by further performing a weighted fusion on the global feature information using the fine-grained feature information. In embodiments of the present disclosure, the fine-grained feature information and the global feature information can include weighting information for respective objects. For example, the weighting information corresponding to the fine-grained feature information can be a first weight, and the weighting information corresponding to the global feature information can be a second weight. In this case, the sum of the weights of the first and second weights can satisfy a target weight value. The target weight value can be a fixed value. The initial weight can represent an initial weight value of the fine-grained feature. The first weight can be obtained by updating the initial weight. The initial weight value can be determined based on actual experience or relevant test results. Using the example of an object to be processed that is an image, after determining the essential feature information (i.e. the target frame information) of the image, the local image feature information (e.g. the image block) can be weighted according to the target frame information in order to obtain the fine-grained image feature. In embodiments of the present disclosure, the local feature information of the image can be extracted by the information recognition layer (i.e., the coding layer) of the multimodal model. For example, an image to be processed can be divided into several image blocks, which can be input into the coding layer of the multimodal model (which has a convolution layer and a local response layer). The features of the image blocks can then be extracted by the convolution layer. A convolution kernel can capture initial local features of the image blocks, such as texture, edges, colors, and the like, through a sliding window operation. The initial local feature can then be abstracted at a high level by the local response layer to obtain the local feature information (e.g., a local object element, shape, or texture information). Using the example of a text object to be processed, after determining the essential feature information (i.e., the text feature information), the local text feature can be weighted using the text feature information to obtain the fine-grained text feature. In embodiments of the present disclosure, the local text feature information can be extracted by the coding layer of the multimodal model. For example, starting from word embeddings, the multimodal model can map words in any text into a high-dimensional vector space. The word vector can capture the semantic information (e.g., semantics, context, and collocation relationships). Then, further processing can be performed by an attention mechanism to allow each word to represent the word's embedding and also to be weighted and adapted according to the context to capture richer local semantic information in order to obtain the local text feature. In embodiments of the present disclosure, determining the target feature information may involve updating the initial weighting of the global feature information using the first weighting, which corresponds to the fine-grained feature information, to obtain the second weighting after updating the global feature information. The target feature information may then be obtained according to the product of the first weighting and the fine-grained feature information, and the product of the second weighting and the global feature information. Depending on the modality of the object to be processed, the target feature information may include the coarse-to-fine-grained image fusion feature (i.e., the target image feature information) or the coarse-to-fine-grained text fusion feature (i.e., the target text feature information). Examples of determining target feature information have already been described above. How to determine fine-grained feature information is described below. In embodiments of the present disclosure, if the object to be processed is an image or a video and the essential feature information is the target frame information, the target frame information can be obtained by performing a detection on the object to be processed. Processing the local feature information based on the essential feature information to obtain the fine-grained feature information can include determining the fine-grained feature information based on the image information overlapping between the local feature information and the target frame information. The target frame information according to embodiments of the present disclosure can represent the detection frame position information and the object category information in the detection frame, obtained by the information detection layer. The target object can be detected in the image by the information detection layer, and the position of the target object in the image (i.e., the coordinates of the detection frame) and the object category can be obtained. The information detection layer can also be an information detection submodel of the multimodal model. For example, if the object to be processed is input into the information detection layer, a tensor can be output that represents the associated information for each detection frame in the image. This includes bounding frame coordinates, a confidence score for each detection frame (i.e., a value indicating the probability that a target object is present in the frame), and a category probability. The bounding frame coordinates can include the center coordinates (x, y) of each frame, as well as the frame's width w and height h. The category probability can indicate the probability of the respective category being present within the frame. In one example, after obtaining the local feature information of the image (such as local object parts, shapes, or texture information), fine-grained feature information of the image can be obtained based on the local object shape, the bounding box coordinates, the object category, and the confidence level. After obtaining this fine-grained feature information, the target feature information (coarsely to finely fused feature information of the image) can be obtained based on the fine-grained image features, local image features, and global image features. With reference to Fig. 3A, the procedure for determining the fine-grained image fusion feature information is described below. Fig. 3A is a schematic diagram showing a determination process of the fine-grained image fusion feature information according to some embodiments of the present disclosure. As shown in Fig. 3A, in 300A the object to be processed is an image 301. The image 301 is detected using an information detection layer 302 in the multimodal model to obtain the target frame information 303. The target frame information includes the target frame coordinates and the category of objects in the detection frame (e.g., dog or bicycle). The image 301 can be recognized using the information recognition layer 304 to obtain the local image feature information 305 and the global image feature information 306 corresponding to the image 301. After obtaining the target frame information 303, the local image feature information 305, and the global image feature information 306, the local image feature information 305 can be weighted using the target frame information 303 to obtain the fine-grained image feature 307. Then, the global image feature information 306 and the fine-grained image feature 307 can be fused to obtain the coarse-fine-grained image fusion feature 308 (i.e., target image feature information). In embodiments of the present disclosure, the local feature can be processed using the target frame information. To improve the model's perception of image details, the multimodal model can be configured to extract fine-grained features in the image. The fine-grained feature enhances performance in image classification and target detection tasks, resulting in higher accuracy and greater expressiveness in multimodal fusion tasks. How to determine the fine-grained feature information is described above as an example, and another example of how to determine the fine-grained feature information is described below. In embodiments of the present disclosure, if the object to be processed is audio or text and the essential feature information is text feature information, processing the local feature information based on the essential feature information to obtain the fine-grained feature information may include determining the fine-grained feature information based on the text information overlapping between the local feature information and the text feature information. In embodiments of the present disclosure, the essential feature information corresponding to the text is text feature information. Obtaining the text feature information can involve performing word segmentation processing on the text to obtain word segmentation features and searching the text. Then, using the information detection layer, detection can be performed on images in the image candidate set to obtain the text information corresponding to the text. Based on this, intersection processing can be performed on the text information and the word segmentation features to determine the overlapping text information. If the weight of the word in the intersection is increased, the text feature information can be obtained. The text information corresponding to the text can be classification information that corresponds to the object in the image obtained through image detection.Once the text feature information has been determined, the text feature information and the local text feature can be weighted and fused to obtain the coarse-grained text fusion feature information. The following section, with reference to Fig. 3B, describes in more detail the method for determining the coarse- and fine-grained text fusion feature information in conjunction with specific embodiments of the present disclosure. Fig. 3B is a schematic diagram showing a determination process of the coarse-grained text fusion feature information according to some embodiments of the present disclosure. As shown in Fig. 3B, in 300B the object to be processed is a text 310. Word segmentation processing is performed on the text 310 to obtain a word segmentation feature 311, and a search is performed on the text 310 to obtain a set of image candidates (i.e., an image library). Then, the information detection layer 302 is configured to detect the image candidates 312 in the image candidate set to obtain the text information 313 corresponding to the text 310. Based on this, the intersection (overlap) of the text information 313 and the word segmentation feature 311 is determined to obtain the text feature information 314. The information recognition layer 304 can be configured to recognize the text 310 in order to obtain the local text feature information 315 and the global text feature information 316 corresponding to the text 310. After obtaining the text feature information 314, the local text feature information 315, and the global text feature information 316, the local text feature information 315 can be weighted using the text feature information 314 to obtain the fine-grained text feature 317. Then, the global text feature information 316 and the fine-grained text feature 317 can be fused to obtain the coarse-fine-grained text fusion feature (i.e., the target text feature information) 318. In embodiments of the present disclosure, the text information overlapping between the local feature information and the text feature information allows the model to capture essential semantic information, enabling it to understand the relationships between different text segments more accurately and to better process complex text tasks. For example, in multi-topic text classification, the model can better understand and distinguish the different topics within the text by combining the local feature and the text feature information. In particular, for texts with multiple or similar topics, the fine-grained feature can improve the classification's level of differentiation. In embodiments of the present disclosure, the method can further include performing word segmentation on the audio or text to obtain multiple word segmentation features and determining from the multiple word segmentation features the word segmentation feature that corresponds to the text information associated with the image candidate, as the text feature information comprises. The image candidate can be obtained based on a text or audio search. In embodiments of the present disclosure, the word segmentation features can be obtained using various methods based on the different classifications of the input information. For example, if the object to be processed is text, obtaining the word segmentation feature may involve performing preprocessing of the text, e.g., removing irrelevant words (a word that occurs frequently is not an important word for analysis), and then performing word segmentation processing on the text using the (multi-language supporting) word segmentation tool to obtain the word segmentation feature. Different word segmentation requirements can be met by selecting different modes (e.g., exact mode, full mode, etc.). If the object to be processed is audio, the procedure for capturing the word segmentation feature may include processing and recognizing the audio information in the early stages (e.g., using Mel-frequency-cepstrum coefficients to extract features from the audio signals and automatic audio recognition techniques) to convert the audio signal into text. Subsequent processes may be similar to text processing processes, which will not be repeated here. In embodiments of the present disclosure, the image candidate can comprise several image candidates obtained by performing recognition and searching of the text for subsequent determination of the text feature information. The reference method and approaches of the image candidate are not limited therein. In one example, if a text search is provided, the search text can be encoded into text vectors. Meanwhile, the system can select some images from a large image set or candidate set to form the initial set of image candidates. This set can be filtered through a prior computation or a simple image search procedure (e.g., quickly filtering metadata and labels based on the image) to obtain the initial set of image candidates. For each initial image candidate, the image-text similarity between the image vector and the search text vector can be calculated. If the image-text similarity meets a predetermined threshold, the initial image with the highest similarity to the search text can be determined as the image candidate. Once the image candidate has been determined, the corresponding text information (e.g., the object classification information in the image) can be generated.Once the text information and the word segmentation feature have been determined, the intersection of these two features can be identified. If the word segmentation lies within this intersection, the feature weighting of the word segmentation can be increased to obtain the text feature information. How to determine the text feature information has been described in detail above, and how to obtain the target feature information is described below. In embodiments of the present disclosure, the fine-grained feature information can comprise a fine-grained feature and a first weight corresponding to the fine-grained feature. Merging the fine-grained feature information and the global feature information to obtain the target feature information can include: updating the initial weight of the global feature information based on the first weight to obtain the second weight corresponding to the global feature information, wherein the sum of the first weight and the second weight satisfies the target weight value. The fine-grained feature and the global feature information can be weighted and merged based on the first weight and the second weight to determine the target feature information. In embodiments of the present disclosure, the second weighting, corresponding to the global feature information, can be obtained by updating the initial weighting of the global feature information. The weighting corresponding to the fine-grained feature can be the first weighting. The target weighting value can be the sum of the first weighting and the second weighting (e.g., a fixed value). The result of multiplying the first weighting and the fine-grained feature can be calculated, as can the result of multiplying the second weighting and the global feature information. Based on this, the target feature information can be determined. In one example, the initial weight corresponding to the global trait information can be α, and the weight corresponding to the fine-grained trait can be the first weight, denoted as β. The values of α and β can be determined based on experimentation or experience. Using the first weight β and the target weight value, the initial weight α can be updated to obtain the updated second weight value α', and the updated α' satisfies α' + β = 1. Based on this, the fine-grained traits (denoted as Flocal) are weighted using β, and the global trait information (denoted as Fglobal) can be weighted using the second weight α' to obtain the target trait information; that is, the target trait information is equal to β · Flocal + α' · Fglobal. In embodiments of the present disclosure, merging the fine-grained feature and the global feature information allows for better capture of local details and global structures of the object being processed. The fine-grained features can focus on the local details of the object being processed (e.g., local shape, texture, etc.), while the global features can focus on the overall structure of the object being processed. The combination of the two can help to accurately identify the result object in situations with complex backgrounds or multiple targets. Examples of determining the target feature information have been described in detail above, and a training process for a deep learning model that utilizes an object processing method is described below. In embodiments of the present disclosure, the object processing method can be performed by a deep learning model comprising an information detection layer and an information recognition layer. The training process of the deep learning model can include inputting a first pattern object to be processed into the information detection layer of the deep learning model to obtain pattern target information, and inputting the first pattern object to be processed and a second pattern object to be processed into the information recognition layer to obtain the first local pattern feature information, the first global pattern feature information, the second local pattern feature information, and the second global pattern feature information. The modality of the first pattern object to be processed can be different from the modality of the second pattern object to be processed. In embodiments of the present disclosure, the deep learning model can be a cross-modality retrieval model, for example, a CLIP (Contrastive Language-Image Pre-training) model. The information detection layer can be configured as a first network layer designed to perform detection on a sample image or a single frame of a sample video to obtain the essential feature information. The information recognition layer can be configured as a second network layer designed to perform feature extraction on the object of the sample image or video to be processed in order to recognize the initial feature information. In embodiments of the present disclosure, the information detection layer may be configured to perform detection on the first pattern object to be processed (e.g., a pattern image or pattern video frame) in order to obtain pattern target frame information, which includes detection frame coordinates (target frame coordinates), a confidence value, and a classification probability. The information recognition layer may be configured to perform feature extraction on the first pattern object to be processed in order to obtain the first local pattern feature information and the first global pattern feature information. The information recognition layer may be configured to perform feature extraction on the second pattern object to be processed (e.g.,to perform a sample text or sample audio) in order to obtain the second local sample feature information and the second global sample feature information. In embodiments of the present disclosure, the deep learning model can be trained using a loss function determined by the following operations: Determining, based on the pattern target frame information, the first essential pattern feature information and the second essential pattern feature information, which correspond to the first pattern object to be processed, respectively.The process involves determining the pattern-target image feature information and the pattern-target text feature information based on the first essential pattern feature information, the second essential pattern feature information, the first local pattern feature information, the first global pattern feature information, the second local pattern feature information, and the second global pattern feature information, in order to obtain pattern feature pairs, and determining the loss function based on the similarity between the pattern feature pairs and a label similarity. The label similarity can be obtained at least based on the initial similarity between the first and second pattern objects to be processed. The local pattern feature information can include the first local pattern feature information and the second local pattern feature information. The global pattern feature information can include the first global pattern feature information and the second global pattern feature information. The first and second patterns can only be used to distinguish between objects to be processed. The specific object represented here is not restricted. In embodiments of the present disclosure, the loss function can be used to measure the degree of similarity between patterns of different modalities in the multimodal model. The label similarity can be determined based on the initial similarity between different pattern objects to be processed. For example, the label similarity can be (0, 1), where 0 represents the lowest degree of similarity between the patterns of different modalities and 1 represents the highest degree of similarity between the patterns of different modalities. Using the example of the first pattern object to be processed as a pattern image and the second as a pattern text, the first essential pattern feature information can be the pattern target frame information corresponding to the pattern image, and the second essential pattern feature information can be the pattern text feature information corresponding to the pattern text. Thus, the pattern target image feature information can be determined based on the pattern target frame information, the local pattern image feature information, and the global pattern image feature information. Based on this, the pattern feature pair can be constructed using the pattern target image feature information and the pattern target text feature information. Then, the similarity between the pattern feature pairs can be calculated using a similarity calculation strategy (e.g.,...).The similarity value (cosine similarity) can be calculated, and the similarity value can be converted into a probability value. This probability value can reflect the likelihood of the pattern image matching the pattern text. For example, the sigmoid function can be used to convert the similarity into the probability, denoted as p. Then, a difference between the predicted probability (p) and the label similarity (0 or 1) can be calculated using a cross-entropy strategy to obtain the loss function by minimizing the error. In embodiments of the present disclosure, the pattern target frame information can represent the pattern detection frame position information corresponding to the first pattern object to be processed and the second pattern object to be processed, and the pattern object categorization information in the pattern detection frame. Fig. 3C is a schematic diagram showing a training process of a deep learning model according to some embodiments of the present disclosure. As shown in Fig. 3C, in 300C the first pattern object to be processed is a pattern image 411 and the second pattern object to be processed is a pattern text 412. An information detection layer 41 in the initial multimodal model is used to perform detection on the pattern image 411 in order to obtain pattern target frame information 413 and text information 418 corresponding to the object in the target frame. The pattern target frame information 413 can include the pattern target frame coordinates and the object categorization information in the pattern detection frame. An information recognition layer 42 in the initial multimodal model is used to recognize the pattern image 411 in order to obtain local pattern image feature information 414 and global pattern image feature information 415 corresponding to the pattern image 411. After obtaining the pattern-target frame information 413, the local pattern-image feature information 414, and the global pattern-image feature information 415, the local pattern-image feature information 414 can be weighted using the pattern-target frame information 413 to obtain the fine-grained pattern-image feature 416. Then, the global pattern-image feature information 415 and the fine-grained pattern-image feature 416 can be fused to obtain the pattern-target-image feature information (i.e., the coarse-fine-grained pattern-image fusion feature) 417. The word segmentation processing can be performed on the sample text 412 to obtain sample word segmentation features 421. The sample word segmentation features 421 and the text information 418 can be intersected to obtain the overlapping text information. If the sample word segmentation feature occurs in the overlapping text information, the weight of the sample word segmentation feature can be increased and serve as the sample text feature information 422. The sample text 412 can be recognized using the information recognition layer 42 to obtain local sample text feature information 423 and the global sample text feature information 424 corresponding to the sample text 412. After obtaining the pattern text feature information 422, the local pattern text feature information 423, and the global pattern text feature information 424, the local pattern text feature information 423 can be weighted using the pattern text feature information 422 to obtain the fine-grained pattern text feature 425. Then, the global pattern text feature information 424 and the fine-grained pattern text feature 425 can be fused to obtain the pattern target text feature information (i.e., the coarse-fine-grained pattern text fusion feature information) 426. After obtaining the pattern-target image feature information 417 and the pattern-target text feature information 426, pattern feature pairs 430 can be obtained. Then the loss function 440 can be determined based on the similarity between the pattern feature pairs 430 and the label similarity, and a trained target multimodal model 450 can be obtained. Based on the object processing method described above, this disclosure also provides an object processing device. The device is described in detail below in conjunction with Fig. 4. Fig. 4 is a schematic block diagram of the object processing device 400 according to embodiments of the present disclosure. As shown in Fig. 4, the object processing device 400 has an information output module 401, a target feature information reference module 402 and a result generation module 403. In embodiments of the present disclosure, the object processing device 400 can be configured to implement the object processing method according to embodiments of the present disclosure. For example, the information output module 401 can perform operation S201, which is used to process the object to be processed, to obtain essential feature information and initial feature information of the object to be processed, where the initial feature information includes local feature information and global feature information. The target feature information reference module 402 can be configured, for example, to perform operation S202 to merge the essential feature information, the local feature information, and the global feature information to obtain the target feature information of the object to be processed. The essential feature information can be target framework information or text feature information based on the object to be processed in different modalities. The result generation module 403 can be configured, for example, to perform operation S203 to determine an object that satisfies the similarity condition regarding the target feature information of the object being processed, as the result object corresponding to the object being processed. The modality of the object being processed can differ from the modality of the result object. The modality can include text, image, video, and / or audio. In embodiments of the present disclosure, based on the information output module 401, the target feature information reference module 402, and the result generation module 403 of the object processing device 400, the target feature information, which contains the coarse-grained fusion information, can be obtained by fusing the essential feature information, the local feature information, and the global feature information of the object to be processed. Then, the result object corresponding to the object to be processed can be determined. Since the target feature information is based on the essential feature information (i.e.,By retaining the original global features, the fine-grained feature for the extracted target frame information or text feature information can be added, and the perception of the fine-grained image and text information during cross-modal retrieval can be improved to enhance the overall feature representation capability and further improve the user experience. The information output module 401, the target feature information reference module 402, and the result generation module 403 can be combined in one module, or the information output module 401, the target feature information reference module 402, and the result generation module 403 can each be divided into several modules. In some other embodiments, at least some functions of the information output module 401, the target feature information reference module 402, and / or the result generation module 403 can be combined with at least some functions of another module, and the combination can be implemented in a single module. In embodiments of the present disclosure, the information output module 401, the target feature information reference module 402, and / or the result generation module 403 can be implemented at least as a hardware circuit, e.g.,The module may be implemented as a field-programmable gate array (FPGA), programmable logic array (PLA), single-chip system, multi-chip module (system-in-package), application-specific integrated circuit (ASIC), or other suitable hardware or firmware for integrating or packaging circuits, or in a suitable combination of software, hardware, and firmware. In some other embodiments, the information output module 401, the target feature information reference module 402, and / or the result generation module 403 may be implemented, at least partially, as a computer program module which, when executed by a computer, causes the computer to perform the functions of the corresponding module. In embodiments of the present disclosure, the target feature information reference module 402 can comprise a submodule for processing local feature information and a fusion processing submodule. The submodule for processing local feature information can be configured to process the local feature information based on the essential feature information in order to obtain the fine-grained feature information. The fusion processing submodule can be configured to fuse the fine-grained feature information and the global feature information in order to obtain the target feature information. In embodiments of the present disclosure, if the object to be processed is an image or video, the essential feature information can be the target frame information. The target frame information can be obtained by performing a detection on the object to be processed. The submodule for processing local feature information can include a unit for determining fine-grained feature information, which is configured to determine the fine-grained feature information based on the image information overlapping between the local feature information and the target frame information. In embodiments of the present disclosure, the essential feature information, if the object to be processed is speech audio or text, can be the text feature information. The submodule for processing local feature information can include a unit for determining fine-grained feature information, which is configured to determine the fine-grained feature information based on the text information overlapping between the local feature information and the text feature information. In embodiments of the present disclosure, the device may further comprise a word segmentation processing module and a text feature information determination module. The word segmentation processing module may be configured to perform word segmentation processing on speech audio or text to obtain multiple word segmentation features. The text feature information determination module may be configured to determine, from the multiple word segmentation features, the word segmentation feature that corresponds to the text information associated with the image candidate, as the text feature information. The image candidate may be obtained based on text or speech audio retrieval. In embodiments of the present disclosure, the fine-grained feature information can comprise the fine-grained features and the first weight corresponding to the fine-grained features. The fusion processing submodule can include a unit for acquiring a second weight and a unit for determining target feature information. The unit for acquiring a second weight can be configured to update the initial weight of the global feature information based on the first weight to obtain the second weight corresponding to the global feature information. The sum of the first weight and the second weight can satisfy a target weight value.The unit for determining target feature information can be set up to determine the target feature information by fusing the fine-grained features and the global feature information based on the first weighting and the second weighting. In embodiments of the present disclosure, the object processing method can be performed by a deep learning model. The deep learning model can have an information detection layer and an information recognition layer. The training process of the deep learning model can include inputting the first pattern object to be processed into the information detection layer of the deep learning model to obtain the pattern target frame information, and inputting the first and second pattern objects to be processed into the information recognition layer to obtain the first local pattern feature information, the first global pattern feature information, the second local pattern feature information, and the second global pattern feature information. The modality of the first pattern object to be processed can be different from the modality of the second pattern object to be processed. In embodiments of the present disclosure, the deep learning model can be trained by the loss function determined by the following operations: Determining, based on the pattern target frame information, the first essential pattern feature information and the second essential pattern feature information, which correspond to the first pattern object to be processed, respectively.The second pattern object to be processed corresponds to the pattern target image feature information and the pattern target text feature information, based on the first essential pattern feature information, the second essential pattern feature information, the first local pattern feature information, the first global pattern feature information, the second local pattern feature information, and the second global pattern feature information, in order to obtain the pattern feature pairs, and the loss function is determined based on the similarity between the pattern feature pairs and the label similarity. The label similarity can be obtained at least based on the initial similarity between the first pattern object to be processed and the second pattern object to be processed. Fig. 5 is a schematic block diagram of an electronic device for implementing the object processing method according to some embodiments of the present disclosure. The electronic device may comprise various types of digital computers, such as laptops, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The electronic device may also comprise various types of mobile devices, such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The components, connections and relationships between the components, and functions of the components described in this disclosure are merely examples and are not intended to limit the description of this disclosure and / or the claimed implementations thereof. As shown in Fig. 5, the device 500 includes an arithmetic unit 501. The arithmetic unit 501 can be configured to perform various suitable actions and processes according to the computer program stored in the read-only memory (ROM) 502 and the computer program loaded into the random-access memory (RAM) 503 by a memory unit 508. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The arithmetic unit 501, the ROM 502, and the RAM 503 are interconnected via a bus 504. An input / output interface (I / O interface) 505 is also connected to the bus 504. Several components in the electronic device 500 are connected to the I / O interface 505, including an input device 506 such as a keyboard, mouse, and the like; an output device 507 such as various types of screens, speakers, and the like; a storage device 508 such as magnetic disks, optical disks, and the like; and a communication device 509 such as a network card, modem, wireless communication transceiver, and the like. The communication device 509 enables the device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunications networks. The Computing Unit 501 can incorporate various general-purpose and / or specialized processing assemblies with processing and computing capabilities. Examples of the Computing Unit 501 include a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) chips, various computation units running machine learning algorithms, a digital signal processor (DSP), and any suitable processor, control unit, microcontroller, and the like. The Computing Unit 501 can perform the various procedures and processing operations described above, such as the avatar control procedure. For example, in some embodiments, the avatar control procedure can be implemented as a computer software program physically embodied on a machine-readable medium such as the Storage Unit 508.In some embodiments, the computer program can be loaded wholly or partially onto and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the processing unit 501, one or more steps of the avatar control procedure described above can be performed. In some other embodiments, the processing unit 501 can be configured to perform the avatar control procedure in another suitable way (for example, by means of firmware). The various embodiments of the systems and procedures described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or a combination thereof. These various embodiments can include implementations in one or more computer programs, wherein the one or more computer programs are executable and / or interpretable on a programmable system comprising at least one programmable processor.The programmable processor may comprise a programmable special-purpose or general-purpose processor capable of receiving data and instructions from a storage system, at least one input device and at least one output device, and of transmitting data and instructions to the storage system, the at least one input device and the at least one output device. The program code used to implement the method according to this disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or control unit of a general-purpose computer, a specialized computer, or another programmable data processing device. When executed by the processor or control unit, the program code then causes the functions / operations specified in the flowcharts and / or block diagrams to be carried out. The program code can be executed entirely on one machine, partially on one machine, partially on one machine as a standalone software package and partially on a remote machine, or entirely on a remote machine or server. In the context of this disclosure, the machine-readable medium can be a physical medium that can contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. The machine-readable medium can be a machine-readable signaling medium or a machine-readable storage medium. The machine-readable medium can include, among other things, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, such a device or apparatus, or a suitable combination thereof.Other specific examples of machine-readable storage media include electrical connections with one or more wires, portable computer drives, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, CD-ROMs, optical storage devices, magnetic storage devices, or any suitable combination thereof. To provide user interaction, the systems and procedures described herein can be implemented on a computer. The computer may have a display device (for example, a CRT or LCD monitor (cathode ray tube or liquid crystal display monitor), configured to show information to the user, as well as a keyboard and a pointing device (for example, a mouse or trackball) through which the user can input information to the computer. Other types of devices may also be configured to provide user interaction. Feedback provided to the user may be any form of sensory feedback (e.g., visual, auditory, or tactile feedback). Input from the user may be received in any form (including auditory, speech, or tactile input). The systems and procedures described herein can be implemented in computer systems that have backend components (for example, data servers), or computer systems that have middleware components (for example, an application server), or computer systems that have frontend components (for example, a user computer with a graphical user interface or a web browser, where the user can interact with implementations of the systems and procedures described herein via the graphical user interface or web browser), or computer systems in any combination of backend, middleware, or frontend components. The system components can be interconnected by digital data communication of any form and medium (for example, a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the internet. The computer system can consist of a client and a server. The client and server are generally remote from each other and typically interact via a communication network. A client-server relationship can be established by running computer programs on the respective computers that have such a client-server relationship. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a type of hosting product in a cloud computing service system designed to address the shortcomings of traditional physical hosts and VPS (Virtual Private Server or simply VPS) services, such as management difficulties and poor commercial scalability. The server can also be a server in a distributed system or a server integrated with a blockchain. A person skilled in the art understands that the features described in various embodiments and / or in the claims of this disclosure can be combined and / or integrated in a variety of ways, even if such combinations or integrations are not expressly described in this disclosure. In particular, the features described in various embodiments and / or in the claims of this disclosure can be combined and / or integrated in a variety of ways without departing from the spirit and teaching of this disclosure. All such combinations and / or integrations fall within the scope of this disclosure. Although the present disclosure has been presented and described with reference to specific embodiments thereof, the person skilled in the art understands that various changes in form and detail can be made to the present disclosure without departing from the spirit and scope of the present disclosure as defined by the appended claims and their equivalents. Therefore, the scope of the present disclosure is not limited to the aforementioned embodiments, but is defined not only by the appended claims, but also by their equivalents. QUOTES INCLUDED IN THE DESCRIPTION This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature CN 202411998148.9
[0001]
Claims
Object processing procedures, comprising: processing an object to be processed to obtain essential feature information and initial feature information of the object to be processed, wherein the initial feature information includes local feature information and global feature information; performing fusion processing on the essential feature information, the local feature information and the global feature information to obtain target feature information of the object to be processed, wherein the essential feature information is target frame information or text feature information of the object to be processed of different modalities;and determining an object that satisfies a similarity condition regarding the target feature information of the object to be processed as a result object that corresponds to the object to be processed, wherein one modality of the object to be processed is different from one modality of the result object and the modality includes text, image, video and / or audio. The method of claim 1, wherein obtaining the target feature information based on the essential feature information, the local feature information, and the global feature information comprises: processing the local feature information based on the essential feature information to obtain fine-grained feature information; and performing fusion processing on the fine-grained feature information and the global feature information to obtain the target feature information. The method of claim 2, wherein: if the object to be processed is an image or a video, the essential feature information is target frame information and the target frame information is obtained by detecting the object to be processed; and the processing of the local feature information based on the essential feature information to obtain fine-grained feature information comprises: determining the fine-grained feature information based on image information overlapping between the local feature information and the target frame information. The method of claim 2, wherein: if the object to be processed is audio or text, the essential feature information is the text feature information; and the processing of the local feature information based on the essential feature information to obtain fine-grained feature information comprises: determining the fine-grained feature information based on text information overlapping between the local feature information and the text feature information. The method of claim 4, further comprising: performing word segmentation processing on the audio or the text to obtain multiple word segmentation features; and determining, from the multiple word segmentation features, a word segmentation feature that is identical to text information corresponding to an image candidate as the text feature information, wherein the image candidate is retrieved based on the text or the audio. The method of claim 2, wherein: the fine-grained feature information comprises a fine-grained feature and a first weight corresponding to the fine-grained feature; performing the fusion processing on the fine-grained feature information and the global feature information to obtain the target feature information comprises: updating an initial weight of the global feature information based on the first weight to obtain a second weight corresponding to the global feature information, wherein a sum of the first weight and the second weight satisfies a target weight value; and performing a weighted fusion on the fine-grained feature and the global feature information based on the first weight and the second weight, respectively, to determine the target feature information. The method of claim 1, wherein: the object processing method is performed by a deep learning model and the deep learning model comprises an information detection layer and an information recognition layer; and a training process of the deep learning model comprises: inputting a first pattern object to be processed into the information detection layer of the deep learning model to obtain pattern target frame information; and inputting the first pattern object to be processed and a second pattern object to be processed into the information recognition layer to obtain first local pattern feature information, first global pattern feature information, second local pattern feature information, and second global pattern feature information, wherein one modality of the first pattern object to be processed is different from one modality of the second pattern object to be processed. The method of claim 7, wherein the deep learning model is trained by a loss function and the determination of the loss function comprises: determining first essential pattern feature information and second essential pattern feature information corresponding to the first pattern object to be processed and the second pattern object to be processed, respectively, based on the pattern target frame information; determining pattern target image feature information and pattern target text feature information based on the first essential pattern feature information, the second essential pattern feature information, the first local pattern feature information, the first global pattern feature information, the second local pattern feature information, and the second global pattern feature information to obtain pattern feature pairs;and determining the loss function based on similarities between the pattern feature pairs and label similarities, wherein the label similarities are obtained at least on the basis of an original similarity between the first pattern object to be processed and the second pattern object to be processed. A computer-readable storage medium on which one or more computer programs are stored which, when executed by one or more processors, cause the one or more processors to: process an object to be processed in order to obtain essential feature information and initial feature information of the object to be processed, wherein the initial feature information includes local feature information and global feature information; perform fusion processing on the essential feature information, the local feature information, and the global feature information to obtain target feature information of the object to be processed, wherein the essential feature information is target frame information or text feature information of the object to be processed of different modalities;to determine an object that satisfies a similarity condition regarding the target feature information of the object to be processed as a result object that corresponds to the object to be processed, wherein one modality of the object to be processed is different from one modality of the result object and the modality includes text, image, video and / or audio. Storage medium according to claim 9, wherein the one or more processors are further configured to: process the local feature information based on the essential feature information in order to obtain fine-grained feature information; and to perform fusion processing on the fine-grained feature information and the global feature information in order to obtain the target feature information. Storage medium according to claim 10, wherein: if the object to be processed is an image or a video, the essential feature information is target frame information and the target frame information is obtained by detecting the object to be processed; and the one or more processors are further configured to determine the fine-grained feature information based on image information overlapping between the local feature information and the target frame information. Storage medium according to claim 10, wherein: if the object to be processed is audio or text, the essential feature information is the text feature information; and the one or more processors are further configured to determine the fine-grained feature information based on text information overlapping between the local feature information and the text feature information. Electronic device comprising: one or more processors; and one or more memories for storing one or more computer programs which, when executed by the one or more processors, cause the one or more processors to: process an object to be processed in order to obtain essential feature information and initial feature information of the object to be processed, wherein the initial feature information includes local feature information and global feature information; perform fusion processing on the essential feature information, the local feature information, and the global feature information to obtain target feature information of the object to be processed, wherein the essential feature information is target frame information or text feature information of the object to be processed of different modalities;to determine an object that satisfies a similarity condition regarding the target feature information of the object to be processed as a result object that corresponds to the object to be processed, wherein one modality of the object to be processed is different from one modality of the result object and the modality includes text, image, video and / or audio. The device according to claim 13, wherein the one or more processors are further configured to: process the local feature information based on the essential feature information to obtain fine-grained feature information; and perform fusion processing on the fine-grained feature information and the global feature information to obtain the target feature information. Device according to claim 14, wherein: if the object to be processed is an image or a video, the essential feature information is target frame information and the target frame information is obtained by detecting the object to be processed; and the one or more processors are further configured to determine the fine-grained feature information based on image information overlapping between the local feature information and the target frame information. Device according to claim 14, wherein: if the object to be processed is audio or text, the essential feature information is the text feature information; and the one or more processors are further configured to determine the fine-grained feature information based on text information overlapping between the local feature information and the text feature information. The device according to claim 16, wherein the one or more processors are further configured to: perform word segmentation processing on the audio or the text to obtain multiple word segmentation features; and determine from the multiple word segmentation features a word segmentation feature that is identical to text information corresponding to an image candidate, as the text feature information, wherein the image candidate is retrieved based on the text or the audio. The device according to claim 14, wherein: the fine-grained feature information comprises a fine-grained feature and a first weight corresponding to the fine-grained feature; the one or more processors are further configured to: update an initial weight of the global feature information based on the first weight to obtain a second weight corresponding to the global feature information, wherein a sum of the first weight and the second weight satisfies a target weight value; and perform a weighted fusion on the fine-grained feature and the global feature information based on the first weight and the second weight, respectively, to determine the target feature information. The device according to claim 13, wherein: the object processing method is performed by a deep learning model and the deep learning model comprises an information detection layer and an information recognition layer; and a training process of the deep learning model comprises: inputting a first pattern object to be processed into the information detection layer of the deep learning model to obtain pattern target frame information; and inputting the first pattern object to be processed and a second pattern object to be processed into the information recognition layer to obtain first local pattern feature information, first global pattern feature information, second local pattern feature information, and second global pattern feature information, wherein one modality of the first pattern object to be processed is different from one modality of the second pattern object to be processed. The device according to claim 19, wherein the deep learning model is trained by a loss function and the determination of the loss function comprises: determining first essential pattern feature information and second essential pattern feature information corresponding to the first pattern object to be processed and the second pattern object to be processed, respectively, based on the pattern target frame information; determining pattern target image feature information and pattern target text feature information based on the first essential pattern feature information, the second essential pattern feature information, the first local pattern feature information, the first global pattern feature information, the second local pattern feature information, and the second global pattern feature information to obtain pattern feature pairs;and determining the loss function based on similarities between the pattern feature pairs and label similarities, wherein the label similarities are obtained at least on the basis of an original similarity between the first pattern object to be processed and the second pattern object to be processed.