Deep learning-based saliency target sorting method and system
By adopting the deep learning dual-branch structure model in the significance target sorting, the competitive relationship between significance targets is learned, and the problem of relying on explicit visual cues in the existing technology is solved, and a more accurate significance target sorting is achieved.
Patent Information
- Application Number
- CN202510270514.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-06
AI Technical Summary
The existing significance target sorting method mainly relies on explicit visual significance cues, neglecting the learning of competitive relationships between significance targets, resulting in insufficient quantification of competition results and affecting the sorting performance.
Using a two-branch structure model based on deep learning, visual branches extract significant features at the visual level, semantic branches extract semantic features of the whole picture and the target, and through semantic knowledge interaction and directional relationship reasoning, the competitive relationship between significant goals is learned to obtain more accurate significance scores.
By integrating multiple features, the model can more accurately quantify the competitive results between significant targets and improve the accuracy of significance target sorting.
Smart Images

Figure CN120107754A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of salient object ranking, and in particular to a salient object ranking method and system based on deep learning. Background Art
[0002] Salient Object Ranking (SOR) is an emerging computer vision task that aims to prioritize objects in a scene according to the order in which human attention shifts. Islam et al. first proposed the concept of saliency ranking and predicted the relative saliency value at the pixel level through a hierarchical representation based on relative saliency and a phased refinement method. However, this method is not completely consistent with human's real visual perception. Subsequently, Siris et al. proposed target-based saliency ranking based on psychology and neuroscience research, designed a bidirectional attention mechanism, and predicted the saliency ranking of objects according to the order in which human attention shifts for the first time. At the same time, it provided a large-scale benchmark dataset, laying the foundation for subsequent research.
[0003] Fang et al. proposed a plug-and-play SOR branch for joint training and designed an attention module for the branch that retains position information to improve the ranking performance. Liu et al. built a graph reasoning module to enhance the competitive interaction between targets and provided a more accurately annotated SOR dataset. Tian et al. combined spatial attention and target-based attention to complete saliency ranking. Recently, Sun et al. proposed a partitioning paradigm-based method that uses dense pyramid transformers to alleviate ambiguity in ranking. Guan et al. introduced human posture cues for saliency ranking and achieved high-level interaction between targets and the surrounding environment.
[0004] Although the above studies have made some progress in the field of SOR, their solutions mainly rely on explicit visual saliency clues. Figure 1 As shown in the figure, A and B represent two different objects, the bidirectional arrows represent the interaction between the two, and the unidirectional arrows represent the characteristics of the objects themselves. Figure 1 The intrinsic features of the target and their dynamic interactions form a salient relationship. In this context, the intrinsic features of the target and its dynamic interaction with the surrounding environment constitute visual saliency cues. The above methods focus more on the superposition of these independent visual cues and often ignore the learning of the competitive relationship between salient targets. This emphasis leads to inaccurate quantification of the competition results, which ultimately affects the performance of salient target sorting. Summary of the invention
[0005] In order to solve the above problems, the present invention proposes a salient target ranking method and system based on deep learning, which widely integrates multiple features and encourages the model to autonomously learn the competitive relationship between salient targets, so as to quantify the competition results more accurately.
[0006] According to some embodiments, the present disclosure adopts the following technical solutions: A salient object ranking method based on deep learning, comprising: Get the image to be processed; The image is input into the trained object ranking model to identify the salient objects and calculate the salient scores, and the final salient object ranking result is obtained; Among them, the target ranking model adopts a dual-branch structure: the visual branch extracts the salient features of the target at the visual level, and different targets obtain their salient scores based on their own salient features; the semantic branch extracts the semantic features of the entire image and the semantic features of the target, and performs semantic knowledge interaction and directional relationship reasoning in turn to obtain the salient scores of each target. Based on the salient scores on the two branches, the final salient targets and ranking results are obtained.
[0007] According to some embodiments, the present disclosure adopts the following technical solutions: A salient object ranking system based on deep learning, comprising: The image acquisition module is configured to: acquire an image to be processed; The object ranking module is configured to: input the image into the trained object ranking model, identify the significant objects and calculate the significant scores, and obtain the final significant object ranking result; Among them, the target ranking model adopts a dual-branch structure: the visual branch extracts the salient features of the target at the visual level, and different targets obtain their salient scores based on their own salient features; the semantic branch extracts the semantic features of the entire image and the semantic features of the target, and performs semantic knowledge interaction and directional relationship reasoning in turn to obtain the salient scores of each target. Based on the salient scores on the two branches, the final salient targets and ranking results are obtained.
[0008] According to some embodiments, the present disclosure adopts the following technical solutions: A computer program product includes a computer program, and when the computer program is executed by a processor, the method for ranking significant objects based on deep learning is implemented.
[0009] According to some embodiments, the present disclosure adopts the following technical solutions: A non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, a method for ranking significant objects based on deep learning is implemented.
[0010] According to some embodiments, the present disclosure adopts the following technical solutions: An electronic device comprises: a processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes the method for ranking significant objects based on deep learning.
[0011] Compared with the prior art, the present invention has the following beneficial effects: The present invention calculates the significance score of each target from the visual level and the semantic level respectively through a dual-branch structure, obtains the final significance target and ranking result based on the significance scores on the two branches, makes full use of the multimodality of images and texts, and improves the accuracy of significance target ranking.
[0012] The present invention designs a global visual feature extractor (GVAE), which integrates multiple saliency information into saliency query using an irregular paradigm from the visual level, focuses on the competitive relationship between salient targets, thereby stimulating the reasoning of competitive relationships and obtaining accurate saliency representation.
[0013] The present invention designs a contextual relation reasoning (CRR) module, which focuses on the semantic level. Based on the semantic representation of the target, a target graph is constructed, the graph is input into a graph convolutional neural network, and the significance level of the target is obtained after a fully connected layer. Its essence is to use text generation inertia to reversely infer the significance relationship of each target in the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings constituting a part of the present disclosure are used to provide a further understanding of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation on the present disclosure.
[0015] Figure 1 A graph of factors that rank significant targets.
[0016] Figure 2 This is a model structure diagram of Example 1.
[0017] Figure 3 This is a structural diagram of the global visual feature extractor (GVAE) of Example 1.
[0018] Figure 4 This is a structural diagram of the contextual relationship reasoning (CRR) module of Example 1.
[0019] Figure 5 This is a diagram of the multimodal guidance method of Example 1. DETAILED DESCRIPTION
[0020] The present disclosure is further described below in conjunction with the accompanying drawings and embodiments.
[0021] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present disclosure belongs.
[0022] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0023] Example 1 In one embodiment of the present disclosure, a method for ranking salient objects based on deep learning is provided, comprising: Step S1: obtaining an image to be processed; Step S2: input the image into the trained object ranking model to identify the significant objects and calculate the significant scores, and obtain the final significant object ranking result; Among them, the target ranking model adopts a dual-branch structure: the visual branch extracts the salient features of the target at the visual level, and different targets obtain their salient scores based on their own salient features; the semantic branch extracts the semantic features of the entire image and the semantic features of the target, and performs semantic knowledge interaction and directional relationship reasoning in turn to obtain the salient scores of each target. Based on the salient scores on the two branches, the final salient targets and ranking results are obtained.
[0024] As an embodiment, a method for ranking significant targets based on deep learning disclosed in the present invention adopts a target ranking model GVCR constructed based on deep learning, widely integrates multiple features, and encourages the model to autonomously learn the competitive relationship between significant targets, so as to more accurately quantify the competition results. The specific implementation process is as follows: As shown in Figure 2, the GVCR architecture includes a visual branch and a semantic branch. In the visual branch, GVAE (global visual feature extractor) extracts information that affects the saliency level at the visual level and integrates it into the saliency query (i.e., global visual query). The yellow part in the lower right corner vividly shows the competition results of the target on the three saliency influencing factors (position, scale, brightness) and the final competition score. In the semantic branch, the CRR (contextual relationship reasoning) module is used to guide the model to reversely reason about the saliency relationship of the target using contextual text information as clues, specifically: 1. Visual Branch The visual branch is used to extract the salient features of the target at the visual level. Different targets obtain their salient scores based on their own salient features, including GLEE, SAM, MS and global visual feature extractor; the GLEE is used to predict the bounding box from the image; the SAM is used to extract the features guided by the image mask and predict the target mask; the MS is used to select the target features from the features guided by the image mask as the input of the global visual feature extractor; the global visual feature extractor integrates the salient information in the target features into the salient query using an irregular paradigm, stimulates the reasoning of competitive relationships, obtains the competitiveness representation of the target, and infers the salient score of the target.
[0025] 1. GLEE predicts the bounding box, specifically: The image to be processed is input into a large-scale trained GLEE model to obtain a high-quality bounding box prediction result. The GLEE model includes a backbone network, a text encoder, a visual prompter and a decoder. The backbone network is used to extract image features, the text encoder and the visual prompter are used to provide text and visual prompts, and the decoder is used to output the bounding box prediction result. During the training process, the decoder simultaneously processes three parts of information from the text encoder, the visual prompter and the image features to predict the bounding box. After large-scale training, this embodiment retains the image backbone network and the decoder part, and its internal parameters will be fixed, so as to effectively obtain high-quality bounding box prediction results: .
[0026] 2. SAM extracts the features guided by the image mask and predicts the target mask, specifically: First, the image is input into the encoder from SAM to extract image mask-guided features, which is called image encoding.
[0027] The decoder then receives the image encoding and uses the bounding box obtained in the previous step as a visual cue to predict a high-precision object mask.
[0028] Similarly, SAM, as a basic model, has also undergone large-scale training, and its parameters can be frozen without affecting the output mask effect. The output is recorded as: Target Mask With bounding box Satisfy between: , that is, for any , each All correspond to the only ,vice versa.
[0029] 3. MS selects target features from the features guided by the image mask, specifically: In order to make the following global visual feature extractor GVAE learn more effectively, this embodiment adds a Mask-guided Selection (MS) operation to select target features for GVAE and define a feature set , used to indicate the The features of the target, this set is encoded by the image generated by the encoder The elements in which only when the mask In Location When the value on is 1, the corresponding feature That is, only when the mask indicates that the position belongs to the target, the relevant features will be included in the feature set, which can be expressed as:
[0030]
[0031] in, represents the image encoding generated by the SAM encoder, Indicates The target mask, Represent the number of channels, height and width respectively, Indicates the selected Characteristics of a target.
[0032] This is different from RoI Pooling (region of interest pooling) or RoI Align (region of interest alignment) commonly used in previous SOR methods. The MS operation of this embodiment does not contain background pollution and can extract pure target-level information, thereby more effectively focusing on the target itself and improving the quality and accuracy of subsequent learning.
[0033] 4. The global visual feature extractor integrates the saliency information in the target features into the saliency query using an irregular paradigm, stimulates the reasoning of competitive relationships, obtains the competitiveness representation of the target, and infers the saliency score of the target, specifically: GVAE aims to extract information that affects the saliency level at the visual level and integrate it into the saliency query to achieve global learning. It learns to integrate multiple saliency information into the saliency query using an irregular paradigm from the visual level, thereby stimulating the reasoning of competitive relationships and obtaining accurate competitiveness representation. This embodiment introduces global visual query as a specific form of saliency query, which can be expressed as , , is a customized height and width; it is worth noting that in this embodiment, a different query is not assigned to each target, but a query is shared. This approach effectively integrates the salient features of different targets and promotes the universal learning of salient information extraction, which is the core of the feature extractor.
[0034] Figure 3 is the structure diagram of GVAE, such as Figure 3 As shown in the figure, GVAE performs cyclic feature extraction on different targets through global visual query, thereby exercising its feature extraction ability in iteration and realizing the function of feature extractor.
[0035] First, in order to facilitate calculation, the shape of the target feature is reshaped. The shape space dimension of , expressed as ,in, is the number of channels of image encoding, It is The length of the target feature.
[0036] Next, we use the cross-attention layer to build a recurrent attention module to fully motivate the feature extractor to learn the competitive relationship between targets. , perform the following operations to effectively extract the characteristics of the target itself, thereby inferring the competitiveness representation of the target, which is expressed as:
[0037]
[0038]
[0039]
[0040] , in, Indicates the initial , as the index increases, Constantly updated, Indicates The result after the update is also the Competitiveness of a target , ρ represents the one-dimensional convolution and tensor splitting operations, Serves as keys and values for cross-attention layers.
[0041] In this work, this embodiment takes The above operations enable global visual query to accurately extract salient information from irregularly aggregated target features at the visual level; global visual query By traversing different targets, we can get the competitiveness representation of different targets, thus achieving significance sorting:
[0042]
[0043] in, It is the recurrent attention module built in the previous step. represents one-dimensional convolution, represents the flattening operation, Represents the fully connected layer, and the final output target The significance score of .
[0044] It is worth noting that It is a global variable from beginning to end, rather than being owned by different targets separately. This method effectively aggregates comprehensive integrated strategies into the global visual query, thereby promoting its learning competition relationship.
[0045] In summary, GVAE can extract the salient features of the target at the visual level, such as scale, brightness, etc., and fully learn the expression of irregular paradigms; however, since the MS operation will lose the spatial position information of the target and it is difficult to establish the dynamic relationship between targets and between targets and backgrounds in the image, it is not enough to accurately judge the saliency level of the target. Therefore, a semantic branch, especially the CRR module, is proposed to jointly guide the sorting of salient targets.
[0046] 2. Semantic Branch The semantic branch extracts the semantic features of the whole image and the target semantic features, performs semantic knowledge interaction and directional relationship reasoning in turn, obtains the significance score of each target, and obtains the final significance target and ranking result based on the significance scores on the two branches, including BLIP, CLIP and contextual relationship reasoning modules; the BLIP is used to extract the semantic features of the whole image; the CLIP is used to extract the semantic features of the target; the contextual relationship reasoning module is used to perform semantic knowledge interaction and directional relationship reasoning to obtain the significance score of each target.
[0047] Specifically, given the superiority of BLIP in perceiving the entire image content and the excellent performance of CLIP in image classification tasks, the image to be processed is first input into the image encoder of BLIP for image encoding, and then the text encoder is used to extract the semantic features of the entire image. At the same time, the bounding box obtained by the visual branch is used to extract the semantic features of the entire image. The original image is cropped at the object level and the cropped result is input into CLIP to extract the semantic features of the object for subsequent relationship reasoning.
[0048] The CRR module focuses on the semantic level, using text to generate inertial reverse reasoning of the salient relationship between objects in the image; the semantic features of the whole image and target semantic features As the input of CRR, sufficient interactive perception and directional reasoning are performed to infer the significant relationship between targets.
[0049] Figure 4 This is the structural diagram of the contextual relation reasoning (CRR) module. Figure 4 (a), (b), and (c) respectively represent the number of targets. The connection status of the nodes when the value is 2, 3, or 4. The critical value at this time is 3, the number of neighbors selected is 2.
[0050] Inspired by the inertia of text generation, the CRR module is designed to use global semantic information and target semantic information to interactively infer relational clues, thereby helping to predict where human attention will turn. The CRR module makes up for the weakness of GVAE in spatial interaction, such as Figure 4 The shown CRR module consists of two parts: semantic knowledge interaction and directional relationship reasoning.
[0051] 1. Semantic knowledge interaction: the target perceives the global semantic information of the image in turn and determines its own semantic position in the image. Specifically: Target features extracted using CLIP , and the full-image features extracted by BLIP Perform a cross-attention operation:
[0052] in, Indicates The semantic features of the target, Represents the semantic features of the entire image, represents a one-dimensional convolution and tensor splitting operation, As the key and value of the cross-attention layer, the output That is the The semantic representation of each target in the image.
[0053] Through semantic knowledge interaction, each target can locate itself based on the global semantic information of the image, thus having a preliminary understanding of its own saliency.
[0054] 2. Directional relationship reasoning: Use graph neural networks to aggregate features on target nodes, thereby promoting dynamic interactions between targets. Specifically: (1) Constructing a target graph based on the semantic representation of the target For the task of ranking significant targets, the relationship between targets is complex and nonlinear, and graph neural networks are excellent in processing non-Euclidean data; design a critical value , when the number of targets is less than When the number of targets is greater than or equal to When using K-nearest neighbors (select Neighbors) promote the prediction of node attributes, which can use prior information to reduce the difficulty of learning and avoid information redundancy; Figure 4 Shown when , The node connection status at that time.
[0055] (2) The graph is input into the graph convolutional neural network. After passing through the fully connected layer, the significance level of the target is obtained, which can be expressed as:
[0056]
[0057]
[0058] in, represents a set of target nodes and , represents the edge set, Representation Node of nearest neighbors, , , They represent graph convolutional neural network, activation function and fully connected layer respectively.
[0059] 3. Based on the significance scores on the two branches, the final significance target and ranking results are obtained The ranking of salient targets is determined jointly by GVAE and CRR as follows:
[0060] in, is a trainable parameter. Finally, the mask of the salient target and the sorting result are combined to get the final output.
[0061] In addition to the model structure of the above-mentioned salient object ranking model GVCR, a variety of multimodal guidance methods are also provided to enable the model to achieve better fusion performance. Figure 5 There are two different multimodal instruction methods from GVCR: The first way, such as Figure 5 As shown in (a), CLIP is used to extract image semantic features and generate semantic queries, which are then enhanced using a recurrent attention module to work in conjunction with the global visual query.
[0062] In this method, only the target features on the GVCR visual branch are used to extract visual features and perceive semantic information respectively, so that significant features can be extracted with emphasis on two different modalities. After adjustment by the fully connected layer, the semantic features extracted by CLIP are used as semantic queries, and then traversed on all targets, and the significant contribution that semantic information can provide on the target is learned through the recurrent attention module. Finally, the enhanced semantic query is spliced with the global visual query, and the significant target ranking is achieved through the fully connected layer.
[0063] The second method is Figure 5 As shown in (b), BLIP and CLIP are still used to obtain the full-image semantic information and target semantic information respectively to achieve semantic knowledge interaction. The obtained semantic representation is fused with the enhanced global visual query and input into GCN to achieve significant target ranking.
[0064] The second method is different from GVCR in that the enhanced global visual query is first fused with the target semantic features that have undergone the semantic knowledge interaction stage, and then they are jointly input into GCN to achieve significant target ranking. Specifically, in the visual branch, GVAE is still used to learn the competitive relationship; in the semantic branch, the target features extracted by CLIP and the full-image features extracted by BLIP are cross-attentioned to achieve semantic knowledge interaction; the target semantic representation and the enhanced global visual query are then passed through the fully connected layer and then residually connected, and then input into GCN to achieve significant target ranking.
[0065] However, the first multimodal guidance method can only provide limited semantic information, and lacks direct relational reasoning guidance like GVCR for ranking salient targets; GVAE emphasizes the learning of the target's own salient characteristics, while semantic knowledge interaction pays more attention to the derivation of global relationships; the direct fusion of the outputs of the two in the second method will cause confusion, which is not conducive to network learning; in summary, the GVCR of this embodiment obtains rich semantic information, and the contextual relationship reasoning module only receives a unique input type, which can obtain more benefits under the multimodal guidance method.
[0066] Example 2 In one embodiment of the present disclosure, a salient object ranking system based on deep learning is provided, comprising: The image acquisition module is configured to: acquire an image to be processed; The object ranking module is configured to: input the image into the trained object ranking model, identify the significant objects and calculate the significant scores, and obtain the final significant object ranking result; Among them, the target ranking model adopts a dual-branch structure: the visual branch extracts the salient features of the target at the visual level, and different targets obtain their salient scores based on their own salient features; the semantic branch extracts the semantic features of the entire image and the semantic features of the target, and performs semantic knowledge interaction and directional relationship reasoning in turn to obtain the salient scores of each target. Based on the salient scores on the two branches, the final salient targets and ranking results are obtained.
[0067] Example 3 In one embodiment of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the method for ranking significant objects based on deep learning is implemented.
[0068] Example 4 In one embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the method for ranking significant objects based on deep learning is implemented.
[0069] Example 5 In one embodiment of the present disclosure, an electronic device is provided, comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the method for ranking significant objects based on deep learning.
[0070] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0071] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0072] Although the above describes the specific implementation methods of the present disclosure in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present disclosure. Technical personnel in the relevant field should understand that on the basis of the technical solution of the present disclosure, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present disclosure.
Claims
1. A method for ranking salient objects based on deep learning, characterized in that: include: Get the image to be processed; The image is input into the trained object ranking model to identify the salient objects and calculate the salient scores, and the final salient object ranking result is obtained; The target ranking model adopts a dual-branch structure: the visual branch extracts the salient features of the target at the visual level, and different targets obtain their salient scores based on their own salient features; The semantic branch extracts the semantic features of the entire image and the target, performs semantic knowledge interaction and directional relationship reasoning in turn, obtains the significance score of each target, and obtains the final significance target and ranking results based on the significance scores on the two branches.
2. A method for ranking salient objects based on deep learning as claimed in claim 1, characterized in that: The visual branch includes GLEE, SAM, MS and a global visual feature extractor; The GLEE is used to predict bounding boxes from images; The SAM is used to extract image mask-guided features and predict the target mask; The MS is used to select target features from the image mask-guided features as input to the global visual feature extractor; The global visual feature extractor integrates the saliency information in the target features into the saliency query using an irregular paradigm, stimulates the reasoning of competitive relationships, obtains the competitiveness representation of the target, and infers the saliency score of the target.
3. A method for ranking salient objects based on deep learning as claimed in claim 2, characterized in that: The global visual feature extractor uses cross attention to construct a recurrent attention module, which encourages the feature extractor to learn the competitive relationship between targets and obtain the competitiveness representation of different targets. Finally, after one-dimensional convolution, flattening operation and full connection layer, the saliency score of the target is output.
4. The method for ranking significant objects based on deep learning according to claim 1, characterized in that: The semantic branch includes BLIP, CLIP and contextual relationship reasoning modules; The BLIP is used to extract semantic features of the entire image; The CLIP is used to extract target semantic features; The contextual relationship reasoning module is used to perform semantic knowledge interaction and directional relationship reasoning to obtain the significance score of each target.
5. A method for ranking salient objects based on deep learning as claimed in claim 4, characterized in that: The semantic knowledge interaction is to perceive the global semantic information of the image in turn through the target and determine the semantic position of the target in the image, specifically: The target semantic features extracted by CLIP and the full-image semantic features extracted by BLIP are cross-attentioned to obtain the semantic representation of the target in the image.
6. A method for ranking salient objects based on deep learning as claimed in claim 4, characterized in that: The directional relationship reasoning uses graph neural networks to perform feature aggregation on target nodes, thereby promoting dynamic interaction between targets. Specifically: Based on the semantic representation of the target, a target graph is constructed; The graph is input into the graph convolutional neural network, and the significance level of the target is obtained after passing through the fully connected layer.
7. A salient object ranking system based on deep learning, characterized in that: include: The image acquisition module is configured to: acquire an image to be processed; The object ranking module is configured to: input the image into the trained object ranking model, identify the significant objects and calculate the significant scores, and obtain the final significant object ranking result; The target ranking model adopts a dual-branch structure: the visual branch extracts the salient features of the target at the visual level, and different targets obtain their salient scores based on their own salient features; The semantic branch extracts the semantic features of the entire image and the target, performs semantic knowledge interaction and directional relationship reasoning in turn, obtains the significance score of each target, and obtains the final significance target and ranking results based on the significance scores on the two branches.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for ranking significant objects based on deep learning as described in any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by the processor, a deep learning-based salient object ranking method as described in any one of claims 1 to 6 is implemented.
10. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes a method for ranking significant objects based on deep learning as described in any one of claims 1 to 6.