Image semantic coding and retrieval system based on visual truly

By constructing a visual semiotics-based image semantic coding and retrieval system, symbol vectors for the signifier, signified, and symbolic layers of images are built, solving the problem that existing image retrieval systems struggle to understand the deep semantics of images and achieving more accurate image retrieval results.

CN121636742APending Publication Date: 2026-03-10TAIZHOU SANXING ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing image retrieval systems struggle to deeply understand the symbolic structures, metaphorical meanings, and deep cultural contexts within images, resulting in insufficient interpretability when retrieving data in complex semantic scenarios. In particular, when faced with abstract concepts or emotional imagery, the retrieval results have a low relevance to the user's intent.

Method used

An image semantic coding and retrieval system based on visual semiotics is adopted. The system performs object detection on images through a three-layer symbol vector generation module, constructs symbol vectors of the signifier layer, signified layer, and symbol layer, and calculates the similarity between images by combining the comprehensive symbol similarity acquisition module to generate retrieval results.

Benefits of technology

It enables multi-level semantic expression of images, improving the accuracy of image retrieval. Especially when faced with images with complex semantics and diverse cultural backgrounds, it can more accurately reflect the physical characteristics, semantic concepts and symbolic meanings of the images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121636742A_ABST
    Figure CN121636742A_ABST
Patent Text Reader

Abstract

The invention discloses an image semantic coding and retrieval system based on visual semifluid, which relates to the technical field of image retrieval, and comprises a three-layer symbol vector generation module, a query image three-layer symbol vector construction module, a comprehensive symbol similarity acquisition module and a retrieval result set acquisition module. According to the system, through mutual cooperation of an image symbol feature extraction module, a three-layer symbol generation module, a symbol vector acquisition module and the like, each image can be converted into a symbol vector of an energy layer, a pointed layer and a symbol layer by the system, then a three-layer symbol vector of each image is obtained, and then the three-layer symbol similarity between a query image and each image is calculated; and finally, the image closest to the semantic of the query image is output, so that the purpose of accurate retrieval is achieved, the multilevel semantic expression of the image is constructed through a Fuliji method, finally, the most relevant retrieval result is returned through the retrieval result set acquisition module, and the accuracy of image retrieval is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image retrieval technology, specifically an image semantic encoding and retrieval system based on visual semiotics. Background Technology

[0002] With the rapid development of artificial intelligence and visual technology, image semantic retrieval has become an important research direction in the field of visual information processing. Current mainstream image retrieval systems typically rely on deep convolutional neural networks or joint vision-language models to encode features of images and text to achieve cross-modal retrieval. However, these models focus on learning semantic co-occurrence or statistical correlation, making it difficult to deeply understand the symbolic structures, metaphorical meanings, and deep cultural contexts within images, resulting in insufficient interpretability when retrieving complex semantic scenarios.

[0003] However, traditional image retrieval methods rely on visual features of images, such as color, texture, and shape, for comparison. These methods often fail to effectively capture the deep semantics in images when dealing with complex semantic hierarchical relationships, especially when faced with complex queries based on abstract concepts such as peace and holiness, or emotional imagery. They lack multi-level structured modeling of visual elements, semantic objects, and symbolic meanings in images, resulting in low relevance of retrieval results to user intent. Based on this, an image semantic encoding and retrieval system based on visual semiotics is proposed. Summary of the Invention

[0004] The purpose of this invention is to provide an image semantic encoding and retrieval system based on visual semiotics to solve the problems mentioned in the background art.

[0005] A visual semiotics-based image semantic encoding and retrieval system includes: The three-layer symbol vector generation module is used to perform object detection on each image in the image library, obtain the image symbol features of each image, and construct three-layer symbol vectors of each image in the image library: signifier layer, signified layer, and symbol layer based on the image symbol features. The query image three-layer symbol vector construction module constructs the three-layer symbol vector of the query image in the same way as obtaining the three-layer symbol vector of each image when the user uploads a query image, thus obtaining the symbol vectors of the signifier layer, signified layer and symbol layer of the query image. The comprehensive symbol similarity acquisition module calculates the three-layer symbol similarity between the query image and each image based on the three-layer symbol vectors of each image in the image library and the three-layer symbol vectors of the query image. It then calculates the comprehensive symbol similarity, where the three-layer symbol similarity is the symbol similarity of the signifier layer, the signified layer, and the symbol layer, respectively. The search result set acquisition module obtains the search result set based on the sorting of three-level symbol similarity, and displays each image in the search result set according to the image order in the search result set.

[0006] As a further aspect of the present invention, the specific method for constructing the symbol vectors of each image indexer layer is as follows: Taking any image in the image library as the target image, the visual symbols A(k) in the target image are identified through object detection. The signifier layer symbol vector of the target image is represented as N(1) = [M(1), M(2), ..., M(Dn)], where M(k) is a non-negative real number, representing the weight value of different visual symbols A(k) in the target image, representing the contribution of different visual symbols to the overall target image. Here, k refers to different visual symbols. Using the same analysis method, the signifier layer symbol vector N(i) corresponding to each image in the image library can be obtained, where i refers to different images in the image library.

[0007] As a further aspect of the present invention: the weight values ​​of different visual symbols A(k) in the image are obtained as follows: The region corresponding to each visual symbol A(k) in the target image is identified by object detection, and the number of pixels in the region corresponding to each visual symbol A(k) is calculated. Then, the proportion of the number of pixels in the region corresponding to each visual symbol A(k) to the total number of pixels in the whole image is calculated, and it is used as the weight value M(k) of each visual symbol A(k) in the target image. Thus, the signifier layer symbol vector of the target image is obtained: N(1) = [M(1), M(2), ..., M(Dn)], where Dn is the number of visual symbols contained in the target image.

[0008] As a further aspect of the present invention, the specific method for constructing the symbol vectors of each image-defined layer is as follows: Identify the different object category symbols L(s) in the target image, and the confidence scores K(s) corresponding to each object category, where s represents different object categories. Label the set of symbols for the indicated layer of the target image as [L(1), L(2), ..., L(d)], where d is the total number of object categories in the indicated layer of the target image, and L(s) represents different object category symbols in the target image. Obtain the sum of the confidence scores K(s) of all object categories in the target image. Use the ratio between the confidence scores K(s) of different object categories in the target image and the sum of the confidence scores K(s) of all different object categories as the confidence weight Q(s) of different object categories in the target image, and then obtain the indicated layer symbol vector of the target image: S(1) = [Q(1), Q(2), ..., Q(d)], where d is the total number of object categories in the target image. Using the same analysis method, the indicated layer symbol vector S(i) corresponding to each image in the image library can be obtained.

[0009] As a further aspect of the present invention, the specific method for constructing the symbol vectors of each image symbol layer is as follows: Based on the target image's reference layer symbol set [L(1), L(2), ..., L(d)], obtain the symbolic layer concept symbol X(s) corresponding to each object category symbol L(s) in the target image. Select a target category symbol from each object category symbol in the target image's reference layer symbol set, and obtain the paths between the target category symbol and its corresponding symbolic layer concept symbol. For any path, obtain each edge that makes up the path and the weight of each edge. Take the product of the weights of each edge as the path weight of the path, and then obtain the path weight corresponding to each path. Based on the number of paths between the target category symbol and its corresponding symbolic layer concept symbol, analyze and obtain the mapping coefficient between the target category symbol and its corresponding symbolic layer concept symbol. Perform the same analysis on other object category symbols in the reference layer symbol set, and then obtain the mapping coefficient Y(s) between each object category symbol in the reference layer symbol set and its corresponding symbolic layer concept symbol. Obtain the ratio T(s) between each mapping coefficient Y(s) and all mapping coefficients Y(s), and then obtain the symbolic layer symbol vector U(1) of the target image. Using the same analysis method, the symbol vector U(i) corresponding to each image in the image library can be obtained, where i represents different images in the image library.

[0010] As a further aspect of the present invention: when there is only one path between the target category symbol and its corresponding symbolic layer concept symbol, the path weight of the corresponding path is used as the mapping coefficient between the target category symbol and its corresponding symbolic layer concept symbol; when there is more than one path between the target category symbol and its corresponding symbolic layer concept symbol, the maximum path weight among the multiple path weights is used as the mapping coefficient between the target category symbol and its corresponding symbolic layer concept symbol.

[0011] As a further aspect of the present invention: the specific method for calculating the three-level symbol similarity between the query image and each of the other images is as follows: According to the similarity calculation formula, for each image in the query image and each image in the image library, the symbol vector similarity between them is calculated at the signifier layer, the signified layer, and the symbol layer, respectively denoted as the signifier layer symbol similarity Z1i, the signified layer symbol similarity Z2i, and the symbol layer symbol similarity Z3i.

[0012] As a further aspect of the present invention, the specific method for obtaining the comprehensive symbol similarity is as follows: The comprehensive symbol similarity symbol vector Vi is calculated by weighted summation, where Vi = β1 × Z1i + β2 × Z2i + β3 × Z3i, and the symbol vectors β1, β2, and β3 are the weight coefficients corresponding to the similarity of each layer, satisfying β1 + β2 + β3 = 1, and β1, β2, and β3 are all greater than or equal to 0.

[0013] As a further aspect of the present invention, the specific method for obtaining the retrieval result set based on the ranking of three-layer symbol similarity is as follows: Iterate through the three-level symbol similarity Vi between the query image and each other, sort the three-level symbol similarity Vi values ​​from largest to smallest, and select the top K images with the highest three-level symbol similarity as the search result set.

[0014] Compared with the prior art, the beneficial effects of the present invention are: This invention, through the coordinated operation of modules such as image symbol feature extraction, three-layer symbol generation, and symbol vector acquisition, enables the system to transform each image into symbol vectors at the signifier, signified, and symbolic layers. Each layer of symbol vectors corresponds to a different semantic dimension of the image, accurately reflecting its physical features, semantic concepts, and symbolic meaning. During a query, the system constructs a three-layer symbol vector for the query image, calculates the three-layer symbol similarity between the query image and images in the image library using a comprehensive symbol similarity acquisition module, and finally returns the most relevant search results through a search result set acquisition module, effectively improving the accuracy of image retrieval. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the system framework structure of the present invention. Detailed Implementation

[0016] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Example 1: Please refer to Figure 1 This application provides an image semantic encoding and retrieval system based on visual semiotics, including: The three-layer symbol vector generation module performs object detection on each image in the image library, acquires the image symbol features of each image, performs object detection on the three-layer symbol features of each image in the image library, generates three-layer symbol vectors for each image based on the three-layer symbols of each image, and provides a structured expression of the visual semantics of the images. Each image constructs a three-layer symbol vector to represent the semantics of the signifier layer, signified layer, and symbol layer of the image, respectively. The three-layer symbol vectors of each image are associated with the corresponding image ID and stored in the database. The three-layer symbols refer to the symbols corresponding to the signifier layer, signified layer, and symbol layer of each image, respectively. The signifier refers to a local visual segment or salient visual feature in an image, that is, the physical features of the specific object, color, shape, etc. in the image, which is output by the detection and segmentation network, such as Faster / Mask R-CNN, SViT-Det, etc. The referent refers to the scene, action, and direct concept constituted by the corresponding semantic concepts or entity categories in the image, such as pigeon, olive branch, and church, which are mapped by a classifier or a standard vocabulary library; Symbolism is an abstract allegory or symbolic meaning based on cultural context, such as peace, holiness, and mourning. It is generated by reasoning from the symbol knowledge graph and rule base of this invention and can be accompanied by confidence level. Object detection and semantic segmentation are performed on images in the image library to extract the main visual elements, such as people, objects, and scenes, forming basic visual symbols. These symbols are extracted using object detection algorithms, such as Faster symbol vectors, R-CNN symbol vectors, or Mask symbol vectors, and serve as the basis for constructing subsequent hierarchical symbol vectors.

[0018] For each image, three layers of symbols are generated. Each image is represented by symbol vectors at the signifier, signified, and symbolic levels. The signifier layer represents the specific physical features in the image, the signified layer represents the semantic concepts in the image, and the symbolic layer maps abstract meanings in conjunction with cultural context. These symbol vectors are ultimately associated with image IDs and stored in a database.

[0019] To obtain the signifier layer symbol vector, a global visual symbol dictionary is first constructed to characterize the main visible elements in the image. This dictionary contains multiple basic visual symbols. Each basic visual symbol corresponds to a main visual element. The main visual element refers to the typical object category, typical color area, typical texture or shape pattern of the image. Typical object categories include people, cars, trees, etc.; typical color areas include blue sky, green grass, etc.; typical texture or shape patterns include water ripples, brick wall texture, etc. Take any image in the image library as the target image, and identify each visual symbol A(k) in the target image through target detection. The symbol vector of the indexer layer of the target image is represented as N(1) = [M(1), M(2), ..., M(Dn)], where M(k) is a non-negative real number, representing the weight value of different visual symbols A(k) in the target image, representing the contribution of different visual symbols to the overall target image, where k refers to different visual symbols, and Dn is the number of visual symbols contained in the target image; The weight values ​​of different visual symbols A(k) in this image are obtained as follows: First, the regions corresponding to each visual symbol A(k) in the target image are identified by object detection or semantic segmentation algorithms, and the number of pixels in the regions corresponding to each visual symbol A(k) is calculated. Then, the proportion of the number of pixels in the regions corresponding to each visual symbol A(k) to the total number of pixels in the whole image is calculated, and this proportion is used as the weight value M(k) of each visual symbol A(k) in the target image. The value range is: 0≤M(k)≤1. Then, the signifier layer symbol vector of the target image is obtained: N(1) = [M(1), M(2), ⋯、M(Dn)]; Using the same analysis method, we can obtain the sign vector N(i) corresponding to each image in the image library, where i represents different images in the image library; By employing object detection and semantic segmentation, the main visual symbols in the image are identified, and a signifier layer symbol vector is calculated based on the weight of each visual symbol in the image. The weight values ​​reflect the degree of contribution of each visual symbol to the overall image. For obtaining the symbol vector of the pointed layer, for the target image, the different object category symbols L(s) in the target image are identified based on object detection, and the confidence K(s) corresponding to each object category are obtained. s represents different object categories, and the confidence K(s) is provided by the detection network, such as Mask R-CNN or YOLO, which represents the model's confidence in whether a certain object category appears in the image. For example, if a pigeon is detected, the object detection network will give a confidence value, such as 0.9, which means that there is 90% confidence in the existence of the pigeon in the image. Similarly, other detected objects, such as olive branches, will also have corresponding confidence values, such as 0.8, which means that there is 80% confidence in the existence of the olive branch. The set of symbols of the pointed layer of the target image is labeled as [L(1), L(2), ..., L(d)], where d is the total number of object categories in the pointed layer of the target image, and L(s) represents different object category symbols in the target image. It's important to note that this confidence score is part of existing object detection techniques. These networks process images using convolutional neural networks, detecting various objects and calculating a confidence score for each. This confidence score can be viewed as a probability value output by the object detection network, representing the reliability of the network's detection result for that object. The confidence score reflects the reliability of an object in the image. If an object has a high confidence score, we consider that object to contribute significantly to the semantics of the image, and therefore assign it a higher weight. In this way, we can ensure that important objects in the image contribute more to the symbol vector of the pointed-to layer, thus affecting the final semantic representation of the image. The sum of the confidence scores K(s) of all object categories in the target image is obtained. The ratio between the confidence scores K(s) of different object categories in the target image and the sum of the confidence scores K(s) of all different object categories is used as the confidence weight Q(s) of different object categories in the target image. Then, the indicated layer symbol vector of the target image is obtained: S(1) = [Q(1), Q(2), ..., Q(d)], where d is the total number of object categories in the target image. Using the same analysis method, we can obtain the layer symbol vector S(i) corresponding to each image in the image library, where i represents different images in the image library; For a target image, the object detection network calculates the indicated layer symbol vector for each object category based on confidence. The higher the confidence value, the greater the semantic contribution of the object to the image, and therefore the greater the weight of the object in the indicated layer symbol vector.

[0020] To obtain the symbol vectors of the symbol layer, a semantic knowledge graph is constructed between the object category symbols of the referent layer and the concept symbols of the symbol layer. This graph is used to represent the mapping relationship from the object category symbols of the referent layer to the concept symbols of the symbol layer, that is, the mapping relationship of abstracting symbolic concepts from specific object category symbols. Based on the set of referent layer symbols of the target image [L(1), L(2), ..., L(d)], we can infer the symbolic concept symbols X(s) corresponding to each object category symbol L(s) in the target image in the symbol semantic knowledge graph. Select a target category symbol from the object category symbols of the symbol set of the target image, and obtain the paths between the target category symbol and its corresponding symbolic layer concept symbol; For any path, obtain each edge that makes up the path and the weight of each edge. A path is generally composed of multiple edges connected in sequence. The product of the weights of each edge is taken as the path weight of the path, and then the path weight Ji corresponding to each path is obtained, where i represents the different paths between the target category symbol and its corresponding symbolic layer concept symbol. When there is only one path between the target category symbol and its corresponding symbolic layer concept symbol, the path weight of the corresponding path is used as the mapping coefficient Y1 between the target category symbol and its corresponding symbolic layer concept symbol, and the corresponding path is recorded as the mapping path between the target category symbol and its corresponding symbolic layer concept symbol; When there is more than one path between a target category symbol and its corresponding symbolic layer concept symbol, the maximum path weight among the multiple path weights Ji is taken as the mapping coefficient Y1 between the target category symbol and its corresponding symbolic layer concept symbol, and the path corresponding to the maximum path weight is recorded as the mapping path between the target category symbol and its corresponding symbolic layer concept symbol; The same analysis is performed on other object category symbols in the symbol set of the indicated layer, and then the mapping coefficient Y(s) between each object category symbol in the symbol set of the indicated layer and its corresponding symbolic layer concept symbol is obtained. The ratio T(s) between each mapping coefficient Y(s) and all mapping coefficients Y(s) is obtained, and then the symbolic layer symbol vector U(1) = [T(1), T(2), ..., T(d)] of the target image is obtained. Using the same analysis method, we can obtain the symbol vector U(i) corresponding to each image in the image library, where i represents different images in the image library; By constructing a semantic knowledge graph, the symbols of the referent layer are mapped to the conceptual symbols of the symbolic layer, and the mapping coefficients are calculated through path weights to generate symbolic layer symbol vectors.

[0021] The query image three-layer symbol vector construction module constructs the three-layer symbol vector of the query image in the same way as obtaining the three-layer symbol vector of each image when the user uploads a query image, thus obtaining the symbol vectors of the signifier layer, signified layer and symbol layer of the query image. The comprehensive symbol similarity acquisition module calculates the three-layer symbol similarity between the query image and each image based on the three-layer symbol vectors of each image in the image library and the three-layer symbol vector of the query image. It then combines these three layers into a comprehensive symbol similarity, which consists of the symbol similarity of the signifier layer, the signified layer, and the symbol layer. Based on the similarity calculation formula, the signifier layer symbol vector of the query image is calculated with the signifier layer symbol vectors of each image to obtain the signifier layer symbol similarity Z1i between the signifier layer symbol vector of the query image and the signifier layer symbol vectors of each image. Similarly, the signified layer symbol similarity Z2i between the signified layer symbol vector of the query image and the signified layer symbol vectors of each image is calculated using the same similarity calculation formula. The symbol similarity is then comprehensively measured by weighting the signifier layer symbol similarity, the signified layer symbol similarity, and the symbol layer symbol similarity to comprehensively measure the three-layer symbol similarity between the query image and each image. That is, by using Vi = β1×Z1i + β2×Z2i + β3×Z3i, we obtain the three-layer symbolic similarity Vi between the query image and each of the other images. Z1i, Z2i, and Z3i are the symbolic similarity of the signifier layer, the signified layer, and the symbolic layer between the query image and each of the other images, respectively. β1, β2, and β3 are the weighting coefficients of the symbolic similarity of the signifier layer, the signified layer, and the symbolic layer, respectively, satisfying β1 + β2 + β3 = 1, and β1, β2, and β3 are all greater than or equal to 0. The weighting coefficients can be dynamically adjusted according to the task requirements. Among them, β1, β2, and β3 are all non-negative weights and can be optimized by the validation set or preset according to the scenario.

[0022] The final comprehensive similarity is obtained by calculating the three-layer symbolic similarity between the query image and each image in the image library, and by combining the symbolic similarity of the signifier layer, the signified layer, and the symbolic layer according to the weighted formula.

[0023] The search result set acquisition module obtains the search result set based on the sorting of three-level symbol similarity, and displays each image according to the image order in the search result set; The images in the entire image library are sorted according to the three-level symbolic similarity Vi between the query image and each other. The top K images with the highest three-level symbolic similarity are selected as the search results. The three-level symbolic similarity Vi of all images in the image library is traversed, and the three-level symbolic similarity Vi between the query image and each other is called. The images are sorted from largest to smallest according to the three-level symbolic similarity Vi between the query image and each other. The top K images with the highest three-level symbolic similarity are selected as the search result set. The images are displayed according to the order of the images in the search result set. The images in the image library are sorted according to the overall symbol similarity, and the top K images with the highest similarity are selected as the search results and displayed in order.

[0024] Through the collaborative efforts of multiple modules, semantic encoding and retrieval of images are achieved. First, the image symbol feature extraction module uses object detection and semantic segmentation techniques to extract the main visual elements of each image in the image library, such as people, objects, and scenes, forming basic visual symbols. These symbols provide the foundation for the subsequent three-layer symbol vector generation. Next, the three-layer symbol vector generation module structures each image into three layers of symbol vectors: the signifier layer, the signified layer, and the symbolic layer, representing the physical features, semantic concepts, and cultural symbols in the image, respectively. Each layer of symbol vectors is generated based on the image content and stored in the database in association with the image ID. The signifier layer symbol vector acquisition module calculates the weight of each visual symbol in the image using an object detection algorithm, characterizing its contribution to the overall image; the signified layer symbol vector acquisition module calculates the confidence of each object category based on the object detection network, generating the signified layer symbol vector. The symbolic layer symbol vector acquisition module maps the signified layer symbols to symbolic layer concept symbols using a semantic knowledge graph, generating the symbolic layer symbol vector. For the query image, the system will construct a corresponding three-layer symbol vector, and then calculate the three-layer symbol similarity between the query image and the image in the library through the comprehensive symbol similarity acquisition module. Finally, the system will return the most relevant search results through the search result set acquisition module.

[0025] This system employs a visual semiotics framework to encode and structure images using three layers of symbols, comprehensively capturing their semantic information. Through the coordinated efforts of modules such as image symbol feature extraction, three-layer symbol generation, and symbol vector acquisition, each image is transformed into a symbol vector at the signifier, signified, and symbolic levels. Each layer of symbol vector corresponds to a different semantic dimension of the image, accurately reflecting its physical features, semantic concepts, and symbolic meaning. During a query, the system constructs a three-layer symbol vector for the query image and compares it with images in the image library, ultimately outputting the image whose semantics are closest to the query image, achieving precise retrieval. This approach not only constructs a multi-layered semantic expression of images using semiotic methods but also combines advanced technologies such as object detection, semantic segmentation, and knowledge graphs to achieve deep semantic mining and precise retrieval of images. This solution effectively improves the accuracy of image retrieval, especially when dealing with semantically complex images with diverse cultural backgrounds. Furthermore, the system can dynamically adjust weighting coefficients according to the needs of different application scenarios, flexibly handling various retrieval tasks and exhibiting strong scalability and adaptability.

[0026] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.

[0027] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A visual semiotics based image semantic coding and retrieval system, characterized in that, The method comprises the following steps: A three-layer symbol vector generation module is used to detect objects in each image in the image library, obtain image symbol features of each image, and construct a three-layer symbol vector of each image in the image library based on the image symbol features, including an enabler layer, a referent layer and a symbol layer; A query image three-layer symbol vector construction module is used to construct a three-layer symbol vector of the query image in the same way as the three-layer symbol vector of each image is obtained when a user uploads a query image, and obtain the symbol vector of the enabler layer, the referent layer and the symbol layer of the query image; A comprehensive symbol similarity acquisition module is used to calculate the three-layer symbol similarity between the query image and each image in the image library based on the three-layer symbol vector of each image in the image library and the three-layer symbol vector of the query image, and further calculate the comprehensive symbol similarity, wherein the three-layer symbol similarity is the symbol similarity of the enabler layer, the referent layer and the symbol layer; A retrieval result set acquisition module is used to obtain a retrieval result set based on the ranking of the three-layer symbol similarity, and display each image in the retrieval result set according to the order of the images in the retrieval result set.

2. The image semantic encoding and retrieval system based on visual semiotics according to claim 1, characterized in that, The specific way to construct the enabler layer symbol vector of each image is as follows: Any image in the image library is taken as a target image, each visual symbol A(k) in the target image is recognized through target detection, and the enabler layer symbol vector of the target image is represented as N(1)= [M(1), M(2)、……、M(Dn)], wherein M(k) is a non-negative real number representing the weight value of different visual symbols A(k) in the target image, which represents the contribution of different visual symbols to the whole target image, wherein k represents different visual symbols, and the same analysis method is used to obtain the enabler layer symbol vector N(i) corresponding to each image in the image library, wherein i represents different images in the image library.

3. The image semantic encoding and retrieval system based on visual semiotics according to claim 2, characterized in that, The weight value of different visual symbols A(k) in the image is obtained in the following way: The area corresponding to each visual symbol A(k) in the target image is recognized through target detection, the pixel number of the area corresponding to each visual symbol A(k) is calculated, then the proportion of the pixel number of the area corresponding to each visual symbol A(k) in the total pixel number of the whole image is calculated, and the proportion is taken as the weight value M(k) of each visual symbol A(k) in the target image, and then the enabler layer symbol vector N(1)= [M(1), M(2)、……、M(Dn)] of the target image is obtained, wherein Dn is the number of visual symbols contained in the target image.

4. The image semantic encoding and retrieval system based on visual semiotics according to claim 2, characterized in that, The specific way to construct the referent layer symbol vector of each image is as follows: Identify different object class symbols L(s) in the target image, and the confidence K(s) of each object class respectively, s is different object class, mark the target layer symbol set of the target image as [L(1), L(2), …, L(d)], wherein d is the total number of object classes in the target image, L(s) is different object class symbol in the target image, obtain the sum of all object class confidence K(s) in the target image, and the ratio between the confidence Q(s) of different object classes in the target image and the sum of all different object class confidence K(s) is the confidence weight of different object classes in the target image, and then obtain the target layer symbol vector of the target image: S(1)= [Q(1), Q(2), …, Q(d)], d is the total number of object classes in the target image; Using the same analysis method, the target layer symbol vector S(i) corresponding to each image in the gallery can be obtained.

5. The image semantic encoding and retrieval system based on visual semiotics according to claim 4, characterized in that, The specific way to construct the symbolic layer symbol vector of each image is: According to the target layer symbol set [L(1), L(2), …, L(d)] of the target image, obtain the symbolic layer concept symbol X(s) corresponding to each object class symbol L(s) in the target image, select a target class symbol from each object class symbol in the target layer symbol set of the target image, obtain each path between the target class symbol and the symbolic layer concept symbol corresponding thereto, for any path, obtain each edge and the weight of each edge, and the product of the weights of each edge is the path weight of the path, and then the path weight corresponding to each path is obtained, and the mapping coefficient between the target class symbol and the symbolic layer concept symbol corresponding thereto is obtained according to the number of each path between the target class symbol and the symbolic layer concept symbol corresponding thereto; The other object class symbols in the target layer symbol set are analyzed in the same way, and then the mapping coefficient Y(s) between each object class symbol in the target layer symbol set and the symbolic layer concept symbol corresponding thereto is obtained, and the ratio T(s) between each mapping coefficient Y(s) and all mapping coefficients Y(s) is obtained, and then the symbolic layer symbol vector U(1)= [T(1), T(2), …, T(d)] of the target image is obtained, and the symbolic layer symbol vector U(i) corresponding to each image in the gallery can be obtained by using the same analysis method, wherein i represents different images in the gallery.

6. The image semantic encoding and retrieval system based on visual semiotics according to claim 5, characterized in that, When there is only one path between the target class symbol and the symbolic layer concept symbol corresponding thereto, the path weight of the corresponding path is taken as the mapping coefficient between the target class symbol and the symbolic layer concept symbol corresponding thereto, and when there is more than one path between the target class symbol and the symbolic layer concept symbol corresponding thereto, the maximum path weight in the multiple path weights is taken as the mapping coefficient between the target class symbol and the symbolic layer concept symbol corresponding thereto.

7. The image semantic encoding and retrieval system based on visual semiotics according to claim 5, characterized in that, The specific way to calculate the three-layer symbol similarity between the query image and each image is: According to the similarity calculation formula, for each image in the image library, the symbol vector similarity of the query image and the image in the image library on the signifier layer, the signified layer and the symbol layer is calculated respectively, and is recorded as the signifier layer symbol similarity Z1i, the signified layer symbol similarity Z2i and the symbol layer symbol similarity Z3i respectively.

8. The image semantic encoding and retrieval system based on visual semiotics according to claim 7, characterized in that, The specific manner of obtaining the comprehensive symbol similarity is as follows: The comprehensive symbol similarity is calculated by weighted summation, and the symbol vector Vi of the symbol layer is Vi=β1×Z1i+β2×Z2i+β3×Z3i, wherein β1, β2 and β3 are weight coefficients corresponding to the symbol vector of the symbol layer, and satisfy β1+β2+β3=1, and β1, β2 and β3 are greater than or equal to 0.

9. The image semantic encoding and retrieval system based on visual semiotics according to claim 8, characterized in that, The specific manner of obtaining the retrieval result set based on the ranking of the three-layer symbol similarity is as follows: The three-layer symbol similarity Vi between the query image and each image is traversed, the three-layer symbol similarity Vi between the query image and each image is ranked from large to small according to the numerical value, and the first K images with the three-layer symbol similarity are selected as the retrieval result set.