Picture library construction method and device based on multi-modal feature fusion
By employing a multimodal feature fusion method, the shooting time, location, identity of the people involved, and background features of the image are obtained. Combined with an interpersonal relationship graph, structured descriptive information is generated, which solves the problem of weak semantic relevance in existing image libraries and achieves efficient image library management and retrieval.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGXI INST OF FASHION TECH
- Filing Date
- 2026-01-21
- Publication Date
- 2026-05-01
AI Technical Summary
Existing image library construction methods rely on simple metadata and basic tags, lacking the mining and integration of deep semantic information of image content. This results in isolated image library content, a single information dimension, and an inability to establish rich semantic connections between images, affecting retrieval and recommendation performance.
A multimodal feature fusion method is adopted to obtain the shooting time, location, identity of people and background features of the image through the feature extraction model. Combined with the interpersonal relationship graph, structured descriptive information is generated to build a local image library.
It enables structured extraction and multi-dimensional association of deep semantic information in image content, improving the semantic richness and retrieval efficiency of the image library.
Smart Images

Figure CN121963255A_ABST
Abstract
Description
A method and apparatus for constructing an image library based on multimodal feature fusion Technical Field
[0001] This application relates to the field of image processing, and in particular to a method and apparatus for constructing an image library based on multimodal feature fusion. Background Technology
[0002] Currently, with the widespread adoption of smart home devices and social applications, the number of digital images is growing exponentially, making traditional image library construction methods insufficient to meet the demands of intelligent management. Most existing image libraries rely on simple metadata (such as time and location) and basic tags for classification and storage, lacking effective mining and integration of deep semantic information about image content. This construction method results in isolated image library content, a single information dimension, and an inability to establish rich semantic relationships between images, thus limiting the effectiveness of subsequent applications such as retrieval and recommendation.
[0003] Specifically, existing technologies typically employ a fragmented feature extraction strategy when building image libraries, such as analyzing faces, scenes, or objects independently, failing to achieve collaborative modeling of multimodal features. Furthermore, the system struggles to effectively represent and integrate complex interpersonal relationships and dynamically changing scene contexts, resulting in image libraries lacking structure, hierarchy, and semantic integrity.
[0004] This application provides a method and apparatus for constructing an image library based on multimodal feature fusion, which solves the problems of low image library retrieval efficiency and insufficient semantic understanding caused by the single dimension of image features and weak semantic relevance in the prior art.
[0005] Firstly, this application provides a method for constructing an image library based on multimodal feature fusion, comprising: acquiring an input image and determining the shooting time and location corresponding to the input image; inputting the input image into a feature extraction model to extract the person's identity features and scene background features corresponding to the input image; inputting the person's identity features into a preset interpersonal relationship model for parsing to determine the interpersonal relationship graph corresponding to the person's identity features; performing multimodal feature fusion on the interpersonal relationship graph, shooting time, shooting location, person's identity features, and scene background features to generate structured descriptive information corresponding to the input image; and constructing a local image library corresponding to the input image based on the structured descriptive information.
[0006] Secondly, this application provides an image library construction device based on multimodal feature fusion, comprising: an input image acquisition module configured to acquire an input image and determine the shooting time and shooting location corresponding to the input image; a feature extraction module configured to input the input image into a feature extraction model to extract the person's identity features and scene background features corresponding to the input image; an interpersonal relationship graph determination module configured to input the person's identity features into a preset interpersonal relationship model for parsing to determine the interpersonal relationship graph corresponding to the person's identity features; a fusion module configured to perform multimodal feature fusion of the interpersonal relationship graph, shooting time, shooting location, person's identity features, and scene background features to generate structured description information corresponding to the input image; and an image library construction module configured to construct a local image library corresponding to the input image based on the structured description information.
[0007] Thirdly, this application provides a readable medium including executable instructions, which, when executed by a processor of an electronic device, cause the electronic device to perform any of the methods described in the first aspect.
[0008] Fourthly, this application provides an electronic device including a processor and a memory storing execution instructions, wherein when the processor executes the execution instructions stored in the memory, the processor performs the method as described in any of the first aspects.
[0009] This application provides a method and apparatus for constructing an image library based on multimodal feature fusion. The method involves acquiring an input image and determining its shooting time and location; inputting the input image into a feature extraction model to extract the corresponding person's identity features and scene background features; parsing the person's identity features using a pre-defined interpersonal relationship model to determine the corresponding interpersonal relationship graph; fusing the interpersonal relationship graph, shooting time, shooting location, person's identity features, and scene background features using multimodal features to generate structured descriptive information corresponding to the input image; and constructing a local image library based on this structured descriptive information. This method achieves structured extraction and multi-dimensional association of deep semantic information from image content, effectively improving the semantic richness and retrieval efficiency of the image library.
[0010] The further effects of the aforementioned non-conventional preferred method will be explained below in conjunction with specific embodiments. Attached Figure Description
[0011] To more clearly illustrate the embodiments of this application or the existing technical solutions, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 is a flowchart illustrating a method for constructing an image library based on multimodal feature fusion according to an embodiment of this application; Figure 2 is a flowchart illustrating another method for constructing an image library based on multimodal feature fusion according to an embodiment of this application; Figure 3 is a flowchart illustrating another method for constructing an image library based on multimodal feature fusion according to an embodiment of this application; Figure 4 is a structural schematic diagram of an image library construction device based on multimodal feature fusion according to an embodiment of this application; Figure 5 is a structural schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] Currently, with the widespread adoption of smart home devices and social applications, the number of digital images is growing exponentially, making traditional image library construction methods insufficient to meet the demands of intelligent management. Most existing image libraries rely on simple metadata (such as time and location) and basic tags for classification and storage, lacking effective mining and integration of deep semantic information about image content. This construction method results in isolated image library content, a single information dimension, and an inability to establish rich semantic relationships between images, thus limiting the effectiveness of subsequent applications such as retrieval and recommendation.
[0015] Specifically, existing technologies typically employ a fragmented feature extraction strategy when building image libraries, such as analyzing faces, scenes, or objects independently, failing to achieve collaborative modeling of multimodal features. Furthermore, the system struggles to effectively represent and integrate complex interpersonal relationships and dynamically changing scene contexts, resulting in image libraries lacking structure, hierarchy, and semantic integrity.
[0016] To address this issue, this application proposes a method for constructing an image library based on multimodal feature fusion. This method aims to solve the problems of low retrieval efficiency and insufficient semantic understanding in existing technologies due to the single dimension of image features and weak semantic relevance. In this embodiment, the method for constructing an image library based on multimodal feature fusion includes: Step 101, acquiring an input image and determining the shooting time and location corresponding to the input image.
[0017] The system needs to read the image files to be processed, i.e., the input images, from the user's local storage device. These input images are typically stored on hard drives, mobile devices, or external storage devices, covering various types of photos taken in the user's daily life. While reading the input images, the system automatically extracts the basic metadata information of the images, the most important of which is determining the time and location of the image capture.
[0018] The capture time is primarily obtained by reading the timestamp information of the image file. This time information records the image's creation or capture time. For images with complete time information, the system can accurately obtain detailed time data such as year, month, day, hour, minute, and second. For images with incomplete or missing time information, the system uses the file's creation time or last modification time as a substitute and marks the source type of that time information.
[0019] Determining the shooting location relies on the geographic location data contained in the image, typically in the form of latitude and longitude coordinates. Cameras with positioning capabilities automatically record geographic location information when taking a picture, and the system analyzes this location data to determine the image's shooting location. For images that do not contain location information, the system marks their location as unknown, reserving a processing interface for possible subsequent location inference or manual annotation.
[0020] Step 102: Input the input image into the feature extraction model to extract the person's identity features and scene background features corresponding to the input image.
[0021] The feature extraction model is responsible for extracting semantically meaningful features from the input image. Employing deep learning techniques, it can simultaneously perform two tasks: person identification and scene background analysis. By using a single network architecture to handle multiple tasks, the model not only improves computational efficiency but also enhances feature sharing and complementarity between different tasks.
[0022] The process of extracting personal identification features first requires locating and detecting facial regions in the input image, and then encoding and identifying the features of each detected face. The system uses a deep neural network to convert facial images into high-dimensional feature vectors, which capture the unique feature information of an individual's face. By calculating and matching similarity with a pre-established family member feature database, the system can accurately identify the specific person appearing in the image and generate a corresponding identity label for each identified person.
[0023] Scene background feature extraction focuses on environmental information and background elements in the image, including various scene elements such as buildings, natural scenery, indoor and outdoor environments, and text signs. The system identifies representative scene features by analyzing the overall composition and local details of the input image. These scene features are encoded into structured feature representations that describe the type, characteristics, and semantic information of the shooting environment, providing crucial support for subsequent image classification and semantic understanding.
[0024] The system's feature extraction goes beyond basic face recognition and scene classification, capturing more detailed feature information. For person feature extraction, the system can identify detailed features such as clothing color, style, text labels, posture, and facial expressions. For scene feature extraction, the system can identify rich background information such as building logos, text content, environmental layout, and weather conditions. This fine-grained feature extraction provides a more comprehensive data foundation for subsequent multimodal fusion.
[0025] Step 103: Input the person's identity characteristics into the preset interpersonal relationship model for analysis, so as to determine the interpersonal relationship map corresponding to the person's identity characteristics.
[0026] Interpersonal relationship models aim to establish a dynamic knowledge representation system capable of accurately describing complex family and social relationships. Unlike traditional simple relationship labeling methods, interpersonal relationship models construct a high-dimensional dynamic relationship network with a time dimension. Modeling interpersonal relationships requires considering the complexity and diversity of real-life interpersonal relationships.
[0027] The many-to-many nature of interpersonal relationships is a crucial consideration in designing interpersonal relationship models. In real life, the relationship between any two people is often not singular but may involve multiple relationship types simultaneously. For example, Zhang San and Li Si may be colleagues and also relatives; if Zhang San's sister marries Li Si, then Li Si becomes Zhang San's brother-in-law. This coexistence of multiple relationships is difficult to accurately describe using traditional relationship representation methods, but it is fully reflected in the interpersonal relationship model of this invention.
[0028] The high-dimensional nature of interpersonal relationships is reflected in the fact that each person may assume multiple social roles, which function within different relationship networks. For example, in family relationships, Zhang San is simultaneously the father of his child, the husband of his spouse, and the son of his father; in social relationships, he may be someone's classmate, someone's colleague, someone's friend, and so on. Each role represents a dimension of relational information, constituting a complex, high-dimensional relationship space.
[0029] Real-world interpersonal relationships are also dynamic and subject to change; the relationship between two people often evolves over time. For example, two people might initially be classmates, then become colleagues after graduation, and eventually marry as their relationship develops. This evolution of relationships is reflected in photos taken at different times, and users' queries may also involve the relationship status at a specific point in time. For instance, a user might ask, "Ten years ago, two people were classmates who went hiking in Huangshan with the whole class. Later, they got married and went to Hong Kong for their honeymoon. Can you find photos of them?" This type of query involves the temporal trajectory of the relationship.
[0030] Based on the above analysis, interpersonal relationships exhibit complex characteristics of being many-to-many, high-dimensional, and dynamically changing. Traditional static knowledge graph methods cannot fully and accurately express this complex relationship structure. Therefore, this invention introduces a time dimension into traditional knowledge graphs, constructing a time-aware dynamic relationship model. In this model, each person's role and the relationships between people are adjusted accordingly over time, forming a four-dimensional relationship representation space (person, relationship type, relationship strength, and time).
[0031] Identify the subject person corresponding to the input image; construct the direct relationship between the subject person and the object person; perform semantic analysis on the direct relationship to determine the indirect relationship between the subject person and the object person in the preset member database; determine the start and end times of the direct and indirect relationships; construct a time-dimensional interpersonal relationship graph based on the subject person, direct relationship, indirect relationship and start and end time, and determine the interpersonal relationship model based on the interpersonal relationship graph.
[0032] When identifying the main figure in an input image, the system selects one or more figures from all the people in the image as the central node for relationship analysis. The main figure is usually the most important or prominent person in the image, or it may be a specific object of the user's attention. Around the main figure, the system begins to build a network of relationships between it and other figures.
[0033] In the process of constructing direct relationships, the system establishes direct connections between subjects and objects based on pre-collected and labeled family member information. These direct relationships include blood relationships such as father and son, mother and daughter, brothers and sisters, marital relationships such as husband and wife, spouses, and social relationships such as colleagues, classmates, friends, and neighbors. Each type of direct relationship has its specific semantic meaning and social attributes, forming the basic links of the interpersonal relationship network.
[0034] The identification of indirect relationships is a crucial manifestation of a system's intelligent reasoning ability. By performing deep semantic analysis on established direct relationships, the system can deduce more complex indirect relationships. For example, when the system knows that Zhang San is Li Si's father, and Li Si is Wang Wu's husband, it can deduce the indirect relationship that Zhang San is Wang Wu's father-in-law. Similarly, the system can deduce complex indirect relationships such as uncle-nephew, cousins, and colleagues' wives. This reasoning ability enables the system to understand and process complex relationship query requests from users.
[0035] The introduction of the time dimension imbues the interpersonal relationship model with dynamic characteristics, accurately reflecting the changes in interpersonal relationships over time. The system records the start and end times of each direct and indirect relationship, forming a complete history of relationship evolution. For example, two people might be classmates in school, become colleagues after graduation, and later develop a marital relationship. By recording the time boundaries of these relationships, the system can accurately determine the state of the relationship between two people at a specific point in time, supporting complex queries with temporal semantics.
[0036] Based on four elements—the main characters, direct relationships, indirect relationships, and start and end times—the system constructs a complete interpersonal relationship graph with a time dimension. This graph not only records the static relationship structure but also dynamically reflects the evolution of relationships, providing rich contextual support for subsequent image semantic understanding and intelligent retrieval.
[0037] Step 104: Perform multimodal feature fusion on the interpersonal relationship map, shooting time, shooting location, personal identity features and scene background features to generate structured description information corresponding to the input image.
[0038] Multimodal feature fusion is a technical step that unifies and integrates the interpersonal relationship graphs, shooting time, shooting location, personal identity features, and scene background features generated by the aforementioned processing steps. This process requires fusing data from different information sources and different representation formats into a unified, structured descriptive system. This fusion is not a simple data splicing, but a deep integration process based on semantic understanding.
[0039] The system integrates the identification results of the person's identity features and the background features of the scene with the previously constructed interpersonal relationship model and the metadata information contained in the photo itself to form a comprehensive analysis of a photo. This comprehensive analysis includes detailed information from multiple dimensions, which can fully describe the semantic content and background situation of the photo.
[0040] For example, regarding facial features, the system first identifies Zhang Moumou as the only person in a photo through interpersonal relationship features. The system further confirms his identity as a first-year high school student in a certain city and obtains his family relationship information through an interpersonal relationship graph, such as his father being Zhang Mou and his mother being Zhou. In clothing feature analysis, the system identifies that the person is wearing a blue T-shirt with the word "Gash" and the number "52" printed on it, carrying a black backpack with decorative patterns, and wearing a red hat with some text or logos. Posture and expression analysis shows that the person is smiling, appearing very happy and confident, holding a bottle of clear mineral water in his right hand, and his left hand hanging naturally. Other detailed features include personalized elements such as small ornaments or pendants hanging from his backpack.
[0041] In terms of background feature analysis, the system performed detailed building identification, confirming that the building's top features red Chinese characters reading "Turpan North Station," indicating it is a train station. The building's design is modern, with an arched roof and a simple, elegant exterior. The scene also includes slogans; a banner on the wall in front of the building reads "All Ethnic Groups are One Family." Environmental feature analysis shows neatly paved ground, appearing clean and tidy, and a clear, well-lit sky, suggesting the shooting likely took place during the day or evening. Based on the building's signage, the system was able to pinpoint the location as Turpan North Station in Turpan City, Xinjiang Uygur Autonomous Region, China.
[0042] During the fusion process, the interpersonal relationship graph provides rich semantic contextual information for the identity characteristics of individuals. The system not only knows which people appear in the image, but also understands the types of relationships between these people, their relationship history, and their specific relationship status at the time of capture. This relational information greatly enriches the semantic content of the image, enabling the system to understand the social context and background of the interactions between the individuals.
[0043] Shooting time and location serve as crucial spatiotemporal background information, complementing and validating scene background features. Time information helps determine the historical context and era of the image, while location information provides geographical environment and spatial background. When this spatiotemporal information is combined with scene features, a complete description of the shooting scene can be constructed, including information from multiple dimensions such as environment type, geographical features, and time characteristics.
[0044] The generation of structured descriptive information follows a unified data model and standard format, ensuring that all images possess the same semantic description. This structured design not only facilitates data storage and management but, more importantly, lays a solid foundation for subsequent intelligent retrieval based on natural language. Each image's structured description includes multiple dimensions such as information about people, relationships, scenes, and spatiotemporal information, forming a complete semantic archive for the image.
[0045] Step 105: Based on the structured description information, construct the local image library corresponding to the input image.
[0046] Based on the generated structured description information, the system establishes a complete, localized intelligent image knowledge base. This image base differs from traditional simple file storage methods; it is a knowledge management system with rich semantic information and intelligent retrieval capabilities.
[0047] The construction of the local image library involves multiple technical aspects, including data organization, index creation, and storage optimization. Based on the characteristics of structured descriptive information, the system establishes a multi-layered, multi-dimensional data organization structure. Each image corresponds to a complete structured record, containing all semantic and metadata information for that image. These records are categorized and indexed according to different dimensions, forming the data foundation that supports complex queries.
[0048] Localized processing is a key feature of this image library; all image analysis, feature extraction, and relationship reasoning are completed on the user's local device. This design not only protects the privacy of users' family photos and avoids the risk of leakage that may result from uploading sensitive data to the cloud, but also improves the system's response speed and ease of use.
[0049] The establishment of a local image library provides data support for subsequent intelligent retrieval and application functions. Through structured descriptive information, the system can support complex queries based on natural language, allowing users to search for specific photos by describing relationships between people, scene features, time, and location. This intelligent retrieval capability significantly improves the efficiency and experience of managing and using family photos.
[0050] As can be seen from the above technical solutions, the beneficial effects of this embodiment are as follows: This application provides a method for constructing an image library based on multimodal feature fusion. The method involves acquiring an input image and determining its shooting time and location; inputting the input image into a feature extraction model to extract the corresponding person's identity features and scene background features; parsing the person's identity features using a preset interpersonal relationship model to determine the interpersonal relationship graph corresponding to the person's identity features; fusing the interpersonal relationship graph, shooting time, shooting location, person's identity features, and scene background features using multimodal features to generate structured descriptive information corresponding to the input image; and constructing a local image library corresponding to the input image based on the structured descriptive information. This achieves structured extraction and multi-dimensional association of deep semantic information in image content, effectively improving the semantic richness and retrieval efficiency of the image library.
[0051] Figure 1 shows only a basic embodiment of an image library construction method based on multimodal feature fusion according to this application. With certain optimizations and extensions, other preferred embodiments of an image library construction method based on multimodal feature fusion can be obtained.
[0052] Figure 2 shows another specific embodiment of an image library construction method based on multimodal feature fusion according to this application.
[0053] In this embodiment, a method for constructing an image library based on multimodal feature fusion includes the following steps: Step 201: Obtain the input image and determine the shooting time and shooting location corresponding to the input image.
[0054] Step 202: Input the input image into the feature extraction model to extract the person's identity features and scene background features corresponding to the input image.
[0055] Step 203: The convolutional neural network uses a pre-trained ResNet50 as the backbone for feature extraction.
[0056] The vital sign extraction model can employ a single-network multi-task learning approach, which offers significant technical advantages. First, it boasts high computational efficiency. By sharing the feature extraction process with the ResNet50 backbone network, the system reduces computation by over 30% compared to using two independent networks. This efficiency improvement is particularly crucial for localized deployment. Second, it demonstrates significant feature reuse. Low-level visual features such as edges, textures, and colors are simultaneously used for both friend / family recognition and background recognition, enhancing the sufficiency and effectiveness of feature utilization. Finally, joint optimization leads to a mutually reinforcing effect between tasks. Background information may aid in friend / family recognition; for example, the feature "wearing glasses" could be a physical characteristic or associated with a specific scene. Joint training of the two tasks allows the model to learn these potential connections.
[0057] We employ a pre-trained ResNet50 network as the backbone architecture for feature extraction. This choice is based on ResNet50's excellent performance and extensive validation in image feature extraction. ResNet50 is a 50-layer deep residual network that addresses the vanishing gradient problem in deep network training by introducing skip connections, effectively extracting multi-level feature information from images. The pre-trained ResNet50 model has been thoroughly trained on large-scale image datasets, learning rich and general visual feature representations, including low-level features such as edges, textures, shapes, and colors, as well as more complex semantic features.
[0058] Another important consideration in choosing ResNet50 as the backbone network is the balance of computational resources. Compared to deeper network architectures, ResNet50 maintains its feature extraction capabilities while having a relatively small number of model parameters and computational complexity, enabling the entire system to run smoothly on ordinary personal computers and even mobile devices. This design philosophy aligns with the core idea of this invention, which emphasizes localized processing, allowing users to enjoy intelligent image management services without relying on high-performance cloud computing resources.
[0059] By employing pre-trained models, the system can fully leverage existing knowledge, avoiding the massive amounts of data and computational resources required to train a network from scratch. Furthermore, the system only needs fine-tuning with relatively little family photo data to adapt to specific tasks of person recognition and scene analysis, significantly reducing the barrier to entry and cost of system deployment.
[0060] Step 204: Add two parallel fully connected branches after the global average pooling layer of ResNet50; the fully connected branches include the first branch and the second branch.
[0061] The global average pooling layer transforms the feature map output from the last convolutional layer of ResNet50 into a fixed-length feature vector containing a high-level semantic feature representation of the input image. By splitting the processing into two independent branches after this layer, the system can perform two different tasks simultaneously based on the same fundamental feature representation.
[0062] Common features are extracted by sharing the ResNet50 backbone network, including basic visual features such as texture, color, and edges of the photo. A dual-branch design is adopted in the fully connected layer part of ResNet50. For example, the first branch identifies relatives and friends through a 150-dimensional fully connected layer, while the second branch classifies the background scene through a 512-dimensional fully connected layer.
[0063] The first branch, with its 150-dimensional design based on Dunbar's number theory, covers key individuals within the user's social circle, ultimately outputting the probability distribution of 150 relatives and friends. For example, for a photo containing the father, the output might be ["Father": 0.95, "Mother": 0.02, "Uncle": 0.01, ...], with the softmax activation function ensuring the sum of all probabilities is 1. The second branch, with its 512-dimensional design, covers a rich variety of scene features, capable of recognizing diverse background environments such as "sunny summer day," "on a bridge," and "indoor living room," outputting probability distributions for various scene features, such as ["Rainy day": 0.88, "Huangshan": 0.76, "Outdoor": 0.92, ...]. The first branch is specifically responsible for person identification, aiming to accurately identify members appearing in images. The second branch focuses on extracting and recognizing scene background features, analyzing environmental information, background elements, and scene types in the image. This dual-branch design allows the system to simultaneously obtain person and scene information in a single forward computation, avoiding the redundant computational overhead of processing two tasks separately.
[0064] The parallel design of the two branches also demonstrates the advantages of feature sharing. Although person recognition and scene recognition focus on different content, they share certain commonalities at the underlying feature level. For example, factors such as lighting conditions, image quality, and shooting angle affect both tasks. By sharing the common features extracted from the ResNet50 backbone, the two tasks can mutually promote each other, improving the overall recognition accuracy and robustness.
[0065] Step 205: The first branch uses the softmax activation function to output the identity probability distribution vector to determine the identity characteristics of the person.
[0066] The first branch receives the general feature vector extracted by ResNet50; it passes the general feature vector through a fully connected layer of the target dimension; the number of dimensions corresponding to the target dimension is the number of members in the member library; it applies the softmax activation function to the output of the fully connected layer to generate an identity probability distribution vector; and it determines the identity features of the person in the input image based on the identity label corresponding to the maximum probability value in the identity probability distribution vector.
[0067] The first branch is specifically optimized for person identification tasks, employing a softmax activation function to output an identity probability distribution vector. It first receives a general feature vector extracted by the ResNet50 backbone network, which contains a high-level semantic representation of the image. Subsequently, the feature vector is input into a fully connected layer of the target dimension for linear transformation; the output dimension of this fully connected layer is equal to the number of members in a pre-defined member database.
[0068] The member database is constructed based on the user's actual family and social circle. According to Dunbar's number theory, a person can maintain stable social relationships with approximately 150 people. Therefore, the system's member database is designed to cover the user's main social contacts, including family members, relatives, friends, colleagues, classmates, and other groups. Each output neuron in the fully connected layer corresponds to a specific member in the member database, and the output value represents the original score of the match between the face in the input image and that member.
[0069] The application of the softmax activation function transforms the original scores into a standardized probability distribution, ensuring that the sum of the probability values for all members equals 1. This probabilistic representation not only provides confidence information for the recognition results but also handles uncertainties during the recognition process. By selecting the maximum value in the probability distribution and its corresponding person label, the system determines the identity features of the person in the input image. This probability-based recognition method also supports multi-person scenes, allowing the system to generate a corresponding identity probability distribution for each detected face in the image.
[0070] Step 206: The second branch uses the sigmoid activation function to output scene feature vectors to determine scene background features.
[0071] The second branch is specifically responsible for extracting and recognizing scene background features, using the sigmoid activation function to output scene feature vectors. Unlike single-selection classification for person identification, scene feature recognition faces a multi-label classification problem because an image may simultaneously contain multiple scene elements and background features. For example, a photo may simultaneously have multiple scene labels such as "outdoors," "sunny," and "park."
[0072] The second branch receives the same ResNet50 general feature vector as input and performs feature transformation through a specially designed fully connected layer. The output dimension of this fully connected layer corresponds to the number of predefined scene feature categories, with each output neuron representing a specific scene feature type. These scene features cover multiple dimensions of information, including environment type, weather conditions, architectural style, natural landscape, and indoor / outdoor distinctions.
[0073] The sigmoid activation function is chosen so that the output value of each scene feature is between 0 and 1, representing the probability of that feature appearing in the current image. Unlike the softmax function, the sigmoid function allows multiple features to have high activation values simultaneously, which aligns with the needs of multi-label classification. By setting an appropriate threshold, the system can identify all salient scene features present in the image, forming a complete scene description vector.
[0074] The first and second branches are jointly trained using a multi-task learning approach; a multi-task loss function is constructed; the multi-task loss function is composed of a weighted sum of the person identification loss of the first branch and the scene recognition loss of the second branch; using the multi-task loss function, the network parameters of the convolutional neural network are optimized through the backpropagation algorithm to determine the feature extraction model.
[0075] Because a single photograph may contain multiple friends and family members, as well as various scene features, the system employs multiple binary cross-entropy loss functions during training to calculate the loss for each person and scene separately. For person recognition, the system calculates an independent binary cross-entropy loss for each possible person; for scene recognition, the system similarly calculates the corresponding binary cross-entropy loss for each scene feature category. This loss function design accurately handles complex annotation situations involving multiple people and multiple scenes, avoiding the limitations of traditional single-label classification methods.
[0076] During the training of the feature extraction model, the system constructs a comprehensive multi-task loss function, which weights and sums the loss from the first branch (person identification) and the loss from the second branch (scene identification). Person identification uses a cross-entropy loss function to measure the difference between the predicted identity probability distribution and the true label. Scene identification uses a binary cross-entropy loss function to calculate the loss between the predicted probability of each scene feature and the true label.
[0077] Through backpropagation, the gradient information of the multi-task loss function is propagated to all parameters of the entire network, achieving end-to-end joint optimization. This training method enables the ResNet50 backbone to learn general feature representations beneficial to both tasks, while the two branches learn task-specific feature transformations respectively. Joint training is generally more efficient than training two independent models separately and achieves better overall performance.
[0078] Step 207: Input the person's identity features into the preset interpersonal relationship model for analysis, so as to determine the interpersonal relationship map corresponding to the person's identity features.
[0079] Step 208: Perform multimodal feature fusion on the interpersonal relationship map, shooting time, shooting location, personal identity features and scene background features to generate structured description information corresponding to the input image.
[0080] Step 209: Based on the structured description information, construct the local image library corresponding to the input image.
[0081] As can be seen from the above technical solution, the beneficial effects of this embodiment are: by using a pre-trained ResNet50 as the feature extraction backbone and designing a dual-branch multi-task learning architecture, efficient and accurate multimodal feature extraction capabilities are achieved. This design enables the system to simultaneously complete two key tasks—person identification and scene feature extraction—in a single forward computation, significantly reducing the system's resource requirements compared to solutions using two independent networks.
[0082] Figure 3 shows another specific embodiment of the image library construction method based on multimodal feature fusion proposed in this application. This embodiment is further described based on the foregoing embodiments.
[0083] In this embodiment, a method for constructing an image library based on multimodal feature fusion includes the following steps: Step 301: Obtain the input image and determine the shooting time and shooting location corresponding to the input image.
[0084] Step 302: Input the input image into the feature extraction model to extract the person's identity features and scene background features corresponding to the input image.
[0085] Step 303: Input the person's identity characteristics into the preset interpersonal relationship model for analysis, so as to determine the interpersonal relationship map corresponding to the person's identity characteristics.
[0086] Step 304: Perform multimodal feature fusion on the interpersonal relationship map, shooting time, shooting location, personal identity features and scene background features to generate structured description information corresponding to the input image.
[0087] Step 305: Based on the structured description information, construct the local image library corresponding to the input image.
[0088] Step 306: Receive query command.
[0089] Receiving query commands is the initial step in user interaction with the intelligent image library system, requiring the system to handle various forms of user input. Users can input query requests via natural language text, such as "Find the grandma in the red dress in last year's Spring Festival family photo" or "Show photos of me and my classmates taken in Huangshan." The system also supports voice input, using speech recognition technology to convert users' spoken descriptions into text-based query commands.
[0090] The process of receiving query commands also includes preliminary preprocessing and standardization of the input information. The system cleans the text entered by the user, removing irrelevant punctuation marks, filler words, and other interfering information, while performing basic grammar checks and error corrections. For voice input, the system uses automatic speech recognition technology to convert the audio signal into text and evaluates the confidence level of the recognition results to ensure the accuracy of the conversion.
[0091] To enhance user experience, the system also supports a multi-turn dialogue interaction mode. Users can refine their search criteria step by step through continuous dialogue, for example, first asking "show photos from last year's Spring Festival," then asking "those with grandma in them," and finally further specifying "those wearing red clothes." The system maintains the context information of the dialogue, integrating information from multiple rounds of interaction to form a complete understanding of the search requirements.
[0092] Step 307: Perform semantic analysis on the query command to extract the query conditions corresponding to the query command; the query conditions include one or more combinations of person identity tags, interpersonal relationship graphs, shooting time, shooting location, and specific scene background.
[0093] Semantic analysis of query commands requires converting the user's natural language description into structured query conditions. This process employs natural language processing techniques, including multiple subtasks such as named entity recognition, relation extraction, and temporal parsing. The system first performs lexical and syntactic analysis on the query command, identifying keywords, phrase structures, and grammatical relationships within the sentence.
[0094] The system needs to identify specific person names or relationship descriptions from query commands. For example, when a user asks for "a photo of Zhang San," the system directly extracts the person tag "Zhang San." When a user asks for "a photo of my father," the system needs to combine the user's identity information to convert "my father" into a specific person tag. For more complex person descriptions, such as "grandmother in red," the system will extract the person role "grandmother" and the additional feature description "wearing red."
[0095] Extracting information from interpersonal relationship graphs involves understanding and parsing complex relationship descriptions. When a user queries an indirect relationship like "my classmate's wife," the system needs to parse out the relationship chain: from the user to the classmate's direct relationship, then from the classmate to his wife, ultimately determining the identity of the target person. The system uses a pre-built interpersonal relationship model to reason, transforming complex relationship descriptions into queryable relationship paths.
[0096] The system parses and processes time information in various forms, including absolute times such as "September 2025," relative times such as "last year's Spring Festival," and fuzzy times such as "when I was a child." It employs time entity recognition technology to convert these natural language time expressions into standard time ranges or points in time. For special time concepts such as holidays and seasons, the system combines historical data and common sense to perform accurate time mapping.
[0097] Location information extraction also requires processing various location representations, including specific place names, relative locations such as "near my home," and location types such as "in the park." The system uses geographic entity recognition and location reasoning technologies to convert these location descriptions into geographic coordinates or regional ranges that can be used for retrieval.
[0098] Recognizing specific scene backgrounds involves extracting information such as environmental descriptions, weather conditions, and activity types. For example, "sunset at the beach" would be parsed as the scene features "beach" and "sunset time," while "indoors on a rainy day" would be identified as scene labels such as "rainy weather" and "indoor environment."
[0099] Step 308: Based on the query conditions, use RAG technology to search the local image library to determine the target image corresponding to the query command.
[0100] Using RAG technology, semantic similarity is calculated in a local image library, and a candidate image set is determined based on the matching degree between the query conditions and the structured description information. A comprehensive matching score is calculated based on the semantic relevance and context matching degree of each candidate image in the candidate image set with the query conditions. Based on the comprehensive matching score, the target image is determined from the candidate image set.
[0101] Based on the extracted query criteria, the system utilizes RAG technology to perform intelligent retrieval within the local image library. RAG technology combines traditional keyword matching retrieval with deep semantic understanding, enabling it to handle complex, multi-dimensional query requirements. The retrieval process first requires calculating the matching degree between the user's query criteria and the structured descriptive information of each image in the image library.
[0102] The system employs semantic vectorization technology to convert query conditions and image descriptions into points in a high-dimensional semantic vector space. By calculating the cosine similarity or Euclidean distance between vectors, the system can quantify the semantic matching degree between the query request and the content of each image. This semantic vector-based matching method can handle the complexities of natural language, such as lexical variations and synonyms.
[0103] During the process of determining the candidate image set, the system sets an appropriate similarity threshold to filter out images with a high degree of matching with the query conditions. This initial filtering process can significantly narrow down the search scope and improve the efficiency of subsequent precise matching. At the same time, the system also considers the completeness and complexity of the query conditions and dynamically adjusts the filtering strategy to ensure that no potentially relevant images are missed.
[0104] The system considers not only semantic relevance but also contextual matching to calculate a comprehensive matching score. Semantic relevance primarily measures the degree of semantic consistency between the image content and the query conditions, while contextual matching considers the consistency of contextual information such as time, location, and relationships between people. For example, when a user queries "family photos from last year's Spring Festival," the system will not only match photos containing the whole family but will also focus on whether the shooting time meets the requirements of the Spring Festival period.
[0105] The dynamic adjustment mechanism of multi-dimensional weights enables the system to optimize matching strategies based on the characteristics of query conditions. When the query mainly involves information about people, the matching score related to people will receive higher weight; when the query emphasizes time or location information, the weight of the corresponding dimension will be increased accordingly. This adaptive weight allocation mechanism ensures that different types of queries can obtain optimal search results.
[0106] The final target image is determined based on the ranking of comprehensive matching scores. The system sorts candidate images from highest to lowest matching score and returns the top N most relevant images according to the user's needs. The system also provides a result explanation function, which explains to the user the reason for the matching of each returned image, enhancing the interpretability of the search results and user trust.
[0107] As can be seen from the above technical solutions, the beneficial effects of this embodiment are as follows: By employing natural language processing technology to perform deep semantic analysis on query commands, it achieves accurate understanding and parsing of complex query needs. By introducing retrieval enhancement generation technology and combining it with multi-dimensional semantic matching algorithms, it achieves high-precision intelligent image retrieval functionality. Compared with traditional keyword matching methods, this solution performs exceptionally well in handling synonyms, near-synonyms, and context-related queries, accurately finding the target image that the user truly needs.
[0108] Figure 4 shows a specific embodiment of an image library construction device based on multimodal feature fusion according to this application. This embodiment of an image library construction device based on multimodal feature fusion is a physical device used to execute the image library construction method based on multimodal feature fusion shown in Figures 1-3. Its technical solution is essentially the same as the above embodiments, and the corresponding descriptions in the above embodiments are also applicable to this embodiment. This embodiment of an image library construction device based on multimodal feature fusion includes: an input image acquisition module 401, configured to acquire an input image and determine the shooting time and shooting location corresponding to the input image; a feature extraction module 402, configured to input the input image into a feature extraction model to extract the person's identity features and scene background features corresponding to the input image; an interpersonal relationship graph determination module 403, configured to input the person's identity features into a preset interpersonal relationship model for parsing to determine the interpersonal relationship graph corresponding to the person's identity features; a fusion module 404, configured to perform multimodal feature fusion of the interpersonal relationship graph, the shooting time, the shooting location, the person's identity features, and the scene background features to generate structured description information corresponding to the input image; and an image library construction module 405, configured to construct a local image library corresponding to the input image based on the structured description information.
[0109] Figure 5 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. The memory may include RAM, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk storage device. Of course, the electronic device may also include other hardware required for other services.
[0110] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only a single bidirectional arrow is used in Figure 5, but this does not indicate that there is only one bus or one type of bus.
[0111] Memory is used to store instructions for execution. Specifically, instructions for execution are computer programs that can be executed. Memory can include main memory and non-volatile memory, and it provides the processor with execution instructions and data.
[0112] In one possible implementation, the processor reads the corresponding execution instructions from non-volatile memory into memory and then runs them. Alternatively, it may obtain the corresponding execution instructions from other devices to logically form an image library construction device based on multimodal feature fusion. The processor executes the execution instructions stored in memory to implement the image library construction method based on multimodal feature fusion provided in any embodiment of this application.
[0113] The method for constructing an image library based on multimodal feature fusion, as provided in the embodiment shown in Figure 4 of this application, can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor.
[0114] The steps of the method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0115] This application also proposes a readable medium that stores execution instructions. When the stored execution instructions are executed by the processor of an electronic device, the electronic device can execute a method for constructing an image library based on multimodal feature fusion provided in any embodiment of this application, and specifically execute the method shown in Figure 1, Figure 2, or Figure 3.
[0116] The electronic devices in the foregoing embodiments may be computers.
[0117] Those skilled in the art will understand that the embodiments of this application can be provided as methods or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or a combination of software and hardware.
[0118] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0119] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0120] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for constructing an image library based on multimodal feature fusion, characterized in that, include: Obtain the input image and determine the shooting time and location corresponding to the input image; input the input image into a feature extraction model to extract the person's identity features and scene background features corresponding to the input image; The identity features of the person are input into a preset interpersonal relationship model for parsing to determine the interpersonal relationship graph corresponding to the identity features of the person; the interpersonal relationship graph, the shooting time, the shooting location, the identity features of the person, and the scene background features are fused using multimodal features to generate structured description information corresponding to the input image; based on the structured description information, a local image library corresponding to the input image is constructed.
2. The method according to claim 1, characterized in that, Also includes: Identify the subject person corresponding to the input image; construct the direct relationship between the subject person and the object person; Perform semantic analysis on the direct relationships to determine the indirect relationships corresponding to the main character in a preset member database; determine the start and end times of the direct and indirect relationships. Based on the main figures, the direct relationships, the indirect relationships, and the start and end times, a time-dimensional interpersonal relationship graph is constructed, and the interpersonal relationship model is determined based on the interpersonal relationship graph.
3. The method according to claim 2, characterized in that, The feature extraction model employs a convolutional neural network, which includes: a pre-trained ResNet50 as the backbone for feature extraction; two parallel fully connected branches added after the global average pooling layer of the ResNet50; the fully connected branches include a first branch and a second branch; the first branch uses a softmax activation function to output an identity probability distribution vector to determine the person's identity features; the second branch uses a sigmoid activation function to output a scene feature vector to determine the scene background features.
4. The method according to claim 3, characterized in that, The first branch uses the softmax activation function to output a probability distribution vector of a person's identity to determine the person's identity features. This includes: the first branch receiving a general feature vector extracted by the ResNet50; passing the general feature vector through a fully connected layer of a target dimension; the number of dimensions corresponding to the target dimension is the number of members in the member library; applying the softmax activation function to the output of the fully connected layer to generate the identity probability distribution vector; and determining the person's identity features in the input image based on the person's identity label corresponding to the maximum probability value in the identity probability distribution vector.
5. The method according to claim 3, characterized in that, Also includes: The first branch and the second branch are jointly trained using a multi-task learning approach; a multi-task loss function is constructed; the multi-task loss function is composed of a weighted sum of the person identification loss of the first branch and the scene recognition loss of the second branch; using the multi-task loss function, the network parameters of the convolutional neural network are optimized through a backpropagation algorithm to determine the feature extraction model.
6. The method according to claim 4, characterized in that, Also includes: Receive a query instruction; perform semantic analysis on the query instruction to extract the query conditions corresponding to the query instruction; The query criteria include one or more combinations of the person's identity tag, the interpersonal relationship map, the shooting time, the shooting location, and a specific scene background; based on the query criteria, the RAG technology is used to search the local image library to determine the target image corresponding to the query command.
7. The method according to claim 6, characterized in that, The step of using RAG technology to search the local image library to determine the target image corresponding to the query instruction includes: using the RAG technology to calculate semantic similarity in the local image library; determining a candidate image set based on the matching degree between the query conditions and the structured description information; calculating a comprehensive matching score based on the semantic relevance and contextual matching degree of each candidate image in the candidate image set with the query conditions; and determining the target image from the candidate image set according to the comprehensive matching score.
8. An image library construction device based on multimodal feature fusion, characterized in that, include: The input image acquisition module is configured to acquire an input image and determine the shooting time and shooting location corresponding to the input image; The feature extraction module is configured to input the input image into the feature extraction model to extract the person's identity features and scene background features corresponding to the input image; The interpersonal relationship graph determination module is configured to input the person's identity features into a preset interpersonal relationship model for parsing, so as to determine the interpersonal relationship graph corresponding to the person's identity features. The fusion module is configured to perform multimodal feature fusion of the interpersonal relationship map, the shooting time, the shooting location, the person's identity features, and the scene background features to generate structured description information corresponding to the input image; The image library building module is configured to build a local image library corresponding to the input image based on the structured description information.
9. A computer-readable storage medium storing a computer program, characterized in that, The computer program is used to execute the image library construction method based on multimodal feature fusion as described in any one of claims 1-7.
10. An electronic device, characterized in that, The electronic device includes: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the image library construction method based on multimodal feature fusion as described in any one of claims 1-7.