Image management method and device and electronic equipment

By extracting the features of the people to be matched in images from chat groups of instant messaging applications and storing them in the photo album, the problem of low image management efficiency is solved, and efficient and accurate image management and retrieval are achieved.

CN121012808APending Publication Date: 2025-11-25VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511260240.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

In instant messaging applications, image management in chat groups is inefficient. Users need to manually scroll through historical messages one by one, making it difficult to retrieve images across dates and multiple groups, and it is easy to miss key content.

Method used

When the number of images received in a target chat group exceeds a threshold, the features of the people to be matched in the images are extracted. Based on the baseline features of the target person, images matching the target person are identified from the images and stored in the album.

Benefits of technology

It eliminates the need for users to manually filter images, improving image management efficiency, enabling precise classification and storage, allowing users to quickly find images related to their target objects, and enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121012808A_ABST
    Figure CN121012808A_ABST
Patent Text Reader

Abstract

The invention discloses an image management method and device and electronic equipment, and belongs to the technical field of electronic equipment, and the method comprises the steps: extracting the to-be-matched character features of each received image under the condition that the number of images received through a target chat group exceeds a first threshold value; determining at least one first image matched with the target object from the received images according to the character features to be matched and reference character features of the target object; and storing the at least one first image to a photo album corresponding to the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of electronic equipment technology, specifically relating to an image management method, apparatus, and electronic equipment. Background Technology

[0002] With the rapid development of mobile internet, instant messaging applications have become the main channel for communication between users. Taking the communication between kindergarten teachers and parents of kindergarten children as an example, teachers often share relevant images of children's activities, classroom performance and daily life moments in instant messaging application chat groups to help parents keep abreast of their children's growth in kindergarten.

[0003] In related technologies, images are embedded in instant messaging chat logs as single images or in scattered forms, lacking systematic classification. For example, parents need to manually scroll through historical messages one by one, which is time-consuming and laborious; especially in scenarios involving cross-date and multi-group chats, the difficulty of image retrieval increases exponentially, and key content is easily missed due to page flipping, resulting in low image management efficiency. Summary of the Invention

[0004] The purpose of this application is to provide an image management method, apparatus, and electronic device that can improve the efficiency of image management in chat groups.

[0005] In a first aspect, embodiments of this application provide an image management method, the method comprising:

[0006] If the number of images received through the target chat group exceeds a first threshold, extract the features of the person to be matched from each of the received images.

[0007] Based on the characteristics of the person to be matched and the baseline characteristics of the target object, at least one first image matching the target object is determined from the received images;

[0008] Store the at least one first image to the album corresponding to the target object.

[0009] Secondly, embodiments of this application provide an image management device, the device comprising:

[0010] The extraction module is used to extract the features of the person to be matched from each of the received images when the number of images received through the target chat group exceeds a first threshold.

[0011] The determining module is used to determine at least one first image that matches the target object from the received images based on the characteristics of the person to be matched and the baseline characteristics of the target object;

[0012] A storage module is used to store the at least one first image into an album corresponding to the target object.

[0013] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the image management method as described in the first aspect.

[0014] Fourthly, embodiments of this application provide a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the image management method as described in the first aspect.

[0015] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the steps of the image management method as described in the first aspect.

[0016] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the steps of the image management method as described in the first aspect.

[0017] In this embodiment, when the number of images received through the target chat group exceeds a first threshold, the features of the person to be matched in each received image are extracted; based on the features of the person to be matched and the baseline features of the target object, at least one first image matching the target object is determined from the received images; and at least one first image is stored in the album corresponding to the target object. On the one hand, this image management process eliminates the need for users to manually filter images one by one, saving users time and effort, enabling them to organize and manage images received in the chat group more efficiently, and improving the overall user experience; on the other hand, this precise classification and storage method allows users to easily and quickly find all images related to the target object in a specific album, avoiding the hassle of searching through a large number of cluttered images, providing users with a more orderly and convenient image viewing and management experience, thereby improving image management efficiency. Attached Figure Description

[0018] Figure 1 This is a flowchart of an image management method provided in an embodiment of this application;

[0019] Figure 2 This is one of the example diagrams of an image management method provided in the embodiments of this application;

[0020] Figure 3 This is a second example diagram of an image management method provided in an embodiment of this application;

[0021] Figure 4 This is the third example diagram of an image management method provided in the embodiments of this application;

[0022] Figure 5 This is the fourth example diagram of an image management method provided in the embodiments of this application;

[0023] Figure 6 This is a structural block diagram of an image management device provided in an embodiment of this application;

[0024] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0025] Figure 8 This is a schematic diagram of the hardware structure of an electronic device that implements the various embodiments of this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0027] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, and the number of objects is not limited; for example, a first object can be one or more. Furthermore, the character " / " generally indicates that the preceding and following objects have an "or" relationship.

[0028] In today's era of highly developed instant messaging, chat groups have become an important platform for people to share information and interact, with image sharing being particularly frequent. For example, in kindergarten parent groups, teachers often share pictures of children's activities at school; in work groups, colleagues also share pictures of project sites, meeting scenes, etc. However, as the number of images in groups continues to increase, users face many difficulties when searching for images of specific individuals, such as their own children, themselves, or specific colleagues. Manually browsing through chat history one by one is not only time-consuming and laborious, but also prone to missing important images among a large number of pictures. In addition, images shared at different times are scattered throughout the chat history, lacking effective classification and organization, which is not conducive to users' centralized management and review of images related to specific individuals, resulting in low efficiency in image management within chat groups.

[0029] To address the aforementioned technical problems, embodiments of this application provide an image management method, apparatus, and electronic device to improve the efficiency of users managing images in chat groups.

[0030] It should be noted that the image management method provided in this application is applicable to electronic devices. In practical applications, such electronic devices include, but are not limited to, mobile terminals such as smartphones, tablets, and personal digital assistants. This application does not limit the scope of these devices.

[0031] Figure 1 This is a flowchart of an image management method provided in an embodiment of this application, such as... Figure 1 As shown, the method may include the following steps: step 101, step 102 and step 103.

[0032] In step 101, if the number of images received through the target chat group exceeds a first threshold, the features of the person to be matched in each of the received images are extracted.

[0033] In this application embodiment, the target chat groups include, but are not limited to: various school parent groups, various training institution parent groups, work groups, hobby groups, life service groups, industry professional groups, regional / cultural groups, and special need groups (such as medical, public welfare, etc.).

[0034] In this embodiment of the application, the user can flexibly set the number of target chat groups according to their own needs. The number can be one or multiple.

[0035] In this embodiment, the target chat group can be set through a user interface provided by the electronic device. Optionally, the user interface displays a list of joined chat groups, and the user selects one or more chat groups as the target chat group by checking boxes. Optionally, the user can also manually enter the name or identifier of the chat group to specify the target chat group.

[0036] In this embodiment, we consider that different users have very different management needs for chat groups in different scenarios. For example, a parent may be in multiple groups, such as a kindergarten class group, an extracurricular class group, and a parent communication group. During specific periods, such as when the child is participating in kindergarten activities, the parent may only focus on the kindergarten class group, and in this case, the kindergarten class group can be set as the target chat group; while when there are important activities in the child's extracurricular class, the extracurricular class group can be set as the target chat group.

[0037] In this embodiment of the application, this flexible setting method can adapt to various complex and ever-changing life scenarios, enabling users to accurately manage the chat groups they need to follow under different circumstances, thus meeting the diverse scenario needs of users.

[0038] As can be seen, in this embodiment of the application, users are allowed to set target chat groups independently, which is applicable to both single chat groups and multiple chat groups, thereby enhancing the flexibility and applicability of the image management method and meeting the image management needs of different users in different scenarios.

[0039] Considering that chat groups with different activity levels have significantly different amounts of information generated and their importance, in this embodiment of the application, the first threshold can be reasonably set according to the activity level of the chat group and the frequency of image reception, such as 10 images, 20 images, etc.

[0040] In this embodiment, for highly active chat groups with frequent image reception, such as photography enthusiast groups, members may share a large number of images in a short period of time. Setting the first threshold to 20 images can filter out some scattered and relatively low-value image information, allowing for focused processing only on image sets that reach a certain number, such as extracting image features and classifying them. This allows electronic devices to focus on more representative images, improving the accuracy and efficiency of information processing.

[0041] In this embodiment, for chat groups with low activity, such as small family chat groups, the image reception frequency is relatively low. If the first threshold is set too high, the electronic device may not be able to trigger the relevant processing flow for a long time, missing some important images; while if it is set too low, such as 5 images, some scattered and unrepresentative images may be unnecessarily processed, causing information overload. Therefore, the first threshold can be set to 10 images, which can ensure timely processing when the number of images reaches a certain scale, while avoiding the inconvenience to users caused by processing too much irrelevant information.

[0042] In this embodiment of the application, when a received image contains multiple faces, the features of the object corresponding to each face can be extracted to obtain the features of the person to be matched in the image.

[0043] In this application embodiment, the human characteristics may include at least one of the following: facial features, clothing features, and body features.

[0044] In this embodiment, facial features may include: the shape of a person's facial features, facial contours, facial expressions, etc. In practical applications, facial recognition technology can be used for facial feature extraction and analysis.

[0045] In this embodiment, clothing features may include: the color, style, and pattern of the clothing worn by the person, as well as accessories such as hats, necklaces, and earrings. In practical applications, image recognition algorithms can be used for clothing feature recognition and extraction.

[0046] In this embodiment, body features may include: a person's height, body shape, and posture such as standing, sitting, and walking. In practical applications, image analysis and model matching can be used to determine body features.

[0047] In this embodiment of the application, considering that different human characteristics have different stability and uniqueness, facial features are highly unique, and each person's facial contours, facial proportions, etc. have subtle differences. However, in some cases, such as when the face is obscured, the angle is not good, or the lighting is insufficient, it may be difficult to accurately identify a person by relying solely on facial features. Clothing features and body features can serve as effective supplements.

[0048] Furthermore, different application scenarios have different requirements and environmental conditions for person recognition. For example, in some formal occasions, such as business meetings or document processing, facial features are the main basis for recognition because the person's face is usually clearly visible and needs to be accurately identified. In some outdoor activities or social gatherings, clothing features and body features may be more prominent and easier to identify. Therefore, by combining facial features, clothing features and body features, it is possible to more accurately determine the person to whom the person belongs.

[0049] In step 102, based on the characteristics of the person to be matched and the baseline characteristics of the target object, at least one first image matching the target object is determined from the received images.

[0050] In this embodiment, image matching is performed based on the extracted features of the person to be matched and the pre-defined baseline features of the target object. The baseline features of the target object include at least one of facial features, clothing features, and body posture features, and these features are specified by the user or automatically collected and stored by the electronic device during the initialization or specific settings phase of the electronic device.

[0051] In this embodiment, a specific matching algorithm can be used to compare and analyze the features of the person to be matched with the baseline features of the target object. The matching algorithm can be selected according to actual needs, such as distance-based algorithms like Euclidean distance or cosine similarity, or machine learning algorithms like support vector machines or neural networks. By calculating the similarity between features, it is determined whether the person in the image matches the target object. When the similarity between the features of the person in the image and the baseline features of the target object exceeds a preset matching threshold, the image is considered to match the target object and is designated as the first image.

[0052] In this embodiment, the matching threshold can be adjusted according to the accuracy and recall requirements of the actual application to flexibly adapt to various application scenarios.

[0053] In the embodiments of this application, the first image typically includes the target object.

[0054] In step 103, at least one first image is stored in the album corresponding to the target object.

[0055] In this embodiment of the application, the album corresponding to the target object can be an electronic album, stored on a local device or a cloud server, for convenient viewing and analysis later.

[0056] In this embodiment of the application, the album corresponding to the target object can be a storage space specially created by the electronic device for each target object, used to centrally store images that match the target object.

[0057] In this embodiment of the application, some of the first images can be stored in the album corresponding to the target object, or all of the first images can be stored in the album corresponding to the target object.

[0058] In this embodiment, a portion of the first images can be automatically filtered and stored in the album corresponding to the target object, or the number of first images can be stored in the album corresponding to the target object according to the number of images preset by the user.

[0059] In this embodiment, users can customize and edit the name of the album corresponding to the target object to meet the user's personalized management needs.

[0060] As can be seen from the above embodiments, in this embodiment, when the number of images received through the target chat group exceeds a first threshold, the features of the person to be matched in each received image are extracted; based on the features of the person to be matched and the baseline features of the target object, at least one first image matching the target object is determined from the received images; and at least one first image is stored in the album corresponding to the target object. On the one hand, this image management process eliminates the need for users to manually filter images one by one, saving users time and effort, enabling them to organize and manage images received in the chat group more efficiently, and improving the overall user experience; on the other hand, this precise classification and storage method allows users to easily and quickly find all images related to the target object in a specific album, avoiding the hassle of searching through a large number of cluttered images, providing users with a more orderly and convenient image viewing and management experience, thereby improving image management efficiency.

[0061] In some embodiments provided in this application, a two-layer comparison architecture of "rapid screening + accurate verification" can be adopted to reduce the computational load while ensuring accuracy. Accordingly, the above step 102 may include the following steps: step 1021 and step 1022.

[0062] In step 1021, a feature comparison algorithm is used to compare the features of the person to be matched with the baseline features of the target object. Based on the comparison results, a set of candidate images similar to the target object is selected from the received images.

[0063] In this application embodiment, the human characteristics may include at least one of the following: facial features, clothing features, and body posture features. Facial features may include: the shape of the person's facial features, facial contours, and facial expressions. Clothing features may include: the color, style, and pattern of the person's clothing, as well as accessories such as hats, necklaces, and earrings. Body posture features may include: the person's height, body type, and posture, such as standing, sitting, and walking postures.

[0064] In this embodiment of the application, a feature comparison algorithm is used to compare the features of the person to be matched with the baseline features of the target object in order to achieve rapid screening. This algorithm is usually a relatively simple and efficient algorithm that can quickly calculate the similarity between the two.

[0065] In the embodiments of this application, the feature matching algorithms include, but are not limited to: feature matching algorithms based on distance metrics, feature matching algorithms based on statistical methods, feature matching algorithms based on machine learning, and feature matching algorithms based on deep learning.

[0066] In this embodiment, although the rapid screening stage uses a relatively simple algorithm, it can initially eliminate a large number of obviously mismatched images, reducing the amount of data that needs to be processed in the subsequent precise verification stage. The precise verification stage then conducts in-depth analysis of the screened candidate images to further eliminate possible misjudgments. This two-layer screening method can effectively reduce the accumulation of errors and improve the accuracy of the final result.

[0067] In step 1022, a deep learning model is used to perform feature comparison processing on the features of the person to be matched in each candidate image in the candidate image set and the baseline features of the target object. Based on the comparison results, at least one first image matching the target object is determined from the candidate image set.

[0068] In this embodiment of the application, a deep learning model is used to compare the features of the person to be matched in each candidate image in the initially selected candidate image set with the baseline features of the target object again, so as to achieve accurate verification.

[0069] In this embodiment, the deep learning model possesses powerful feature extraction and learning capabilities, enabling it to more accurately capture subtle differences in human characteristics. This precise verification step improves matching accuracy and reduces misjudgments.

[0070] In this application embodiment, the deep learning model includes, but is not limited to: convolutional neural network model, recurrent neural network model, attention mechanism model, etc.

[0071] In this embodiment, a deep learning model is used in the precise verification stage. This model can automatically learn complex patterns and high-level abstract representations of human features. Compared with traditional feature matching algorithms, deep learning models can better handle various complex situations, such as changes in lighting, pose changes, and occlusion. For example, in face recognition, even if part of the target object's face is occluded, the deep learning model can still accurately match it using other learned facial feature information, thereby ensuring a high degree of match between the final determined first image and the target object, and improving the accuracy of the entire comparison process.

[0072] For example, a feature comparison algorithm based on distance metric is used in the rapid screening stage, and a ResNet-34 deep learning model is used in the accurate verification stage.

[0073] As can be seen, in this embodiment of the application, a two-layer comparison architecture of "rapid screening + precise verification" is adopted. This architecture divides the human feature comparison process into two stages: first, preliminary screening is carried out, and then precise matching is carried out. This ensures the final matching accuracy while reducing unnecessary calculations and improving the overall processing efficiency.

[0074] In some embodiments provided in this application, the provided image management method can also be used in... Figure 1 Based on the illustrated embodiment, the following step is added: Step 100;

[0075] In step 100, the human features of the visual reference material are extracted, and the extracted human features are determined as the baseline human features of the target object.

[0076] In this embodiment of the application, the visual reference material may include at least one of the following: a single person reference image, a first labeled area of ​​a single person reference image, a second labeled area of ​​a group photo reference image, a reference video, and a non-realistic rendered image.

[0077] In this embodiment of the application, the single-person reference image, the first annotation area, the second annotation area, and the reference video all include the target object.

[0078] In this embodiment of the application, the non-realistic rendered image may include at least one of the following: cartoon image, emoticon image, sketch image, oil painting image.

[0079] In this embodiment, the first annotation region of the single-person reference image and the second annotation region of the group photo reference image allow users or electronic devices to specify key feature regions of the target object.

[0080] In this embodiment, considering that there may be interfering information in the image in some cases, such as cluttered background or partial occlusion by other people, these interferences can be eliminated by marking the area, allowing the feature extraction algorithm to focus more on the key parts of the target object, such as the face or specific clothing patterns, thereby improving the accuracy of feature extraction.

[0081] In this embodiment of the application, considering that in a group photo of multiple people, the target object is partially obscured by other people, but by marking the unobscured facial area of ​​the target object, the feature extraction algorithm can accurately extract facial features and avoid the influence of the obscured part.

[0082] In this embodiment of the application, the reference video contains dynamic information of the target object at different times, which can provide richer human characteristics compared to static images.

[0083] In some cases, it may be difficult to obtain a real image of the target object, or the number of real images may be limited. Non-realistic rendered images, such as cartoon images and emoji images, can supplement real images and provide additional information about the person's features.

[0084] As can be seen, in this embodiment, the baseline characteristics of the target object are established by extracting features from various types of visual reference materials. These visual reference materials cover a wide range, and different types of visual reference materials can showcase the characteristics of the target object from multiple dimensions, improving the comprehensiveness and accuracy of feature extraction.

[0085] In some embodiments provided in this application, the provided image management method can also be used in... Figure 1 Based on the embodiment shown, after step 103 above, the following step is added: step 104.

[0086] In step 104, a non-realistic rendered image and / or video containing the target object is generated based on at least one first image.

[0087] In this embodiment of the application, when the baseline human features of the target object are derived from a non-realistic rendering image, the human features of the first image are highly similar to those of the non-realistic rendering image. In this case, generating a non-realistic rendering image containing the target object based on the first image can enrich the quantity and content of non-realistic rendering images and enhance their appeal.

[0088] For example, if the target chat group is a kindergarten parent group, and the target is a child named "Miaomiao," and Miaomiao's baseline characteristics are derived from a cute emoji, then after filtering out the first image matching Miaomiao from the images received from the kindergarten parent group, an emoji for Miaomiao is generated based on at least one of the first images. Furthermore, Miaomiao's emoji can be added to instant messaging tools for parents to easily save and share.

[0089] In this embodiment of the application, when the baseline features of the target object are derived from a video, the features of the first image are highly similar to those of the video. In this case, generating a video containing the target object based on the first image can enrich the number and content of the video and increase its interest.

[0090] For example, if the target chat group is a kindergarten parent group and the target is a child named "TuTu", and the baseline characteristics of TuTu are derived from a video, then after filtering out the first image that matches TuTu from the images received from the kindergarten parent group, a video clip containing TuTu is obtained by splicing together at least one of the first images, so that parents can save it.

[0091] As can be seen, in the embodiments of this application, the target object can be presented in an artistic, stylized and dynamic way by generating non-realistic rendered images, videos and other forms, which can bring a new visual experience to the user.

[0092] In some embodiments provided in this application, step 101 above may include the following step: step 1011.

[0093] In step 1011, if the number of images received through the target chat group within a preset time window exceeds a first threshold, the features of the person to be matched in each received image are extracted.

[0094] In this embodiment, the preset time window sets a clear time range for image reception. Only images received within this specific time range will be included in the subsequent processing flow. This helps to exclude irrelevant images outside the time window and avoid interference from images received during long periods of accumulation or unrelated time periods. This allows for focusing on image data closely related to the target task within a specific time period, thereby improving the accuracy and relevance of data processing.

[0095] In this embodiment of the application, considering that images from different time periods may have different feature distributions and correlations, by setting a preset time window, it can be ensured that the extracted features of the person to be matched are based on images within the same time period. These images may be more similar in terms of shooting environment, person status, etc., thereby making the extracted features more consistent and timely, which is conducive to more accurate matching of person features in the future.

[0096] In this embodiment, the preset time window can be flexibly set according to the user's actual needs, enabling the user to quickly obtain matching images within the time period of interest, meeting the user's image search and management needs in specific scenarios, and improving user satisfaction.

[0097] In this embodiment of the application, different users may have different preferences for setting the time window. For example, some users may want to focus on images from the most recent hour, while others may focus more on images from the same day. By allowing users to customize the preset time window, more personalized image processing services can be provided to users.

[0098] For example, the preset time windows are 30 minutes, 1 hour, 3 hours, 6 hours, 12 hours, 24 hours, etc.

[0099] In this embodiment of the application, for chat groups with strong immediacy, such as work discussion groups or emergency coordination groups, information updates are rapid and communication is frequent. The preset time window can be set to 30 minutes or 1 hour, so as to capture a large amount of information interaction in a short period of time within the group in a timely manner.

[0100] In this embodiment of the application, for daily chat groups of friends or interest groups, the pace of communication is relatively slow. The preset time window can be set to 3 hours or 6 hours, which can cover the normal communication of the group within a certain period of time, and will not cause frequent triggering of monitoring and feature extraction operations due to the time window being too short, thus increasing the burden on electronic devices.

[0101] In this embodiment of the application, for groups that require long-term information accumulation, such as academic research groups or industry information sharing groups, the importance and value of the information may take a long time to accumulate and be realized. The preset time window can be set to 12 hours or 24 hours to collect all images of the group within a day for comprehensive management.

[0102] As can be seen in this embodiment, a large number of images may flood into a chat group in a short period of time. Without filtering and feature extraction, these images will accumulate into a massive dataset, making them difficult to manage and utilize effectively. Extracting the features of the people to be matched when the number of images received within a preset time window exceeds a first threshold can transform the images into a more structured and analyzable data format. For example, in a work group for a large event, photos from the event are constantly being uploaded. By extracting the features of the people involved, the images can be quickly classified and organized, archived according to different people, and easily retrieved and used later.

[0103] In some embodiments provided in this application, step 1011 may include the following steps: step 10111 and step 10112;

[0104] In step 10111, when the number of target chat groups is 1 and the number of images received through the target chat groups within a preset time window exceeds a first threshold, the features of the people to be matched for each image received by the target chat groups within the preset time window are extracted.

[0105] In this embodiment of the application, when the number of target chat groups is 1, the electronic device continuously monitors the chat group. If it is detected that the number of images received by the chat group within a preset time window exceeds a preset first threshold, the features of the people to be matched in each image received by the chat group within the preset time window are extracted.

[0106] In this embodiment of the application, some users may only be interested in images in a specific chat group. For example, a user may participate in a kindergarten parent group and only want to manage the images of their own child shared in the group. In this case, the electronic device continuously monitors this chat group and extracts features when the number of images received within a preset time window exceeds a first threshold. This can accurately meet the user's needs for image management in this specific chat group and avoid unnecessary processing of other unrelated chat groups.

[0107] In step 10112, if the number of target chat groups is greater than 1 and the sum of the number of images received through all target chat groups within a preset time window exceeds a first threshold, the features of the people to be matched in each image received by all target chat groups within the preset time window are extracted.

[0108] In this embodiment, when the number of target chat groups is greater than one, the electronic device simultaneously monitors all the set target chat groups. If the sum of the number of images received by all target chat groups within a preset time window exceeds a first threshold, the features of the individuals to be matched in each image received by all target chat groups within the preset time window are extracted.

[0109] In this embodiment, some users may participate in multiple related chat groups simultaneously. For example, a parent may have joined multiple groups, such as a kindergarten class group and an extracurricular activity group, and wish to comprehensively manage images related to their child in these groups. When the sum of the number of images received by all the set target chat groups within a preset time window exceeds a first threshold, features are extracted. This allows for the comprehensive collection of images related to the target object, ensuring that no important images in any target group are missed, guaranteeing the integrity of the image data, and meeting the user's need for centralized management of images from multiple related target groups.

[0110] As can be seen, in this embodiment, when the number of target chat groups is 1 and the number of received images exceeds the first threshold within a preset time window, only the features of the people to be matched in the images within that group are extracted, which can concentrate processing effort on the communication content of a specific chat group. When the number of target chat groups is greater than 1 and the total number of received images exceeds the first threshold, the features of the people to be matched in the images of all groups are extracted, which can achieve comprehensive integration of information from multiple groups.

[0111] In some embodiments provided in this application, the provided image management method can also be used in... Figure 1 Based on the embodiment shown, after step 103 above, the following step is added: step 105.

[0112] In step 105, the reception time and source group information of each first image in the album corresponding to the target object are marked.

[0113] In this embodiment of the application, while storing the first image, the electronic device automatically labels the reception time and source group information of each first image.

[0114] In this embodiment of the application, the reception time of the first image can be accurate to a specific time point, such as year-month-day-hour-minute-second; in addition, it can also support the generation of a visual timeline browsing interface to facilitate users to view the image.

[0115] In this embodiment of the application, the source group information of the first image clearly shows which target chat group the first image was received from. If the first image comes from multiple target chat groups, all relevant group information is marked to facilitate subsequent query and management by the user.

[0116] As can be seen, in this embodiment of the application, when annotating image information, not only is the receiving time recorded, but also the source group information is annotated in detail. Especially for images from multiple groups, users can clearly understand the source of the image, providing users with a more convenient and comprehensive image search and usage experience.

[0117] In some embodiments provided in this application, step 103 above may include the following steps: step 1031.

[0118] In step 1031, if the number of first images is greater than 1, the action similarity of all first images is determined, and the top N first images with the largest action differences are stored in the album corresponding to the target object; wherein, the lower the action similarity, the greater the action difference, and N≥1.

[0119] In this embodiment, a feature extraction method can be used to extract action-related features from each first image. For example, a deep learning model, such as a convolutional neural network combined with optical flow, can be used to extract action features. Specifically, a pre-trained CNN model is used to extract static features of the image, while optical flow is used to calculate the motion information of pixels in the image sequence. The static features and motion information are then fused together as action features.

[0120] In this embodiment, a similarity measurement method can be used to calculate the action similarity between any two first images. This similarity measurement method includes, but is not limited to, cosine similarity and Euclidean distance. Taking cosine similarity as an example, the extracted action features are represented as vectors, and their similarity is measured by calculating the cosine of the angle between two vectors. The closer the cosine value is to 1, the more similar the actions are; the closer it is to -1, the greater the difference in actions.

[0121] In this embodiment, the top N first images with the greatest action differences can be identified based on the calculated action similarity. Specifically, this can be achieved by sorting the action similarity between all pairs of images, or by using a more efficient algorithm (such as clustering algorithm combined with similarity threshold filtering).

[0122] In this embodiment, the top N images with the largest differences in actions are stored in the album corresponding to the target object. This album can be an electronic album, stored on a local device or a cloud server, for convenient subsequent viewing and analysis.

[0123] As can be seen, in this embodiment of the application, the first image with a large difference in action is selected as the representative image for storage, thereby achieving intelligent deduplication. On the one hand, this makes the image set of the target object more diverse and avoids storing a large number of images with similar actions, saving storage space while better reflecting the behavioral characteristics and personality of the target object. On the other hand, when it is necessary to find a specific action image of the target object, since the album stores images with large differences in action, the required action type can be quickly located, reducing retrieval time and workload.

[0124] For ease of understanding, combined with Figures 2 to 5 The four examples shown illustrate the above method implementations.

[0125] Example 1: Duoduo's parents can select 3-5 individual photos of Duoduo from their local photo album, and mark key areas such as the face and clothing by drawing circles. In this case, only the features of the person within that area will be extracted. If no areas are marked, the entire image will be used as the area for feature extraction. Figure 2As shown, taking a single-person image as an example, firstly, the electronic device 20 can extract the features of Duoduo in the single-person image 21 as the baseline features for subsequent image comparison. Then, when the kindergarten parent group 22 receives 10 images within 10 minutes, an automatic comparison process is triggered, and the comparison progress is notified to the user via a pop-up window 23. For example, the first round of rapid comparison filters out 5 suspected images, and the second round of detailed comparison confirms 3 matching images. After the comparison is completed, a pop-up window 24 can notify the user that the matched images will be automatically saved to the "2025-08-01_Duoduo Exclusive" album. In addition, information such as the image's reception time and source can also be labeled.

[0126] Example 2: If Youyou's parents provide a group photo containing Youyou, then the image area where Youyou is located can be marked in the group photo, and only the features of people in that image area can be extracted. For example... Figure 3 As shown, firstly, the electronic device 30 can extract the features of the people in the image region 311 where Youyou is located in the group photo image 31 as the baseline features for subsequent image comparison. Then, when the kindergarten parent group 32 and the dance class extracurricular group 33 receive 23 images within 3 minutes, an automatic comparison process is triggered, and the comparison progress is notified to the user via a pop-up window 34. For example, the first round of rapid comparison filters out 5 suspected images, and the second round of detailed comparison confirms 3 matching images. After the comparison is completed, a pop-up window 35 can notify the user that the matched images will be automatically saved to the "2025-08-01_Youyou Exclusive" album. In addition, information such as the image's reception time and source can also be labeled.

[0127] Example 3: Duoduo's parents can select a cartoon-style or emoji-like non-realistic rendered image, such as... Figure 4 As shown, firstly, the electronic device 40 can extract the human features from the non-realistic rendered image 41 as baseline human features for subsequent image comparison. Then, when the kindergarten parent group 42 receives 10 images within 10 minutes, an automatic comparison process is triggered, and the user is notified of the comparison progress via a pop-up window 43. For example, the first round of rapid comparison filters out 5 suspected images, and the second round of detailed comparison confirms 3 matching images. After the comparison is completed, a pop-up window 44 notifies the user that the matched images will be automatically saved to the "2025-08-01_Duoduo Exclusive" album. In addition to annotating the image's reception time and source, cartoon images or emoticons can also be generated based on the images.

[0128] Example 4: Duoduo's parents can select a video containing Duoduo from their local photo album, such as... Figure 5As shown, firstly, electronic device 50 can extract the facial features of Duoduo in video 51 as baseline facial features for subsequent image comparison. Then, when the kindergarten parent group 52 receives 10 images within 10 minutes, an automatic comparison process is triggered, and the comparison progress is notified to the user via pop-up window 53. For example, the first round of rapid comparison filters out 5 suspected images, and the second round of detailed comparison confirms 3 matching images. After the comparison is completed, a pop-up window 54 notifies the user that the matched images will be automatically saved to the "2025-08-01_Duoduo Exclusive" album. In addition to annotating the image reception time and source, video clips containing Duoduo can also be generated based on the images.

[0129] As can be seen, the image management method in this application embodiment is applied to a chat group scenario. When the number of images received by a specific chat group exceeds a set threshold, the features of people in the images are automatically extracted, matched with the features of the target object, and the filtered images are stored in the album corresponding to the target object, so as to achieve effective management and organization of images.

[0130] The image management method provided in this application can be executed by an image management device. This application uses an image management device executing the image management method as an example to illustrate the image management device provided in this application.

[0131] Figure 6 This is a structural block diagram of an image management device provided in an embodiment of this application, such as... Figure 6 As shown, the image management device 600 may include: an extraction module 601, a determination module 602, and a storage module 603;

[0132] The extraction module 601 is used to extract the features of the person to be matched in each of the received images when the number of images received through the target chat group exceeds a first threshold.

[0133] The determining module 602 is used to determine at least one first image that matches the target object from the received images based on the characteristics of the person to be matched and the baseline characteristics of the target object;

[0134] The storage module 603 is used to store the at least one first image into the album corresponding to the target object.

[0135] As can be seen from the above embodiments, in this embodiment, when the number of images received through the target chat group exceeds a first threshold, the features of the person to be matched in each received image are extracted; based on the features of the person to be matched and the baseline features of the target object, at least one first image matching the target object is determined from the received images; and at least one first image is stored in the album corresponding to the target object. On the one hand, this image management process eliminates the need for users to manually filter images one by one, saving users time and effort, enabling them to organize and manage images received in the chat group more efficiently, and improving the overall user experience; on the other hand, this precise classification and storage method allows users to easily and quickly find all images related to the target object in a specific album, avoiding the hassle of searching through a large number of cluttered images, providing users with a more orderly and convenient image viewing and management experience, thereby improving image management efficiency.

[0136] Optionally, as an embodiment, the person's characteristics include at least one of the following: facial features, clothing features, and body features;

[0137] The determining module 602 is specifically used to perform feature comparison processing on the features of the person to be matched and the baseline features of the target object through a feature comparison algorithm, and to filter out a set of candidate images similar to the target object from the received images based on the comparison results; and to perform feature comparison processing on the features of the person to be matched of each candidate image in the candidate image set and the baseline features of the target object through a deep learning model, and to determine at least one first image matching the target object from the candidate image set based on the comparison results.

[0138] Optionally, as an embodiment, the extraction module 601 is further configured to extract the human features of the visual reference material and determine the extracted human features as the baseline human features of the target object; wherein, the visual reference material includes at least one of the following: a single-person reference image, a first labeled area of ​​a single-person reference image, a second labeled area of ​​a group photo reference image, a reference video, and a non-realistic rendering image; the single-person reference image, the first labeled area, the second labeled area, and the reference video all include the target object; the non-realistic rendering image includes at least one of the following: a cartoon image, an emoticon image, a sketch image, and an oil painting image.

[0139] Optionally, as an embodiment, the image management device 600 may further include:

[0140] A generation module is used to generate a non-realistic rendered image and / or video containing the target object based on at least one first image.

[0141] Optionally, as an embodiment, the extraction module 601 is specifically used to extract the features of the person to be matched from each image received by the target chat group within the preset time window when the number of target chat groups is 1 and the number of images received through the target chat group within the preset time window exceeds a first threshold; and to extract the features of the person to be matched from each image received by all the target chat groups within the preset time window when the number of target chat groups is greater than 1 and the sum of the number of images received through all the target chat groups within the preset time window exceeds a first threshold.

[0142] Optionally, as an embodiment, the image management device 600 may further include:

[0143] The annotation module is used to annotate the reception time and source group information of each of the first images in the album.

[0144] Optionally, as an embodiment, the storage module 603 is specifically used to determine the action similarity of all the first images when the number of the first images is greater than 1, and store the top N first images with the largest action differences into the album corresponding to the target object; wherein, the lower the action similarity, the greater the action difference, and N≥1.

[0145] The image management device in this application embodiment can be a stand-alone electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0146] The image management device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0147] The image management device provided in this application embodiment can achieve... Figure 1 To avoid repetition, the various processes implemented in the method embodiment shown will not be described again here.

[0148] Optionally, such as Figure 7 As shown, this application embodiment also provides an electronic device 700, including a processor 701 and a memory 702. The memory 702 stores a program or instructions that can run on the processor 701. When the program or instructions are executed by the processor 701, they implement the various steps of the above-described image management method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0149] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0150] Figure 8 This is a schematic diagram of the hardware structure of an electronic device that implements the various embodiments of this application.

[0151] The electronic device 800 includes, but is not limited to, components such as: radio frequency unit 801, network module 802, audio output unit 803, input unit 804, sensor 805, display unit 806, user input unit 807, interface unit 808, memory 809, and processor 810.

[0152] Those skilled in the art will understand that the electronic device 800 may also include a power supply (e.g., a battery) for supplying power to various components. The power supply may be logically connected to the processor 810 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 8 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0153] The processor 810 is configured to, when the number of images received through a target chat group exceeds a first threshold, extract the features of the person to be matched from each of the received images; determine at least one first image that matches the target object from the received images based on the features of the person to be matched and the baseline features of the target object; and store the at least one first image in the album corresponding to the target object.

[0154] As can be seen, in this embodiment of the application, on the one hand, this image management process does not require users to manually filter images one by one, saving users time and effort, enabling users to organize and manage images received in chat groups more efficiently, and improving the overall user experience; on the other hand, this precise classification and storage method allows users to easily and quickly find all images related to the target object in a specific album, avoiding the trouble of searching in a large number of messy images, providing users with a more orderly and convenient image viewing and management experience, thereby improving image management efficiency.

[0155] Optionally, as an embodiment, the person's characteristics include at least one of the following: facial features, clothing features, and body features;

[0156] The processor 810 is specifically configured to perform feature comparison processing on the features of the person to be matched and the baseline features of the target object using a feature comparison algorithm, and to filter out a set of candidate images similar to the target object from the received images based on the comparison results; and to perform feature comparison processing on the features of the person to be matched in each candidate image in the candidate image set and the baseline features of the target object using a deep learning model, and to determine at least one first image matching the target object from the candidate image set based on the comparison results.

[0157] Optionally, as an embodiment, the processor 810 is further configured to extract human features from visual reference materials and determine the extracted human features as the baseline human features of the target object; wherein, the visual reference materials include at least one of the following: a single-person reference image, a first labeled area of ​​a single-person reference image, a second labeled area of ​​a group photo reference image, a reference video, and a non-realistic rendering image; the single-person reference image, the first labeled area, the second labeled area, and the reference video all include the target object; the non-realistic rendering image includes at least one of the following: a cartoon image, an emoji image, a sketch image, and an oil painting image.

[0158] Optionally, as an embodiment, the processor 810 is also configured to generate a non-realistic rendered image and / or video containing the target object based on the at least one first image.

[0159] Optionally, as an embodiment, the processor 810 is specifically configured to extract the features of the person to be matched from each image received by the target chat group within the preset time window when the number of target chat groups is 1 and the number of images received through the target chat group within the preset time window exceeds a first threshold; and to extract the features of the person to be matched from each image received by all the target chat groups within the preset time window when the number of target chat groups is greater than 1 and the sum of the number of images received through all the target chat groups within the preset time window exceeds a first threshold.

[0160] Optionally, as an embodiment, the processor 810 is also used to annotate the reception time and source group information of each of the first images in the album.

[0161] Optionally, as an embodiment, the processor 810 is specifically configured to, when the number of the first images is greater than 1, determine the action similarity of all the first images, and store the top N first images with the largest action differences into the album corresponding to the target object; wherein, the lower the action similarity, the greater the action difference, and N≥1

[0162] It should be understood that, in this embodiment, the input unit 804 may include a graphics processing unit (GPU) 8041 and a microphone 8042. The GPU 8041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 806 may include a display panel 8061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 807 includes at least one of a touch panel 8071 and other input devices 8072. The touch panel 8071 is also called a touch screen. The touch panel 8071 may include a touch detection device and a touch controller. Other input devices 8072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0163] The memory 809 can be used to store software programs and various data. The memory 809 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (e.g., sound playback function, image playback function, etc.). Furthermore, the memory 809 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 809 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0164] Processor 810 may include one or more processing units; optionally, processor 810 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 810.

[0165] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image management method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0166] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0167] This application also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described image management method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0168] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0169] This application also provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described image management method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0170] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0171] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (e.g., ROM / RAM, magnetic disk, optical disk), including several instructions to cause a terminal (e.g., a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0172] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An image management method, characterized in that, The method includes: If the number of images received through the target chat group exceeds a first threshold, extract the features of the person to be matched from each of the received images. Based on the characteristics of the person to be matched and the baseline characteristics of the target object, at least one first image matching the target object is determined from the received images; Store the at least one first image to the album corresponding to the target object.

2. The method according to claim 1, characterized in that, The character features include at least one of the following: facial features, clothing features, and body features; The step of determining at least one first image matching the target object from the received images based on the characteristics of the person to be matched and the baseline characteristics of the target object includes: The feature comparison algorithm is used to compare the features of the person to be matched with the baseline features of the target object. Based on the comparison results, a set of candidate images similar to the target object is selected from the received images. Using a deep learning model, the features of the person to be matched in each candidate image in the candidate image set are compared with the baseline features of the target object. Based on the comparison results, at least one first image matching the target object is determined from the candidate image set.

3. The method according to claim 1, characterized in that, Before determining at least one first image matching the target object from the received images based on the features of the person to be matched and the baseline features of the target object, the method further includes: Extract the character features from the visual reference material, and determine the extracted character features as the baseline character features of the target object; The visual reference materials include at least one of the following: a single person reference image, a first labeled area of ​​a single person reference image, a second labeled area of ​​a group photo reference image, a reference video, and a non-realistic rendered image; The single-person reference image, the first labeled area, the second labeled area, and the reference video include the target object; The non-realistic rendered images include at least one of the following: cartoon images, emoji images, sketch images, and oil painting images.

4. The method according to claim 3, characterized in that, After determining at least one first image matching the target object from the received images based on the characteristics of the person to be matched and the baseline characteristics of the target object, the method further includes: Based on the at least one first image, generate a non-realistic rendered image and / or video containing the target object.

5. The method according to claim 1, characterized in that, When the number of images received through the target chat group exceeds a first threshold, the feature extraction of the person to be matched from each received image includes: If the number of target chat groups is 1 and the number of images received through the target chat group within a preset time window exceeds a first threshold, extract the features of the person to be matched for each image received by the target chat group within the preset time window. If the number of target chat groups is greater than 1 and the sum of the number of images received through all the target chat groups within the preset time window exceeds a first threshold, extract the features of the person to be matched from each image received by all the target chat groups within the preset time window.

6. The method according to claim 1, characterized in that, The method further includes: The receiving time and source group information of each of the first images in the album are marked.

7. The method according to claim 1, characterized in that, The step of storing the at least one first image to the album corresponding to the target object includes: If the number of the first images is greater than 1, determine the action similarity of all the first images, and store the top N first images with the largest action differences into the album corresponding to the target object; wherein, the lower the action similarity, the greater the action difference, and N≥1.

8. An image management device, characterized in that, The device includes: The extraction module is used to extract the features of the person to be matched from each of the received images when the number of images received through the target chat group exceeds a first threshold. The determining module is used to determine at least one first image that matches the target object from the received images based on the characteristics of the person to be matched and the baseline characteristics of the target object; A storage module is used to store the at least one first image into an album corresponding to the target object.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing programs or instructions that can run on the processor, the programs or instructions being executed by the processor to implement the steps of the image management method as described in any one of claims 1 to 7.

10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the image management method as described in any one of claims 1 to 7.