Image description method and electronic equipment

By recognizing and processing images in the image library to generate personalized descriptions, the problem of users finding it difficult to locate target images in a massive image library is solved, enabling fast and accurate image search and improving the user experience.

CN121904644APending Publication Date: 2026-04-21HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2024-10-21
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, users need to spend a lot of time and effort to find target images in a massive image library. Existing image search functions cannot quickly and accurately locate target images, which affects the user experience.

Method used

By performing image recognition processing, a bounding box for the target person's face is determined, and personalized, specific descriptive information is generated. Image search is then performed based on this descriptive information, improving search efficiency and accuracy.

Benefits of technology

It enables fast and accurate image search based on personalized description information, thus improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121904644A_ABST
    Figure CN121904644A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image description method and electronic equipment, relates to the technical field of image processing, and aims to describe an image, generate personalized and specific description information of the image and search the image based on the personalized description information, so that the search efficiency and accuracy can be improved, and the user experience can be improved. The method comprises the following steps: acquiring a photo album to be processed, wherein the photo album to be processed comprises at least one image to be described; performing recognition processing on each to-be-described image, and determining a portrait detection frame of a target person in each to-be-described image; and performing image description according to the portrait detection frame of the target person in each to-be-described image to obtain description information of each to-be-described image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing technology, and more particularly to an image description method and electronic device. Background Technology

[0002] With the continuous development of computer technology, smart devices offer increasingly richer functions, and users are demanding a faster, more efficient, and more innovative user experience. Nowadays, users frequently use smartphones to take photos and record their lives, often accumulating thousands upon thousands of photos in their albums. Searching for specific images within this vast library requires significant time and effort, impacting the user experience. Summary of the Invention

[0003] The embodiments of this disclosure provide an image description method and electronic device that can generate personalized and specific descriptive information for images. Image search based on the personalized descriptive information can improve search efficiency and accuracy.

[0004] To achieve the above objectives, the embodiments of this disclosure adopt the following technical solutions:

[0005] Firstly, embodiments of this disclosure provide an image description method applied to an electronic device. The method includes: firstly, acquiring a photo album to be processed, which includes at least one image to be described; then, performing recognition processing on each image to be described to determine a bounding box for a target person in each image; and finally, performing image description based on the bounding boxes for the target person in each image to obtain description information for each image. By describing images using this image description method, personalized and specific description information is generated. During image search, searches can be performed based on this personalized and specific description information, enabling quick and accurate retrieval of target images, improving search efficiency and accuracy, and enhancing the user experience.

[0006] In conjunction with the first aspect, another possible implementation involves obtaining the album to be processed, including: acquiring multiple images to be processed; performing recognition processing on each image to obtain the object corresponding to each image; classifying the multiple images according to the object corresponding to each image to obtain at least one album; wherein multiple images corresponding to the same object are grouped into the same album; and determining the album to be processed based on at least one album. Based on the above possible implementations, electronic devices can classify images according to the objects in the images, grouping images of the same object into the same album, thus achieving intelligent album creation. Image description based on the created albums can quickly identify the target object to be described, helping to improve the efficiency of generating image description information.

[0007] In conjunction with the first aspect, another possible implementation involves performing recognition processing on each image to be described to determine the face detection bounding box of the target person in each image. This includes: performing recognition processing on each image to be described to determine the face detection bounding box, face features, and face detection bounding box of each person in each image; determining the target person based on the face detection bounding box and face features of each person in each image; and determining the face detection bounding box of the target person in each image based on the face detection bounding box of the target person and the face detection bounding boxes of each person. Based on the above possible implementation, the electronic device determines the person appearing most frequently in the images to be described as the target person based on the face features of each person in the multiple images to be described included in the album to be processed, and then determines the face detection bounding box of the target person in the image to be described. Based on the face detection bounding box of the target person, the segmented image corresponding to the target person can be segmented during image description, and the remaining parts of the image that do not need to be described can be cropped. This improves the efficiency of generating image description information while reducing interference, which helps to improve the accuracy of image description information.

[0008] In conjunction with the first aspect, another possible implementation involves determining the target person based on the face detection bounding boxes and facial features of each person in each image to be described. This includes: determining a first image to be described with the fewest people from all the images to be described; determining the similarity distance between the faces in the first image to be described based on the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described (excluding the first image to be described); and determining the target person based on the similarity distance between the faces in the first image to be described. Based on this possible implementation, since the target person appears in every image to be described in the album to be processed, selecting the image with the fewest people as the first image to be described reduces the number of facial similarity distances that need to be calculated, which helps to quickly determine the target person and thus improves the efficiency of generating image description information.

[0009] In conjunction with the first aspect, another possible implementation involves determining the similarity distance of each face in the first image to be described based on the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described (excluding the first image to be described). This includes: determining the distance between the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described based on the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described based on the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described based on the distance between the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described; and determining the similarity distance of each face in the first image to be described based on the distance between the faces in the first image to be described and the second images to be described. Based on the above possible implementation methods, the similarity distance of each face in the first image to be described is determined according to the distance between the facial features of each person in the image to be described. The smaller the similarity distance, the more similar the face is to the faces in other images to be described, thus providing data support for subsequent determination of the target person based on the similarity distance.

[0010] In conjunction with the first aspect, another possible implementation involves determining the image detection box of the target person in each image to be described based on the face detection box and the image detection box of each person. This includes: determining candidate image detection boxes for the target person in each image to be described based on the face detection box and the image detection boxes of each person; wherein the face detection box of the target person is located within the candidate image detection box; determining the distance between the face detection box and the candidate image detection box in each image to be described; and determining the image detection box of the target person in each image to be described based on the distance between the face detection box and the candidate image detection box. Based on the above possible implementations, considering the constraint that the face detection box is located within the image detection box, under this constraint, the image detection box closest to the face detection box is taken as the image detection box of the target person. This eliminates the possibility of mismatch between the face detection box and the image detection box caused by special shooting postures, thereby improving the accuracy of the determined image detection box of the target person and achieving fast and accurate matching of the image detection box that matches the face detection box of the target person.

[0011] In conjunction with the first aspect, another possible implementation involves image description based on the bounding boxes of the target person in each image to be described, obtaining descriptive information for each image. This includes: acquiring a pre-trained visual model for image description; determining segmented images of the target person in each image based on the bounding boxes of the target person in each image; and using the visual model to describe the segmented images of the target person in each image, thereby obtaining descriptive information for each image. Based on the above possible implementation, the image corresponding to the target person is described using a pre-trained visual model for image description, generating personalized and specific descriptive information for that target person.

[0012] In conjunction with the first aspect, another possible implementation of the image description method further includes: responding to a first user operation by displaying an image search interface, the image search interface including a search box; receiving search information entered by the user in the search box, the search information being used to search for a target image; determining the target image among the images to be described based on the search information and the description information of each image to be described; and finally displaying the target image in the image search interface. Based on the above possible implementations, when a user enters search information for the target image they wish to search for in the search box, and searches are performed based on the search information and personalized, specific description information of the image, an accurate target image can be quickly determined, improving search efficiency and accuracy, and enhancing the user experience.

[0013] Secondly, embodiments of this disclosure provide an image description apparatus that can be applied to an electronic device to implement the image description method described in the first aspect. The functions of this image description apparatus can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned functions, such as an acquisition module, a recognition module, and a description module.

[0014] The acquisition module is configured to acquire a photo album to be processed, which includes at least one image to be described.

[0015] The recognition module is configured to perform recognition processing on each of the images to be described, and determine the portrait detection box of the target person in each of the images to be described.

[0016] The description module is configured to perform image description based on the portrait detection bounding box of the target person in each of the images to be described, and obtain description information for each of the images to be described.

[0017] In conjunction with the second aspect, in one possible implementation, the acquisition module is further configured to acquire multiple images to be processed; perform recognition processing on each of the images to be processed to obtain the object corresponding to each image to be processed; classify the multiple images to be processed according to the object corresponding to each image to be processed to obtain at least one album; wherein multiple images to be processed corresponding to the same object are classified into the same album; and determine the album to be processed according to the at least one album.

[0018] In conjunction with the second aspect, in one possible implementation, the recognition module is further configured to perform recognition processing on each of the images to be described, determine the face detection box, face features, and portrait detection box of each person in each of the images to be described; determine the target person based on the face detection box and face features of each person in each of the images to be described; and determine the portrait detection box of the target person in each of the images to be described based on the face detection box of the target person and the portrait detection box of each person in each of the images to be described.

[0019] In conjunction with the second aspect, in one possible implementation, the recognition module is further configured to determine a first image to be described with the fewest people from among the images to be described; determine the similarity distance of each face in the first image to be described based on the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described (excluding the first image to be described); and determine the target person based on the similarity distance of each face in the first image to be described.

[0020] In conjunction with the second aspect, in one possible implementation, the recognition module is further configured to: determine the distance between the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described, based on the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described (excluding the first image to be described); determine the distance between each face in the first image to be described and each of the second images to be described, based on the distance between the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described; and determine the similarity distance between each face in the first image to be described and each of the second images to be described, based on the distance between the faces in the first image to be described and each of the second images to be described.

[0021] In conjunction with the second aspect, in one possible implementation, the recognition module is further configured to: determine a candidate image detection box for the target person in each of the images to be described based on the face detection box and the image detection box for each person in each image to be described; wherein the face detection box of the target person is located within the candidate image detection box; determine the distance between the face detection box and the candidate image detection box in each of the images to be described; and determine the image detection box of the target person in each of the images to be described based on the distance between the face detection box and the candidate image detection box in each of the images to be described.

[0022] In conjunction with the second aspect, in one possible implementation, the description module is further configured to acquire a pre-trained visual model for describing images; determine a segmented image of the target person in each of the images to be described based on the portrait detection box of the target person in each of the images to be described; and use the visual model to perform image description on the segmented image of the target person in each of the images to be described, thereby obtaining description information for each of the images to be described.

[0023] In conjunction with the second aspect, in one possible implementation, the image description device further includes: a display module, a receiving module, and a search module;

[0024] The display module is configured to display an image search interface in response to a first user operation, the image search interface including a search box;

[0025] The receiving module is configured to receive search information entered by the user in the search box, the search information being used to search for a target image;

[0026] The search module is configured to determine the target image among the images to be described based on the search information and the description information of each image to be described;

[0027] The display module is also configured to display the target image in the image search interface.

[0028] Thirdly, this disclosure provides an electronic device, including: a memory, a display screen, and one or more processors; the memory, the display screen, and the processors are coupled. The memory stores computer program code, which includes computer instructions; when the electronic device is running, the processor executes one or more computer instructions stored in the memory to cause the electronic device to perform an image description method as described in any of the first aspects above.

[0029] Fourthly, this disclosure provides a computer storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform an image description method as described in any of the first aspects.

[0030] Fifthly, this disclosure provides a computer program product that, when run on an electronic device, causes the electronic device to execute the image description method as described in any of the first aspects.

[0031] In a sixth aspect, an apparatus (e.g., a system-on-a-chip) is provided, comprising a processor for supporting an electronic device in performing the functions described in the first aspect above. In one possible design, the apparatus further comprises a memory for storing program instructions and data necessary for the electronic device. When the apparatus is a system-on-a-chip, it may be composed of chips or may include chips and other discrete devices.

[0032] It should be understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0033] Figure 1 A schematic diagram of an interface for a gallery application provided by existing technology;

[0034] Figure 2 This is a schematic diagram of another interface for a gallery application provided in the prior art;

[0035] Figure 3 A schematic diagram of the shooting trajectory created in a gallery application provided in the prior art;

[0036] Figure 4 A schematic diagram of a people's photo album for recognition in existing image library applications;

[0037] Figure 5 A schematic diagram of the module interaction architecture of the gallery application provided in this embodiment of the disclosure;

[0038] Figure 6 A schematic diagram of an interface illustrating an example of image description provided in this disclosure;

[0039] Figure 7 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this disclosure;

[0040] Figure 8 A schematic diagram of the software structure of an electronic device provided in an embodiment of this disclosure;

[0041] Figure 9 A schematic diagram illustrating the implementation flow of an image description method provided in this embodiment of the disclosure;

[0042] Figure 10 A schematic diagram illustrating an example of an image description method provided in this disclosure;

[0043] Figure 11 A schematic diagram illustrating the implementation flow of another image description method provided in this embodiment of the disclosure;

[0044] Figure 12 A schematic diagram illustrating the implementation flow of another image description method provided in this embodiment of the disclosure;

[0045] Figure 13 This is a schematic diagram of the structure of a chip system provided in an embodiment of this disclosure. Detailed Implementation

[0046] The technical solutions of the embodiments of this disclosure will be described below with reference to the accompanying drawings. In the description of this disclosure, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can represent A or B. "And / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. Furthermore, in the description of this disclosure, unless otherwise stated, "multiple" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. Furthermore, to facilitate a clear description of the technical solutions of the embodiments of this disclosure, the terms "first" and "second" are used in the embodiments of this disclosure to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that "first" and "second" are not necessarily different. Meanwhile, in the embodiments of this disclosure, words such as "exemplary" or "for example" are used to indicate that something is being used as an example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this disclosure should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner for ease of understanding.

[0047] Furthermore, the network architecture and business scenarios described in the embodiments of this disclosure are for the purpose of more clearly illustrating the technical solutions of the embodiments of this disclosure, and do not constitute a limitation on the technical solutions provided in the embodiments of this disclosure. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this disclosure are also applicable to similar technical problems.

[0048] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the relevant concepts or technologies is given first:

[0049] 1. Natural Language Processing (NLP) refers to the technology of using natural language, which is used by humans to communicate, to interact with machines. It is an important direction in the fields of computer science and artificial intelligence.

[0050] 2. Image captioning is an interdisciplinary task in the fields of computer vision and natural language processing. It aims to enable computers to automatically generate a descriptive text based on an input image. This text usually contains key information such as the main objects, actions, and scenes in the image.

[0051] 3. The Hungarian Algorithm, also known as the KM Algorithm, is a classic algorithm for solving the task assignment problem in polynomial time. The task assignment problem involves n tasks and n people, each with different costs or efficiencies in completing different tasks. The goal is to allocate tasks to minimize the total cost or maximize the total efficiency. The core idea of ​​the Hungarian Algorithm is to perform a series of row and column transformations on the efficiency matrix so that each row and column contains a zero element. The optimal solution is then found by identifying the allocation scheme with the most zero elements.

[0052] With the continuous development of computer technology, electronic devices offer users a richer functional experience. Currently, users frequently use electronic devices (such as smartphones) to take photos and record their lives, and the photo albums on these devices often contain thousands or even tens of thousands of images. When users need to find the desired image among these massive amounts of images, they need to spend a lot of time and effort manually searching through them.

[0053] To facilitate user searches and improve search efficiency, electronic device photo albums offer image search functions. Existing image search functions mainly include the following search methods:

[0054] Keyword-based image search: Image library apps have recognition and tagging capabilities, or users can manually tag images beforehand. Users can enter keywords (such as items, plants, animals, text on images, etc.) in the search box, and the image library app will search and display relevant images. This search method requires pre-recognition of images and automatic or manual pre-tags. If the user's keywords are inaccurate or if images are not pre-recognized or tagged, the target image may not be found, or a category of images may be found, requiring the user to further filter the search results for the target image. For example, Figure 1 A schematic diagram of an interface for a gallery application provided by existing technology, such as Figure 1 As shown in (a), the search interface of the image library application includes a search box 101 and multiple image frames 102. Users can enter keywords such as "flower" into the search box 101 to search for images. The electronic device then searches the image library for images containing the keyword "flower." Figure 1 As shown in (b) above, among the multiple "flower" images 103, the user needs to further select the target image 104 they actually want to search from among the multiple "flower" images 103. After receiving the user's click on the target image 104, the electronic device enlarges the target image 104 to the size shown. Figure 1 The full-screen page display shown in (c) is shown in the image.

[0055] Time-based image search: The gallery app records the timestamps of each image and displays them in order of increasing time. Users can slide the timeline to quickly find photos from a specific period. This search method requires users to remember the image's date; otherwise, it's not quick to find the image. Furthermore, when multiple images are recorded for the same date, users need to further filter through the results for that date. For example, Figure 2 This is an alternative interface diagram for a gallery application provided in the prior art, such as... Figure 2 As shown in (a), multiple images are displayed in chronological order, from closest to furthest. When the user swipes up or down, the gallery interface displays the following: Figure 2 As shown in (b) of the diagram, the timeline 201 allows users to quickly locate the date corresponding to the target image they wish to search for. Then, when the electronic device receives a user's click on the target image, it enlarges the image to full-screen display.

[0056] Location-based image search: The gallery app records the shooting location of each photo. Photos from different locations can be categorized into different albums, or a shooting trajectory map can be created based on the shooting location. Users can search for photos and videos taken at a selected location by clicking on an album or a shooting trajectory map. This search method can only search for images whose shooting location is recorded, and if multiple images are found for the selected location, users need to further filter the search results for their target images. Figure 3 A schematic diagram of the shooting trajectory created in a gallery application provided in the prior art, such as Figure 3As shown, multiple images are categorized into different locations on the shooting trajectory map according to their shooting location. Users can click on a location on the shooting trajectory map based on the shooting location of the photo they want to search for. The electronic device determines and displays the image corresponding to the clicked location based on the user's click. When the selected location includes multiple images, the user needs to further select the target image they want to search for from among the multiple images.

[0057] Image search based on people: The photo library app features people recognition, automatically identifying and categorizing photos of different people into different albums. When a user clicks on an album containing a specific person, they can find all photos of that person. If an album contains multiple images of that person, the user needs to further filter through those images to find their desired result. Figure 4 An illustration of a people album for recognition in existing image library applications, such as... Figure 4 As shown, different people are categorized into different albums. When a user wants to search for a specific photo of a person, they click on the album of that person, but still need to select the image they want to search for from the multiple images included in that album.

[0058] Based on the above analysis, it can be seen that the image search function provided in the existing technology requires users to perform complex search operations and cannot quickly and accurately find the target image, thus failing to meet users' query needs and affecting the user experience.

[0059] To address the aforementioned technical problems, this disclosure provides an image library application that offers an image search function based on image descriptions. By processing the images, personalized and specific descriptive information is generated. Users can perform image searches based on this descriptive information, enabling them to quickly and accurately find target images and improve the user experience.

[0060] See Figure 5 , Figure 5 This is a schematic diagram of the module interaction architecture of the gallery application provided in the embodiments of this disclosure, such as... Figure 5 As shown, the image library application includes an image description device 501 and an image search device 502. The image description device 501 first acquires a photo album to be processed, which includes at least one image to be described; then it performs recognition processing on each image to be described to determine the portrait detection box of the target person in each image to be described; finally, it performs image description based on the portrait detection box of the target person in each image to be described, and obtains the description information of each image to be described.

[0061] For example, such as Figure 5As shown, the electronic device can store the image description information in a data file of the local file system 503. When a user needs to search for a target image, they enter the search information for the target image in the search box of the image search interface. The image search device 502 receives the search information entered by the user, loads the image description information from the local file system 503, performs a search and match based on the search information and the description information of each image, and determines the target image from all images.

[0062] The gallery application of this disclosure provides an image search function based on image description. The image description device describes the images included in the album of the gallery application to generate personalized and specific description information for the images. The image search device performs image search based on the personalized description information, which can improve search efficiency and accuracy and enhance the user experience.

[0063] For example, images in an electronic device's image library application are processed using image description methods to generate personalized and specific descriptive information for each image. This image description information may include, but is not limited to, information describing the object, its characteristics, and its behavior.

[0064] See Figure 6 ,by Figure 6 Taking the image shown in (a) as an example, when describing the image, the image is identified and described. The description information of the image can include the person xxx, xxx's clothing and behavior. For example, the description information of the generated image can be "xxx is running while wearing a white shirt and black pants".

[0065] The electronic device stores the descriptive information of a given image, such as "xxx is running while wearing a white shirt and black trousers".

[0066] When a user wants to search for an image, they can describe the target image they want to find using natural language. For example, a user can enter the natural language description of the target image they want to find in the search box of a gallery application. The electronic device processes the natural language input by the user to obtain search information. The electronic device then searches for target description information that matches the user's search information in the pre-stored description information of each image, and outputs the image described in the target description information as the target image.

[0067] For example, in the case of Figure 6When searching for the image shown in (a), the user can enter "photo of xxx running" in the search box. The electronic device determines the search information as "xxx" and "running" based on the natural language input by the user. The target description information matched by the search information "xxx" and "running" is "xxx is running while wearing a white shirt and black pants". Figure 6 The image shown in (a) is identified as the target image, and the search interface outputs the image as shown in (a). Figure 6 The search results shown in (b) are as follows. When an electronic device receives a user click... Figure 6 When operating on the target image shown in (b), the target image is enlarged to full screen display, such as... Figure 6 As shown in (c) in the figure.

[0068] In this embodiment, when performing image search based on image description, the image is described in detail and specifically using an image description method, generating personalized image description information. When performing image search based on personalized description information, the user only needs to enter the natural language of the image they want to search for in the search box to quickly and accurately find the target image. Compared to existing image search methods based on people, locations, times, keywords, etc., this method improves search efficiency and accuracy. Furthermore, in this embodiment, the information entered by the user in the search box can be natural language, making the search method more user-friendly and enhancing the user experience. Image search based on personalized descriptions can quickly and accurately find target images, improving search efficiency and accuracy, and enhancing the user experience.

[0069] Based on the above embodiments, this disclosure provides an image description method applied to electronic devices. By processing images using this method, personalized, detailed, and accurate descriptive information can be generated. Image searching based on this descriptive information can quickly and accurately locate target images, improving search efficiency and accuracy, and enhancing the user experience.

[0070] For example, the electronic device in this disclosure may be a portable computer (such as a mobile phone), a tablet computer, a wearable device, an in-vehicle device, an augmented reality (AR) / virtual reality (VR) device, a desktop / laptop / handheld computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, as well as a cellular phone, a personal digital assistant (PDA), a media player, and other devices capable of describing images. This disclosure does not impose any special limitations on the specific form of the electronic device.

[0071] Please refer to Figure 7 , Figure 7 A schematic diagram of a possible hardware structure for an electronic device is shown:

[0072] like Figure 7 As shown, the electronic device 700 may include: a processor 710, an external memory interface 720, an internal memory 721, a universal serial bus (USB) interface 730, a charging management module 740, a power management module 741, a battery 742, an antenna 1, an antenna 2, a mobile communication module 750, a wireless communication module 760, an audio module 770, a speaker 770A, a receiver 770B, a microphone 770C, a headphone jack 770D, a sensor module 780, buttons 790, a motor 791, an indicator 792, a camera 793, a display screen 794, and a subscriber identification module (SIM) card interface 795, etc.

[0073] The aforementioned sensor module 780 may include sensors such as pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, and bone conduction sensors.

[0074] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 700. In other embodiments, the electronic device 700 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0075] The processor 710 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.

[0076] The controller can be the nerve center and command center of the electronic device 700. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0077] The processor 710 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 710 is a cache memory. This memory can store instructions or data that the processor 710 has just used or that are used repeatedly. If the processor 710 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 710, and thus improves the efficiency of the system.

[0078] In some embodiments, the processor 710 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I4C) interface, an inter-integrated circuit sound (I4S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0079] It is understood that the interface connection relationships between the modules illustrated in this embodiment are merely illustrative and do not constitute a structural limitation on the electronic device 700. In other embodiments, the electronic device 700 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0080] The charging management module 740 receives charging input from a charger, which can be a wireless charger or a wired charger. While charging the battery 742, the charging management module 740 can also supply power to electronic devices via the power management module 741.

[0081] The power management module 741 connects the battery 742, the charging management module 740, and the processor 710. The power management module 741 receives input from the battery 742 and / or the charging management module 740, and supplies power to the processor 710, internal memory 721, external memory, display 794, camera 793, and wireless communication module 760, etc. In some embodiments, the power management module 741 and the charging management module 740 may also be housed in the same device.

[0082] The wireless communication function of the electronic device 700 can be implemented through antenna 1, antenna 2, mobile communication module 750, wireless communication module 760, modem processor, and baseband processor. In some embodiments, antenna 1 of the electronic device 700 is coupled to mobile communication module 750, and antenna 2 is coupled to wireless communication module 760, enabling the electronic device 700 to communicate with networks and other devices through wireless communication technology.

[0083] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 700 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.

[0084] The mobile communication module 750 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 700. The mobile communication module 750 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 750 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation.

[0085] The mobile communication module 750 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via the antenna 1. In some embodiments, at least some functional modules of the mobile communication module 750 can be housed in the processor 710. In some embodiments, at least some functional modules of the mobile communication module 750 and at least some modules of the processor 710 can be housed in the same device.

[0086] The wireless communication module 760 can provide solutions for wireless communication applications on electronic devices 700, including WLAN (such as wireless fidelity, Wi-Fi) networks, Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc.

[0087] The wireless communication module 760 can be one or more devices integrating at least one communication processing module. The wireless communication module 760 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signal, and sends the processed signal to processor 710. The wireless communication module 760 can also receive signals to be transmitted from processor 710, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0088] Electronic device 700 implements display functions through a GPU, a display screen 794, and an application processor. The GPU is a microprocessor for image processing, connecting the display screen 794 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 710 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0089] The display screen 794 is used to display images, videos, etc. The display screen 794 includes a display panel.

[0090] The electronic device 700 can implement its shooting function through an ISP, a camera 793, a video codec, a GPU, a display 794, and an application processor. The ISP is used to process the data fed back by the camera 793. The camera 793 is used to capture still images or videos. In some embodiments, the electronic device 700 may include one or N cameras 793, where N is a positive integer greater than 1.

[0091] The external storage interface 720 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the electronic device 700. The external storage card communicates with the processor 710 through the external storage interface 720 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.

[0092] Internal memory 721 can be used to store computer executable program code, which includes instructions. Processor 710 executes various functional applications and data processing of electronic device 700 by running the instructions stored in internal memory 721. For example, in embodiments of this disclosure, processor 710 can execute instructions stored in internal memory 721, which may include a program storage area and a data storage area.

[0093] The program storage area can store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.). The data storage area can store data created during the use of the electronic device 700 (such as audio data, phonebook, etc.). Furthermore, the internal memory 721 can include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0094] Electronic device 700 can implement audio functions such as music playback and recording through audio module 770, speaker 770A, receiver 770B, microphone 770C, headphone jack 770D, and application processor.

[0095] Buttons 790 include a power button, volume buttons, etc. Buttons 790 can be mechanical buttons or touch-sensitive buttons. A motor 791 can generate vibration alerts. Motor 791 can be used for incoming call vibration alerts or for touch vibration feedback. An indicator 792 can be an indicator light, used to indicate charging status, battery level changes, messages, missed calls, notifications, etc. A SIM card interface 795 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 795 to achieve contact and separation with the electronic device 700. The electronic device 700 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 795 can support Nano SIM cards, Micro SIM cards, and other SIM cards.

[0096] In addition, an operating system, such as HarmonyOS, iOS, Android, or Windows, runs on the aforementioned components. Applications can be installed and run on this operating system. In other embodiments, the electronic device may run multiple operating systems.

[0097] It should be understood that Figure 7The hardware modules included in the illustrated electronic device are merely illustrative and do not limit the specific structure of the electronic device. In fact, the electronic device provided in this disclosure may also include other hardware modules that interact with the hardware modules illustrated in the figures; these are not specifically limited here. For example, the electronic device may also include a flash, a miniature projector, etc. Furthermore, if the electronic device is a personal computer (PC), it may also include components such as a keyboard and a mouse.

[0098] The software system of the aforementioned electronic device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This disclosure embodiment uses a layered architecture. Taking the system as an example, the software structure of the electronic device is illustrated.

[0099] Figure 8 This is a software architecture block diagram of an electronic device according to an embodiment of the present disclosure. The layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, [the following is omitted as the text is incomplete and requires further context]. The system is divided into four layers, from top to bottom: application layer, application framework layer, Android runtime, system library, and kernel layer.

[0100] The application layer can include a series of application packages. For example, these application packages may include applications such as camera, gallery, calendar, phone, map, navigation, WLAN, Bluetooth, SystemUI, themes, and wallpaper applications. The systemUI is used to display the interface of the electronic device, such as the gallery application's interface. For example, the gallery may include photos and videos taken by the camera, as well as screenshots, screen recordings, downloaded images and videos, etc. The gallery application's search interface includes a search box and a search button. Users can enter natural language to describe the images they want to search for in the search box and click the search button to search for the desired images in the gallery based on the user's input.

[0101] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0102] like Figure 8 As shown, the application framework layer may include an activity manager, a window manager, a content provider, a view system, a resource manager, a notification manager, etc., and this disclosure embodiment does not impose any limitations on this.

[0103] Activity Manager: Used to manage the lifecycle of each application. Applications typically run in the operating system as Activities. For each Activity, the Activity Manager maintains a corresponding application record (ActivityRecord), which records the state of the application's Activities. The Activity Manager can use this ActivityRecord as an identifier to schedule the application's Activity processes.

[0104] WindowManagerService: Used to manage graphical user interface (GUI) resources used on the screen. Specifically, it can be used for: getting screen size, creating and destroying windows, showing and hiding windows, window layout, focus management, and input method management.

[0105] Content providers store and retrieve data, making that data accessible to applications. This data can include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc. Resource managers provide applications with various resources, such as localized strings, icons, images, layout files, video files, etc.

[0106] The application framework layer may also include a power manager, a phone window manager, an image captioning service, a wallpaper service, and an image compositor (surfaceflinger). The power manager is used to shut down unnecessary hardware components, effectively reducing power consumption. The phone window manager handles phone-related functions. The wallpaper service handles wallpaper-related content. The image compositor is a system service used to handle layer compositing. The image captioning service handles content related to image captioning tasks.

[0107] The Android Runtime comprises the core libraries and the virtual machine. The Android Runtime is responsible for scheduling and managing the Android system. The core libraries consist of two parts: one part contains the functionalities that Java calls, and the other part contains the core Android libraries. The application layer and application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0108] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0109] The Surface Manager manages the display subsystem and provides 2D and 3D layer blending for multiple applications. The Media Library supports playback and recording of various common audio and video formats, as well as still image files. The Media Library supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG. OpenGL ES is used for 3D graphics drawing, image rendering, compositing, and layer processing. SGL is a 2D graphics engine.

[0110] The kernel layer is the layer between hardware and software. The kernel layer includes at least display drivers, camera drivers, audio drivers, sensor drivers, etc., but this disclosure does not impose any limitations on these components.

[0111] The methods described in the following embodiments can all be implemented in electronic devices having the above-described hardware or software structures.

[0112] The image description method provided by this disclosure is described in detail below with reference to the accompanying drawings. This method can be applied to the aforementioned electronic device. Figure 9 This is a flowchart illustrating an image description method provided in an embodiment of this disclosure, such as... Figure 9 As shown, the method specifically includes the following steps:

[0113] Step 901: Obtain multiple images to be processed.

[0114] Taking a mobile phone as an example, the image description method provided in this disclosure will be described. For instance, the image to be processed may include static images and dynamic images. Static images may include photos taken by the user, downloaded images, and screenshots, while dynamic images may include videos taken by the user, GIFs (such as Graphics Interchange Format GIFs, live GIFs, etc.), and screen recordings.

[0115] Among the multiple images to be processed, there is at least one image to be described, and each image to be described includes an object to be described.

[0116] Step 902: Perform recognition and classification processing on the image to be processed to determine at least one album.

[0117] The process involves recognizing and grouping images of the same object into a single album, resulting in at least one album. Specifically, after acquiring at least one image, the electronic device performs recognition processing to identify objects within each image. Then, based on the similarity of objects within different images, the images are grouped together, with images of the same object grouped into the same album. For example, images containing person A are grouped into the "Person A" album, images containing person B into the "Person B" album, images containing flowers and plants into the "Flowers and Plants" album, images containing cars into the "Cars" album, and so on.

[0118] For example, when determining an object in an image to be processed, if the image is a static image, the object can be directly identified by performing recognition processing on the static image. If the image is a dynamic image, the frame to be recognized can be determined based on the dynamic image, and then the frame to be recognized can be processed to obtain the object in the image.

[0119] For example, a single frame in a dynamic image (such as the first frame, keyframe, last frame, or other specific frame) can be used as the frame to be identified and processed. The identification result of this frame is then used as the identification result of the image to be processed, and the objects in this frame are identified as the objects in the image to be processed. Alternatively, multiple frames in a dynamic image (such as the first three frames, multiple keyframes, subsequent frames, or other specific frames) can be used as frames to be identified and processed separately, with objects meeting preset conditions identified as the objects in the image to be processed. Furthermore, all frames in a dynamic image can be used as frames to be identified and processed separately, with objects meeting preset conditions identified as the objects in the image to be processed. For example, preset conditions could include those with the highest frequency of occurrence or the largest number of pixels occupied.

[0120] It should be noted that the images to be described within the same album may include not only the same object to be described, but also other objects, which can be objects to be described from other albums. In other words, the same image containing multiple objects can be categorized into multiple albums. For example, a photo of person A and person B together can be categorized into both person A's album and person B's album.

[0121] Step 903: In response to the user's operation to determine the album to be processed, obtain the album to be processed.

[0122] The interface of the electronic device gallery application allows users to view various albums and receive user actions to select albums for processing. For example, after a user clicks on an album, the images included in that album are displayed. When personalized descriptions of images in an album are needed, the user selects the album containing the image to be described as the album to be processed.

[0123] The album to be processed can be an album in a mobile photo gallery application, containing at least one image to be described. All images in the album to be described include the same object to be described. For example, the album to be processed can be a people album, where each image includes at least the target person to be described. In addition to the target person, other people may also be included in the images to be described. Figure 10 As shown in (a), the album to be processed can be the album of person A. Each image in the album includes person A. In addition to person A, some images in the album may also include other people such as person B and person C.

[0124] In some embodiments, the album to be processed can also be the album of other objects, such as the album of animal D, in which case the object to be described is animal D. This embodiment uses the example of a people's album to be processed and a people's album to be described, but this is not intended to limit the album to be processed or the object to be described.

[0125] Step 904: Perform recognition processing on each image to be described in the album to obtain the face detection box, face features and portrait detection box of each person in each image to be described.

[0126] For example, an electronic device stores a pre-trained image processing model. This model is used to determine the location of faces, facial features, and portraits of individuals in an image, obtaining face detection boxes, facial features, and portrait detection boxes for each face. When image recognition processing is required, the pre-trained image processing model is used to perform face recognition and portrait recognition processing on each image to be described in the image to be described, obtaining face detection boxes, facial features, and portrait detection boxes for each individual in each image. Figure 10 As shown in (b), the face detection box and the portrait detection box can be represented as rectangles, which respectively represent the location of the face and the location of the portrait in the image to be described.

[0127] Step 905: Determine the target face based on the face detection bounding box and face features of each person in each image to be described.

[0128] Based on the facial features of each person in each image to be described, identify the face present in each image to be described, and take it as the face of the target person, and take the target person as the owner of the album to be processed.

[0129] If the album contains images of a single person (i.e., only one face is identified in an image), then that single face is identified as the target face. If all images in the album contain images of multiple people (i.e., at least two faces are identified in all images), then the target face can be determined using the following steps:

[0130] Step 9051: Determine the first image to be described from all the images to be described, which has the fewest people.

[0131] The number of people is equal to the number of face detection boxes. The number of people in each image to be described in the album is determined, and the image with the fewest people is selected as the first image to be described from all images to be described.

[0132] It should be noted that when there are two or more images to be described that have the same and minimum number of people, one of these images can be randomly selected as the first image to be described. For example... Figure 10 As shown in (b), the number of people in the image to be described 1001 and the image to be described 1002 are equal and the minimum, and the image to be described 1001 is randomly selected as the first image to be described.

[0133] Step 9052: Calculate the distance between the facial features of each person in the first image to be described and the facial features of each person in the second image to be described.

[0134] The second image to be described is any image in the album to be described, excluding the first image. The distance between the first facial feature of the first person in the first image to be described and the second facial feature of the second person in the second image to be described represents the similarity between the faces of the first and second people, and can be calculated using the Euclidean distance formula.

[0135] like Figure 10 As shown in (b), the distances between each face in the image to be described 1001 and each face in the image to be described 1002 and each face in the image to be described 1003 are calculated respectively, and the distances between the face in the face detection box 10012 in the image to be described and each face in the image to be described 1002 and each face in the image to be described 1003 are calculated respectively.

[0136] Step 9053: Determine the similarity distance of each face in the first image to be described based on the distance between the facial features of each person in the first image to be described and the facial features of each person in the second image to be described.

[0137] Based on the distances between the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described, the minimum distance between the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described is determined. This minimum distance is taken as the distance between each face in the first image to be described and each face in the second image to be described. The sum of the distances between each face in the first image to be described and all the second images to be described is calculated to obtain the total distance corresponding to each face in the first image to be described, which is the similarity distance of each face in the first image to be described.

[0138] like Figure 10 As shown in (b), the first distance between the features of the face in the face detection box 10011 of the image to be described 1001 and the features of the face in the face detection box 10021 of the image to be described 1002, calculated in step 9052, and the second distance between the features of the face in the face detection box 10011 and the features of the face in the face detection box 10022 of the image to be described 1002 are compared. The smaller of the first and second distances is taken as the distance between the face in the face detection box 10011 and the image to be described 1002. Similarly, the distances between the face in the face detection box 10011 and the image to be described 1003, the distances between the face in the face detection box 10012 and the image to be described 1002, and the distances between the face in the face detection box 10012 and the image to be described 1003 are determined. The total distance corresponding to a face in face detection box 10011 is equal to the sum of the distance between the face in face detection box 10011 and the image to be described 1002 and the distance between the face in face detection box 10011 and the image to be described 1003. The total distance corresponding to a face in face detection box 10012 is equal to the sum of the distance between the face in face detection box 10012 and the image to be described 1002 and the distance between the face in face detection box 10012 and the image to be described 1003.

[0139] Step 9054: Determine the target face based on the total distances corresponding to each face in the first image to be described.

[0140] For example, the face with the smallest total distance among all faces in the first image to be described can be taken as the target face.

[0141] like Figure 10 As shown in (b), the face in face detection box 10011 appears in both image 1002 and image 1003 to be described, while the face in face detection box 10012 appears in image 1002 to be described but does not appear in image 1003 to be described. Therefore, the total distance corresponding to the face in face detection box 10011 is less than the total distance corresponding to the face in face detection box 10012, and the face in face detection box 10011 is taken as the target face.

[0142] Since the target face appears in every image in the album to be processed, the image with the fewest people is first selected as the first image to be described. The total distances corresponding to all faces in this first image to be described are calculated, which reduces the number of people that need to be calculated and helps to quickly find the target face. Then, the distances between the facial features of each person in the first image to be described and the facial features of each person in the other images to be described are calculated. The minimum distance is found to determine the total distance corresponding to each face. Then, the minimum total distance is determined from the total distances corresponding to each face, and the target face is found to exist in all the images to be described.

[0143] Step 906: Determine the target face's image detection box based on the face detection box of the target face and the image detection box of each person in each image to be described.

[0144] After identifying the target face, the distance between the target face's bounding box and all other face detection boxes in each image to be described is calculated. For example, the distance between the midpoint of the top border of the target face detection box and the midpoint of the top border of each face detection box can be calculated, and this distance can be used as the distance between the target face's bounding box and all other face detection boxes.

[0145] In general photo poses, the closest distance between the face detection bounding box and the image detection bounding box of the same person is the minimum. Based on this, the image detection bounding box closest to the target face is selected from all image detection bounding boxes, thus obtaining the target person's image bounding box in the image. However, for special photo poses such as lying on one's side or tilting the body, the image detection bounding box closest to the target person's face may not necessarily be the target person's image detection bounding box. Matching based on the closest distance might result in matching the face detection bounding box of person 1 to the image detection bounding box of person 2.

[0146] To address the aforementioned issue of mismatch between face detection boxes and portrait detection boxes, this embodiment considers the constraint that "the center point of the face detection box is located within the portrait detection box." Under this constraint, the portrait detection box closest to the face detection box is used as the portrait detection box for the target person. This eliminates the possibility of mismatches caused by special shooting poses, improving matching accuracy and enabling fast and accurate matching of the portrait detection box that matches the target person's face detection box. By adding the constraint that "the center point of the face detection box is located within the portrait detection box," the matching accuracy between the determined portrait detection box and the face detection box is improved, thereby ensuring that the description information described by the portrait detection box is the description information of the target person.

[0147] In one implementation, when determining the best-matching face detection bounding box for the target person using the Hungarian algorithm, the input to the Hungarian algorithm is an M×N distance matrix D:

[0148]

[0149] Where, d ij This represents the distance between the i-th face detection box and the j-th portrait detection box (e.g., the distance between the midpoints of the top borders of the two detection boxes). M is the number of face detection boxes in the image to be described, and N is the number of portrait detection boxes in the image to be described. For example, the output of the Hungarian algorithm could be such that the distance... The smallest combination of human detection bounding boxes [j1,j2,…,j n ].

[0150] The constraint matrix C corresponding to the constraint condition "the center point of the face detection box is located in the face detection box" is shown in formula (2):

[0151]

[0152] Where, when the center point of the i-th face detection bounding box is located within the j-th image detection bounding box, c ij =0; when the center point of the i-th face detection box is outside the j-th face detection box, c ij =+∞.

[0153] Based on the distance matrix D, and considering the constraint matrix C, determine the optimal distance matrix D. ′ For example, the input distance matrix D of the Hungarian algorithm can be added to the constraint matrix C to obtain the optimized distance matrix D. ′ D ′ =D+C. Replace the distance matrix D in the Hungarian algorithm with the optimized distance matrix D. ′ Perform calculations to ensure that the output human detection bounding boxes are a combination of [j1,j2,…,j...]. n The accuracy is higher.

[0154] Step 907: Perform image description based on the face detection bounding box of the target face in the image to be described, and obtain the image description information.

[0155] For example, when determining the descriptive information of an image, a pre-trained visual model can be obtained first. This visual model can be an existing model such as InternVL 2.0, which will not be described in detail in this embodiment. Then, based on the portrait detection bounding boxes of the target person in each image to be described, a segmented image of the target person in each image to be described is determined. This segmented image is the image that needs to be described, after removing other non-target persons, background, and other redundant information. Finally, the visual model is used to perform image description on the segmented images of the target person in each image to be described, thereby obtaining the descriptive information of each image to be described.

[0156] For example, the descriptive information of an image may include information such as describing the object, the object's characteristics, and the object's behavior. The object's characteristics may include features such as clothing or accessories, and the object's behavior may include information such as the object's own actions and interactive behaviors.

[0157] The image description method provided in this disclosure can be applied to various products that require image searching. By describing images using this method, personalized and specific descriptive information is generated. During image searches, the search can be performed based on this personalized and specific descriptive information, enabling quick and accurate retrieval of target images, improving search efficiency and accuracy, and enhancing the user experience.

[0158] Based on the above embodiments, this disclosure further provides an image description method. Figure 11 This is a flowchart illustrating another image description method provided in this disclosure, which is applied to electronic devices, such as mobile phones. Figure 11 As shown, the method includes the following steps:

[0159] Step 1101: Obtain the album to be processed.

[0160] The album to be processed includes at least one image to be described. For example, the album to be processed can be obtained by an electronic device categorizing multiple images to be processed, or it can be an album pre-categorized by other devices, which is not limited in this embodiment.

[0161] In one implementation, the album to be processed can be obtained through the following steps:

[0162] Step 11011: Obtain multiple images to be processed.

[0163] For example, the image to be processed may include static images and dynamic images. Static images may include photos taken by the user, downloaded images, and screenshots, while dynamic images may include videos taken by the user, GIFs (such as Graphics Interchange Format GIFs, live GIFs, etc.), and screen recordings.

[0164] Among the multiple images to be processed, there is at least one image to be described, and each image to be described includes an object to be described.

[0165] Step 11012: Perform recognition processing on each image to be processed to obtain the object corresponding to each image.

[0166] The image to be processed is then identified to determine the objects contained within it. One image can correspond to one object or multiple objects. These objects can include people, animals, plants, etc. Different objects possess different characteristics. For example, facial features can be used as the characteristics of a person object.

[0167] For example, if no object is successfully identified in the image to be processed, the image can correspond to 0 objects. Images where no objects are successfully identified can be marked as images where objects were not successfully identified.

[0168] Step 11013: Based on the objects corresponding to each image to be processed, classify and process multiple images to be processed to obtain at least one album.

[0169] Based on the identified objects, images of the same object are categorized into one album, resulting in at least one album. Multiple images of the same object are grouped into the same album. Specifically, after acquiring at least one image to be processed, the electronic device performs recognition processing to identify objects within each image. Then, based on the similarity of objects in different images, the images are categorized, with images of the same object grouped into the same album. For example, images containing person A are categorized into the album for person A, images containing person B are categorized into the album for person B, images containing flowers and plants are categorized into the album for flowers and plants, images containing cars are categorized into the album for cars, and so on.

[0170] For example, when classifying images marked as unidentified objects, images that were not successfully identified can be left unclassified or classified into an album of unidentified objects.

[0171] It should be noted that the images to be described within the same album may include not only the same object to be described, but also other objects, which can be objects to be described from other albums. In other words, the same image containing multiple objects can be categorized into multiple albums. For example, a photo of person A and person B together can be categorized into both person A's album and person B's album.

[0172] Step 11014: Determine the album to be processed based on at least one album.

[0173] For example, an electronic device can select a photo album chosen by the user as the album to be processed from at least one album based on the user's selection operation. Alternatively, the electronic device can sequentially select each album in the gallery as the album to be processed according to the naming order of the albums, and sequentially perform image descriptions on the images in each album of the gallery.

[0174] Step 1102: Perform recognition processing on each image to be described, and determine the portrait detection box of the target person in each image to be described.

[0175] In some embodiments, the target person's portrait detection bounding box can be determined through the following steps:

[0176] Step 11021: Perform recognition processing on each image to be described to determine the face detection box, face features and portrait detection box of each person in each image to be described.

[0177] For example, an electronic device stores a pre-trained image processing model. This model is used to determine the location of faces, facial features, and portraits of individuals in an image, resulting in face detection boxes, facial features, and portrait detection boxes for each face. When image recognition processing is required, the pre-trained image processing model is used to perform face recognition and portrait recognition processing on each image to be described in the image to be described, resulting in face detection boxes, facial features, and portrait detection boxes for each individual in each image.

[0178] Step 11022: Determine the target person based on the face detection bounding box and face features of each person in each image to be described.

[0179] Based on the facial features of each person in each image to be described, the person present in each image to be described is identified as the target person, and this target person is the owner of the album to be processed. In this embodiment of the disclosure, the target person refers to the object to be described corresponding to the current album to be processed.

[0180] If the album contains images of a single person (i.e., only one face is identified in an image), then the person corresponding to that single face is identified as the target person. If all images in the album contain images of multiple people (i.e., at least two faces are identified in all images), then the target person can be determined using the following steps:

[0181] Step 110221: Determine the first image to be described from all the images to be described, which has the fewest people.

[0182] The number of people is equal to the number of face detection boxes. The number of people in each image to be described in the album is determined, and the image with the fewest people is selected as the first image to be described from all images to be described.

[0183] Step 110222: Based on the facial features of each person in the first image to be described and the facial features of each person in each second image to be described (excluding the first image to be described), determine the similarity distance of each face in the first image to be described.

[0184] When determining the similarity distance between faces in the first image to be described, the distance between the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described (excluding the first image to be described) is first determined based on the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described. The second images to be described are all images to be described in the album to be processed, excluding the first image to be described. The distance between the first facial feature of the first person in the first image to be described and the second facial feature of the second person in the second image to be described represents the similarity between the faces of the first person and the second person, and can be calculated using the Euclidean distance formula.

[0185] Then, based on the distances between the facial features of each person in the first image to be described and the facial features of each person in the second image to be described, the distances between each face in the first image to be described and the distances between each face in the second image to be described are determined. Based on these distances, the minimum distance between the facial features of each person in the first image to be described and the facial features of each person in the second image to be described is determined. This minimum distance is then taken as the distance between each face in the first image to be described and the distances between each face in the second image to be described.

[0186] Finally, based on the distances between each face in the first image to be described and each of the second images to be described, the similarity distance of each face in the first image to be described is determined. The sum of the distances between each face in the first image to be described and all the second images to be described is calculated to obtain the similarity distance corresponding to each face in the first image to be described.

[0187] Step 110223: Determine the target person based on the similarity distance of each face in the first image to be described.

[0188] For example, the person whose face has the smallest total distance among all faces in the first image to be described can be identified as the target person.

[0189] Since the target person appears in every image in the album to be processed, the image with the fewest people is first selected as the first image to be described. The total distances corresponding to all faces in this first image to be described are calculated, which reduces the number of people that need to be calculated and helps to quickly find the target person. Then, the distances between the facial features of each person in the first image to be described and the facial features of each person in the other images to be described are calculated. The minimum distance is found to determine the total distance corresponding to each face. Then, the minimum total distance is determined from the total distances corresponding to each face, thus identifying the target person who exists in all the images to be described.

[0190] Step 11023: Determine the image detection box of the target person in each image to be described based on the face detection box and the portrait detection box of each person in each image to be described.

[0191] For example, when determining the bounding box of a target person in each image to be described, candidate bounding boxes for the target person can first be determined based on the face bounding box of the target person in each image to be described and the bounding boxes of each person. After the target person is determined, the bounding box of the target person in each image to be described is determined based on the positional relationship between the face bounding box of the target person in each image to be described and the bounding boxes of each person. For example, the positional relationship between the center point of the face bounding box of the target person in each image to be described and all the bounding boxes can be calculated, and candidate bounding boxes can be selected from all the bounding boxes. The candidate bounding box includes the center point of the target person's face bounding box, meaning the target person's face bounding box is located within the candidate bounding box.

[0192] Then, the distance between the face detection bounding box of the target person and the candidate image detection bounding boxes in each image to be described is determined. For example, the distance between the midpoint of the upper border of the face detection bounding box and the midpoint of the upper border of each candidate image detection bounding box can be calculated, and this distance can be used as the distance between the face detection bounding box of the target person and each candidate image detection bounding box.

[0193] Finally, based on the distance between the target person's face detection bounding box and the candidate face detection bounding boxes in each image to be described, the face detection bounding box of the target person in each image to be described is determined. For example, the face detection bounding box with the smallest distance between the target person's face detection bounding box and the candidate face detection bounding boxes in each image to be described can be selected as the face detection bounding box of the target person.

[0194] In this embodiment, the constraint "the center point of the face detection box is located within the portrait detection box" is considered. Under this constraint, the portrait detection box closest to the face detection box is used as the portrait detection box of the target person. This eliminates the possibility of mismatch between the face detection box and the portrait detection box caused by special shooting postures, thereby improving matching accuracy and enabling fast and accurate matching of the portrait detection box that matches the face detection box of the target person. By adding the constraint "the center point of the face detection box is located within the portrait detection box," the matching accuracy between the determined portrait detection box and the face detection box can be improved, thereby ensuring that the description information described by the portrait detection box is the description information of the target person.

[0195] In some embodiments, when the center point of the target person in the image to be described is outside all the image detection boxes, i.e. there are no candidate image detection boxes, the distance between the midpoint of the upper border of the target person's face detection box and the midpoint of the upper border of each candidate image detection box can be calculated, and the image detection box with the smallest distance between the target person's face detection box and the candidate image detection boxes in each image to be described can be selected as the target person's image detection box.

[0196] Step 1103: Perform image description based on the human figure detection bounding box of the target person in each image to be described, and obtain the description information of each image to be described.

[0197] In some embodiments, the descriptive information of an image can be determined according to the following steps:

[0198] Step 11031: Obtain a pre-trained visual model for describing the image.

[0199] The visual model can be a pre-trained model used to describe image information. For example, the visual model can be an InternVL 2.0 model.

[0200] Step 11032: Based on the image detection bounding box of the target person in each image to be described, determine the segmented image of the target person in each image to be described.

[0201] Step 11033: Use a visual model to perform image description on the segmented images of the target person in each image to be described, and obtain the description information of each image to be described.

[0202] The image search method provided in this disclosure describes images using image description methods, generating personalized and specific descriptive information for the images. When performing image searches, the search can be based on this personalized and specific descriptive information, enabling quick and accurate retrieval of target images, improving search efficiency and accuracy, and enhancing the user experience.

[0203] In some embodiments, image search can be further performed based on the above embodiments, and the image description method may further include... Figure 12 Steps 1201 to 1204 are shown below:

[0204] Step 1201: In response to user operation, display the image search interface.

[0205] The image search interface includes a search box. The image search interface can be like... Figure 1 As shown in (a) in the figure.

[0206] Step 1202: Receive the search information entered by the user in the search box.

[0207] For example, a user can enter natural language in a search box, and the electronic device performs natural language processing on the user's input to obtain search information, which is then used to search for the target image.

[0208] Step 1203: Based on the search information and the description information of each image to be described, determine the target image among the images to be described.

[0209] Based on the search information, a search and matching process is performed on the image description information. The description information with the highest matching degree is taken as the target description information, and the image corresponding to the target description information is determined as the target image.

[0210] Step 1204: Display the target image in the image search interface.

[0211] In this embodiment of the disclosure, an image description method is used to describe the image, generating personalized and specific descriptive information. When performing an image search, the user enters the search information for the target image they wish to search for in the search box. The search is then performed based on this search information and the personalized and specific descriptive information, enabling the rapid identification of the accurate target image. This improves search efficiency and accuracy, enhancing the user experience.

[0212] It should be understood that the steps in the above-described method embodiments provided in this disclosure can be implemented by integrated logic circuits in the processor hardware or by instructions in software form. The method steps disclosed in conjunction with the embodiments of this disclosure can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.

[0213] In one example, the unit in the above device may be one or more integrated circuits configured to implement the above methods, such as one or more ASICs, or one or more DSPs, or one or more FPGAs, or a combination of at least two of these integrated circuit forms.

[0214] For example, when the units in the device can be implemented through a processing element scheduler, the processing element can be a general-purpose processor, such as a CPU or other processor capable of calling programs. Alternatively, these units can be integrated together to form a system-on-a-chip (SoC).

[0215] In one implementation, the units that implement the corresponding steps in the above methods can be implemented in the form of a processing element scheduler. For example, the device may include a processing element and a storage element, wherein the processing element calls a program stored in the storage element to execute the methods of the above method embodiments. The storage element may be a storage element located on the same chip as the processing element, i.e., an on-chip storage element.

[0216] In another implementation, the program used to perform the above methods can be located on a storage element on a different chip than the processing element, i.e., an off-chip storage element. In this case, the processing element calls or loads the program from the off-chip storage element onto the on-chip storage element to call and execute the methods of the above method embodiments.

[0217] For example, embodiments of this disclosure may also provide an apparatus, such as an electronic device, which may include a processor and a memory for storing processor-executable instructions. When the processor is configured to execute the aforementioned instructions, it causes the electronic device to implement the image description method as described in the foregoing embodiments. The memory may be located within or outside the electronic device, and the processor may include one or more processors.

[0218] In another implementation, the unit implementing each step of the above method can be configured as one or more processing elements, which can be disposed on the corresponding electronic device described above. These processing elements can be integrated circuits, such as one or more ASICs, one or more DSPs, one or more FPGAs, or combinations of these types of integrated circuits. These integrated circuits can be integrated together to form a chip.

[0219] For example, embodiments of this disclosure also provide a chip, such as Figure 13 As shown, the chip system includes at least one processor 1301 and at least one interface circuit 1302. The processor 1301 and the interface circuit 1302 are interconnected via lines. For example, the interface circuit 1302 can be used to receive signals from other devices. As another example, the interface circuit 1302 can be used to send signals to other devices (e.g., the processor 1301).

[0220] For example, interface circuit 1302 can read instructions stored in the device's memory and send those instructions to processor 1301. When the instructions are executed by processor 1301, they enable electronic devices (such as...) Figure 7The electronic device 700 shown performs the steps in the above embodiments. Of course, the chip system may also include other discrete devices, and this disclosure does not specifically limit this.

[0221] This disclosure also provides a computer-readable storage medium storing computer program instructions thereon. When the computer program instructions are executed by an electronic device, the electronic device can implement the image description method described above.

[0222] This disclosure also provides a computer program product, including computer instructions for operation of the electronic device described above. When the computer instructions are executed in the electronic device, the electronic device enables the image description method described above. Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0223] In the several embodiments provided in this disclosure, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0224] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0225] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0226] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product, such as a program. This software product is stored in a program product, such as a computer-readable storage medium, and includes several instructions to cause a terminal device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0227] For example, embodiments of this disclosure may also provide a computer-readable storage medium storing computer program instructions thereon. When the computer program instructions are executed by an electronic device, the electronic device causes the electronic device to implement the image description method as described in the foregoing method embodiments.

[0228] The above are merely specific embodiments of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any changes or substitutions within the technical scope disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. An image description method, characterized in that, Applied to electronic devices, the method includes: Obtain the album to be processed, which includes at least one image to be described; The images to be described are processed for recognition, and the image detection box of the target person in each image to be described is determined. Image description is performed based on the human figure detection bounding box of the target person in each of the images to be described, and description information of each image to be described is obtained.

2. The method according to claim 1, characterized in that, The process of obtaining the album to be processed includes: Acquire multiple images to be processed; The images to be processed are identified to obtain the objects corresponding to each image to be processed. Based on the objects corresponding to each of the images to be processed, the multiple images to be processed are classified to obtain at least one album; wherein, multiple images to be processed corresponding to the same object are classified into the same album; Based on the at least one album, determine the album to be processed.

3. The method according to claim 1, characterized in that, The step of performing recognition processing on each of the images to be described, and determining the portrait detection box of the target person in each of the images to be described, includes: The images to be described are processed for recognition to determine the face detection box, face features and portrait detection box of each person in each image to be described; The target person is determined based on the face detection bounding box and face features of each person in the image to be described. Based on the face detection bounding box and the portrait detection bounding box of the target person in each of the images to be described, the portrait detection bounding box of the target person in each of the images to be described is determined.

4. The method according to claim 3, characterized in that, The step of determining the target person based on the face detection bounding boxes and facial features of each person in each of the images to be described includes: Determine the first image to be described from among the images to be described, which has the fewest people. Based on the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described (excluding the first image to be described), the similarity distance of each face in the first image to be described is determined. The target person is determined based on the similarity distance of each face in the first image to be described.

5. The method according to claim 4, characterized in that, The step of determining the similarity distance of each face in the first image to be described based on the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described (excluding the first image to be described) includes: Based on the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described (excluding the first image to be described), the distance between the facial features of each person in the first image to be described and the facial features of each person in each of the second images to be described is determined. The distance between each face in the first image to be described and each face in the second image to be described is determined based on the distance between the facial features of each person in the first image to be described and the facial features of each person in the second image to be described. The similarity distance of each face in the first image to be described is determined based on the distance between each face in the first image to be described and each of the second images to be described.

6. The method according to claim 3, characterized in that, The step of determining the image detection box of the target person in each of the images to be described based on the face detection box and the image detection box of each person in each image to be described includes: Based on the face detection bounding box and the portrait detection bounding box of the target person in each of the images to be described, a candidate portrait detection bounding box of the target person in each of the images to be described is determined; wherein, the face detection bounding box of the target person is located within the candidate portrait detection bounding box; Determine the distance between the face detection box of the target person and the candidate image detection box in each of the images to be described; The face detection box of the target person in each of the images to be described is determined based on the distance between the face detection box of the target person and the candidate face detection box in each of the images to be described.

7. The method according to claim 1, characterized in that, The step of performing image description based on the portrait detection bounding box of the target person in each of the images to be described, to obtain the description information of each image to be described, includes: Obtain a pre-trained visual model for describing the image; Based on the portrait detection bounding box of the target person in each of the images to be described, a segmented image of the target person in each of the images to be described is determined; The visual model is used to perform image description on the segmented images of the target person in each of the images to be described, thereby obtaining the description information of each of the images to be described.

8. The method according to claim 1, characterized in that, The method further includes: In response to user operation, an image search interface is displayed, which includes a search box; Receive search information entered by the user in the search box, the search information being used to search for a target image; Based on the search information and the description information of each of the images to be described, the target image is determined among the images to be described; The target image is displayed in the image search interface.

9. An electronic device, characterized in that, include: Processor, memory, bus, and communication interface; The memory is used to store computer execution instructions. The processor is connected to the memory via the bus. When the electronic device is running, the processor executes the computer execution instructions stored in the memory to cause the electronic device to perform the image description method as described in any one of claims 1-8.

10. A computer-readable storage medium storing instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the image description method as described in any one of claims 1-8.