Image acquisition method and device, electronic equipment, storage medium and program product
By performing semantic parsing and cross-modal matching on the descriptive information, the problem of low efficiency in photo search is solved, and fast and accurate image acquisition and album generation are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NEW H3C INTELLIGENCE TERMINAL CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-10
AI Technical Summary
Current technologies have low efficiency in photo searching, requiring users to manually categorize or search for photos based on shooting time, resulting in low search efficiency.
By performing semantic parsing on the descriptive information to obtain image metadata and semantic information, and then performing matching based on metadata and semantic vectors, cross-modal semantic matching is achieved, improving search efficiency and accuracy.
It enables quick and accurate retrieval of target images based on descriptive information, improving the efficiency and accuracy of photo searching and simplifying the process of image classification and album generation.
Smart Images

Figure CN121833987A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer processing, and particularly relates to an image acquisition method and device, electronic equipment, a storage medium and a program product. BACKGROUND
[0002] At present, with the popularity of smart phones, digital cameras and other devices, the acquisition of photos is more and more convenient, and the accumulation of photos in a single device is also more and more.
[0003] In the related art, a user manually classifies a plurality of photos stored in a device in advance, and when a specific photo is searched for, the storage location of the photo is determined through the previous manual classification, and then the specific photo is acquired. If the user has not previously manually classified the photos, the storage location of the photos can be determined according to the shooting time of the photos, and then the specific photo is acquired.
[0004] However, in the above related art, the photo searching efficiency is low. SUMMARY
[0005] The present application provides an image acquisition method and device, electronic equipment, a storage medium and a program product to solve the problem of low photo searching efficiency.
[0006] In a first aspect, the present application provides an image acquisition method, which comprises: preprocessing a candidate image to obtain metadata and a semantic vector of the candidate image; receiving an image acquisition instruction, wherein the image acquisition instruction comprises description information used for image acquisition; performing semantic analysis on the description information to obtain image metadata and image semantic information; matching the image metadata with metadata of each candidate image to obtain a first image set; matching the image semantic information with a semantic vector of each candidate image to obtain a second image set; performing image screening based on the intersection of the first image set and the second image set to obtain a target image matched with the description information.
[0007] In a second aspect, the present application provides an image acquisition device, which comprises: an image processing module configured to preprocess a candidate image to obtain metadata and a semantic vector of the candidate image; an instruction receiving module configured to receive an image acquisition instruction, wherein the image acquisition instruction comprises description information used for image acquisition; a description analyzing module configured to perform semantic analysis on the description information to obtain image metadata and image semantic information; The data matching module is used to match the image metadata and the metadata of each candidate image to obtain a first image set; A semantic matching module is used to match the semantic information of the image with the semantic vectors of each candidate image to obtain a second image set; The image acquisition module is used to perform image filtering based on the intersection of the first image set and the second image set to obtain a target image that matches the description information.
[0008] Thirdly, this application provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the image acquisition method described in the first aspect or any corresponding embodiment.
[0009] Fourthly, this application provides a computer-readable storage medium storing computer instructions for causing a computer to perform the image acquisition method described in the first aspect or any corresponding embodiment thereof.
[0010] Fifthly, this application provides a computer program product, including computer instructions for causing a computer to execute the image acquisition method described in the first aspect or any corresponding embodiment thereof.
[0011] The image acquisition method provided in this embodiment obtains image metadata and image semantic information by semantically parsing descriptive information. Then, it performs matching based on the image metadata and image semantic information to obtain a target image that matches the descriptive information, achieving cross-modal semantic matching. By matching the descriptive information with candidate images, it achieves the search from descriptive information to target images, improving the efficiency of image retrieval. Furthermore, considering the characteristic that descriptive information contains a lot of information, it parses the descriptive information into image metadata and image semantic information. The image metadata is matched with the metadata of the candidate images, and the image semantic information is matched with the semantic vector of the candidate images. This helps to improve the matching degree between descriptive information and target images, thereby improving the overall accuracy of the image acquisition method that obtains target images through descriptive information. Attached Figure Description To more clearly illustrate the technical solutions in the specific embodiments or related technologies of this application, the drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram illustrating an application scenario according to an embodiment of this application; Figure 2 An illustrative diagram of an application scenario is shown below; Figure 3 This is a schematic flowchart of a first method for acquiring an image according to an embodiment of this application; Figure 4 An exemplary schematic diagram of an image acquisition method is shown; Figure 5 This is a schematic diagram of a second process for an image acquisition method according to an embodiment of this application; Figure 6 This is a schematic diagram of a third process of an image acquisition method according to an embodiment of this application; Figure 7 An exemplary diagram illustrates a method for generating photo albums based on descriptive information; Figure 8 An exemplary schematic diagram of an image preprocessing method is shown; Figure 9 This is a schematic diagram of the fourth process of the image acquisition method according to the embodiments of this application; Figure 10 This is a structural block diagram of an image acquisition device according to an embodiment of this application; Figure 11 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] It is understood that before using the technical solutions disclosed in the various embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0015] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0016] As one optional application scenario in the embodiments of this application, such as Figure 1 As shown, the image acquisition system may include at least one terminal device and at least one server. Figure 1 The system is illustrated in the example, which includes a computer 101, a mobile terminal 102, and a server 103, and the terminal devices such as the computer 101 and the mobile terminal 102 are connected to the server 103 through a network 110.
[0017] Specifically, the terminal device can be a smartphone, tablet, laptop, PDA, desktop computer, game console, smart TV, smart wearable device, in-vehicle terminal, VR (Virtual Reality) device, AR (Augmented Reality) device, etc. Server 103 can be a standalone physical server, a server cluster, a distributed system, or a cloud server providing cloud services. Network 110 can be a wired or wireless network, examples of which include, but are not limited to, the Internet, corporate intranet, local area network, wide area network, mobile communication network, and combinations thereof.
[0018] For example, in this embodiment of the application, an application is installed on the terminal device. This application can be any application with image search functionality. Optionally, the application can be an application that requires downloading and installation, or it can be a mini-program that can be used instantly; this embodiment of the application does not limit this.
[0019] For example, the server is the backend server of the aforementioned application. Figure 2 As shown, the application displays an image display interface, which has descriptive information acquisition and image display functions; the server has image preprocessing services, image search services, and automatic album generation services, and information processing models, semantic extraction models, face recognition technology, semantic databases, and relational databases provide underlying support for the various services in the server.
[0020] For image preprocessing services, after the server obtains the candidate images uploaded or scanned by the user through the application, it uses a semantic extraction model to extract semantic vectors from the candidate images and then uses face recognition technology to obtain face tags for the candidate images. Further, the semantic vectors are stored in a semantic database, and metadata is stored in a relational database. The metadata includes face tags, the acquisition time of the candidate images, and the acquisition location of the candidate images.
[0021] For image search services, after the server obtains the descriptive information for image acquisition through the application, it performs semantic parsing of the descriptive information using an information processing model to obtain image metadata and image semantic information. It then matches the image metadata with the metadata of each candidate image using stored data in a relational database to obtain a first image set. Finally, it matches the image semantic information with the semantic vectors of each candidate image using stored data in a semantic database to obtain a second image set. Furthermore, it performs image filtering based on the intersection of the first and second image sets to obtain target images that match the descriptive information.
[0022] For the automatic album generation service, the semantic vectors of each image to be processed are obtained through the semantic extraction model or the stored data in the semantic database. Then, based on the semantic vectors of each image to be processed, the images to be processed are clustered to obtain one or more clustered image sets. After that, a clustered image set is used as a second album, and the second album theme of each second album is generated through the information processing model.
[0023] For example, the information processing model is a pre-trained MLLM (Multimodal Large Language Model), and the semantic extraction model is a pre-trained CLIP (Contrastive Language–Image Pretraining). For example, the face recognition technology can be a pre-trained face recognition model or a pre-written face recognition program.
[0024] It should be noted that the above description of the terminal device and server is merely illustrative and explanatory. In practical applications, the terminal device and server can be flexibly configured and adjusted. For example, if the terminal device's computing power is sufficient, the image preprocessing service, image search service, and automatic album generation service can be executed directly by the terminal device's processor; or, to reduce the storage pressure on the terminal device, while the terminal device's processor executes the image preprocessing service, image search service, and automatic album generation service, the semantic database and relational database can be stored on a separate server; and so on.
[0025] According to an embodiment of this application, an image acquisition method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0026] This embodiment provides an image acquisition method that can be used in the aforementioned terminal device or server (hereinafter referred to as electronic device). Figure 3 This is a flowchart of an image acquisition method according to an embodiment of this application, such as... Figure 3 As shown, the process includes the following steps: Step S301: Preprocess the image to be selected to obtain the metadata and semantic vector of the image to be selected.
[0027] The candidate image refers to an image pre-stored in an electronic device. In the embodiments of this application, before acquiring the image, the electronic device preprocesses the candidate image to obtain the metadata and semantic vector of the candidate image.
[0028] For example, the metadata of the candidate image includes the face tag to which the candidate image belongs, the acquisition time of the candidate image, and the acquisition location of the candidate image. The acquisition time of the candidate image refers to the generation time of the candidate image; the acquisition location of the candidate image refers to the location of the acquisition device when the candidate image was generated. For example, if the candidate image is a photograph, then the acquisition time of the candidate image is the time the photograph was taken, and the acquisition location of the candidate image is the camera's positioning information at the time the photograph was taken.
[0029] Step S302: Receive image acquisition instruction.
[0030] In this embodiment of the application, when acquiring an image, the electronic device receives an image acquisition instruction. The image acquisition instruction includes descriptive information for image acquisition.
[0031] Optionally, the description information is text information.
[0032] In one possible implementation, the electronic device obtains descriptive information through text input. For example, the application obtains text input information through a user's text input operation and identifies this text input information as descriptive information; correspondingly, the electronic device obtains this descriptive information. Here, "user" refers to the user of the electronic device.
[0033] In another possible implementation, the electronic device obtains descriptive information through voice input. For example, the application described above obtains voice input information through a user's voice input operation; correspondingly, the electronic device obtains this voice input information and converts it into text information to obtain descriptive information.
[0034] Optionally, once the electronic device determines that the description information input is complete, it generates an image acquisition instruction based on the description information.
[0035] Step S303: Perform semantic parsing on the description information to obtain image metadata and image semantic information.
[0036] In this embodiment, after obtaining the image acquisition instruction, the electronic device performs semantic parsing on the descriptive information contained in the image acquisition instruction to obtain image metadata and image semantic information. The image metadata is used to characterize the metadata of the target image described by the descriptive information, and the image semantic information is used to characterize the semantic vector contained in the target image described by the descriptive information.
[0037] For example, the metadata of the target image includes the face tag to which the target image belongs, the acquisition time of the target image, and the acquisition location of the target image. The acquisition time of the target image refers to the generation time of the target image; the acquisition location of the target image refers to the location of the device acquiring the target image at the time of its generation. For example, if the target image is a photograph, then the acquisition time of the target image is the time the photograph was taken, and the acquisition location of the target image is the camera's positioning information at the time the photograph was taken.
[0038] For example, such as Figure 4 As shown, the electronic device performs semantic parsing of the descriptive information using the aforementioned information processing model. The electronic device inputs the semantic parsing instructions and descriptive information into the information processing model, thereby obtaining the image metadata and image semantic information output by the model. The semantic parsing instructions are used to control the information processing model to perform semantic parsing of the descriptive information.
[0039] Step S304: Match the image metadata with the metadata of each candidate image to obtain the first image set.
[0040] In this embodiment of the application, after obtaining the aforementioned image metadata, the electronic device matches the image information with the metadata of each candidate image to obtain a first image set. The candidate images refer to images pre-stored in the electronic device, and the first image set includes at least one candidate image. Optionally, the candidate images included in the first image set are referred to as first images.
[0041] For example, such as Figure 4 As shown, the electronic device matches image metadata with the metadata of each candidate image using the aforementioned relational database. The relational database stores the metadata of each candidate image.
[0042] Step S305: Match the semantic information of the images with the semantic vectors of each candidate image to obtain the second image set.
[0043] In this embodiment of the application, after obtaining the aforementioned image semantic information, the electronic device matches the image semantic information with the semantic vectors of each candidate image to obtain a second image set. The second image set includes at least one candidate image. Optionally, the candidate images included in the second image set are referred to as second images.
[0044] For example, such as Figure 4 As shown, the electronic device matches the semantic information of the image with the semantic vectors of each candidate image using the aforementioned semantic database. The semantic database stores the semantic vectors of each candidate image.
[0045] Step S306: Based on the intersection of the first image set and the second image set, perform image filtering to obtain target images that match the description information.
[0046] In this embodiment of the application, after obtaining the first image set and the second image set, the electronic device performs image filtering based on the intersection of the first image set and the second image set to obtain a target image that matches the above-described information.
[0047] The image acquisition method provided in this embodiment obtains image metadata and image semantic information by semantically parsing the descriptive information. Then, it performs matching based on the image metadata and image semantic information to obtain a target image that matches the descriptive information, realizing cross-modal semantic matching. By matching the descriptive information with the candidate image, it achieves the search from the descriptive information to the target image, improving the efficiency of image search. Moreover, considering that the descriptive information contains a lot of information, it parses the descriptive information into image metadata and image semantic information. The image metadata is matched with the metadata of the candidate image, and the image semantic information is matched with the semantic vector of the candidate image. This helps to improve the matching degree between the descriptive information and the target image, thereby improving the accuracy of the overall image acquisition method of obtaining the target image through the descriptive information.
[0048] This embodiment provides an image acquisition method, which can be used in the aforementioned terminal device or server. Figure 5 This is a flowchart of an image acquisition method according to an embodiment of this application, such as... Figure 5 As shown, the process includes the following steps: Step S501: Preprocess the image to be selected to obtain its metadata and semantic vector. For details, please refer to [link to relevant documentation]. Figure 3 Step S301 of the illustrated embodiment will not be described again here.
[0049] Step S502: Receive image acquisition command. See details below. Figure 3 Step S302 of the illustrated embodiment will not be described again here.
[0050] Step S503: Semantic parsing is performed on the description information to obtain image metadata and image semantic information. For details, please refer to [link to relevant documentation]. Figure 3 Step S303 of the illustrated embodiment will not be described again here.
[0051] Step S504: Match the image metadata with the metadata of each candidate image to obtain the first image set. For details, please refer to [link to relevant documentation]. Figure 3 Step S304 of the illustrated embodiment will not be described again here.
[0052] Step S505 involves matching the semantic information of the images with the semantic vectors of each candidate image to obtain the second image set. For details, please refer to [link to relevant documentation]. Figure 3 Step S305 of the illustrated embodiment will not be described again here.
[0053] Step S506: Based on the intersection of the first image set and the second image set, image filtering is performed to obtain the target image that matches the description information.
[0054] Specifically, step S506 includes: Step S5061: Perform intersection processing on the first image set and the second image set to obtain the third image set.
[0055] In this embodiment of the application, after acquiring the first image set and the second image set, the electronic device performs intersection processing on the first image set and the second image set to obtain a third image set. The third image set includes at least one candidate image. Optionally, the candidate image included in the third image set is referred to as the third image.
[0056] For example, the electronic device performs intersection analysis on the first image set and the second image set to obtain an intermediate image; further, it performs similar image filtering on the intermediate image set to remove similar images from the intermediate image set, resulting in a third image set. The intermediate image set includes at least one candidate image. Optionally, the candidate images included in the intermediate image set are referred to as intermediate images. For example, a similar image refers to an image whose overlap is greater than a similarity threshold. This similarity threshold can be any value and can be flexibly set and adjusted according to actual conditions; this embodiment does not limit this.
[0057] Step S5062: Determine the score of each third image based on the degree of matching between each third image in the third image set and the description information.
[0058] In this embodiment, after acquiring the aforementioned third image set, the electronic device determines a score for each third image based on the degree of matching between each third image in the third image set and the descriptive information. For example, the degree of matching between the third image and the descriptive information is used to characterize the degree of matching between the third image and the input intent of the descriptive information.
[0059] For example, the electronic device scores the third image based on the degree of matching between the third image and the descriptive information using the aforementioned information processing model. The electronic device inputs a scoring instruction, scoring indicators, the third image, and the aforementioned descriptive information into the information processing model to obtain the score of the third image output by the information processing model. The scoring instruction controls the information processing model to score the third image, and the scoring indicators include whether the image meets the requirements described in the descriptive information.
[0060] Optionally, in this embodiment of the application, in order to improve the accuracy of the scoring, the electronic device determines the score of each third image based on the degree of matching between each third image and the description information, as well as the image quality of each third image.
[0061] For example, image quality includes image sharpness and the rationality of image composition. For example, the above scoring indicators also include whether the image is sharp and whether the image composition is reasonable.
[0062] Step S5063: The third image with a score greater than the score threshold is determined as the fourth image.
[0063] In this embodiment of the application, after obtaining the score of the third image, the electronic device determines the third image with a score greater than the score threshold as the fourth image.
[0064] For example, the scoring threshold can be any value, and the scoring threshold can be flexibly set and adjusted according to the actual situation. This application embodiment does not limit this.
[0065] Optionally, when the number of fourth images is greater than one, the set of multiple fourth images is called the fourth image set.
[0066] Step S5064: If the number of fourth images is greater than the number threshold, the fourth image is determined as the target image.
[0067] In this embodiment of the application, after acquiring the fourth image, the electronic device determines the fourth image as the target image if the number of fourth images is greater than or equal to the number threshold.
[0068] For example, the quantity threshold can be any value, and can be flexibly set and adjusted according to the actual situation, such as 10, 50, 100, etc., and this application embodiment does not limit it in this way. It should be noted that there is a special case where the quantity threshold is 1, that is, as long as the fourth image exists, the fourth image is determined as the target image.
[0069] Step S5065: If the number of fourth images is less than the number threshold, obtain the target image from the candidate images based on the keywords of the fourth images.
[0070] In this embodiment of the application, after acquiring the fourth image, if the number of fourth images is less than a threshold, the electronic device acquires the target image from the candidate images based on the keywords of the fourth image.
[0071] Specifically, step S5065 includes: Step a1: The n fourth images with the highest scores are selected as the fifth images.
[0072] For example, in an embodiment of this application, when the number of fourth images is less than a threshold, the electronic device determines the n highest-scoring fourth images as the fifth images. Here, n is a positive integer.
[0073] For example, n can be any value. The value of n can be flexibly set and adjusted according to the actual situation, such as 3, 5, 6, etc. This application embodiment does not limit this.
[0074] Step a2: Extract keywords from the fifth image to obtain the keywords of the fifth image.
[0075] For example, in this embodiment of the application, after acquiring the fifth image, the electronic device extracts keywords from the fifth image to obtain the keywords of the fifth image. Optionally, when the number of fifth images is greater than one, the set of multiple fifth images is called the fifth image set.
[0076] For example, the electronic device extracts keywords from the fifth image using the aforementioned information processing model. The electronic device inputs the keyword extraction instruction and the fifth image into the information processing model to obtain the keywords of the fifth image output by the information processing model.
[0077] Step a3: Match the keywords of each fifth image with the semantic vectors of each candidate image to obtain the sixth image set.
[0078] For example, in this embodiment of the application, after obtaining the keywords of the fifth image, the electronic device matches the keywords of each fifth image with the semantic vectors of each candidate image to obtain a sixth image set. The sixth image set includes at least one candidate image. Optionally, the candidate images included in the sixth image set are referred to as the sixth image.
[0079] Step a4: Based on the intersection of the first image set and the sixth image set, perform image filtering to obtain the target image.
[0080] For example, in this embodiment of the application, after obtaining the sixth image set, the electronic device performs image filtering based on the intersection of the first image set and the sixth image set to obtain the target image. The intersection of the first image set and the sixth image set, and the image filtering, are similar to the intersection of the first image set and the second image set, as described above, and will not be repeated here.
[0081] The image acquisition method provided in this embodiment obtains a third image set by intersecting the first and second image sets. Then, target images are selected based on the matching degree between each third image and the descriptive information. After metadata matching and semantic matching, the score of the third image is determined based on the matching degree between the third image and the descriptive information. Target images are then selected based on the scores of the third images, further improving the matching degree between the descriptive information and the target images, thereby improving the overall accuracy of the image acquisition method. Furthermore, if the number of fourth images obtained after metadata matching, semantic matching, and matching degree selection is less than a threshold, the matching range is further expanded based on the determination that the fourth image matches the descriptive information. The target images are obtained by matching keywords of the fourth image, ensuring that a sufficient number of target images can be obtained from the descriptive information, thereby indirectly improving the image search capability of the descriptive information.
[0082] Furthermore, determining the score of the third image based on the degree of matching between the third image and the descriptive information, combined with the image quality of the third image, helps to improve the accuracy of the score.
[0083] In addition, the n highest-rated fourth images are identified as the fifth images. The target images are obtained by matching the keywords of the fifth images. While expanding the matching range, more accurate keywords are obtained from the images with higher matching degrees. The matching is supplemented only based on the more accurate keywords. This improves the ability to search for images based on descriptive information while saving time as much as possible, which is conducive to improving the efficiency of searching for images based on descriptive information.
[0084] This embodiment provides an image acquisition method, which can be used in the aforementioned terminal device or server.Figure 6 This is a flowchart of an image acquisition method according to an embodiment of this application, such as... Figure 6 As shown, the process includes the following steps: Step S601: Preprocess the image to be selected to obtain its metadata and semantic vector. For details, please refer to [link to relevant documentation]. Figure 3 Step S301 of the illustrated embodiment will not be described again here.
[0085] Step S602: Receive image acquisition instruction. See details below. Figure 3 Step S302 of the illustrated embodiment will not be described again here.
[0086] Step S603: Perform semantic parsing on the description information to obtain image metadata and image semantic information. For details, please refer to [link to relevant documentation]. Figure 3 Step S303 of the illustrated embodiment will not be described again here.
[0087] Step S604: Match the image metadata with the metadata of each candidate image to obtain the first image set. For details, please refer to [link to relevant documentation]. Figure 3 Step S304 of the illustrated embodiment will not be described again here.
[0088] Step S605 involves matching the semantic information of the images with the semantic vectors of each candidate image to obtain the second image set. For details, please refer to [link to relevant documentation]. Figure 3 Step S305 of the illustrated embodiment will not be described again here.
[0089] Step S606: Image filtering is performed based on the intersection of the first image set and the second image set to obtain target images that match the description information. For details, please refer to [link to details]. Figure 5 Step S506 of the illustrated embodiment will not be described again here.
[0090] Step S607: Based on the target image, generate text information for the target image.
[0091] In this embodiment of the application, after acquiring the target image, the electronic device generates text information about the target image based on the target image. This text information describes the content contained in the target image.
[0092] For example, the electronic device generates text information of the target image using the aforementioned information processing model. The electronic device inputs a description instruction and the target image into the information processing model, and obtains the text information of the target image output by the information processing model. The description instruction is used to control the information processing model to generate the text information of the target image.
[0093] Optionally, in this embodiment of the application, the number of target images is multiple.
[0094] Step S608: Combine multiple target images into a first album, and generate the first album theme based on the text information of each target image.
[0095] In this embodiment of the application, after acquiring the text information of the target images, the electronic device treats multiple target images as a first album and generates a first album theme based on the text information of each target image.
[0096] For example, the electronic device generates a first album theme using the aforementioned information processing model. The electronic device inputs the theme generation instruction and the text information of each target image into the information processing model to obtain the first album theme output by the information processing model.
[0097] The image acquisition method provided in this embodiment generates a first album theme shared by different target images through the text information of the target images. That is, this application provides a method for generating albums based on description information. The image is automatically and quickly classified through description information, and multiple target images are combined into a first album to generate a first album theme. The operation is simple, the image classification efficiency is high, and the generation efficiency of the first album is improved.
[0098] For example, in conjunction with the reference Figure 7 This paper provides a complete overview of the image acquisition method of this application from the perspective of generating a first album based on descriptive information. Specifically, it includes: Step S701: Preprocess the image to be selected to obtain the metadata and semantic vector of the image to be selected.
[0099] Step S702: Receive an image acquisition instruction, which includes descriptive information for image acquisition.
[0100] Step S703: Perform semantic parsing on the description information to obtain image metadata and image semantic information.
[0101] Step S704: Based on image metadata and image semantic information, a third image set is obtained by matching from the candidate images.
[0102] Step S705: Determine the score of each third image based on the degree of matching between each third image and the description information, as well as the image quality of each third image.
[0103] Step S706: The third image with a score greater than the score threshold is determined as the fourth image.
[0104] Step S707: Determine whether the number of fourth images is greater than or equal to the number threshold. If the number of fourth images is greater than or equal to the number threshold, proceed to step S711; if the number of fourth images is less than the number threshold, proceed to steps S708-710.
[0105] Step S708: The n fourth images with the highest scores are determined as the fifth images, where n is a positive integer.
[0106] Step S709: Extract keywords from the fifth image to obtain the keywords of the fifth image.
[0107] Step S710: The keywords of the fifth image are used as new image semantic information, and the process is repeated starting from step S703 above.
[0108] Step S711: The fourth image is determined as the target image.
[0109] Step S712: Generate text information for the target image based on the target image.
[0110] Step S713: Based on the text information of each target image, the multiple target images are combined into a first album, and the first album theme of the first album is generated. For example, such as Figure 8 As shown, step S701 above includes the following steps: Step S801: Semantic extraction is performed on the image to be selected to obtain the semantic vector of the image to be selected.
[0111] In this embodiment of the application, after acquiring the candidate image, the electronic device performs semantic extraction on the candidate image to obtain the semantic vector of the candidate image.
[0112] For example, the electronic device performs semantic extraction on the candidate image using the aforementioned semantic extraction model. The electronic device inputs the candidate image into the semantic extraction model and obtains the semantic vector of the candidate image output by the semantic extraction model.
[0113] Step S802: Store the semantic vector and the association between the semantic vector and the candidate image.
[0114] In this embodiment of the application, after obtaining the above semantic vector, the electronic device stores the semantic vector and the association between the semantic vector and the candidate image.
[0115] For example, such as Figure 8 As shown, semantic vectors, and the relationships between semantic vectors and candidate images, are stored in the aforementioned semantic database.
[0116] Step S803: Perform face recognition on the selected image to obtain the face label of the selected image.
[0117] In this embodiment, after acquiring a candidate image, the electronic device performs face recognition on the candidate image to obtain a face tag to which the candidate image belongs. This face tag is used to characterize whether the candidate image contains a face, and the relationship between the contained face and the aforementioned user. For example, this relationship can be user-defined, such as "son," "daughter," or "friend"; or the relationship can be automatically distinguished by the electronic device, such as "relationship 1," "relationship 2," or "relationship 3," where the same relationship represents the same face.
[0118] For example, the electronic device performs face recognition on the selected image using the aforementioned face recognition technology.
[0119] Step S804: The face tag, the acquisition time of the candidate image, and the acquisition location of the candidate image are determined as the metadata of the candidate image.
[0120] In this embodiment of the application, after the electronic device acquires the face tag, it determines the face tag, the acquisition time of the candidate image, and the acquisition location of the candidate image as the metadata of the candidate image.
[0121] The acquisition time of the candidate image refers to the generation time of the candidate image. For example, if the candidate image is a photograph, then the acquisition time of the candidate image is the time when the photograph was taken.
[0122] The acquisition location of the candidate image refers to the location of the acquisition device when the candidate image is generated. For example, if the candidate image is a photograph, the acquisition location of the candidate image is the camera's positioning information when the photograph was taken.
[0123] Step S805: Store metadata and the association between the metadata and the candidate images.
[0124] In this embodiment of the application, after obtaining the aforementioned metadata, the electronic device stores the metadata and the association between the metadata and the candidate image.
[0125] For example, such as Figure 8 As shown, metadata, and the association between metadata and candidate images, are stored in the aforementioned relational database.
[0126] The image acquisition method provided in this embodiment extracts semantic vectors from the candidate image and stores them, and performs face recognition on the candidate image to store metadata. The pre-storage of semantic vectors and metadata helps to improve the subsequent matching speed, thereby improving the acquisition efficiency of the target image.
[0127] The above section introduced the method of generating albums based on descriptive information. The following section introduces the method of automatically generating albums.
[0128] This embodiment provides an image acquisition method, which can be used in the aforementioned terminal device or server. Figure 9 This is a flowchart of an image acquisition method according to an embodiment of this application, such as... Figure 9 As shown, the process includes the following steps: Step S901: Obtain more than one image to be processed from the candidate images.
[0129] In this embodiment of the application, after acquiring the candidate images, the electronic device acquires more than one image to be processed from the candidate images.
[0130] In one possible implementation, the electronic device identifies all candidate images as images to be processed.
[0131] In another possible implementation, the electronic device determines a portion of the candidate images as images to be processed. Specifically, step S901 above includes: Step S9011: Obtain browsing information for each candidate image.
[0132] In this embodiment of the application, after acquiring the candidate images, the electronic device acquires browsing information for each candidate image. The browsing information includes at least one of the following: browsing count and browsing duration.
[0133] Step S9012: Select images that meet the set conditions for browsing information as images to be processed.
[0134] In this embodiment of the application, after obtaining the above-mentioned browsing information, the electronic device determines the candidate images whose browsing information meets the set conditions as images to be processed.
[0135] For example, the set conditions include at least one of the following: the number of views is greater than a number threshold, and the viewing duration is greater than a duration threshold. The number of views threshold and the duration threshold can be any values, and can be flexibly set and adjusted according to actual circumstances. This application embodiment does not limit this.
[0136] Step S902: Based on the semantic vectors of each image to be processed, perform clustering processing on more than one image to be processed to obtain p clustered image sets.
[0137] In this embodiment, after acquiring the images to be processed, the electronic device performs clustering processing on more than one image to be processed based on the semantic vector of each image to be processed, resulting in p clustered image sets. Here, p is a positive integer, and each clustered image set includes at least one image to be processed. Optionally, the images to be processed in the clustered image sets are referred to as clustered images.
[0138] For example, the semantic vector of the image to be processed is based on the above.Figure 8 The semantic vector obtained and stored by image preprocessing in the embodiment.
[0139] For example, electronic devices cluster images to be processed based on the distance between semantic vectors.
[0140] Step S903: Take a clustered image set as a second album and generate the second album theme for each second album.
[0141] In this embodiment of the application, after obtaining the above-mentioned clustered image set, a clustered image set is used as a second album, and a second album theme is generated for each second album.
[0142] For example, the electronic device generates a second album theme using the aforementioned information processing model. The electronic device inputs a description instruction and clustered images into the information processing model to obtain the text information of the clustered images output by the information processing model; further, it inputs a theme generation instruction and the text information of each clustered image into the information processing model to obtain the second album theme output by the information processing model.
[0143] It should be noted that the above description of the automatic album generation method is only exemplary and illustrative. In practical applications, albums related to holidays can also be automatically generated based on holidays.
[0144] The image acquisition method provided in this embodiment obtains multiple clustered image sets by clustering the semantic vectors of the images to be processed, and uses one clustered image set as an album to generate a second album theme for each album. That is, this application provides an automatic album generation method, which is simple to operate and, compared with the method of manually selecting images to generate albums in related technologies, is conducive to improving the efficiency of album generation.
[0145] In addition, by using the browsing information of the candidate images to determine the images to be processed, and incorporating the browsing information when the album is automatically generated, the automatically generated album conforms to the user's browsing habits. This makes it easier for users to quickly retrieve frequently viewed images from the already generated album, thereby reducing the number of searches for the same image and thus helping to reduce the burden of searching for images.
[0146] This embodiment also provides an image acquisition device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0147] This embodiment provides an image acquisition device, such as... Figure 10As shown, it includes: Image processing module 1001 is used to preprocess the image to be selected to obtain the metadata and semantic vector of the image to be selected; The instruction receiving module 1002 is used to receive an image acquisition instruction, wherein the image acquisition instruction includes descriptive information for image acquisition; The description parsing module 1003 is used to perform semantic parsing on the description information to obtain image metadata and image semantic information; The data matching module 1004 is used to match the image metadata and the metadata of each candidate image to obtain the first image set; The semantic matching module 1005 is used to match the semantic information of the image with the semantic vectors of each candidate image to obtain the second image set; The image acquisition module 1006 is used to perform image filtering based on the intersection of the first image set and the second image set to obtain a target image that matches the description information.
[0148] In some optional implementations, the image acquisition module 1005 includes: The intersection processing unit is used to perform intersection processing on the first image set and the second image set to obtain the third image set; The image filtering unit is used to filter images from the third image set based on a preset strategy to obtain the target image.
[0149] In some optional implementations, the image filtering unit includes: The image scoring subunit is used to determine the score of each third image based on the degree of matching between each third image in the third image set and the descriptive information, as well as the image quality of each third image. The image determination subunit is used to determine the third image, which has a score greater than the score threshold, as the fourth image. The image acquisition subunit is used to determine the fourth image as the target image when the number of fourth images is greater than the number threshold. The key matching subunit is used to obtain the target image from the candidate images based on the keywords of the fourth image when the number of fourth images is less than the number threshold.
[0150] In some alternative implementations, the image key matching subunit is used for: The n highest-rated fourth images are selected as the fifth images, where n is a positive integer. Keyword extraction is performed on the fifth image to obtain the keywords of the fifth image; The keywords of each fifth image are matched with the semantic vectors of each candidate image to obtain the sixth image set; The target image is obtained by filtering images based on the intersection of the first image set and the sixth image set.
[0151] In some alternative implementations, the number of target images is multiple, and the apparatus further includes: The text generation module is used to generate text information for the target image based on the target image. The first generation module is used to treat multiple target images as a first album and generate the first album theme based on the text information of each target image. In some alternative implementations, the image processing module 1001 includes: The semantic extraction unit is used to extract semantics from the image to be selected, and obtain the semantic vector of the image to be selected. Vector storage unit, used to store semantic vectors and the relationship between semantic vectors and candidate images; The face recognition unit is used to perform face recognition on the selected image and obtain the face label of the selected image; The data acquisition unit is used to determine the acquisition time and location of the candidate image as the metadata of the candidate image; The data storage unit is used to store metadata and the association between the metadata and the candidate images.
[0152] In some alternative embodiments, the apparatus further includes: The image selection module is used to obtain more than one image to be processed from the candidate images; The image clustering module is used to cluster more than one image to be processed based on the semantic vector of each image to be processed, so as to obtain p clustered image sets; where p is a positive integer. The second generation module is used to take a clustered image set as a second album and generate the second album theme for each second album.
[0153] In some alternative implementations, the image selection module includes: The browsing acquisition unit is used to acquire browsing information for each candidate image, including at least one of browsing count and browsing duration. The image selection unit is used to identify candidate images that meet the set conditions of the browsing information as images to be processed.
[0154] The image acquisition apparatus provided in this application can execute the image acquisition method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0155] Figure 11This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0156] The following is a detailed reference. Figure 11 The diagram illustrates a structural schematic suitable for implementing the electronic device described in the embodiments of this application. The electronic device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 1101, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1102 or a program loaded from memory 1108 into random access memory (RAM) 1103. The RAM 1103 also stores various programs and data required for the operation of the electronic device. The processor 1101, ROM 1102, and RAM 1103 are interconnected via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0157] Typically, the following devices can be connected to I / O interface 1105: input devices 1106 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 1107 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 1108 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1109. Communication device 1109 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 11 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0158] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1109, or installed from memory 1108, or installed from ROM 1102. When the computer program is executed by processor 1101, it performs the functions defined in the image acquisition method of embodiments of this application.
[0159] Figure 11 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0160] This application also provides a computer-readable storage medium. The methods described in this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the image acquisition method shown in the above embodiments is implemented.
[0161] A portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0162] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and all such modifications and variations fall within the scope defined by the appended claims.
Claims
1. An image acquisition method, characterized in that, The method comprises: preprocessing the candidate image to obtain metadata and a semantic vector of the candidate image; receiving an image acquisition instruction, wherein the image acquisition instruction comprises description information for image acquisition; performing semantic analysis on the description information to obtain image metadata and image semantic information; matching the image metadata with the metadata of each candidate image to obtain a first image set; matching the image semantic information with the semantic vector of each candidate image to obtain a second image set; performing image screening based on the intersection of the first image set and the second image set to obtain a target image matching the description information.
2. The method of claim 1, wherein, The image screening based on the intersection of the first image set and the second image set to obtain a target image matching the description information comprises: performing intersection processing on the first image set and the second image set to obtain a third image set; determining the score of each third image based on the matching degree between each third image in the third image set and the description information; determining a fourth image as a third image with a score greater than a score threshold; in a case where the number of fourth images is greater than a number threshold, determining the fourth image as the target image; in a case where the number of fourth images is less than the number threshold, obtaining the target image from the candidate image based on the keyword of the fourth image.
3. The method of claim 2, wherein, The determination of the score of each third image based on the matching degree between each third image and the description information comprises: determining the score of each third image based on the matching degree between each third image and the description information and the image quality of each third image.
4. The method of claim 2, wherein, The obtaining of the target image from the candidate image based on the keyword of the fourth image comprises: determining the top n fourth images with the highest scores as fifth images, n being a positive integer; extracting the keyword of the fifth image to obtain the keyword of the fifth image; matching the keyword of each fifth image with the semantic vector of each candidate image to obtain a sixth image set; performing image screening based on the intersection of the first image set and the sixth image set to obtain the target image.
5. The method of claim 1, wherein, The number of target images is multiple, and the method further comprises: generating text information of the target image based on the target image; regarding multiple target images as a first album, and generating a first album theme of the first album based on the text information of each target image.
6. The method of claim 1, wherein, The preprocessing of the candidate image to obtain metadata and a semantic vector of the candidate image comprises: performing semantic extraction on the candidate image to obtain a semantic vector of the candidate image; storing the semantic vector and the association relationship between the semantic vector and the candidate image; performing face recognition on the candidate image to obtain a face label to which the candidate image belongs; determining the face label, the acquisition time of the candidate image, and the acquisition location of the candidate image as metadata of the candidate image; store the metadata and an association between the metadata and the candidate image.
7. The method according to any one of claims 1 to 5, characterized in that, The method further comprises: obtaining more than one to-be-processed image from the candidate image; performing clustering processing on the more than one to-be-processed image based on the semantic vector of each to-be-processed image, to obtain p clustered image sets; wherein p is a positive integer; taking one of the clustered image sets as one second album, and generating a second album theme of each second album.
8. The method of claim 7, wherein, The obtaining of the more than one to-be-processed image from the candidate image comprises: obtaining browsing information of each candidate image, the browsing information including at least one of browsing times and browsing duration; determining a candidate image whose browsing information meets a set condition as the to-be-processed image.
9. An image acquisition device, characterized in that The apparatus comprises: an image processing module configured to perform preprocessing on a candidate image, to obtain metadata and a semantic vector of the candidate image; an instruction receiving module configured to receive an image obtaining instruction, the image obtaining instruction including description information used for image obtaining; a description analyzing module configured to perform semantic analysis on the description information, to obtain image metadata and image semantic information; a data matching module configured to match the image metadata and the metadata of each candidate image, to obtain a first image set; a semantic matching module configured to match the image semantic information and the semantic vector of each candidate image, to obtain a second image set; an image obtaining module configured to perform image screening based on an intersection of the first image set and the second image set, to obtain a target image matched with the description information.
10. An electronic device, comprising: comprise: a memory and a processor, which are communicatively connected, and the memory stores computer instructions, and the processor executes the computer instructions to perform the image obtaining method in any one of claims 1 to 8.
11. A computer readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, and the computer instructions are used to make a computer execute the image obtaining method in any one of claims 1 to 8.
12. A computer program product, characterised in that, comprise computer instructions, and the computer instructions are used to make a computer execute the image obtaining method in any one of claims 1 to 8.