Image pushing method and device, electronic equipment and computer readable medium

By using convolutional neural network and semantic recognition module to automatically find matching images in the target interface, the problem of inefficient selection of images by users manually is solved, and fast and accurate image push and personalized recommendation are achieved.

CN120541250APending Publication Date: 2025-08-26GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510698472.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the prior art, users need to manually operate when selecting images, resulting in inefficiency, difficulty in ensuring matching, lack of personalized recommendations, and unable to meet the diverse needs of users.

Method used

The target content is obtained through the target interface based on the screen display, and the convolutional neural network and semantic recognition module are used to find matching images in the set of selected images, and display them on the screen and confirm them by the user, and finally display the selected image in the specified area of ​​the target interface.

Benefits of technology

This enables users to quickly find matching images and insert target interfaces without searching one by one in a large number of images, reducing operation complexity and improving matching degree and personalized recommendation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541250A_ABST
    Figure CN120541250A_ABST
Patent Text Reader

Abstract

The invention discloses an image pushing method and device, electronic equipment and a computer readable medium. Target content is acquired based on a target interface displayed on a screen of the electronic equipment; searching a to-be-selected image matched with the target content as a first image in a to-be-selected image set, the to-be-selected image set comprising a plurality of to-be-selected images; the first image is displayed on the screen; determining a second image selected by the user based on the first image; and displaying the second image in a specified area of the target interface. Therefore, through the content obtained based on the target interface, the first image matched with the content can be automatically found in the to-be-selected image set and pushed to the user, so that the user can quickly find the second image needing to be used from the first image and insert the second image into the target interface; the situation that a user needs to search proper images one by one in the multiple to-be-selected images of the to-be-selected image set is avoided, and the operation complexity of the user is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and more specifically, to an image push method, device, electronic device, and computer-readable medium. Background Art

[0002] Currently, when users need to send images in interactive scenarios, they rely on manual operation. For example, if a user wants to add an image, they need to open the album, filter out one or more suitable images from a large number of them, and then manually upload them. This operation method requires users to find the right image among a large number of images, which is too cumbersome. Summary of the Invention

[0003] This application proposes an image push method, device, electronic device and computer-readable medium to improve the above-mentioned defects.

[0004] In a first aspect, the present application provides an image push method, which is applied to an electronic device, and the method includes: obtaining target content based on a target interface displayed on the screen of the electronic device; searching for a candidate image matching the target content as a first image in a set of candidate images, wherein the set of candidate images includes multiple candidate images; displaying the first image on the screen; determining a second image selected by the user based on the first image; and displaying the second image in a specified area of ​​the target interface.

[0005] In a second aspect, the present application also provides an image pushing device, characterized in that the device includes: an acquisition unit, a search unit, a push unit, a determination unit and a display unit. The acquisition unit is used to acquire target content based on a target interface displayed on the screen of the electronic device. The search unit is used to search for a candidate image matching the target content as a first image in a candidate image set, wherein the candidate image set includes multiple candidate images. The push unit is used to display the first image on the screen. The determination unit is used to determine a second image selected by the user based on the first image. The display unit is used to display the second image in a specified area of ​​the target interface.

[0006] In a third aspect, the present application also provides an electronic device comprising: one or more processors; a memory; and one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to execute the above method.

[0007] In a fourth aspect, the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a program code executable by a processor, and when the program code is executed by the processor, the processor executes the above method.

[0008] The image push method, device, electronic device, and computer-readable medium provided by the present application obtain target content based on a target interface displayed on the screen of the electronic device; search for a candidate image matching the target content as a first image in a candidate image set, wherein the candidate image set includes multiple candidate images; display the first image on the screen; determine a second image selected by the user based on the first image; and display the second image in a designated area of ​​the target interface. Therefore, by obtaining content based on the target interface, a first image matching the content can be automatically found in the candidate image set and pushed to the user, so that the user can quickly find the second image to be used from the first image and insert it into the target interface, avoiding the need for the user to search for a suitable image one by one among multiple candidate images in the candidate image set, thereby reducing the complexity of the user's operation.

[0009] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0011] Figure 1 A schematic diagram of an image selection interface provided in an embodiment of the present application is shown; Figure 2 A flowchart of an image push method provided by an embodiment of the present application is shown; Figure 3 A schematic diagram of a target interface provided by an embodiment of the present application is shown; Figure 4 A flowchart of an image push method provided by another embodiment of the present application is shown; Figure 5 A schematic diagram of a target interface provided by another embodiment of the present application is shown; Figure 6 A schematic diagram of a target interface provided by another embodiment of the present application is shown; Figure 7 A flowchart of an image push method provided by another embodiment of the present application is shown; Figure 8 A schematic diagram of a target interface provided by yet another embodiment of the present application is shown; Figure 9 A schematic diagram of a target interface provided by yet another embodiment of the present application is shown; Figure 10 A flowchart for image push of input content provided by an embodiment of the present application is shown; Figure 11 A schematic diagram of a target interface provided by yet another embodiment of the present application is shown; Figure 12 A schematic diagram of a target interface provided by yet another embodiment of the present application is shown; Figure 13 A flowchart of an image push method provided by another embodiment of the present application is shown; Figure 14 A module block diagram of an image push device provided by an embodiment of the present application is shown; Figure 15 A structural block diagram of an electronic device provided in an embodiment of the present application is shown; Figure 16 A storage unit for storing or carrying program codes for implementing the method according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0012] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for which protection is claimed, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work fall within the scope of protection of the present application.

[0013] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.

[0014] Currently, in user interaction scenarios, users' demand for image materials is becoming increasingly frequent and diverse, covering multiple scenarios from e-commerce reviews, social media content creation to work material submission. However, when supporting users to obtain and use images, current mainstream products still rely on users to manually select images. Figure 1 As shown in Figure 1, users need to manually select the appropriate image from multiple images. However, this approach has problems such as low efficiency and fragmented experience.

[0015] Taking the review scenario as an example, if a user wants to upload an image of a product or food after shopping or enjoying a meal, they must go through the following processes: clicking the upload button, entering the system album, searching by time or folder, filtering through multiple directories, manually selecting the image, and confirming the upload. It usually takes tens of seconds to several minutes for a user to complete the selection of an image. If multiple images are combined or searched across albums, the time cost can increase exponentially. In scenarios such as short video creation and editing long articles with multiple images, users may even need to repeatedly exit the content editing interface to search for images, interrupting the creative process.

[0016] On the one hand, users rely on vague memories (such as shooting time and location) for retrieval, but most users cannot accurately recall the storage location of images that are a certain distance away. On the other hand, the system's album classification only supports basic timelines, geographic locations, or rough AI-generated tags (such as "people" and "food"), lacking in-depth analysis of the image content semantics (such as "blue dress" and "conference room whiteboard"). As a result, users need to spend a lot of time and effort to find the right image among a large number of images.

[0017] Therefore, the inventors have discovered through research on the above technologies that the current user interaction scenarios have the following shortcomings: 1. The operation is cumbersome and time-consuming: Users need to spend a lot of time searching for suitable images in the album. Especially when there are a large number of images in the album, the screening process may take several minutes or even longer, which greatly reduces the efficiency of users posting content or filling out forms.

[0018] 2. Difficulty in ensuring matching: Because image selection relies entirely on user subjective judgment, there may be semantic mismatches between the image and the text. For example, when writing a food review, a user might mistakenly select a landscape image, affecting the quality and effectiveness of the content.

[0019] 3. Lack of personalized recommendations: The system cannot provide personalized image recommendations based on users' historical behavior and preferences. Regardless of the user's interests and usage habits, the image acquisition method and process are the same, failing to meet the diverse needs of users.

[0020] Therefore, in order to overcome the above-mentioned drawbacks, the embodiment of the present application provides an image push method, which is applied to an electronic device, which can be a device capable of running an application, such as a smartphone, a tablet computer, an e-book, etc. Specifically, the method includes: S201 to S205.

[0021] S201: Acquire target content based on a target interface displayed on the screen of the electronic device.

[0022] In an embodiment of the present application, a user can insert or send an image within the target interface. In other words, the target interface refers to an interface with the function of sending images. For example, the target interface can be a multimedia messaging service interface, an instant messaging interface, a storage interface, a review interface, a content publishing interface, and a form interface. Among them, the multimedia messaging service interface has the function of editing content. Images can be inserted or text can be entered within the content editing area, and then the edited content can be sent to the other party. For example, the multimedia messaging service interface can be the SMS boundary interface of a text messaging application, the content query interface of a query application, etc. The instant messaging interface can refer to a chat interface, which also has a content editing area. Within this area, users can select and send images, or send images through certain gestures, such as dragging an image into the chat interface. The storage interface has an image selection control, and users can select images based on this control and upload them to the cloud. The review interface is used to edit and send review content for an item, which allows users to insert and send images within the review interface. The form interface can usually display a standard ID template (such as a social security card that requires the chip area to be exposed), error demonstration prompts (such as a "finger blocking information" animated warning), and a description of the image that needs to be uploaded. Users upload the image they indicate through this form interface.

[0023] It should be noted that the target interface refers to an interface that can support the user to send an image, and the image may refer to an image selected from multiple images. The specific type of the target interface is not limited here.

[0024] In the embodiments of the present application, the target content can refer to the content entered by the user on the interface or the content displayed on the target interface. It is understood that the content entered by the user can reflect the content currently of interest to the user and can serve as a reference for the image desired by the user. In addition, the content displayed on the target interface can also serve as a reference for the image desired by the user. Therefore, the target content can be used to predict the image desired by the user so that the image can be pushed to the user.

[0025] S202: Searching for a candidate image matching the target content as a first image in a candidate image set, wherein the candidate image set includes a plurality of candidate images.

[0026] In one embodiment, the image set includes multiple images to be selected. The images to be selected may be images captured by a user through a camera or downloaded from a network, without limitation. For example, the image set may be an album consisting of multiple images to be selected stored by a camera application of an electronic device. The stored images may be stored locally or in the cloud, without limitation.

[0027] In the embodiments of the present application, the selected image corresponds to image information, which may include entity information of the image, wherein the entity information may be the main object of the image. Specifically, an image may typically contain multiple target objects. For example, a landscape photo may include mountains, houses, grass, trees, flowers, animals, white clouds, etc. For another example, a portrait photo may include the person, items used by the person, and items around the person (e.g., fixtures, vehicles, etc.). By identifying the target image, the main object in the target image can be determined.

[0028] For example, the subject of an image generally refers to the most prominent, prominent, or important part or object in an image. The subject is typically the focus or core of the image; it is the primary object or content of the user's attention. In other words, the subject generally refers to the most important target or object in the image. For example, given a street view image, a "car," "person," or "building" might be considered the subject. Object detection and segmentation algorithms can be used to identify the subject object in an image.

[0029] In the embodiments of the present application, an index can be pre-established for each of the multiple candidate images in the candidate image set. Specifically, corresponding image information can be assigned to each candidate image. This image information can include the image's subject type, shooting time, location, shooting equipment, image type (food image, landscape image, portrait image), and so on. It should be noted that the image type refers to the type of the entire image, while the image's subject type refers to the type of the subject object in the image. For example, an image with a hamburger as its subject object would have an image type of food, and a subject type of hamburger.

[0030] As an implementation, feature extraction is performed for each candidate image in the album. Using image feature extraction models based on convolutional neural networks (CNNs), such as ResNet and VGG, rich visual features are extracted from the images. These visual features include, for example, color, texture, shape, and object recognition features. Simultaneously, combined with metadata such as the image's shooting time, location, and device, detailed index information is constructed for each candidate image and stored in an image database for rapid retrieval and matching. This index information is the aforementioned image information.

[0031] Specifically, convolutional neural networks are used to extract image features such as color, texture, shape, and object recognition. Object recognition features are used to determine the image's main subject and image type. For example, image classification models are used to predict image categories (e.g., "landscape," "person," or "architecture"). Object detection algorithms (e.g., YOLO and Faster R-CNN) are used to extract the location and category of multiple objects in the image. Image metadata (e.g., shooting time, location, and shooting device) is typically stored in the image's EXIF ​​(Exchangeable Image File Format) information.

[0032] As an implementation method, through analysis of the target content, the key information corresponding to the target content may include entity information and auxiliary information corresponding to the entity information. The entity information refers to the core subject of the target content (such as product name, location, person name, etc.), and the auxiliary information may refer to related information of the entity information. For example, the auxiliary information may include attribute information, quantity information and relationship information, etc., wherein the attribute information may be an attribute used to describe the entity information, such as the color, price, function, etc. of the product; the time information may be time content related to the entity content, such as the time of the event, the scheduled date, etc.; the quantity information may be specific numerical information such as quantity, size, capacity, etc.; and the relationship information may be a description of the relationship between the query objects.

[0033] In an embodiment of the present application, the key information corresponding to the target content may include the semantics of the target content or the type of the target content, and the target content may be the content entered by the user in the interface, the interface itself, or the content displayed in the interface.

[0034] As can be seen, since the candidate images correspond to image information, and the target content also corresponds to key information, semantics, or types, the two can be matched. For example, if the target content is Sichuan cuisine, then if the candidate image corresponds to the same type of Sichuan cuisine, the two will have a higher degree of match. However, if the candidate image is Shandong cuisine, the match will be lower. Therefore, by matching the target content with the candidate images, it is possible to find the candidate image that matches the target content from the candidate image set and use it as the first image.

[0035] S203: Display the first image on the screen.

[0036] As an embodiment, the first image may be displayed on the screen by displaying the first image in the target interface. Specifically, the first image may be displayed in the editing area of ​​the target interface. For example, an editing area may be provided in the target interface, and a virtual keyboard may be provided in the editing area so that the user can input content through the virtual keyboard. Alternatively, a plurality of controls may be displayed in the editing area, including an image insertion control. When the user clicks the insertion control, an image selection interface may be displayed. The image selection interface may display the selected images in the set of selected images. Therefore, after determining the first image, the first image may be displayed in the editing area of ​​the target interface. For example, the content displayed in the editing area may be changed to the first image, that is, the first images may be adjusted to an appropriate size and displayed one by one in the editing interface.

[0037] like Figure 3 As shown in (a) and (b), the target interface shown is the instant messaging interface, such as Figure 3 As shown in (a), the virtual keyboard is displayed in the editing area 301. After the first image is determined, the multiple first images 302 are displayed in the editing area 301.

[0038] As another implementation, the first image may be displayed in a floating manner on the target interface, for example, the first image may be displayed through a floating window on the target interface.

[0039] As another embodiment, the first image may be displayed in an image selection interface. Specifically, the image selection interface may be Figure 1 In the interface shown, if the first image has not been determined, the images displayed in the image selection interface are the various images in the candidate set. For example, more images can be browsed by swiping up and down in the image selection interface. However, if the first image has been determined, when the user calls up the image selection interface in the target interface, the image displayed in the image selection interface is the first image. This allows the user to maintain the same operating habits as before the first image was displayed, but the content browsed is the image recommended by the system, namely the first image.

[0040] S204: Determine a second image selected by the user based on the first image.

[0041] In one embodiment, the user can select at least some of the multiple first images as the second image. For example, the first image can be selected by clicking on the first image, selecting an image through a voice command, or dragging the first image, for example, dragging the first image to a designated area of ​​the target interface. The specific selection method is not limited herein.

[0042] S205: Displaying the second image in a designated area of ​​the target interface.

[0043] Exemplarily, the designated area can be an area within the target interface for displaying images inserted by the user. Of course, it can also be an area corresponding to the target content, for example, the area where the target content is located or the adjacent area of ​​the display position corresponding to the target content, for example, the area above or below the display position corresponding to the target content, without specific limitation.

[0044] It should be noted that the designated area can be the area where the user currently wants to insert an image in the target interface. If the target interface has multiple areas where images can be inserted, namely image display areas, then the designated area can be one of the multiple image display areas. For example, it can be any area of ​​the multiple image display areas where no image is displayed, or it can be an image display area corresponding to the target content. For example, it can be a display area above or below the display position corresponding to the target content.

[0045] Therefore, in an embodiment of the present application, target content is obtained based on a target interface displayed on the screen of the electronic device; a candidate image matching the target content is searched for as a first image in a candidate image set, wherein the candidate image set includes multiple candidate images; the first image is displayed on the screen; a second image selected by the user based on the first image is determined; and the second image is displayed in a designated area of ​​the target interface. Therefore, by obtaining content based on the target interface, a first image matching the content can be automatically found in the candidate image set and pushed to the user, so that the user can quickly find the second image to be used from the first image and insert it into the target interface, avoiding the need for the user to search for a suitable image one by one among the multiple candidate images in the candidate image set, thereby reducing the complexity of the user's operation.

[0046] See also Figure 4 , Figure 4 An image push method provided by an embodiment of the present application is shown, which is applied to an electronic device and specifically includes: S401 to S405.

[0047] S401: In a target interface displayed on a screen of the electronic device, target content is obtained based on image indication content, wherein the image indication content is used to indicate an image that a user needs to upload through the target interface.

[0048] As an embodiment, the target interface is a form interface, that is, the interface clearly indicates the image that the user needs to insert in the interface. Typically, the interface has at least one image upload area, and image indication content is displayed near the image upload area.

[0049] like Figure 5 As shown, the target interface displays image indication content 501 and an image upload area 502 corresponding to the image indication content 501. It can be seen that the image indication content is used to indicate the image that the user needs to upload through the target interface. It should be noted that when the image indication content 501 corresponds to the image upload area 502, the image indication content is used to represent the image that the user needs to upload in the image upload area corresponding to the image indication content. Moreover, when there are multiple image indication contents on the interface, and each image indication content corresponds to at least one image upload area, each image indication content is used to represent the image that the user needs to upload in the image upload area corresponding to the image indication content.

[0050] Therefore, when there are multiple image indication contents on the target interface, at least one image indication content can be used as the target content. As an embodiment, the image type to be inserted corresponding to each image indication content can be identified, and the image indication content corresponding to the image type that meets the preset type can be used as the target content, wherein the preset type can be a pre-set image type that the user rarely takes or obtains. For example, the preset type can be an ID type or a certification document type. Since users do not often take or save these types of images, the time difference between the saved time of such images and the current time may be greater than the specified difference, making it cumbersome for users to search. Therefore, setting this type of image as a preset type can help users quickly find images of this type.

[0051] like Figure 5 As shown, the image indication content 501 is used to indicate that the image uploaded by the user through the target interface is an ID photo. Therefore, the implementation method of obtaining the target content based on the image indication content may be to use the image indication content as the target content. Of course, it may also be possible to determine at least one image content as the target content from all the image indication contents corresponding to the target interface when there are multiple image indication contents on the target interface. The determination method may refer to the aforementioned content and will not be repeated here.

[0052] S402: Searching for a candidate image matching the target content as a first image in a candidate image set, wherein the candidate image set includes a plurality of candidate images.

[0053] As an embodiment, searching for a candidate image that matches the target content in the candidate image set may be to determine the degree of matching between the target content and each candidate image in the candidate image set, and determine the matching candidate image corresponding to the target content based on the matching degree of each image.

[0054] For example, the method for determining the match between the target content and the selected image may be to perform a matching operation between the content description information of the target content and the image information of the selected image.

[0055] As an embodiment, the content description information may include at least one of key information and semantic information. Key information, also known as keywords, is usually a single word or phrase extracted from the text, representing the core content of the text. These keywords are usually direct, surface-level words that help quickly identify the subject of the text. Semantic information refers to the deeper meaning in the text. It not only focuses on the appearance of individual words, but also understands the relationship between these words and the overall meaning of the text. Semantic information involves understanding context, emotion, intent, etc.

[0056] It can be understood that the implementation method of determining the matching degree between the image information of the selected images in the set of selected images and the target content is to determine the content description information corresponding to the target content, wherein the content description information includes at least one of the key information and semantic information corresponding to the target content; determine the matching degree between the image information of the selected images in the set of selected images and the content description information of the target content, and obtain the matching degree of each of the selected images.

[0057] That is to say, based on the key information of the target content, the matching degree between the target content and the image information of the to-be-selected image in the to-be-selected image set can be determined, and the first image can be found; based on the semantic signal of the target content, the matching degree between the target content and the image information of the to-be-selected image in the to-be-selected image set can be determined, and the first image can be found; or the matching degree between the target content and the image information of the to-be-selected image in the to-be-selected image set can be determined in combination with the key information and semantic information of the target content, and the first image can be found. For example, the first matching degree between the key information and the image information of the to-be-selected image can be determined, and the second matching degree between the semantic information and the image information of the to-be-selected image can be determined, and the matching degree between the target content and the to-be-selected image can be obtained based on the first matching degree and the second matching degree.

[0058] For example, an electronic device includes a semantic recognition module capable of obtaining key information and semantic information about target content. The target content is input into the module, which initiates a data preprocessing process. This process first normalizes the input text, including removing special characters, correcting spelling errors, and standardizing capitalization, to ensure the accuracy of subsequent semantic analysis.

[0059] The semantic recognition module uses a pre-trained language model based on the Transformer architecture, such as the GPT series or BERT. This pre-trained language model performs deep semantic analysis on pre-processed text. The model segmentes the input text, converts it into word vector representations, and then uses a multi-layer Transformer encoder to extract features and understand semantics from the word vectors.

[0060] Specifically, the model tokenizes the input text, breaking it down into individual words or subwords (depending on the tokenization algorithm used, this may be further broken down into subwords, characters, and so on). Each token (or subword) is converted into a word vector. These vectors are learned from a large amount of text data and capture the semantics and context of the word. Word vectors are typically represented using word embeddings (such as Word2Vec, GloVe, or the model's built-in embeddings). Next, the tokenized and tokenized input is passed to the multi-layer encoder of the Transformer model. The Transformer architecture effectively captures relationships between words and handles dependencies within text. Each encoder layer extracts features from the input word vectors, gradually understanding the contextual details. For example, if the target content is "I went to an amazing Italian restaurant last week. The pasta and pizza were both delicious," the model can recognize keywords such as "Italian restaurant," "pasta," "pizza," and "delicious food," and understand the positive sentiment expressed in the sentence and the specific context of the dining experience. In other words, keywords such as "Italian restaurant," "pasta," "pizza," and "food" can be considered the aforementioned key information, representing the entity information corresponding to the target content. This entity information can be the core processing of the target content, that is, the main content of the target content. The emotion and scene can then be considered auxiliary information corresponding to this entity information. Semantic information can also be obtained. For example, this semantic information can include emotion, context, and intent. In this example, the emotion is positive and refers to a preference for food. The context is the user describing a dining experience at a restaurant. The intent is that the user may be sharing a dining experience or expressing a recommendation for a particular restaurant or food.

[0061] Therefore, it can be seen that the semantic information of the target content can be obtained based on the target content and its context. That is, the content description information includes semantic information. Therefore, the implementation method for determining the content description information corresponding to the target content is to obtain the context content corresponding to the target content; and then obtain the semantic information corresponding to the target content based on the target content and the context content.

[0062] As an implementation method, the contextual content may be the browsing content of the user within a preset time period. The browsing content is not limited to the content that the user pays attention to in the target interface. The content of interest may include input content, comment content, and shared content.

[0063] For example, if the target content represents uploading an ID card picture, the contextual content may be the user identity information corresponding to the form interface. The user identity information may be determined by determining the function type corresponding to the control selected by the user before entering the form interface. If the function type is the first type, it is determined that the user needs to create their own information. The contextual content is the identity information of the currently logged-in user, and the corresponding semantic information indicates that the currently logged-in user's ID card picture needs to be uploaded. If the function type is the second type, it is determined that the user needs to create someone else's information. The contextual content is the identity information of another user, and the corresponding semantic information indicates that the user needs to upload another user's ID card picture.

[0064] Furthermore, the contextual content can also be other content outside of the target content within the form interface. This other content can serve as auxiliary information for the target content. This auxiliary information can include information such as the time, location, or posture of the image to be uploaded. The corresponding voice message can then indicate that an ID card image that meets this auxiliary information needs to be uploaded. Therefore, after obtaining the semantic information of the target content and the image information of the candidate image, a matching operation can be performed on the target content and the candidate image based on the semantic information of the target content and the image information of the candidate image.

[0065] Specifically, the image matching and recommendation module searches and matches images in the image database based on the key information and semantic information extracted by the semantic recognition module. Using algorithms such as cosine similarity and Euclidean distance, it calculates the similarity between image features and semantic features, screening images with a high semantic relevance to the user's input text. For example, when a user enters a review about food, the system retrieves images from the photo album containing elements such as food and restaurant scenes and sorts them by similarity from highest to lowest.

[0066] It is understood that the semantic recognition module analyzes the text entered by the user and extracts keywords and semantic features. For example, when a user enters a review about delicious food, the semantic recognition module will identify keywords such as "delicious food," "food," and "restaurant." The image matching and recommendation module first needs to extract features from the images in the database. Deep learning models (such as convolutional neural networks (CNNs)) are typically used to extract the visual features, or image information, of each image. These features include visual information such as color, shape, texture, objects, and scenes in the image. Once the semantic information of the user-entered text and the visual features of the image are obtained, the system can use a similarity calculation algorithm to compare the similarity between the two. For example, this can be cosine similarity or Euclidean distance.

[0067] Cosine similarity calculates the angle between two vectors; the closer the value is to 1, the more similar the two vectors are. Here, cosine similarity is used to measure the similarity between the semantic vector of the input text and the visual feature vector of the image. Euclidean distance measures the straight-line distance between two vectors; smaller distances indicate greater similarity. In the space of images and semantic features, Euclidean distance can be used to measure similarity. Similarity here can be understood as the degree of match between the target content and the pre-set image.

[0068] Then, for a certain target content, according to the matching degree between each candidate image and the target content, a candidate image matching the target content is obtained, thereby obtaining a first image.

[0069] As an implementation method, the implementation method of searching for a candidate image matching the target content as the first image in the candidate image set may be to determine the matching degree between the image information of the candidate image in the candidate image set and the target content; and determine the first image based on the matching degree of each candidate image.

[0070] For example, the candidate images with a matching degree greater than a specified threshold may be selected as the first image. Alternatively, a certain number of candidate images ranked high may be selected as the first image. For example, determining the first image based on the matching degree of each candidate image may be performed by selecting a preset number of candidate images as the first image based on the matching degrees of all the candidate images, in descending order.

[0071] As an embodiment, the preset number can be a value set based on actual needs, or it can be determined based on the target interface. Exemplarily, if a user uploads an image through the target interface, the target interface corresponds to the number of images that the user needs to upload. For example, the target interface corresponds to image indication content, and the image indication content includes images that indicate that the user needs to upload through the target interface. Based on the image, the lower limit of the number of images corresponding to the image indication content can be determined. For example, if the image indication content is to upload the front and back photos of the ID card, the lower limit of the number of images corresponding to the image indication content can be determined to be 2. Therefore, based on the image indication content of the target interface, the lower limit of the number of images of the image indication content can be determined, and thus based on the lower limit of the number of images, the preset number can be determined, that is, the preset number is not less than the lower limit of the number of images.

[0072] Furthermore, in addition to recommending images to users based on the matching between target content and candidate images, the accuracy of recommendations can also be improved by integrating user behavior. Specifically, determining the first image based on the matching degree of each candidate image can include obtaining user preference information; determining a rating for the candidate image based on the matching degree of the candidate image and the user preference information; and selecting a preset number of candidate images as the first image in descending order of rating.

[0073] As an implementation method, based on the matching degree of all the candidate images, the candidate images are sorted in descending order to obtain a first image sequence; based on user preference information, the sorting of the candidate images in the first image sequence is adjusted to obtain a second image sequence; from the second image sequence, the top N candidate images are selected as the first image, where N is a positive integer.

[0074] For example, user preference information may be obtained based on at least one of a user's selection history for multiple recommended images, browsing habits, and interest level. The selection history may include the type, category, and / or style of images selected by the user each time; browsing habits may include images viewed by the user, the duration of viewing, the number of clicks, and the like; and interest level may include the user's preferred shooting style, color tone, and theme.

[0075] As an implementation, a personalized recommendation model can be trained based on the user's preference information, and the recommendation model can provide a recommendation score for each preset image. For example, the recommendation model can be a model built based on a collaborative filtering algorithm, a deep learning algorithm, or other architecture.

[0076] Collaborative filtering is used to recommend content that users might like. In this application, this collaborative filtering algorithm can be item-based collaborative filtering. Item-based collaborative filtering generates recommendations by calculating the similarity between images. When a user selects an image, the system can recommend other similar images. For example, if a user selects a food image, the system will recommend other food images with similar styles, tones, and themes based on the similarity between that image and other images.

[0077] Therefore, based on the user's selection record, browsing habits and / or interest, the user's images of interest can be obtained, and the preset images are compared with the images of interest. Based on the similarity, the recommendation score of each candidate image can be determined.

[0078] In addition, for deep learning algorithms, images of interest determined based on the user's selection records, browsing habits and / or interest level can be used as sample images to train the recommendation model, so that the recommendation model can learn and extract the visual features of each sample image (such as color, texture, shooting angle, scene, etc.), and use these features to calculate the recommendation score of the selected image.

[0079] Therefore, based on the matching degree of all the candidate images, the candidate images are sorted in descending order to obtain a first image sequence; then, based on the recommendation model, the recommendation score of each image in the first image sequence is obtained, and the images with a recommendation score greater than the score threshold are adjusted to the front position of the first image sequence to obtain a second image sequence.

[0080] As another embodiment, the second image sequence may be obtained by determining the degree of match between the target content and each image to be selected, and then determining the recommendation score of each image to be selected, and determining the score value of each image to be selected based on the recommendation score and the degree of match, and based on the score value, obtaining an image sequence in descending order of the score value, and taking the top N images as the first image.

[0081] Exemplarily, based on the first weight, the second weight, the matching degree and the recommended score, a score value of the selected image is obtained, wherein the first weight is the weight corresponding to the matching degree, and the second weight is the weight corresponding to the recommended score. Then, a first product between the first weight and the matching degree is obtained, and a second product between the second weight and the recommended score is obtained. The first product and the second product are summed to obtain the score value.

[0082] Therefore, embodiments of the present application can also incorporate the results of user behavior analysis and learning modules to personalize recommended images. Based on the user's historical selection records, browsing habits, interests, and other data, collaborative filtering algorithms and deep learning algorithms are used to generate a personalized recommendation model for each user. For example, if a user frequently chooses to take pictures of food in a unique style, the system will prioritize pictures with similar styles when recommending, thereby improving the accuracy of recommendations and user satisfaction.

[0083] In other words, a hybrid recommendation approach combines collaborative filtering, deep learning, and content-based recommendations to generate personalized recommendations. The system prioritizes images that match the user's interests. For example, if a user frequently selects food photos with a unique style, other images with similar styles will be prioritized.

[0084] S403: Display the first image on the screen.

[0085] In the embodiment of the present application, the first image can be displayed on the target interface in the form of a floating window, and the display position of the floating window can be within a preset range corresponding to the display area of ​​the image indication content, or within a preset range corresponding to the image upload area corresponding to the image indication content, which is not limited here.

[0086] It can be understood that the implementation method of displaying the first image on the screen can be to display a schematic diagram corresponding to the first image on the screen. The schematic diagram can be a thumbnail of the first image, that is, it can be an image with reduced size and / or image quality. Through this schematic diagram, not only can the user understand the image content of the first image, but also it can avoid directly displaying the original image of the first image on the screen, which may cause other content on the screen to be blocked.

[0087] As an implementation method, after the first image is displayed on the screen, the user can interact with the first image. For example, the user can like, collect, mark, and other operations on the first image to express how much he likes the image. Then, the system of the electronic device will collect these interaction data of the user to serve the aforementioned recommendation model. In other words, the system provides rich user interaction functions to facilitate users to operate and manage recommended images. Users can like, collect, mark, and other operations on recommended images to express how much they like the image. At the same time, the system will collect user interaction behavior data, such as the number of clicks, dwell time, selection results, etc.

[0088] S404: Determine a second image selected by the user based on the first image.

[0089] like Figure 6As shown, the target interface is a form interface. After the first image is determined, the first image is displayed on the interface via a floating window 601. In one embodiment, the first image in the floating window can be slid into the floating window to be displayed. The user can select the first image in the floating window, and the selected first image will be used as the second image.

[0090] For example, the first images may be selected by operating a selection control corresponding to each first image, such as Figure 6 As shown, the user can select the corresponding first image by selecting the control 603. In addition, all the first images can also be selected, that is, all the first images are used as the second image, such as Figure 6 As shown, the floating window 601 displays a select-all control “One-click upload”, and the user operates (eg, clicks) the select-all control to select all first images with one click.

[0091] S405: Display the second image in a designated area of ​​the target interface.

[0092] The designated area may be the image upload area corresponding to the aforementioned image indication content.

[0093] Combined with reference Figure 5 and Figure 6 After the user selects the second image from the first image in the floating window 601 , the second image may be automatically displayed in the image upload area 502 .

[0094] It should be noted that when there are multiple target contents, the designated area corresponding to each target content in the interface can be automatically identified, and the second image corresponding to each target content selected by the user can be inserted into the designated area corresponding to the target content.

[0095] As an implementation, when the target interface is a form interface, the user needs to upload some images through the form interface. These images typically have certain standards, for example, parameters such as the image size must meet certain conditions. Therefore, the image indication content may include image parameter information, which may include parameters such as the image size, resolution, and color information. For example, the size may be required to be 1 inch, and the color information may be required to be a blue background. Therefore, an implementation method for displaying the second image in a designated area of ​​the target interface may be to adjust the second image based on the image parameter information; and then display the adjusted second image in the image upload area of ​​the target interface corresponding to the image indication content.

[0096] Therefore, in a form-filling scenario, the system automatically identifies the first image from the set of candidate images based on the image indicator displayed within the form interface and displays it on the interface. The user can then select an appropriate image, eliminating the need to search for an image within a pre-set image set (e.g., a photo album). Furthermore, when the user clicks on a recommended image, the system automatically populates the image upload area specified in the form, adjusting the image's size, format, and other parameters based on the form's requirements, i.e., the image parameter information in the image indicator.

[0097] See also Figure 7 , Figure 7 An image push method provided by an embodiment of the present application is shown, which is applied to an electronic device and specifically includes: S701 to S705.

[0098] S701: Obtain target content based on content input by a user in a target interface displayed on a screen of the electronic device.

[0099] In the embodiment of the present application, the user can enter content in the target interface, and of course, can also insert or send pictures in the interface. Exemplarily, the target interface can be a multimedia messaging service interface, an instant messaging interface, an evaluation interface, and a content publishing interface.

[0100] As an implementation method, the content currently input by the user in the target interface is obtained to obtain the target content. For example, the content input by the user can be directly used as the target content. Figure 8 As shown, Figure 8 The target interface shown is a review interface, which is used to review a certain business. The interface includes a content editing area 801, in which the user can enter content and insert pictures and videos. Therefore, the target content can be the text content entered by the user in the content editing area.

[0101] In addition, if Figure 9 As shown, Figure 9The target interface shown is an instant messaging interface, which includes a content editing area 901 and a message display area 902. The content editing area 901 can display the content currently input by the user and not sent, and the message display area 902 can display the content that the user has sent or received. For example, the target content can be the content currently input by the user displayed in the content editing area 901. In addition, the target content can also be the content displayed in the message display area 902. Specifically, it can be the content recently sent by the target user in the message display area 902, that is, the last content corresponding to the target user in the message display area 902. The target user can be the user who is currently logged into the application corresponding to the target interface, and can also be understood as the user who is currently logged into the electronic device.

[0102] Therefore, when the target interface is running on the screen of the electronic device, the content input by the user in the interface can be obtained. Specifically, it can be the currently input content or the most recently published content. However, since the published content also needs to be input through the target interface, the most recently published content also belongs to the content input by the user in the interface.

[0103] S702: Searching for a candidate image matching the target content as a first image in a candidate image set, wherein the candidate image set includes a plurality of candidate images.

[0104] S703: Display the first image on the screen.

[0105] S704: Determine a second image selected by the user based on the first image.

[0106] S705: Display the second image in a designated area of ​​the target interface.

[0107] like Figure 10 As shown, for the text content input by the user in the target interface, the text information corresponding to the target content is determined, and then a matching operation is performed. The specific matching process can be referred to the previous embodiment and will not be repeated here. Then, an image recommended to the user, i.e., a first image, is obtained and displayed on the screen.

[0108] As an embodiment, the first image may be displayed on the screen in a floating window on the target interface, and the display position of the first image may be adjacent to the display area of ​​the target content.

[0109] In addition, since the designated area of ​​the target interface can be the content editing area of ​​the target interface or the message display area. Specifically, in some embodiments, the logic for the user to send an image in the target interface is that the user inserts the image in the content editing area of ​​the target interface, and then sends the image after detecting the user's confirmation instruction. Therefore, in this case, the designated area is the content editing area of ​​the target interface. Figure 11 As shown, the control Figure 8 It can be seen that the first image 1101 recommended to the user is displayed in the area corresponding to the content editing area of ​​the interface. The user can select an image, namely the second image, from the first image. The second image will be inserted into the content editing area. The user can send the image within the content editing area to complete the evaluation.

[0110] In other embodiments, the logic for the user to send an image on the target interface is that after the user selects an image on the target interface, the image is directly sent and displayed in the message display area of ​​the interface. In this case, the designated area is the message display area of ​​the interface. Figure 12 As shown, the first image 1101 recommended to the user is displayed in the area corresponding to the message display area of ​​the interface. The user can select an image, namely a second image, from the first image. The second image will be directly sent and displayed in the message display area of ​​the target interface. It can be seen that after the user sends the content "I saw flowers yesterday", this content is used as the target content. Through analysis of this target content, it can be extracted that the subject is flowers. In other words, the keyword corresponding to the target content is determined to be flowers, so that the image related to "flowers", namely the first image, can be automatically pushed to the user.

[0111] As an embodiment, the content description information corresponding to the target content may include semantic information. In addition to being obtained from the target content itself, the semantic information may also be obtained in combination with the contextual content corresponding to the target content. In other words, the semantic information of the target content may be obtained in combination with the contextual content.

[0112] Specifically, an implementation method for obtaining target content based on content input by a user within a target interface displayed on the screen of the electronic device may include determining the designated content input by the user within the target interface displayed on the screen of the electronic device; determining contextual content associated with the designated content; and then the designated content is the target content. The contextual content includes at least one of historical content, content browsed by the user within the target interface, and content corresponding to a reply to the designated content. The historical content includes content input by the user within the target interface prior to inputting the designated content.

[0113] In other words, the semantic recognition module also has the ability to understand the context. It not only analyzes the currently input text, but also combines the previous input content to comprehensively understand the user's expression intention. For example, when a user enters multiple messages in succession, the system can accurately understand the meaning of subsequent messages based on the topics mentioned previously. At the same time, for some semantically ambiguous or incomplete expressions, the model will use its knowledge reserves and semantic understanding capabilities to complete the semantics and provide more accurate semantic information for subsequent image matching. Specifically, pre-trained language models such as GPT can help the system reason about incomplete sentences. For example, when a user enters "I like that kind of blue-toned landscape", if the system knows that the user has previously mentioned "mountain scenery", it can complete the user's intention as "I like blue-toned mountain scenery".

[0114] As an implementation, if the user is dissatisfied with the recommended image or continues to enter text, the system will update the recommended images in real time based on new semantic information and user feedback. That is, after the first image is displayed on the screen, the system determines whether the user has selected at least one first image. If not, the electronic device will capture the content that the user continues to enter within the target interface as the new target content and return to executing the operation of searching for a candidate image matching the target content in the candidate image set as the first image, as well as subsequent operations.

[0115] Alternatively, regardless of whether the user selects the first image, as long as new content is detected, a recommendation operation based on the new content will be performed. For example, taking the target interface as a review interface, when filling out the review, the user first enters content about the food, and the system recommends corresponding food images. When the user continues to enter "the environment is also very good", the system immediately retrieves and recommends images that include elements of the restaurant environment, providing the user with a more accurate recommendation service.

[0116] See also Figure 13 , Figure 13 An image push method provided by an embodiment of the present application is shown, which is applied to an electronic device and specifically includes: S1301 to S1305.

[0117] S1301: Determine the interface type of the target interface based on display content corresponding to the target interface displayed on the screen of the electronic device, wherein the interface type serves as the target content.

[0118] As an implementation manner, in addition to using the above-mentioned input content and image indication content to obtain the target content, the interface type of the target interface can also be used as the target content.

[0119] Exemplarily, an interface screenshot of the target interface can be obtained, and the interface type of the target interface can be determined based on the interface screenshot. That is to say, by identifying the interface screenshot to obtain the main elements of the interface, the interface type can usually be determined by analyzing the main elements it contains. Different types of interfaces will have different elements and component structures. For example: the evaluation interface usually contains star ratings, text input boxes, submit buttons, evaluation labels (such as "good reviews", "bad reviews"), instant messaging interfaces usually include chat boxes, contact lists, message input boxes, send buttons, chat records, etc., payment interfaces: usually have payment buttons, order details, amount display, payment method selection (such as credit card, XXAPP payment, etc.) and payment confirmation buttons.

[0120] In addition, the type of the interface may also be determined based on the type of the application. Specifically, the application corresponding to the target interface may be obtained, and the type of the application may be used as the type of the target interface.

[0121] S1302: Searching for a candidate image matching the target content as a first image in a candidate image set, wherein the candidate image set includes a plurality of candidate images.

[0122] S1303: Display the first image on the screen.

[0123] S1304: Determine a second image selected by the user based on the first image.

[0124] S1305: Display the second image in a designated area of ​​the target interface.

[0125] As an implementation method, in an embodiment of the present application, the interface type of the target interface can be used to determine a set of images to be selected. Specifically, the album application of the electronic device can be provided with multiple preset images, and the image set corresponding to the multiple preset images is a preset image set. The electronic device searches for images that match the interface type from the preset image set based on the interface type of the target interface as the images to be selected. Then, an operation is performed to search for an image to be selected that matches the target content as the first image in the image set to be selected. The target content can be determined based on a method other than the interface type, for example, based on the content input by the user in the target interface or the content indicated by the image.

[0126] See also Figure 14 , which shows a structural block diagram of an image pushing device 1400 provided in an embodiment of the present application. The device may include: an acquisition unit 1401, a search unit 1402, a pushing unit 1403, a determination unit 1404 and a display unit 1405.

[0127] The acquisition unit 1401 is configured to acquire target content based on a target interface displayed on the screen of the electronic device.

[0128] Furthermore, the acquisition unit 1401 is further configured to obtain target content based on image indication content in the target interface displayed on the screen of the electronic device, wherein the image indication content is configured to indicate an image that the user needs to upload through the target interface.

[0129] Furthermore, the acquiring unit 1401 is further configured to adjust the second image based on the image parameter information; and display the adjusted second image in an image upload area corresponding to the image indication content in the target interface.

[0130] Furthermore, the acquisition unit 1401 is further configured to obtain target content based on content input by the user in the target interface displayed on the screen of the electronic device.

[0131] Furthermore, the acquisition unit 1401 is further configured to determine an interface type of the target interface based on display content corresponding to the target interface displayed on the screen of the electronic device, wherein the interface type serves as the target content.

[0132] The searching unit 1402 is configured to search for a candidate image matching the target content as a first image in a candidate image set, wherein the candidate image set includes a plurality of candidate images.

[0133] Furthermore, the search unit 1402 is further configured to determine a degree of matching between image information of the candidate images in the candidate image set and the target content; and determine the first image based on the matching degree of each candidate image.

[0134] Furthermore, the searching unit 1402 is further configured to select a preset number of candidate images as the first images based on the matching degrees of all the candidate images in descending order.

[0135] Furthermore, the search unit 1402 is also used to obtain user preference information; determine the score value of the candidate image based on the matching degree of the candidate image and the user preference information; and select a preset number of candidate images as the first image in descending order of the score value.

[0136] The pushing unit 1403 is configured to display the first image on the screen.

[0137] The determining unit 1404 is configured to determine the second image selected by the user based on the first image.

[0138] The display unit 1405 is configured to display the second image in a designated area of ​​the target interface.

[0139] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0140] In several embodiments provided in this application, the coupling between modules may be electrical, mechanical or other forms of coupling.

[0141] In addition, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional modules.

[0142] Please refer to Figure 15 , which shows a structural block diagram of an electronic device provided in an embodiment of the present application. The electronic device 100 can be an electronic device capable of running applications, such as a smartphone, a tablet computer, an e-book, etc. The electronic device 100 in the present application may include one or more of the following components: a processor 110, a memory 120, and one or more applications, wherein the one or more applications may be stored in the memory 120 and configured to be executed by one or more processors 110, and one or more programs are configured to execute the method described in the aforementioned method embodiment.

[0143] The processor 110 may include one or more processing cores. The processor 110 utilizes various interfaces and circuits to connect various components within the electronic device 100. It executes instructions, programs, code sets, or instruction sets stored in the memory 120, and accesses data stored in the memory 120 to perform various functions and process data within the electronic device 100. Optionally, the processor 110 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 110 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may also be implemented independently of the processor 110 via a separate communications chip.

[0144] The memory 120 may include random access memory (RAM) or read-only memory (ROM). The memory 120 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the various method embodiments described below, and the like. The data storage area may also store data created by the electronic device 100 during use (such as a phone book, audio and video data, and chat history data).

[0145] Please refer to Figure 16 , which shows a block diagram of a computer-readable medium provided in an embodiment of the present application. The computer-readable medium 1600 stores program code, which can be called by a processor to execute the method described in the above method embodiment.

[0146] Computer-readable medium 1600 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or ROM. Alternatively, computer-readable medium 1600 may include non-transitory computer-readable storage media. Computer-readable medium 1600 has storage space for program code 1610 for executing any of the method steps described above. This program code can be read from or written to one or more computer program products. Program code 1610 may be compressed, for example, in a suitable format.

[0147] Therefore, the embodiment of the present application obtains target content based on the target interface displayed on the screen of the electronic device; searches for a candidate image that matches the target content as the first image in the candidate image set, wherein the candidate image set includes multiple candidate images; displays the first image on the screen; determines the second image selected by the user based on the first image; and displays the second image in a designated area of ​​the target interface. Therefore, by obtaining content based on the target interface, the first image that matches the content can be automatically found in the candidate image set and pushed to the user, so that the user can quickly find the second image to be used from the first image and insert it into the target interface, avoiding the need for the user to search for a suitable image one by one among the multiple candidate images in the candidate image set, thereby reducing the complexity of the user's operation.

[0148] Users no longer need to manually search for images in their photo albums; the system automatically recommends relevant images, significantly reducing the time required for content creation and form completion, and improving the fluidity and convenience of user operations. Through intelligent semantic matching, recommended images are highly semantically aligned with the text entered by the user, enhancing the expressiveness and communication of the content, making the published content more engaging and professional. Based on users' historical choices and behavioral data, the system continuously learns and optimizes its recommendation algorithm, providing each user with personalized image recommendations that align with their interests and usage habits, thereby improving user satisfaction and loyalty.

[0149] In addition to recommended images, the system provides easy-to-use image editing and synthesis functions. Users can crop, add filters, and annotate recommended images. They can also combine multiple recommended images to create more creative and expressive content. Integrated voice recognition technology supports voice input. The system can also recommend and fill in images based on voice semantics, providing users with a more convenient interaction method, especially suitable for scenarios where hands are inconvenient or quick operation is required.

[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An image push method, characterized in that: Applied to electronic equipment, the method includes: Acquiring target content based on a target interface displayed on the screen of the electronic device; Searching for a candidate image matching the target content in a candidate image set as a first image, wherein the candidate image set includes a plurality of candidate images; displaying the first image on the screen; determining a second image selected by a user based on the first image; The second image is displayed in a designated area of ​​the target interface.

2. The method according to claim 1, characterized in that The acquiring target content based on the target interface displayed on the screen of the electronic device includes: In the target interface displayed on the screen of the electronic device, target content is obtained based on image indication content, wherein the image indication content is used to indicate an image that the user needs to upload through the target interface.

3. The method according to claim 2, characterized in that The designated area is the image upload area corresponding to the image indication content.

4. The method according to claim 3, characterized in that The image indication content includes image parameter information, and the displaying of the second image in a designated area of ​​the target interface includes: adjusting the second image based on the image parameter information; The adjusted second image is displayed in the image upload area corresponding to the image indication content in the target interface.

5. The method according to claim 1, characterized in that The acquiring target content based on the target interface displayed on the screen of the electronic device includes: The target content is obtained based on the content input by the user in the target interface displayed on the screen of the electronic device.

6. The method according to claim 1, characterized in that The acquiring target content based on the target interface displayed on the screen of the electronic device includes: Based on the display content corresponding to the target interface displayed on the screen of the electronic device, the interface type of the target interface is determined, wherein the interface type serves as the target content.

7. The method according to claim 1, characterized in that The image to be selected in the image set to be selected corresponds to image information, and searching the image to be selected in the image set to be selected for matching the target content as the first image includes: Determining a degree of matching between image information of a candidate image in the candidate image set and the target content; The first image is determined based on the matching degree of each of the candidate images.

8. The method according to claim 7, characterized in that The determining the first image based on the matching degree of each of the candidate images includes: Based on the matching degrees of all the candidate images, a preset number of candidate images are selected as first images in descending order.

9. The method according to claim 7, characterized in that The determining the first image based on the matching degree of each of the candidate images includes: Obtain user preference information; Determining a score for the image to be selected based on the matching degree of the image to be selected and the user preference information; A preset number of candidate images are selected as first images in descending order of the score values.

10. The method according to claim 7, characterized in that Determining the degree of matching between the image information of the candidate images in the candidate image set and the target content includes: Determining content description information corresponding to the target content, where the content description information includes at least one of key information and semantic information corresponding to the target content; The matching degree between the image information of the candidate images in the candidate image set and the content description information of the target content is determined to obtain the matching degree of each candidate image.

11. The method according to claim 10, characterized in that The content description information includes semantic information, and determining the content description information corresponding to the target content includes: Acquire context content corresponding to the target content; Based on the target content and the context content, semantic information corresponding to the target content is obtained.

12. An image pushing device, characterized in that: Applied to electronic equipment, the device comprises: an acquiring unit, configured to acquire target content based on a target interface displayed on a screen of the electronic device; a searching unit, configured to search a candidate image matching the target content as a first image in a candidate image set, wherein the candidate image set includes a plurality of candidate images; a pushing unit, configured to display the first image on the screen; a determining unit, configured to determine a second image selected by the user based on the first image; A display unit is configured to display the second image in a designated area of ​​the target interface.

13. An electronic device, characterized in that: include: one or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to execute the method according to any one of claims 1 to 11.

14. A computer-readable medium, characterized in that The computer-readable medium stores a program code executable by a processor, and when the program code is executed by the processor, the processor executes the method according to any one of claims 1 to 11.