Image search method and device
Through the image recognition model, the object area and category are identified, the selection control is generated, and the image segmentation and uploading is performed in response to user operations, which solves the problems of cumbersome user operations and high resource consumption in the prior art, and achieves user experience improvement and resource saving.
Patent Information
- Application Number
- CN201910943975.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-09-30
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2039-09-30
AI Technical Summary
In the prior art, users need to manually adjust the selection box to select items of interest, which is cumbersome and consumes a lot of network resources and back-end service storage resources.
The image recognition model is used to identify the area and category of the item in the image, generate selection controls, and image segmentation and uploading in response to user operations, reducing the difficulty of user operations and optimizing resource utilization.
It simplifies user operations, saves network resources and storage resources for back-end services, and improves user experience.
Smart Images

Figure CN111782848B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, specifically to the field of image processing technology, and more particularly to an image search method and device. Background Art
[0002] In the prior art, a user obtains an image containing an item of interest through a terminal device. If the user wants to obtain relevant information about the item of interest, the terminal device needs to provide a selection box. The user manually adjusts the box to select the item of interest in the image. Then, based on the user's operation, the terminal device sends the image containing the item of interest and the coordinates of the item of interest selected by the user to the server. The server extracts features based on the received information and returns relevant information about the item of interest to the user. Summary of the Invention
[0003] The embodiments of the present application provide an image search method and device.
[0004] In a first aspect, the present application provides an image search method, which includes: in response to acquiring an image to be identified, using an image recognition model to identify the image to be identified to obtain a first image, wherein the first image includes information for indicating an image area where an object contained in the image to be identified is located and an object category; presenting the first image and generating a control for receiving a user selection operation; in response to receiving a user selection operation for an object in the first image, obtaining a second image based on the image area where the object selected by the user is located; uploading the second image and the information on the object category selected by the user to a server; receiving and presenting object information returned by the server that matches the object selected by the user.
[0005] In some embodiments, the control for receiving a user selection operation includes a selection control, and presenting a first image and generating a control for receiving a user selection operation includes: presenting a first image and generating a selection control for receiving a user selection operation corresponding to an image area where an object in the image to be identified is located.
[0006] In some embodiments, in response to acquiring the image to be recognized, performing image recognition on the image to be recognized to obtain the first image includes: in response to acquiring the image to be recognized, calling a MobileNet model to recognize the image to be recognized to obtain the first image.
[0007] In some embodiments, the MobileNet model is obtained by background downloading.
[0008] In some embodiments, in response to receiving a user selection operation on an object in the first image, obtaining a second image based on the image area where the object selected by the user is located includes: in response to receiving a user selection operation on the object in the first image, expanding the image area where the object selected by the user is located outward by a preset number of pixels and then performing image segmentation to obtain the second image.
[0009] In a second aspect, the present application provides an image search device, which includes: an identification module, configured to, in response to obtaining an image to be identified, use an image recognition model to identify the image to be identified to obtain a first image, wherein the first image includes information indicating the image area where the item contained in the image to be identified is located and the item category; a presentation module, configured to present the first image and generate a control for receiving a user selection operation; a selection module, configured to, in response to receiving a user selection operation for an item in the first image, obtain a second image based on the image area where the item selected by the user is located; a sending module, configured to upload the second image and information on the item category selected by the user to a server; and a display module, configured to receive and present item information returned by the server that matches the item selected by the user.
[0010] In some embodiments, the control that receives the user selection operation includes a selection control, and the presentation module is further configured to: present the first image and generate a selection control that receives the user selection operation and corresponds to the image area where the object in the image to be identified is located.
[0011] In some embodiments, the recognition module is further configured to: in response to acquiring an image to be recognized, call a MobileNet model to perform image recognition on the image to be recognized to obtain a first image.
[0012] In some embodiments, the MobileNet model is obtained by background downloading.
[0013] In some embodiments, the selection module is further configured to: in response to receiving a user selection operation on an object in the first image, expand the image area where the user-selected object is located outward by a preset number of pixels and then perform image segmentation to obtain a second image.
[0014] In a third aspect, the present application provides an electronic device comprising one or more processors; a storage device storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement an image search method.
[0015] In a fourth aspect, the present application provides a computer-readable medium having a computer program stored thereon, which implements an image search method when executed by a processor.
[0016] The image search method and device provided by the present application obtain a first image by using an image recognition model to identify the image to be identified in response to obtaining an image to be identified, wherein the first image includes information indicating the image area where the object contained in the image to be identified is located and the category of the object; presents the first image and generates a control for receiving a user selection operation; in response to receiving a user selection operation for an object in the first image, obtains a second image based on the image area where the object selected by the user is located; uploads the second image and the information about the category of the object selected by the user to a server; receives and presents the object information returned by the server that matches the object selected by the user, thereby reducing the difficulty of the user in obtaining information related to the object to be searched, improving the user experience, and saving the user's network resources and the storage resources of the back-end service. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;
[0018] Figure 2 is a flowchart of an embodiment of an image search method according to the present application;
[0019] Figure 3 is a schematic diagram of an application scenario of the image search method according to the present application;
[0020] Figure 4 is a flowchart of another embodiment of the image search method according to the present application;
[0021] Figure 5 is a schematic diagram of an embodiment of an image search device according to the present application;
[0022] Figure 6 It is a structural diagram of a computer system suitable for implementing a server in an embodiment of the present application. DETAILED DESCRIPTION
[0023] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the relevant invention are shown in the accompanying drawings.
[0024] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0025] Figure 1 An exemplary system architecture 100 is shown to which embodiments of the image search method of the present application may be applied.
[0026] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103 running a warehousing method, a network 104, and a server 105. Network 104 is a medium for providing a communication link between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0027] The terminal devices 101, 102, and 103 can be hardware or software. When the terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with camera devices and display devices, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers. When the terminal devices 101, 102, and 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules (for example, software or software modules for providing distributed services), or they can be implemented as a single software or software module. No specific limitation is made here.
[0028] The terminal devices 101 , 102 , and 103 are configured to acquire and identify an image to be identified to obtain a first image and to perform image segmentation in response to receiving a user's selection operation for an item of interest to obtain a second image.
[0029] The terminal devices 101 , 102 , and 103 interact with the server 105 via the network 104 to submit the second image and information about the category of the item selected by the user, and receive and present relevant information about the item selected by the user returned by the server.
[0030] The server 105 may be a server that provides various services, such as a background image processing server that supports the images displayed on the terminal devices 101, 102, and 103. The background image processing server may analyze and process the received images, and feed back the processing results (e.g., searched information related to the item selected by the user) to the terminal device or store them in the server 105.
[0031] It should be noted that the image search method provided in the embodiment of the present application is mainly executed by a terminal device. Accordingly, the image search device is also mainly provided in the terminal device.
[0032] It should be noted that the server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules (e.g., software or software modules for providing distributed services), or as a single software or software module. No specific limitations are given here.
[0033] It should be understood that Figure 1 The number of terminal devices 101, 102, 103, network 104 and server 105 is only illustrative. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0034] Figure 2 A flowchart 200 of an embodiment of an image search method applicable to the present application is shown. The image search method includes the following steps:
[0035] Step 201: In response to obtaining an image to be identified, an image recognition model is used to identify the image to be identified to obtain a first image, wherein the first image includes information indicating an image region where an object contained in the image to be identified is located and an object category.
[0036] In this embodiment, the execution subject of the image search method (eg Figure 1 The terminal devices 101, 102, 103 shown in the figure can obtain the image to be identified by shooting on-site with a camera device, or can directly call the image to be identified from a local storage device, which is not limited in this application.
[0037] In a specific embodiment, the image to be identified may be a picture taken by a camera, or may be a picture found in a computer picture library.
[0038] Here, the execution entity may use an image recognition model from existing technologies or future development technologies to perform image recognition on the image to be recognized to obtain the first image, such as image recognition models such as Inception, SqueezeNet, and MobileNet, which are not limited in this application. The image recognition model may be obtained by training a preset convolutional neural network model using a sample image dataset. Here, the sample image dataset includes training data consisting of object images and training labels consisting of object categories.
[0039] It's important to note that convolutional neural networks primarily consist of convolutional layers, pooling layers, activation layers, and fully connected layers. The convolutional layer is the most important layer in a convolutional neural network, primarily used to extract local features from sample images. A convolution operation is a mathematical operation that multiplies corresponding positions of two matrices and then adds them together. In image processing, the original sample image serves as an input matrix, while the other matrix is the convolution kernel. The kernel is empirically chosen to achieve specific functions, such as smoothing and sharpening.
[0040] The pooling layer is primarily used to retain the main features after the previous convolution, reducing the number of parameters and computation required for the next layer, preventing overfitting, and maintaining certain invariance properties of the sample image, such as translation, rotation, and scaling. Pooling operations mainly include mean pooling and maximum pooling.
[0041] The excitation layer is used to perform nonlinear mapping on the output of the previous layer.
[0042] The fully connected layer is usually located at the end of the convolutional neural network and has weighted connections with the neurons in the previous layer.
[0043] The execution subject uses an image recognition model to recognize the acquired image to be recognized based on the acquired image to be recognized, and outputs the image area where the object in the image to be recognized is located and information about the category of the object.
[0044] In some optional implementations, in step 201 , a MobileNet model may be called to identify the image to be identified to obtain a first image.
[0045] In these optional implementations, the execution entity may use the MobileNet model preset in the APP (Aplication) to recognize the image to be recognized, or may use the MobileNet model downloaded from the background to recognize the image to be recognized. This application does not limit this.
[0046] The MobileNet model is a new deep neural network designed for mobile devices. It replaces traditional convolutional layers with depthwise separable convolutional layers. These layers primarily consist of depthwise and pointwise convolutional layers. Depthwise convolutional layers use a different convolution kernel for each input channel, with one kernel corresponding to each input channel. This applies a single filter to each input channel, whereas standard convolution kernels are applied to all input channels. Pointwise convolutional layers differ from conventional convolutions in that they use a 1x1 kernel. A standard depthwise separable convolution involves filtering each input channel individually through a depthwise convolutional layer, then combining them using a pointwise convolutional layer to form a new output. By splitting the depthwise separable convolutional layer into two layers—one for filtering and one for combining—this avoids the computational overhead of standard convolution operations, where each operation involves combining features from all input channels simultaneously while filtering, thereby generating new features. This significantly reduces computational effort and model size.
[0047] In a specific implementation, MobileNet uses 3x3 depth-separable convolution to reduce the computational complexity by 8-9 times compared to standard convolution, while maintaining the accuracy of image recognition to the greatest extent.
[0048] In addition, it should be noted that based on different platforms, such as Android or iOS, the call to the MobileNet model can be implemented through the mobile terminal Lib (library), that is, a static link library. Among them, the mobile terminal Lib can include the implementation code of various functions.
[0049] In a specific implementation, if the current platform is the Android platform, since the Android platform is mainly implemented based on the Java language, the MobileNet model can be encapsulated as a mobile-side Lib with the suffix .jar for calling; if the current platform is the iOS platform, since the iOS platform is mainly implemented based on the C language, the MobileNet model can be encapsulated as a mobile-side Lib with the suffix .a for calling.
[0050] This implementation, by adopting the MobileNet model, can fully utilize the limited resources of mobile devices and embedded applications, effectively maximize the accuracy of the model, and solve the problem that commonly used image recognition models cannot be applied to scenarios such as mobile or embedded devices that require low latency and fast response speed.
[0051] In some optional methods, the MobileNet model is obtained through background download.
[0052] In this implementation, the MobileNet model is not pre-installed in the app but downloaded in the background, which does not increase the size of the initial package and reduces memory usage.
[0053] Step 202: Present a first image and generate a control for receiving a user selection operation.
[0054] In this embodiment, the execution entity presents a first image and generates a control for receiving a user's selection operation for an item of interest, wherein the control for receiving the user's selection operation can be an encapsulation of data or methods for receiving user selection operations in existing technologies or future development methods, such as a selection control, an input control, etc., and this application does not limit this.
[0055] In some optional implementations, the control for receiving a user selection operation includes a selection control, and presenting the first image and generating a control for receiving a user selection operation includes: presenting the first image and generating a selection control for receiving a user selection operation corresponding to an image area where an object in the image to be identified is located.
[0056] In these implementations, the selection control may be an encapsulation of data or methods for receiving user selection operations in existing technologies or future development technologies, such as buttons, list boxes, drop-down lists, etc., and this application does not limit this.
[0057] In a specific implementation, the selection control may be a button corresponding to the image region where the object to be identified is located. The button may be in a regular shape, such as a square or circle, or may be in an irregular shape that matches the outline of the image region where the identified object is located, and this application does not limit this.
[0058] By setting up selection controls, users can select items of interest simply by clicking, which reduces the difficulty of operation and improves the user experience.
[0059] Step 203 : In response to receiving a user selection operation on an object in the first image, a second image is obtained based on the image region where the object selected by the user is located.
[0060] In this embodiment, in response to receiving a user's selection operation on an object in the first image, the execution entity may segment the image region where the object selected by the user is located by calling an image segmentation model to obtain a second image.
[0061] The image segmentation model may adopt an image segmentation network model in existing technology or future development technology, such as FCN (Fully Convolutional Networks for Semantic Segmentation, deep learning image segmentation), U-net, etc., which is not limited in this application.
[0062] In one specific implementation, a U-net model can be used to segment the image region containing the user-selected item to obtain a second image. The U-net model is a deep convolutional network suitable for low-data-volume image segmentation. It consists of two main paths: a downsampling / encoding path on the left and an upsampling / encoding path on the right. It can run on embedded devices, effectively improving computing speed and reducing device power consumption.
[0063] For example, in one application scenario, the first image obtained by the executing entity includes items such as a desktop fan, a water cup, and a TV. Among them, the item that the user is interested in is the water cup. After receiving the user's selection operation for the water cup, the executing entity calls an image segmentation model, such as U-net, to segment the image area where the water cup is located, and obtains an image that only includes the image area where the water cup is located as the second image.
[0064] Step 204: Upload the second image and the information of the item category selected by the user to the server.
[0065] In this embodiment, the execution entity uploads the second image and the information of the item category selected by the user to the server via a wired or wireless method.
[0066] The wireless connection methods may include but are not limited to 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or to be developed in the future.
[0067] In one application scenario, if the user-selected item included in the second image is a desktop fan, the item category information is fan. Through the user's selection operation on the image area where the desktop fan is located, the execution entity can send the second image including the image area where the desktop fan is located and the category information fan of the desktop fan to the server.
[0068] Step 205: Receive and present the item information returned by the server that matches the item selected by the user.
[0069] In this embodiment, after the server receives the second image including the image area where the item selected by the user is located and the item category information selected by the user, it extracts the image features of the item, that is, extracts the feature vector of the item, where the feature vector may include feature vectors such as color, texture, and shape.
[0070] The server searches the database for information related to the item based on the item's image features and item category information. This information may include the item's description, price, a hyperlink to a purchase address, and the like.
[0071] Finally, the server sends the found matching images and information about related products to the execution entity, which displays the information after receiving it.
[0072] Continue to see Figure 3 , Figure 3 FIG. 4 is a schematic diagram of an application scenario of the image search method according to this embodiment.
[0073] exist Figure 3In an application scenario, user 301 can obtain an image to be identified through terminal device 302. For example, the image can be captured by a camera on terminal device 302, or it can be an image found in a computer image library. The image to be identified can include multiple items, for example, a mobile phone, a laptop computer, and a printer. In response to receiving the image to be identified, terminal device 302 identifies the image to be identified and obtains a first image 303. First image 303 includes the image region where the identified items are located and information about the item categories. For example, first image 303 may include image regions of a mobile phone, a laptop computer, and a printer, as well as information about the categories of the three items: mobile phone, computer, and printer. Terminal device 302 presents the identified first image 303 and generates a control for receiving a user selection operation. Here, the control for receiving a user click operation can be a button corresponding to the image region where the identified item is located. If the item of interest to user 301 is a mobile phone, upon receiving the user's click operation on the mobile phone, terminal device 302 invokes an image segmentation model, for example, a U-net model, to segment the image and obtain a second image 304 based on the image region where the mobile phone is located. The terminal device 302 may send the second image 304 and the category information of the mobile phone to the server 305 , and present the image after receiving the information returned by the server 305 .
[0074] The image search method provided by the embodiments of the present disclosure obtains an image to be identified by performing image recognition on the image to be identified to obtain a first image, wherein the first image includes information indicating an image region where an identified object contained in the image to be identified is located and an object category, presents the first image and generates a control for receiving a user's selection operation on the identified object, and in response to receiving a user's selection operation on the identified object in the first image, calls an image segmentation model to segment the image region where the object selected by the user is located to obtain a second image, uploads the second image and the information on the object category selected by the user to a server, receives and presents object information returned by the server that matches the object selected by the user, reduces the difficulty of the user in obtaining information related to the object to be searched, improves the user experience, and saves the user's network resources and storage resources of the back-end service.
[0075] Further references Figure 4 , which shows a process 400 of another embodiment of an image search method. The process 400 of the image search method includes the following steps:
[0076] Step 401: In response to obtaining an image to be identified, an image recognition model is used to identify the image to be identified to obtain a first image, wherein the first image includes information indicating an image region where an object contained in the image to be identified is located and an object category.
[0077] In this embodiment, the implementation details and technical effects of step 401 can be found in the description of step 201 and will not be repeated here.
[0078] Step 402: Present a first image and generate a control for receiving a user selection operation.
[0079] In this embodiment, the implementation details and technical effects of step 402 can be found in the description of step 202 and will not be repeated here.
[0080] Step 403 : In response to receiving a user selection operation for an object in the first image, the image region where the object selected by the user is located is expanded outward by a preset number of pixels and then image segmented to obtain a second image.
[0081] In this embodiment, the execution entity can expand outward by a preset number of pixels, such as 20, 30, etc., based on the first image area where the item selected by the user is located, to expand the image area where the item selected by the user is located, obtain the second image area where the item selected by the user is located, and then call the image segmentation model to perform image segmentation to obtain the second image.
[0082] For example, in one application scenario, the U-net image segmentation model is called to segment the image area where the user-selected item is located. Since the convolutional layer in the U-net network will cause a certain degree of edge information loss when processing the image, the input image needs to be expanded to a certain extent.
[0083] If the image area where the object is located, identified based on the object recognition model, is marked with a square frame, then after receiving the user's selection operation for the object of interest, the execution entity first determines the four corner coordinates of the square frame marking the user's object of interest, wherein the image area contained in the square frame of the user's object of interest is the first image area; then the four corner coordinates are expanded outward by twenty pixels to obtain a new square frame, wherein the image area contained in the new square frame is the second image area; finally, the U-net image segmentation model is called for segmentation, and the second image area is used as the second image.
[0084] Step 404: Upload the second image and the information of the item category selected by the user to the server.
[0085] In this embodiment, the implementation details and technical effects of step 404 can be found in the description of step 204 and will not be repeated here.
[0086] Step 405: Receive and present the item information returned by the server that matches the item selected by the user.
[0087] In this embodiment, the implementation details and technical effects of step 405 can be found in the description of step 205 and will not be repeated here.
[0088] The above-mentioned embodiment of the present application obtains a first image by acquiring an image to be identified and identifying the image to be identified, presents the first image and generates a control for receiving a user selection operation, expands the image area where the item selected by the user is located outward by a preset number of pixels to obtain a second image, and uploads the second image and the category information of the item selected by the user to the server, so that the execution entity can better utilize the image segmentation model to process edge pixels, and effectively avoids the loss of key information caused by deviation when the image recognition model identifies the image area where the item is located.
[0089] Further references Figure 5 As an implementation of the methods shown in the above figures, the present application provides an embodiment of an image search device, which is similar to Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0090] like Figure 5 As shown, the image search device 500 of this embodiment includes: a recognition module 501, a presentation module 502, a selection module 503, a sending module 504, and a display module 505. The recognition module 501 is configured to, in response to acquiring an image to be recognized, perform image recognition on the image to be recognized using an image recognition model to obtain a first image, wherein the first image includes information indicating an image region where an item contained in the image to be recognized is located and an item category; the presentation module 502 is configured to present the first image and generate a control for receiving a user selection operation; the selection module 503 is configured to, in response to receiving a user selection operation for an item in the first image, obtain a second image based on the image region where the item selected by the user is located; the sending module 504 is configured to upload the second image and the item category information selected by the user to a server; and the receiving and presentation module 505 is configured to receive and present item information returned by the server that matches the item selected by the user.
[0091] In this embodiment, the recognition module 501 in the image search apparatus 500 may use an image recognition model in existing technology or future development technology to perform image recognition on the image to be recognized to obtain the first image.
[0092] In this embodiment, the above-mentioned presentation module 502 presents the first image and generates a control for receiving the user's selection operation on the identified object, wherein the control for receiving the user's selection operation can be an encapsulation of data or methods for receiving user selection operations in existing technologies or future development methods, such as a selection control, an input control, etc.
[0093] In this embodiment, in response to receiving a user's selection operation on an object in the first image, the selection module 503 obtains the second image by image segmentation based on the image region where the object selected by the user is located.
[0094] In this embodiment, the sending module 504 uploads the second image and the information of the item category selected by the user to the server via wired or wireless means.
[0095] In this embodiment, the receiving and presenting module 505 receives and presents the item information that matches the item selected by the user and is returned by the server based on the second image and the information of the item category selected by the user.
[0096] In some optional implementations of this embodiment, the recognition module 603 is further configured to, in response to acquiring an image to be recognized, call a MobileNet model to perform image recognition on the image to be recognized to obtain a first image.
[0097] In some optional implementations of this embodiment, the control that receives the user selection operation includes a selection control, and the presentation module is further configured to: present the first image and generate a selection control that receives the user selection operation and corresponds to the image area where the object in the image to be identified is located.
[0098] Reference below Figure 6 , which shows a structural diagram of a computer system 600 suitable for implementing a client device or server of an embodiment of the present application.
[0099] like Figure 6 As shown, the computer system 600 includes a processor (e.g., a central processing unit (CPU)) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage portion 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the system 600 are also stored in the RAM 603. The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0100] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, and the like; an output section 607 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a LAN card or a modem. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 610 as needed, so that computer programs read therefrom can be installed into the storage section 608 as needed.
[0101] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program comprising program code for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication portion 609 and / or installed from a removable medium 611.
[0102] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0103] The modules involved in the embodiments described in this application can be implemented by software or hardware. The modules described can also be set in a processor. For example, it can be described as: a processor includes a recognition module, a presentation module, a segmentation module, a sending module and a receiving and display module. Among them, the names of these modules do not constitute a limitation of the modules themselves in some cases. For example, the recognition module can also be described as "a module that recognizes the image to be recognized to obtain a first image in response to obtaining the image to be recognized."
[0104] As another aspect, the present application also provides a non-volatile computer storage medium, which may be the non-volatile computer storage medium included in the apparatus of the above embodiment, or a standalone non-volatile computer storage medium not incorporated into a client device. The non-volatile computer storage medium stores one or more programs, which, when executed by a device, causes the device to: in response to acquiring an image to be recognized, perform image recognition on the image to be recognized to obtain a first image, wherein the first image includes information indicating an image region where a recognized object contained in the image to be recognized is located and an object category; present the first image and generate a control for receiving a user's selection operation for the recognized object; in response to receiving a user's selection operation for the recognized object in the first image, invoke an image segmentation model to segment the image region where the user-selected object is located to obtain a second image; upload the second image and the information about the object category selected by the user to a server; and receive and present object information matching the user-selected object returned by the server.
[0105] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention herein is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features having similar functions disclosed in this application.
Claims
1. An image search method, comprising: In response to acquiring an image to be identified, using an image recognition model to identify the image to be identified to obtain a first image, wherein the first image includes information indicating an image region where each of at least two items included in the image to be identified is located and an item category; Presenting the first image and generating a selection control corresponding to an image region where an object in the image to be identified is located for receiving a user selection operation; In response to receiving a user selection operation of an object in the first image through the selection control, obtaining a second image based on the image area where the object selected by the user is located; Uploading the second image and information about the item category selected by the user to a server; Receive and present item information returned by the server that matches the item selected by the user.
2. The method according to claim 1, wherein, in response to acquiring the image to be recognized, performing image recognition on the image to be recognized to obtain the first image comprises: In response to acquiring the image to be recognized, calling the MobileNet model to recognize the image to be recognized to obtain a first image.
3. According to the method of claim 2, the MobileNet model is obtained by background downloading.
4. The method according to claim 1, wherein, in response to receiving a user selection operation on an object in the first image, obtaining a second image based on the image region where the object selected by the user is located comprises: In response to receiving a user selection operation on an object in the first image, the image area where the object selected by the user is located is expanded outward by a preset number of pixels and then image segmentation is performed to obtain a second image.
5. An image search device, comprising: a recognition module configured to, in response to acquiring an image to be recognized, recognize the image to be recognized using an image recognition model to obtain a first image, wherein the first image includes information indicating an image region where each of at least two objects included in the image to be recognized is located and an object category; a presentation module configured to present the first image and generate a selection control corresponding to an image region where an object in the image to be identified is located for receiving a user selection operation; a selection module configured to, in response to receiving a user selection operation of an object in the first image through the selection control, obtain a second image based on the image area where the object selected by the user is located; a sending module configured to upload the second image and information about the item category selected by the user to a server; The display module is configured to receive and present item information returned by the server that matches the item selected by the user.
6. The apparatus according to claim 5, wherein the identification module is further configured to: In response to acquiring the image to be recognized, calling the MobileNet model to perform image recognition on the image to be recognized to obtain a first image.
7. The device according to claim 6, wherein the MobileNet model is obtained by background downloading.
8. The apparatus according to claim 5, wherein the selection module is further configured to: In response to receiving a user selection operation on an object in the first image, the image area where the object selected by the user is located is expanded outward by a preset number of pixels and then image segmentation is performed to obtain a second image.
9. An electronic device comprising: one or more processors; A storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.
10. A computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Interactive type image search system and method
CN101216841A
Detection method and device for objects in images
CN106780612A