Search method for target object and search method for image

CN115408562BActive Publication Date: 2026-09-08ALIBABA INNOVATION PRIVATE LIMITED
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110606200.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-26
Publication Date
2026-09-08
Estimated Expiration
2041-05-26

AI Technical Summary

Technical Problem

[0005]本发明实施例提供了一种目标对象的搜索方法和图像的搜索方法,以至少解决现有技术中电商平台难以实现定制化的图像搜索的技术问题

Benefits of technology

[0011] In this embodiment of the invention, by displaying original multimedia content on an interactive interface, wherein the original multimedia content includes at least one target object to be edited, responding to editing instructions sensed in the interactive interface, determining the editing interaction mode corresponding to the editing instructions, and generating an edited target image according to the determined editing interaction mode, responding to search instructions, performing a search based on the target image, searching for at least one search result matching the target image from the image library, and displaying at least one search result on the interactive interface, the invention enables users to edit the searched image when using images for searching, and to perform customized image searches based on the edited image, thus solving the technical problem of difficulty in achieving customized image searches in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115408562B_ABST
    Figure CN115408562B_ABST
Patent Text Reader

Abstract

The application discloses a search method of a target object and a search method of an image. The method comprises the following steps: displaying original multimedia content on an interactive interface, wherein the original multimedia content comprises at least one target object to be edited; in response to an editing instruction sensed in the interactive interface, determining an editing interactive mode corresponding to the editing instruction, and generating an edited target image according to the determined editing interactive mode, wherein the editing instruction is used for editing display content in at least one display area in the target object, and the interactive interface provides multiple editing interactive modes; in response to a search instruction, searching at least one search result matched with the target image from an image library based on the target image; and displaying the at least one search result on the interactive interface. The application solves the technical problem that it is difficult to realize customized image search in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically, to a method for searching target objects and a method for searching images. Background Technology

[0002] In e-commerce platforms, users can search for similar products by uploading product images. Image search can be achieved through vectorized retrieval based on image visual content, which involves converting image content into vectorized features through a model and establishing an image feature library. Image retrieval is then achieved through image feature matching. However, this image retrieval method based on visual content and vectorized features only supports fixed image inputs. Users cannot achieve customized image retrieval. For example, a user might search for a pair of similar shoes by image, but only like the pattern and upper design of the shoes and want to modify the stiletto heel in the image to a chunky heel. Existing vectorized retrieval based on image visual content cannot directly achieve customized image retrieval.

[0003] In related technologies, to achieve customized and interactive image retrieval, a candidate set of image feature libraries can be retrieved first using image search, and then the desired text content can be input for secondary filtering. However, this method relies on the construction of labels in the image feature library, and can only be recalled if there are corresponding text descriptions in the image feature library, resulting in a poor user search experience.

[0004] There is currently no effective solution to the technical problem of difficulty in achieving customized image search in the existing technologies mentioned above. Summary of the Invention

[0005] This invention provides a method for searching target objects and a method for searching images, in order to at least solve the technical problem that e-commerce platforms in the prior art find it difficult to achieve customized image searches.

[0006] According to one aspect of the present invention, a method for searching a target object is provided, comprising: displaying original multimedia content on an interactive interface, wherein the original multimedia content includes at least one target object to be edited; responding to an editing instruction sensed in the interactive interface, determining an editing interaction mode corresponding to the editing instruction, and generating an edited target image according to the determined editing interaction mode, wherein the editing instruction is used to edit display content in at least one display area of ​​the target object, and the interactive interface provides multiple editing interaction modes; responding to a search instruction, performing a retrieval based on the target image, and searching for at least one search result matching the target image from an image library; and displaying at least one search result on the interactive interface.

[0007] According to one aspect of the present invention, an image search method is provided, comprising: displaying a live video to be identified in a live streaming interface, wherein the live video is used to play product information, the live video is composed of multiple frames of images, and at least one frame of images displays a product; extracting the frame images displaying the product to obtain original multimedia content, wherein the product displayed in the original multimedia content is the target object to be edited; generating an edited target image in response to an editing instruction sensed in the live streaming interface, wherein the editing instruction is used to edit the display content in at least one display area of ​​the target object; performing a search based on the target image in response to a search instruction, and searching for at least one search result matching the target image from a video library; and displaying at least one search result on the live streaming interface.

[0008] According to one aspect of the present invention, an image search method is provided, comprising: playing a teaching video in a teaching interface and obtaining original courseware content to be edited in the teaching video, wherein the original courseware content displays different types of teaching content; generating an edited target courseware in response to an editing instruction sensed in the teaching interface, wherein the editing instruction is used to edit courseware content in at least one display area of ​​the original courseware content; performing a search based on the target courseware in response to a search instruction, and searching for at least one search result matching the target courseware from a courseware library; and displaying at least one search result on the teaching interface.

[0009] According to one aspect of the present invention, an image search method is provided, comprising: acquiring a physiological video of a site to be examined using a medical device; displaying the physiological video on an examination interface of the medical device, wherein the physiological video displays at least one lesion object to be edited; generating an edited target pathological image in response to an editing instruction sensed in the examination interface of the medical device, wherein the editing instruction is used to edit the display content in at least one display area of ​​the lesion object; performing a search based on the target pathological image in response to a search instruction, and searching for at least one search result matching the target pathological image from an image library; filling the search result into text recording case information to obtain structured filled case text; and displaying the structured filled case text on the examination interface.

[0010] According to one aspect of the present invention, a method for searching a target object is provided, comprising: a cloud server receiving an editing message from a client, wherein the editing message carries identification information characterizing a video to be identified; the cloud server obtaining original multimedia content based on the identification information, wherein the original multimedia content includes at least one target object to be edited; the cloud server responding to an editing instruction sensed in an interactive interface and generating an edited target image, wherein the editing instruction is used to edit the content of at least one part of the target object; the cloud server responding to a search instruction and performing a retrieval based on the target image, searching for at least one search result matching the target image from an image library; and the cloud server feeding back at least one search result to the client, wherein the client displays at least one search result on the interactive interface.

[0011] In this embodiment of the invention, by displaying original multimedia content on an interactive interface, wherein the original multimedia content includes at least one target object to be edited, responding to editing instructions sensed in the interactive interface, determining the editing interaction mode corresponding to the editing instructions, and generating an edited target image according to the determined editing interaction mode, responding to search instructions, performing a search based on the target image, searching for at least one search result matching the target image from the image library, and displaying at least one search result on the interactive interface, the invention enables users to edit the searched image when using images for searching, and to perform customized image searches based on the edited image, thus solving the technical problem of difficulty in achieving customized image searches in the prior art. Attached Figure Description

[0012] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0013] Figure 1 A hardware structure block diagram of a computing device for implementing a search method for a target object is shown;

[0014] Figure 2 A flowchart is provided for a method of searching for a target object according to an embodiment of this application;

[0015] Figure 3 A schematic diagram of an optional target object search method according to an embodiment of this application is provided;

[0016] Figure 4 A schematic diagram of an optional editing interaction mode according to an embodiment of this application is provided;

[0017] Figure 5 A schematic diagram of an optional editing interaction mode according to an embodiment of this application is provided;

[0018] Figure 6 A schematic diagram of an optional editing interaction mode according to an embodiment of this application is provided;

[0019] Figure 7 A schematic diagram of an optional editing interaction mode according to an embodiment of this application is provided;

[0020] Figure 8 A schematic diagram of an optional editing interaction mode according to an embodiment of this application is provided;

[0021] Figure 9 A schematic diagram of an optional image generation algorithm according to an embodiment of this application is provided;

[0022] Figure 10 A flowchart of an image search method according to an embodiment of this application is provided;

[0023] Figure 11 A flowchart of an image search method according to an embodiment of this application is provided;

[0024] Figure 12 A flowchart of an image search method according to an embodiment of this application is provided;

[0025] Figure 13 A flowchart is provided for a method of searching for a target object according to an embodiment of this application;

[0026] Figure 14 A schematic diagram of a target object search device according to an embodiment of this application is provided;

[0027] Figure 15 A schematic diagram of an image search device according to an embodiment of this application is provided;

[0028] Figure 16 A schematic diagram of an image search device according to an embodiment of this application is provided;

[0029] Figure 17 A schematic diagram of an image search device according to an embodiment of this application is provided;

[0030] Figure 18 A schematic diagram of a target object search device according to an embodiment of this application is provided;

[0031] Figure 19 This is a structural block diagram of a computer terminal according to an embodiment of this application;

[0032] Figure 20 This is a schematic diagram of an optional target object search method according to an embodiment of this application. Detailed Implementation

[0033] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0035] Example 1

[0036] According to an embodiment of the present invention, a method embodiment for searching a target object is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0037] The method embodiment provided in Embodiment 1 of this application can be executed in a mobile terminal, computing device or similar computing device. Figure 1 A hardware block diagram of a computing device (or mobile device) for implementing a method for searching for target objects is shown. Figure 1 As shown, the computing device 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computing device 10 may also include a... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0038] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computing device 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as processor control (e.g., selection of a variable resistor termination path connected to an interface).

[0039] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the target object search method in this embodiment of the invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby implementing the above-mentioned application vulnerability detection method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computing device 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0040] The transmission module 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computing device 10. In one example, the transmission module 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission module 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0041] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computing device 10 (or mobile device).

[0042] It should be noted here that, in some embodiments, the above... Figure 1The computing device (or mobile device) shown has a touch display (also referred to as a "touchscreen" or "touch display screen"). In some embodiments, the above... Figure 1 The computing device (or mobile device) shown has a graphical user interface (GUI), through which users can interact with the GUI by touching and / or gesturing on a touch-sensitive surface. The human-computer interaction functions here may include the following interactions: creating web pages, drawing, word processing, creating electronic documents, playing games, video conferencing, instant messaging, sending and receiving emails, call interface, playing digital video, playing digital music, and / or web browsing, etc. Executable instructions for performing the above human-computer interaction functions are configured / stored in one or more processor-executable computer program products or readable storage media.

[0043] Figure 2 This is a flowchart of a target object search method according to Embodiment 1 of the present invention, as follows: Figure 2 As shown, the method includes:

[0044] Step S202: Display the original multimedia content on the interactive interface, wherein the original multimedia content includes at least one target object to be edited.

[0045] The aforementioned interactive interface can be the human-computer interaction interface displayed on the client side of a smart device (such as a mobile phone, smart tablet, and computer). Specifically, the smart device has a touch screen, and users can interact with the smart device through touch operation on the touch screen. The original multimedia content can be displayed on the touch screen interface of the smart device.

[0046] The aforementioned original multimedia content includes, but is not limited to, images, videos, and text files, and includes the target object to be edited. The original multimedia content can be images uploaded by the user through the interactive interface. For example, in e-commerce shopping applications, users can upload images to search for through the interactive interface of the e-commerce shopping platform's client; the images contained in the searched images constitute the original multimedia content. Original multimedia content can also be images selected by the user from displayed images in the interactive interface. For example, when browsing product images in the interactive interface of an e-commerce shopping platform's client, a user can select an image as the image to search for. Original multimedia content can also be videos uploaded or selected by the user, or a frame from a video. For example, in e-commerce shopping applications, a video containing products can be directly used as the original multimedia content, or a frame from the video can be selected as the original multimedia content.

[0047] The target object mentioned above can be an image in the original multimedia content that needs to be edited. Specifically, the target object can be a complete image in the original multimedia content, or it can be a part of the original multimedia content. For example, in the interactive interface of an e-commerce shopping platform's client, when a user uploads an image of a product to be searched, the target object can be the complete product image in the image, or it can be a part of the product image.

[0048] Step S204: Respond to the editing command sensed in the interactive interface, determine the editing interaction mode corresponding to the editing command, and generate the edited target image according to the determined editing interaction mode. The editing command is used to edit the display content in at least one display area of ​​the target object, and the interactive interface provides multiple editing interaction modes.

[0049] Specifically, the aforementioned editing commands can be triggered by the user through touch operations in the interactive interface. For example, the user can click the "Edit" button in the interactive interface to trigger the editing command, or the user can directly perform touch operations on the original multimedia content according to preset gestures to trigger the editing command.

[0050] The target object can be divided into multiple display areas, each with different display content. Users can edit the content displayed in different display areas to modify a specific component of the target object. For example, in the interactive interface of an e-commerce shopping platform's client, if the target object in the original multimedia content is a shoe, then the shoe has multiple display areas, such as the heel area, the front surface area, and the side area. Users can modify the content displayed in at least one of these areas.

[0051] It should be noted that the same display area within a target object can include multiple modifiable display contents. For example, if the target object is a shoe, the display contents in the display area on the front surface of the shoe can include the shoe's color, the pattern on the shoe's surface, etc. Users can edit multiple display contents within a single display area at the same time.

[0052] The above-mentioned editing interaction mode is a way to edit the displayed content of the target object. The image editing interface can have a variety of editing interaction modes. For example, the displayed content can be modified by editing text attributes, or by using template replacement to modify the displayed content.

[0053] The target image mentioned above is a new image regenerated from the original multimedia content after editing, and can be used for image search. The target image can be the entire image of the modified target object, or it can be only a modified portion of the target object. For example, if the original multimedia content is an image containing shoes, and the user edits the heel, changing a stiletto heel to a chunky heel, the resulting image of a shoe with a chunky heel is the target image.

[0054] Step S206: In response to the search command, perform a retrieval based on the target image and search the image library for at least one search result that matches the target image.

[0055] The search command can be triggered by the user through touch operations in the interactive interface, for example, by clicking the "Search" button in the interactive interface. The search results can be a list of objects that match the modified object in the target image.

[0056] In one optional embodiment, the image library is an image feature library built based on the vectorized features of the image. After obtaining the target image edited by the user, the vectorized features of the target image can be extracted and matched with the vectorized features in the image feature library to obtain the search results.

[0057] Step S208: Display at least one search result on the interactive interface.

[0058] Specifically, search results can be displayed in a list format on the interactive interface. For example, in an e-commerce shopping platform application, the interactive interface includes the client interface of the e-commerce platform, where users can select products to purchase from the search results. In an optional embodiment, search results can be sorted and displayed according to their degree of matching with the target image.

[0059] In one alternative embodiment, Figure 3 A schematic diagram of an optional image search method according to an embodiment of this application is provided, such as... Figure 3 As shown, the interactive interface is the client-side interface of the e-commerce platform. The aforementioned original multimedia content can be the original image containing the product. Customized image search methods for users include: users can take photos of the product to obtain an image containing the original image and upload it to the e-commerce platform. The client-side interactive interface includes an original image display interface 31, which can display the original image uploaded by the user. In the client-side interactive interface of the e-commerce platform, users can directly use the original image for searching, or click the "Edit" button to enter the image editing interface 32 to edit certain parts of the target object in the original image (e.g., ...). Figure 3As shown, the target object is a shoe (the heel can be edited). After editing, click the "Generate" button to generate the edited target image. The target image is previewed in preview interface 33, and an image search is performed based on the target image. The search results are displayed on the display interface (i.e., image search results display interface 34), where users can browse and purchase products. If the current search results do not meet the user's needs, they can return to the image editing interface 32 to continue editing the image and repeat the above image search process.

[0060] It should be noted that the original multimedia content can be multimedia content from any application field. For example, the original multimedia content mentioned above can be courseware content in teaching videos, physiological videos or images collected by medical equipment, product shopping videos in e-commerce live streaming videos, and film and television videos in cultural and entertainment applications. For example, when the original multimedia content is a film and television video played on an interactive interface, the target object can be the clothing, accessories, or props of the characters in the film and television video. When watching the film and television video, the user can obtain an image containing the target object by pausing or taking a screenshot and edit it. The edited image can be used for searching, and the search results will be returned to the user.

[0061] In this embodiment, by displaying original multimedia content on an interactive interface, wherein the original multimedia content includes at least one target object to be edited, responding to editing instructions sensed in the interactive interface, determining the editing interaction mode corresponding to the editing instructions, and generating an edited target image according to the determined editing interaction mode, responding to search instructions, performing a search based on the target image, searching for at least one search result matching the target image from the image library, and displaying at least one search result on the interactive interface, the user can edit the searched image when using image search, and perform customized image search based on the edited image, solving the technical problem of difficulty in achieving customized image search in the prior art.

[0062] As an optional embodiment, generating an edited target image in response to an editing command sensed in the interactive interface includes: generating an editing command by triggering an editing function and entering an image editing interface on the interactive interface; executing at least one editing interaction mode in the image editing interface based on the editing command, wherein the editing interaction mode is used to determine an editing method for editing the display content of at least one display area of ​​the target object in the original multimedia content; and editing the display content of at least one display area of ​​the target object in the original multimedia content using the editing interaction mode to generate the edited target image.

[0063] The editing function can be triggered by the user's touch operation on the interactive interface. For example, after uploading the original multimedia content on the interactive interface, the user can generate an image editing interface by clicking the "Edit" button on the interactive interface.

[0064] In one optional embodiment, an image generation algorithm can be invoked to render the edited content obtained through an interactive editing mode. This algorithm generates a target image based on the original multimedia content and the edited result. The rendered target image is more realistic and plausible than an image directly edited by the user, providing a more authentic visual experience. For example, in an e-commerce shopping platform application where the target product is shoes, a user edits the heel of the shoes in the image editing interface, changing a thin heel to a thick heel. By invoking the image generation algorithm, a realistic image of the shoes with a thick heel is generated, providing the user with a better visual experience.

[0065] As an optional embodiment, after generating the edited target image, the method further includes: triggering a preview function to preview the edited target image and obtain a preview result; receiving a re-editing instruction and editing the preview result according to the re-editing instruction to generate a re-edited target image; responding to a re-search instruction, performing a search based on the re-edited target image, and searching the image library for at least one search result that matches the re-edited target image; and displaying at least one search result on the interactive interface.

[0066] Specifically, the preview function can be triggered by user touch operations on the interactive interface. For example, clicking the "Generate" button generates a preview interface that displays the generated target image. For example, ... Figure 3 As shown, in the client display interface of the e-commerce platform, after the user completes the editing of the target object in the image editing interface 32, the user can generate a preview interface 33 by clicking the "Generate" button. The "Preview" button generates a preview interface to display the edited target image.

[0067] By previewing the edited target image, users can intuitively judge whether they are satisfied with the editing result. If they are not satisfied with the preview result, they can return to the step of editing the target object and continue to edit and search the target object multiple times to make the search results more in line with the user's actual needs.

[0068] The aforementioned re-editing command can be triggered in the same way as the aforementioned editing command, and the aforementioned re-search command can be triggered in the same way as the aforementioned search command. For example, the preview interface can display the edited target image. Users can click the "Edit" or "Return" button in the preview interface to re-enter the image editing interface to edit the image and generate a new target image for preview. Users can then search based on the new target image to obtain search results. If the search results are still unsatisfactory, users can return to the image editing interface to edit and search the modified image again.

[0069] As an optional embodiment, editing the display content of at least one display area of ​​a target object in the original multimedia content using an editing interaction mode includes: obtaining multiple image attributes of the target object, wherein the image attributes characterize the display content of the target object in the corresponding display area displayed on the interactive interface; using the editing interaction mode, determining the editable image attributes among the multiple image attributes; editing the attribute values ​​of the editable image attributes to generate edited image attributes, wherein the edited image attributes constitute the editing result.

[0070] Multiple image attributes can include editable and non-editable image attributes. After obtaining multiple image attributes, it is necessary to determine the editable image attributes. The attribute values ​​mentioned above are the specific content contained in the editable image attributes. For example, if the target object is a shoe, the image attribute can be color, and the attribute value can be a specific color such as black or red.

[0071] Figure 4 A schematic diagram of an optional editing interaction mode according to an embodiment of this application is provided, such as... Figure 4 As shown, in the display interface of the e-commerce shopping platform, the user uploads original multimedia content 41. The target object in the original multimedia content 41 is a pair of shoes. Multiple original multimedia content attributes 43 of the shoes are obtained through an attribute recognition algorithm: toe style (pointed toe), heel style (stiletto heel), opening depth (shallow opening), and color (black). Among these, toe style, heel style, opening depth, and color are all editable image attributes. The image attributes also include multiple candidate attribute values. For example, for toe style, candidate attribute values ​​include "closed toe," "toe loop," and "peep toe." The user can edit the original image attribute value "pointed toe" to any one of the multiple candidate attribute values ​​"closed toe," "toe loop," and "peep toe." The user can modify the attribute values ​​of the image attributes through touch operations on the interactive interface to obtain the modified image attributes 44. For example, as shown... Figure 4As shown, the "thin heel" selected by the box in "Heel Style" is the original image attribute value. When the user wants to modify the "Heel Style," they can select the candidate attribute value "thick heel" to complete the modification of the image attribute. The modified "Heel Style" attribute value box will then show "thick heel," which is the modified attribute value. The user can then click the "Generate" button to generate the edited target image 42 using an image generation algorithm. Furthermore, users can combine and modify multiple image attributes; for example, they can simultaneously modify the toe style and heel style to edit multiple image attributes.

[0072] It should be noted that, Figure 4 The display effect and touch operation of the image attribute editing method in the editing interaction mode are for illustrative purposes only. The actual display effect and touch operation may be different, and there are no restrictions here.

[0073] In an optional embodiment, before obtaining multiple image attributes of the target object, the method further includes: using an image recognition algorithm to identify the original multimedia content and identify multiple image attributes of the target object in the original multimedia content, wherein a deep learning algorithm is used to train image samples to generate the image recognition algorithm.

[0074] Image samples can include original multimedia content and corresponding annotations. For example, in e-commerce shopping platform applications, product images and annotations of product image attributes can be used as image samples. The initial model of the image recognition algorithm is trained using a deep learning algorithm to obtain the trained image recognition algorithm.

[0075] Image recognition algorithms can output multiple corresponding image attributes based on the original multimedia content provided by the user. For example, ... Figure 4 As shown, in the display interface of the e-commerce shopping platform, the original multimedia content can be the original image 41 of the product. The target object in the original image 41 uploaded by the user is the product shoe. Through the image recognition algorithm, multiple original image attributes 43 of the shoe and their corresponding attribute values ​​are identified: the toe style is pointed, the heel style is thin heel, the opening depth is shallow, and the color is black. The identified multiple original image attributes 43 are displayed on the display interface. Users can edit the attribute values ​​of the original image attributes 43 through touch operation on the display interface.

[0076] In this embodiment, the original multimedia content is identified by an image recognition algorithm and the attribute values ​​of multiple image attributes are automatically output. Users do not need to manually input or select the attribute values ​​of the image attributes. On the one hand, this can improve the user's operating experience, and on the other hand, it can avoid inaccurate determination of the original multimedia content attributes due to different user perceptions.

[0077] As an optional embodiment, editing the display content of at least one display area of ​​a target object in the original multimedia content using an editing interaction mode includes: using the editing interaction mode, determining the content to be edited selected in the image editing interface, wherein the content to be edited is at least one image attribute contained in the target object in the original multimedia content, and the image attribute characterizes the display content of the target object in the corresponding display area displayed on the interaction interface; replacing the content to be edited with selected template content, wherein the template content is displayed in the template library in the image editing interface; and generating the edited image attributes based on the replacement result, wherein the edited image attributes constitute the editing result.

[0078] In the image editing interface, the selected content to be edited corresponds to at least one display area of ​​the target object. Users can select the display content of a certain display area of ​​the target object as the content to be edited through touch operation, so as to edit one or more image attributes.

[0079] It should be noted that the template library contains pre-stored template content corresponding to image attributes. This allows users to select multiple templates in the image editing interface after selecting the content to be edited. For example, if the target object is a pair of shoes, the image attributes of the shoes can include the toe style and the heel style. The template library stores multiple toe templates corresponding to the toe styles and multiple heel templates corresponding to the heel styles.

[0080] Figure 5 A schematic diagram of an optional editing interaction mode according to an embodiment of this application is provided, such as... Figure 5 As shown, in the application of an e-commerce shopping platform, the original multimedia content can be the original image of the product, and the aforementioned interactive interface can be an image editing interface 500. The user uploads an original image 51, and the target object in the original image 51 is a shoe. The user selects the heel portion of the shoe through touch operation as the content to be edited 511. "The heel style is a stiletto heel" is one of the image attributes of the shoe. Based on the selected heel image, multiple template contents corresponding to the heel are retrieved from the template library and displayed in the image editing interface 500, such as... Figure 5 As shown, the template content includes two areas: "Template Selection" and "Template Fine-tuning". The "Template Selection" area displays multiple heel templates for selection. By selecting a heel template image in the "Template Selection" area, the user can replace the heel image in the original image 51. For example, the thin heel in the original image 51 can be replaced with a block heel, and the image attribute corresponding to the heel style in the original image 51 can be modified to "block heel". Based on the editing result of "block heel", a target image of a shoe with "block heel" is generated, thus realizing the editing of the original image.

[0081] In one optional embodiment, the image editing interface displays at least one of the following types of editing areas: a template editing area for displaying at least one template library, wherein different template libraries correspond to different parts of the target object, and the template library includes multiple template contents for the corresponding parts; a template content editing area for displaying multiple attribute parameters of the template content in the template library, wherein the template content is adjusted by editing the attribute parameters; and a historical image area for displaying editing results generated within a predetermined period, wherein the editing results displayed in the historical image area can be directly selected and directly replace the target object in the original multimedia content.

[0082] It should be noted that in the template editing area, users can select multiple parts of the target object at the same time as the content to be edited. Correspondingly, the template editing area will display multiple template libraries corresponding to the multiple parts.

[0083] In the template content editing area, users can select one or more attribute parameters to fine-tune the template content. The attribute parameters match the corresponding template content; different attribute parameters can be set for different template content.

[0084] For the historical image area, the preset period can be set according to the user's needs. For example, setting the preset period to 20 minutes will retain only the editing results of the most recent 20 minutes in the historical image area. The editing results displayed in the historical image area can be only the template content used by the user, or images generated based on the replacement results. The user can select any image in the historical image area as the target image for image retrieval. Since directly replacing parts of the target object with template content may result in a replacement image that does not match the original multimedia content, the user can use touch operations to generate a new image based on an image generation algorithm to regenerate the edited parts and generate a realistic and reasonable image. Multiple images generated by the user can be displayed in the historical image area.

[0085] For example, such as Figure 5As shown, in the application of an e-commerce shopping platform, the original multimedia content can be the original image of the product. The aforementioned interactive interface can be an image editing interface 500. The image editing interface 500 includes a display area 51 for the original image, which displays the product shoe as the target object and the selected content to be edited 511. The image editing interface 500 also includes a template selection area 52 (i.e., the aforementioned template editing area). When the content to be edited 511 is a shoe heel style, the template selection area 52 displays various shoe heel templates corresponding to the shoe heel style. Users can simultaneously select both the shoe heel style and the shoe upper style as the content to be edited 511. Correspondingly, the template selection area 52 can simultaneously display multiple templates corresponding to the shoe heel style and multiple templates corresponding to the shoe upper style. Users can select from multiple template libraries to replace the content to be edited 511 in the original image. The image editing interface 500 also includes a template fine-tuning area 53 (i.e., the template content editing area mentioned above). When the template content corresponding to the content to be edited 511 is a shoe heel style, the template fine-tuning area 53 can use "color," "material," and "pattern" as multiple attribute parameters corresponding to the shoe heel style. Users can adjust the details of the "color," "material," and "pattern" of the shoe heel. The image editing interface 500 also includes a generation history area 54 (i.e., the historical image area mentioned above). The generation history area 54 displays multiple images of product shoes generated by the user replacing templates for different parts of the shoe. Users can select any image displayed in the generation history area 54 as the target image for image retrieval.

[0086] As an optional embodiment, editing the display content of at least one display area of ​​the target object in the original multimedia content using an editing interaction mode includes: using the editing interaction mode to obtain a line drawing image of the target object; displaying the line drawing image of the target object in an image editing interface and calling the drawing board tool; using the drawing board tool to edit the line drawing image of the target object and generate an editing result.

[0087] The simplified image of the target object can be obtained by recognizing the original multimedia content using a preset algorithm, or it can be drawn by the user on the drawing board based on the outline of the original multimedia content. It should be noted that the image attributes to be edited in the simplified image of the target object should be consistent with the image attributes in the original multimedia content. For example, if the original multimedia content can be the original image of a product, and the heel of the shoe in the original image is a stiletto heel, then the heel of the shoe in the simplified image should also be a stiletto heel.

[0088] The above editing result is the edited line drawing image. Based on the edited line drawing image, the target image can be generated by calling the preset image generation algorithm.

[0089] In one alternative embodiment, the SKETCH detection algorithm (i.e., sketch detection algorithm) can be used to detect the original multimedia content, identify the target object, and generate a sketch image of the target object.

[0090] Figure 6 A schematic diagram of an optional editing interaction mode according to an embodiment of this application is provided, such as... Figure 6 As shown, in the application of an e-commerce shopping platform, the original multimedia content can be the original image 61 of the product, with the target object being shoes. The user uploads the original image 61, which is then recognized using the SKETCH detection algorithm to obtain a simplified sketch image 62 of the shoe. The image attributes of the simplified sketch image 62—"pointed toe," "stiletto heel," and "shallow opening"—are consistent with the original image 61. After obtaining the simplified sketch image 62, a drawing tool 600 can be generated. The drawing tool 600 displays the recognized simplified sketch image 62 and includes tools such as a brush, eraser, pencil, magnifying glass, and color palette 601. These tools allow for image editing of the simplified sketch image 62. For example, the eraser can be used to remove the stiletto heel, and a thicker heel can be drawn with the pencil, thus modifying the image attributes of the target object and obtaining the edited result 63. After modification, the user can use touch operations (e.g., clicking the "Generate" button) to call the target image 64 generated by the image generation algorithm. Figure 6 As shown, the target image 64 is consistent with the edited result 63, and has the image attribute of thick lines. However, in terms of visual effect, compared with the line drawing effect of the edited result 63, the target image 64 has a more realistic and reasonable visual effect.

[0091] As an optional embodiment, an editing interaction mode is used to edit the display content of at least one display area of ​​a target object in the original multimedia content, including: marking the editable content to be edited from the image editing interface, wherein the editable content is key point information of multiple key points of the target object displayed in the original multimedia content, and the multiple key points constitute the outer contour shape of the target object; if a drag command for any one or more key points is detected, responding to the drag command, obtaining the drag result after executing the drag command, wherein the drag command indicates the target position after the key point that was dragged is moved; and generating an editing result based on the drag result.

[0092] The aforementioned key point information can be the location information of the key points, such as the coordinates of the key points in the image editing interface. Multiple key points of the target object can be automatically marked in the original multimedia content using a preset detection algorithm, or the user can manually mark them on the target object. In one optional embodiment, a preset number of key points can be automatically marked in the original multimedia content using a preset detection algorithm. In locations where no key points are marked, the user can manually add key points according to editing needs. The number of key points is determined based on the size of the target object and the user's settings, and is not limited here.

[0093] The aforementioned drag-and-drop command can be triggered by the user's drag-and-drop touch operation on the image editing interface. For example, the user selects any one or more key points and triggers the drag-and-drop command by long-pressing and moving them on the touchscreen.

[0094] It should be noted that the target position of the key point after dragging can be considered as the key point position of the editing result. Based on the key point position of the editing result, the outer contour shape of the edited target object can be formed, and then the target image can be obtained through image generation algorithm.

[0095] In one optional embodiment, a keypoint detection algorithm is used to detect the original multimedia content, detect keypoint information of multiple keypoints of the target object, and mark each keypoint displayed in the original multimedia content.

[0096] Figure 7 A schematic diagram of an optional editing interaction mode according to an embodiment of this application is provided, such as... Figure 7 As shown, in the application of an e-commerce shopping platform, the original multimedia content can be the original image 71 of the product. The target object is a shoe. The user uploads the original image 71, and the original image 71 is detected by a keypoint detection algorithm to obtain image 72 marked with keypoints 721. Keypoints can be marked with star-shaped symbols. It should be noted that image 72 only shows some keypoints, and the number and position of keypoints may vary in reality. When the user wants to change the heel of the shoe from a thin heel to a thick heel, they can select the keypoints of the heel area in image 72 and drag them. Image 73 shows the dragged heel keypoints 731, which are the hollow star-shaped keypoint markers in image 73. Compared with the keypoint positions of the heel area in image 72, dragging to both sides thickens the heel. After completing the dragging of the keypoints, a reasonable and realistic target image 74 is generated by calling an image generation algorithm for subsequent image search.

[0097] It should be noted that, Figure 7The star-shaped marker used is only one optional marker form for key points. In actual applications, different forms of key point markers can also be used, and there are no restrictions here.

[0098] As an optional embodiment, editing the display content of at least one display area of ​​a target object in the original multimedia content is performed using an editing interaction mode, including: based on the editing interaction mode, responding to an interactive editing command and invoking an image conversion algorithm; using the image conversion algorithm to convert the original multimedia content into a three-dimensional model image, wherein the original multimedia content is a two-dimensional model image, and the target object in the three-dimensional model image is displayed as a stereoscopic image; mapping the three-dimensional model image to obtain parameters of multiple editable image attributes in the stereoscopic image, wherein the image attributes characterize the display content of the target object in the corresponding display area on the three-dimensional model image; displaying the parameters of the multiple editable image attributes on the image editing interface, and displaying adjustment controls for each parameter; if an adjustment command is detected to be executed on the adjustment controls of any one or more parameters, obtaining the adjustment result after executing the adjustment command, wherein the adjustment command is used to drag the progress bar of the parameter adjustment control to the target position; and generating an editing result based on the adjustment result.

[0099] The image editing interface described above displays multiple editable image attribute parameters. These parameters can be set according to the image attributes of the target object. Different target objects or different image attributes can correspond to multiple editable parameters. By mapping the 3D model image, parameters with editable and adjustable quantized indices are obtained. Users can edit the image attributes by editing the quantized indices of these parameters. For example, for "block heel" and "stiletto heel" in shoe heel styles, mapping yields the quantized heel width, which serves as an editable parameter for modifying the heel style.

[0100] The different positions of the progress bar correspond to the values ​​of different image attribute parameters. By modifying the position of the progress bar marker, the value of the image attribute parameter can be modified. The target position of the progress bar corresponds to the quantitative index of the parameter required by the user.

[0101] Figure 8 A schematic diagram of an optional editing interaction mode according to an embodiment of this application is provided, such as... Figure 8As shown, in the application of an e-commerce shopping platform, the original multimedia content can be the original image 81 of the product. The target object is a shoe. The user uploads the original image 81, and the image conversion algorithm is a 2D-3D modeling algorithm. The 2D-3D modeling algorithm converts the two-dimensional original image 81 into a three-dimensional model image 82. The image editing interface 800 contains several editable image attribute parameters such as "heel width," "heel height," "opening depth," and "toe length." Adjustment controls 801 are set at corresponding positions for each image attribute parameter. These controls have adjustable progress bars. For example, when the user needs to change the heel from a stiletto heel to a chunky heel, they can adjust the position of the square indicator on the progress bar in the "heel width" adjustment control 801 to increase the heel width value. After modifying the parameters, the user can click the "Finish" and "Save" buttons in sequence. Based on the image generation algorithm, a target image 83 is generated according to the original image and the edited parameter information. The target image is a two-dimensional image and can be used for subsequent image searches. It should be noted that... Figure 8 The adjustment control 801 shown is only one optional adjustment method. In actual applications, the adjustment method of the adjustment control 801 can be different.

[0102] In an optional embodiment, after obtaining the adjustment result after executing the adjustment instruction, the above method further includes: triggering a three-dimensional preview function to preview and display the adjustment result in a three-dimensional mode.

[0103] The 3D preview function can be automatically triggered after the user completes the editing of the above parameters. For example, each time the user completes a parameter modification, the change in the 3D model image corresponding to this parameter editing will be displayed. The 3D preview function can also be manually triggered by the user. For example, after the user completes the editing of multiple parameters, clicking the "Finish" button will display the final 3D model image after editing multiple parameters in the image editing interface.

[0104] like Figure 8 As shown, the image editing interface 800 includes a 3D preview display area to show the edited 3D model image, allowing users to see the changes in the corresponding parameters of the 3D model image in real time. For example, by dragging the adjustment control 801 corresponding to "heel width," users can see the changes in the thickness of the heel width in real time in the 3D preview display area.

[0105] As an optional embodiment, editing the display content of at least one display area of ​​a target object in the original multimedia content using an editing interactive mode to generate an edited target image may include the following steps:

[0106] Step S2041: Edit the content displayed in at least one display area of ​​the target object in the original multimedia content using the editing interaction mode to obtain the editing result;

[0107] Step S2042: Call the image generation algorithm to render the edited result and generate the edited target image.

[0108] Specifically, the display area where editing was performed can be located from the editing results; based on the location results, the area to be rendered in the original multimedia content can be obtained, where the area to be rendered corresponds to the display area where editing was performed; the content of the area to be rendered is input into the image generation algorithm to generate the rendering result, where the image generation algorithm is a training model, and the encoder and decoder included in the training model both adopt convolutional neural network models; based on the rendering result, the edited target image is generated.

[0109] The edited display area can be located from the edit results. This can be done using an object detection model or by manual input from the user. The content of the area to be rendered varies depending on the editing method used on the target object. For example, when editing the target object using image attribute values, the content of the area to be rendered can include attribute value encoding vectors, which are vectors corresponding to the attribute values ​​(e.g., text vectors for shoe heel style, color, texture, etc.). When editing using template content replacement, line drawing image modification, keypoint dragging, or 3D model image parameter adjustment methods, the corresponding template content image, line drawing image, keypoint information, and 3D model image parameters in the edit results need to be converted into feature vectors before being used as the content of the area to be rendered.

[0110] In an optional embodiment, the training model described above can adopt an autoencoder structure. The objective function of the training model is the mean squared error loss function (MSE Loss function) between the restored target image and the unlabeled original multimedia content. By training and optimizing the training model, the visual realism and rationality of the edited target image can be improved.

[0111] Figure 9 A schematic diagram of an optional image generation algorithm according to an embodiment of this application is provided, such as... Figure 9As shown, the original multimedia content can be the original image 91 of the product. After editing the target object, the editing result is obtained. The editing result includes any one of the following: template image (obtained by editing the target object through the above-mentioned template content replacement method), sketch image (obtained by editing the target object through the above-mentioned simple sketch image modification method), key point information (obtained by editing the target object through the above-mentioned key point dragging method), and three-dimensional parameters (obtained by adjusting the parameters of the above-mentioned three-dimensional model image). The editing result is input into the attribute / template positioning module 95 to locate the display area where the editing has been performed, and the display area where the editing has been performed is marked in the original image 91 using a position box 911. The original image 91 with the locally marked position box is input into the encoder 92 and converted into a latent variable. When editing the target object using the method of editing the attribute values ​​of image attributes, the attribute value encoding vector (e.g., such as...) is used. Figure 9 As shown, the text vector information of the shoe heel (which is a thick heel) is input and converted into latent variables. The latent variables corresponding to the locally labeled original image 91 and the latent variables corresponding to the attribute value encoding vectors are input into the decoder 93 to regenerate the target image 96. When editing using methods such as template content replacement, simple drawing image modification, key point dragging, and 3D model image parameter adjustment, any one of the template content image, simple drawing image, key point information, or 3D model image parameters in the editing result is first input into the encoder 94, converted into feature vectors, and then converted into latent variables. The latent variables corresponding to the locally labeled original image 91 and the latent variables corresponding to the above feature vectors are input into the decoder 93 to regenerate the target image 96. The decoder 92, encoder 93, and encoder 94 are all implemented using convolutional neural networks (CNNs).

[0112] In one alternative embodiment, the display area where editing was performed is located by at least one of the following methods: locating the display area where editing was performed from the editing result using an object detection model, or locating the display area where editing was performed from the editing result by receiving an external selection command.

[0113] For example, such as Figure 9 As shown, the attribute / template localization module 95 can be an object detection model. The target object is a pair of shoes. The user has modified the style of the shoe heel. Based on any one of the parameters of the template content image, the line drawing image, the key point information, and the 3D model image in the editing result, the module locates the display area that has been edited in the original image and marks the above-mentioned display area that has been edited with the position box 911. The module determines that the image of the editing result corresponds to the modified display area in the original image and obtains the localization result.

[0114] The aforementioned external selection command can be triggered by the user through selection operations on the interactive interface. For example, the user can directly drag out a position box in the image of the edited result and select the display area that has been edited to obtain the positioning result.

[0115] As an optional embodiment, a retrieval is performed based on the target image, and multiple search results matching the target image are searched from the image library. The multiple search results are then displayed on an interactive interface, including at least one of the following: classifying the multiple search results to obtain classification results, and displaying the multiple search results on the interactive interface based on the classification results; obtaining the similarity between the multiple search results and the target image, and displaying the multiple search results on the interactive interface based on the similarity, wherein search results with similarity below a preset value are marked.

[0116] Since multiple search results are displayed on the interactive interface, they can be categorized to help users find the results that meet their needs. For example, in the interactive interface of an e-commerce platform's client, multiple product search results obtained by searching for the target image are displayed. These results can be categorized according to the price range of the products and displayed according to different price ranges, allowing users to quickly find search results that meet their budget.

[0117] The search results can also be displayed based on their similarity to the target image. In one optional embodiment, the search results can be sorted according to their similarity values, so that the search results with high similarity are displayed at the top of the interactive interface, making it easier for users to find search results that meet their requirements.

[0118] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0119] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0120] Example 2

[0121] According to an embodiment of the present invention, an embodiment of an image search method is also provided. Figure 10 This is a flowchart of an image search method according to Embodiment 2 of the present invention, as shown below. Figure 10 As shown, the method includes:

[0122] Step S1002: Display the live video to be identified in the live streaming interface. The live video is used to play product information. The live video consists of multiple frames of images, and at least one frame of images displays the product.

[0123] The aforementioned live streaming interface can be a human-computer interaction interface displayed on the client side of a smart device (such as a mobile phone, smart tablet, and computer). Specifically, the smart device has a touch screen, and users can interact with the smart device through touch operation on the touch screen. The live video can be displayed on the touch screen interface of the smart device.

[0124] Live video can be a live video on an e-commerce shopping platform, where users can watch live videos of products through the e-commerce shopping platform's client.

[0125] Step S1004: Extract the frame image displaying the product to obtain the original multimedia content, wherein the product displayed in the original multimedia content is the target object to be edited.

[0126] The aforementioned original multimedia content includes, but is not limited to, images, videos, and text files, and includes the target object to be edited. Specifically, the original multimedia content can be the original image. When a user watches a live stream, they can obtain the original image by selecting a frame containing the product, or by using the detection algorithm of an e-commerce shopping platform to detect the product in the live stream. The original multimedia content can also be a video clip extracted from the live stream, which includes multiple frames showcasing the product.

[0127] The target object mentioned above refers to the object in the original multimedia content that needs to be edited. It can be a complete image in the original multimedia content or a part of the original multimedia content. For example, in the live broadcast interface of an e-commerce shopping platform, after obtaining the original multimedia content, the target object can be a complete product image in the original multimedia content or a part of the product image.

[0128] Step S1006: Respond to the editing command sensed in the live broadcast interface and generate the edited target image, wherein the editing command is used to edit the display content in at least one display area of ​​the target object.

[0129] Specifically, the aforementioned editing commands can be triggered by the user through touch operations in the live streaming interface. For example, the user can click the "Edit" button in the interactive interface to trigger the editing command, or the user can directly perform touch operations on the original multimedia content according to preset gestures to trigger the editing command.

[0130] The target object can be divided into multiple display areas, each with different display content. Users can edit the content displayed in different display areas to modify a specific component of the target object. For example, in a live streaming interface of an e-commerce shopping platform, if the target object in the original multimedia content is a shoe, then the shoe has multiple display areas, such as the heel area, the front surface area, and the side area. Users can modify the content displayed in at least one of these areas.

[0131] It should be noted that the same display area within a target object can include multiple modifiable display contents. For example, if the target object is a shoe, the display contents in the display area on the front surface of the shoe can include the shoe's color, the pattern on the shoe's surface, etc. Users can edit multiple display contents within a single display area at the same time.

[0132] The target image mentioned above is a new image regenerated from the original multimedia content after editing, and can be used for image search. The target image can be the entire image of the modified target object, or it can be only a modified portion of the target object. For example, if the original multimedia content is an image containing shoes, and the user edits the heel, changing a stiletto heel to a chunky heel, the resulting image of a shoe with a chunky heel is the target image.

[0133] Step S1008: In response to the search instruction, perform a retrieval based on the target image and search the video library for at least one search result that matches the target image.

[0134] The search command can be triggered by the user through touch operations within the live streaming interface; for example, the user can trigger the search command by clicking the "Search" button on the interactive interface. The search results can be a list of objects that match the modified object in the target image.

[0135] In one optional embodiment, the image library is an image feature library built based on the vectorized features of the image. After obtaining the target image edited by the user, the vectorized features of the target image can be extracted and matched with the vectorized features in the image feature library to obtain the search results.

[0136] Step S1010: Display at least one search result on the live streaming interface.

[0137] Specifically, search results can be displayed in a list format on the interactive interface. For example, in an e-commerce shopping platform application, the interactive interface is the client interface of the e-commerce platform, where users can select products to purchase from the search results. In an optional embodiment, search results can be sorted and displayed according to their degree of matching with the target image.

[0138] In this embodiment, by displaying the live video to be identified in the live streaming interface, extracting the frame images of the displayed product, obtaining the original multimedia content, responding to the editing instructions sensed in the live streaming interface, generating the edited target image, responding to the search instructions, performing a search based on the target image, searching for at least one search result matching the target image from the video library, and displaying at least one search result on the live streaming interface, the system enables users to edit a frame image when performing an image search using a specific frame image in the live streaming video, and perform customized image searches based on the edited image, solving the technical problem of difficulty in achieving customized image searches in the prior art.

[0139] As an optional embodiment, responding to an editing command sensed in the live streaming interface to generate an edited target image includes: generating an editing command by triggering an editing function and entering an image editing interface on the live streaming interface; executing at least one editing interaction mode in the image editing interface based on the editing command; using the editing interaction mode to edit the display content in at least one display area of ​​the target object in the original multimedia content to generate an editing result; and calling an image generation algorithm to render the editing result to generate the edited target image.

[0140] The editing function can be triggered by the user's touch operation on the live broadcast interface. For example, the user can generate an image editing interface by clicking the "Edit" button on the live broadcast interface.

[0141] The editing interaction mode is a way to edit the displayed content of the target object. The image editing interface can have a variety of editing interaction modes. For example, the displayed content can be modified by editing text attributes, or by using template replacement to modify the displayed content.

[0142] Image generation algorithms are used to generate target images based on original multimedia content and editing results. The target images after image rendering are more realistic and reasonable than images directly edited by users, and can bring users a realistic visual effect.

[0143] In an optional embodiment, after generating the edited target image, the method further includes: triggering a preview function to preview and display the edited target image.

[0144] Specifically, the preview function can be triggered by user touch operations on the live streaming interface. For example, clicking the "Generate" button generates a preview interface that displays the generated target image. By previewing the edited target image, users can intuitively judge whether they are satisfied with the editing result. If they are not satisfied with the preview result, they can return to the step of editing the target object and continue editing it.

[0145] As an optional embodiment, the editing interaction mode is used to edit the display content of at least one display area of ​​the target object in the original multimedia content to generate an editing result, including: obtaining multiple image attributes of the target object, wherein the image attributes represent the display content of the target object in the corresponding display area displayed on the live broadcast interface; using the editing interaction mode, determining the editable image attributes among the multiple image attributes; editing the attribute values ​​of the editable image attributes to generate edited image attributes, wherein the edited image attributes constitute the editing result.

[0146] Multiple image attributes can include editable and non-editable image attributes. After obtaining multiple image attributes, it is necessary to determine the editable image attributes. The attribute values ​​mentioned above are the specific content contained in the editable image attributes. For example, if the target object is a shoe, the image attribute can be color, and the attribute value can be a specific color such as black or red.

[0147] In an optional embodiment, before obtaining multiple image attributes of the target object, the method further includes: using an image recognition algorithm to identify the original multimedia content and identify multiple image attributes of the target object in the original multimedia content, wherein a deep learning algorithm is used to train image samples to generate the image recognition algorithm.

[0148] Image samples can include images from the original multimedia content and their corresponding annotations. For example, in e-commerce shopping platform applications, product images from live videos, along with annotations of the product's image attributes, can be used as image samples. A deep learning algorithm is then used to train the initial model of the image recognition algorithm, resulting in a trained image recognition algorithm. The image recognition algorithm can then output multiple corresponding image attributes based on the original multimedia content provided by the user.

[0149] As an optional embodiment, the editing interaction mode is used to edit the display content of at least one display area of ​​the target object in the original multimedia content to generate an editing result, including: using the editing interaction mode, determining the content to be edited selected in the image editing interface, wherein the content to be edited is at least one image attribute contained in the target object in the original multimedia content, and the image attribute represents the display content of the target object in the corresponding display area displayed on the live broadcast interface; replacing the content to be edited with selected template content, wherein the template content is displayed in the template library in the image editing interface; and generating the edited image attributes based on the replacement result, wherein the edited image attributes constitute the editing result.

[0150] In the image editing interface, the selected content to be edited corresponds to at least one display area of ​​the target object. Users can select the display content of a certain display area of ​​the target object as the content to be edited through touch operation, so as to edit one or more image attributes.

[0151] It should be noted that the template library contains pre-stored template content corresponding to image attributes. This allows users to select multiple templates in the image editing interface after selecting the content to be edited. For example, if the target object is a pair of shoes, the image attributes of the shoes can include the toe style and the heel style. The template library stores multiple toe templates corresponding to the toe styles and multiple heel templates corresponding to the heel styles.

[0152] In one optional embodiment, the image editing interface displays at least one of the following types of editing areas: a template editing area for displaying at least one template library, wherein different template libraries correspond to different parts of the target object, and the template library includes multiple template contents for the corresponding parts; a template content editing area for displaying multiple attribute parameters of the template content in the template library, wherein the template content is adjusted by editing the attribute parameters; and a historical image area for displaying editing results generated within a predetermined period, wherein the editing results displayed in the historical image area can be directly selected and directly replace the target object in the original multimedia content.

[0153] It should be noted that in the template editing area, users can select multiple parts of the target object at the same time as the content to be edited. Correspondingly, the template editing area will display multiple template libraries corresponding to the multiple parts.

[0154] In the template content editing area, users can select one or more attribute parameters to fine-tune the template content. The attribute parameters match the corresponding template content; different attribute parameters can be set for different template content.

[0155] As an optional embodiment, the editing interaction mode is used to edit the display content of at least one display area of ​​the target object in the original multimedia content to generate an editing result, including: using the editing interaction mode to obtain a line drawing image of the target object; displaying the line drawing image of the target object in the image editing interface and calling the drawing board tool; using the drawing board tool to edit the line drawing image of the target object to generate an editing result.

[0156] The simplified image of the target object can be obtained by recognizing the original multimedia content using a preset algorithm, or it can be drawn by the user on the drawing board based on the outline of the original multimedia content. It should be noted that the image attributes to be edited in the simplified image of the target object should be consistent with the image attributes in the original multimedia content. For example, if the heel of the shoe in the original multimedia content is a stiletto heel, then the heel of the shoe in the simplified image should also be a stiletto heel.

[0157] The above editing result is the edited line drawing image. Based on the edited line drawing image, the target image can be generated by calling the preset image generation algorithm.

[0158] In one optional embodiment, the SKETCH detection algorithm (i.e., sketch detection algorithm) is used to detect the original multimedia content, identify the target object, and generate a sketch image of the target object.

[0159] As an optional embodiment, an editing interaction mode is used to edit the display content of at least one display area of ​​a target object in the original multimedia content, generating an editing result. This includes: marking the editable content from the image editing interface, wherein the editable content is key point information of multiple key points of the target object displayed in the original multimedia content, and the multiple key points constitute the outer contour shape of the target object; if a drag command for any one or more key points is detected, responding to the drag command, obtaining the drag result after executing the drag command, wherein the drag command indicates the target position after the key point that was dragged is moved; and generating an editing result based on the drag result.

[0160] The aforementioned key point information can be the location information of the key points, such as the coordinates of the key points in the image editing interface. Multiple key points of the target object can be automatically marked in the original multimedia content using a preset detection algorithm, or the user can manually mark them on the target object. In one optional embodiment, a preset number of key points can be automatically marked in the original multimedia content using a preset detection algorithm. In locations where no key points are marked, the user can manually add key points according to editing needs. The number of key points is determined based on the size of the target object and the user's settings, and is not limited here.

[0161] The aforementioned drag-and-drop command can be triggered by the user's drag-and-drop touch operation on the image editing interface. For example, the user selects any one or more key points and triggers the drag-and-drop command by long-pressing and moving them on the touchscreen.

[0162] It should be noted that the target position of the key point after dragging can be considered as the key point position of the editing result. Based on the key point position of the editing result, the outer contour shape of the edited target object can be formed, and then the target image can be obtained through image generation algorithm.

[0163] In one optional embodiment, a keypoint detection algorithm is used to detect the original multimedia content, detect keypoint information of multiple keypoints of the target object, and mark each keypoint displayed in the original multimedia content.

[0164] As an optional embodiment, an editing interaction mode is used to edit the display content of at least one display area of ​​a target object in the original multimedia content, generating an editing result. This includes: based on the editing interaction mode, responding to an interactive editing command, and invoking an image conversion algorithm; using the image conversion algorithm to convert the original multimedia content into a three-dimensional model image, wherein the original multimedia content is a two-dimensional model image, and the target object in the three-dimensional model image is displayed as a stereoscopic image; mapping the three-dimensional model image to obtain parameters of multiple editable image attributes in the stereoscopic image, wherein the image attributes characterize the display content of the target object in the corresponding display area on the three-dimensional model image; displaying the parameters of the multiple editable image attributes on the image editing interface, and displaying adjustment controls for each parameter; if an adjustment command is detected to be executed on the adjustment controls of any one or more parameters, obtaining the adjustment result after executing the adjustment command, wherein the adjustment command is used to drag the progress bar of the parameter adjustment control to the target position; and generating an editing result based on the adjustment result.

[0165] The image editing interface described above displays multiple editable image attribute parameters. These parameters can be set according to the image attributes of the target object. Different target objects or different image attributes can correspond to multiple editable parameters. By mapping the 3D model image, parameters with editable and adjustable quantized indices are obtained. Users can edit the image attributes by editing the quantized indices of these parameters. For example, for "block heel" and "stiletto heel" in shoe heel styles, mapping yields the quantized heel width, which serves as an editable parameter for modifying the heel style.

[0166] The different positions of the progress bar correspond to the values ​​of different image attribute parameters. By modifying the position of the progress bar marker, the value of the image attribute parameter can be modified. The target position of the progress bar corresponds to the quantitative index of the parameter required by the user.

[0167] In an alternative implementation, after obtaining the adjustment result after executing the adjustment instruction, the method further includes: triggering a three-dimensional preview function to preview and display the adjustment result in a three-dimensional mode.

[0168] As an optional embodiment, an image generation algorithm is invoked to render the edited result and generate an edited target image, including: locating the display area where the editing was performed from the edited result; obtaining the area to be rendered in the original multimedia content based on the location result, wherein the area to be rendered corresponds to the display area where the editing was performed; inputting the content of the area to be rendered into the image generation algorithm to generate a rendering result, wherein the image generation algorithm is a training model, and the encoder and decoder included in the training model both adopt convolutional neural network models; and generating the edited target image based on the rendering result.

[0169] Specifically, the edited display area is located from the editing results. This can be done using an object detection model or by manual input from the user. The content of the area to be rendered varies depending on the editing method used on the target object. For example, when editing the target object using image attribute values, the content of the area to be rendered can include attribute value encoding vectors, which are vectors corresponding to the attribute values ​​(e.g., text vectors for shoe heel style, color, texture, etc.). When editing using template content replacement, line drawing image modification, key point dragging, or 3D model image parameter adjustment methods, the corresponding template content image, line drawing image, key point information, and 3D model image parameters in the editing results need to be converted into feature vectors before being used as the content of the area to be rendered.

[0170] In an optional embodiment, the training model described above can adopt an autoencoder structure. The objective function of the training model is the mean squared error loss function (MSE Loss function) between the restored target image and the unlabeled original multimedia content. By training and optimizing the training model, the visual realism and rationality of the edited target image can be improved.

[0171] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant description in Embodiment 1, and will not be repeated here.

[0172] Example 3

[0173] According to an embodiment of the present invention, an embodiment of an image search method is also provided. Figure 11 This is a flowchart of an image search method according to Embodiment 3 of the present invention, as follows: Figure 11 As shown, the method includes:

[0174] Step S1102: Play the teaching video in the teaching interface and obtain the original courseware content to be edited in the teaching video, wherein the original courseware content displays different types of teaching content.

[0175] The instructional videos can be from various online educational platforms, and users can watch them on the client's interface. These different types of instructional content can cover different subjects (e.g., mathematics, English), or different topics or chapters within the same subject.

[0176] The original courseware content to be edited is the content contained in a single frame of the instructional video.

[0177] Step S1104: Respond to the editing command sensed in the teaching interface and generate the edited target courseware, wherein the editing command is used to edit the courseware content in at least one display area of ​​the original courseware content.

[0178] Specifically, the aforementioned editing commands can be triggered by the user through touch operations in the teaching interface. For example, the user can click the "Edit" button in the teaching interface to trigger the editing command, or the user can directly perform touch operations on the original courseware content to be edited according to preset gestures to trigger the editing command.

[0179] The original courseware content can be divided into multiple display areas, each containing different courseware content. Users can edit the content within different display areas to modify specific parts of the original courseware content. The same display area within the target object can contain multiple editable pieces of courseware content, just as the same display area within the original courseware content can contain multiple editable pieces of courseware content.

[0180] The target courseware mentioned above is a new image regenerated from a single frame containing the content of the original courseware, after editing, and can be used for image searching. The target courseware can be the entire image of the modified courseware content, or it can be only a portion of the courseware content that has been modified.

[0181] Step S1106: In response to the search command, perform a search based on the target courseware and find at least one search result that matches the target courseware from the courseware library.

[0182] The above search commands can be triggered by the user through touch operations in the teaching interface. For example, the user can click the "Search" button in the teaching interface to trigger the search command.

[0183] The search results above can be a list of courseware that matches the modified content of the target courseware.

[0184] Step S1108: Display at least one search result on the teaching interface.

[0185] After completing the search for the target courseware, a list of courseware that matches the modified courseware content can be displayed in the teaching interface, and users can select from the displayed courseware list.

[0186] In this embodiment, by playing a teaching video in the teaching interface and obtaining the original courseware content to be edited in the teaching video, responding to the editing command sensed in the teaching interface, generating the edited target courseware, responding to the search command, performing a search based on the target courseware, searching for at least one search result matching the target courseware from the courseware library, and displaying at least one search result on the teaching interface, the user can edit a frame of image containing the original courseware content in the teaching video when performing image search using the teaching video, and perform customized image search based on the edited image, solving the technical problem of difficulty in achieving customized image search in the prior art.

[0187] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant description in Embodiment 1, and will not be repeated here.

[0188] Example 4

[0189] According to an embodiment of the present invention, an embodiment of an image search method is also provided. Figure 12 This is a flowchart of an image search method according to Embodiment 4 of the present invention, as follows: Figure 12 As shown, the method includes:

[0190] Step S1202: Collect physiological video of the area to be tested using medical equipment.

[0191] The physiological video of the area to be tested can be an image video taken by medical testing equipment such as MRI, CT (Computed Tomography), or Doppler ultrasound, which contains an image of the area to be tested.

[0192] Step S1204: Display a physiological video on the examination interface of the medical device, wherein the physiological video shows at least one lesion object to be edited.

[0193] The lesion to be edited can be a pathological image of the area to be detected presented in a physiological video. For example, in an image video captured by a Doppler ultrasound device, the pathological image can be a shadow. The corresponding disease can be determined based on the location of the shadow in the area to be detected, the size of the shadow, and the shape of the shadow.

[0194] Step S1206: In response to the editing command sensed in the examination interface of the medical device, an edited target pathological image is generated, wherein the editing command is used to edit the display content in at least one display area of ​​the lesion object.

[0195] Specifically, the aforementioned editing instructions can be triggered by the doctor through touch operation on the examination interface of the medical device. For example, the doctor can click the "Edit" button on the examination interface to trigger the editing instructions, or the doctor can directly perform touch operation on the physiological video according to preset gestures to trigger the editing instructions.

[0196] In physiological videos, lesion objects can be divided into multiple display areas, each with different content. Doctors can edit the content within different display areas to modify the specific content corresponding to a particular part of the lesion object. It's important to note that the same display area within a lesion object in a physiological video can include multiple modifiable display contents. Doctors can edit multiple contents within a single display area simultaneously. For example, if a lesion object in a physiological video contains pathological pattern A and pathological pattern B, the doctor can trigger an editing command by editing pathological pattern A. The medical device can then regenerate a modified target pathological image based on the doctor's edits to pathological pattern A.

[0197] The target pathological image mentioned above is a newly generated image rewritten from a frame containing the lesion in the physiological video, and can be used for image searching. The target pathological image can be a complete image of the modified lesion, or only a modified portion of the lesion.

[0198] Step S1208: In response to the search command, perform a search based on the target pathological image and search the image library for at least one search result that matches the target pathological image.

[0199] The search command can be triggered by the doctor through touch operation in the examination interface. For example, the doctor can click the "Search" button in the interactive interface to trigger the search command.

[0200] The search results can be a list of lesion objects that match the modified lesion objects in the target pathological image. In an optional embodiment, the image library is a pathological image feature library built based on the vectorized features of the pathological image. After obtaining the target pathological image edited by the doctor, the vectorized features of the target pathological image can be extracted and matched with the vectorized features in the image feature library to obtain search results. The search results can be used to indicate the disease information corresponding to the target pathological image.

[0201] Step S1210: Fill the search results into the text containing the case information to obtain the structured case text.

[0202] Specifically, the search results can include vector features of the lesion pattern of the target pathological image and the corresponding disease information. For example, the search results include pattern information such as the location of the lesion pattern to be detected, the size of the shadow, and the shape of the shadow, as well as the corresponding disease information. The pattern information and disease information of the lesion pattern can be filled into the text that records the case information.

[0203] Step S1212: Display the structured and filled case text on the inspection interface.

[0204] In this embodiment, a physiological video of the area to be tested is acquired by a medical device, and the physiological video is displayed on the examination interface of the medical device. Responding to editing commands sensed on the examination interface, an edited target pathological image is generated. Responding to search commands, a search is performed based on the target pathological image, and at least one search result matching the target pathological image is found in the image library. The search result is then filled into a text containing case information to obtain a structured case text, which is displayed on the examination interface. This allows doctors to edit a frame containing a lesion object in the physiological video when using the physiological video for image search, and to perform customized image search based on the edited lesion image, solving the technical problem of difficulty in achieving customized image search in the prior art.

[0205] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant description in Embodiment 1, and will not be repeated here.

[0206] Example 5

[0207] According to an embodiment of the present invention, a method for searching a target object is also provided. Figure 13This is a flowchart of the target object search method according to Embodiment 5 of the present invention, as follows: Figure 13 As shown, the method includes:

[0208] In step S1302, the cloud server receives an editing message from the client, wherein the editing message carries identification information representing the identified video.

[0209] The aforementioned cloud server is used to host various algorithm models for image processing, including image generation algorithm models and detection models for recognizing original multimedia content in videos.

[0210] The aforementioned client can be video software or application on the user's smart device (such as a mobile phone, smart tablet, or computer), allowing the user to watch videos on the internet. The smart device has a human-computer interaction interface, through which the user performs touch operations to trigger editing messages.

[0211] The aforementioned video for image search is the video to be identified. This video can be any type of video, such as live streams from e-commerce platforms, physiological videos captured by medical devices, or educational videos. The aforementioned identification information is used to identify the position of the original multimedia content within the video; for example, the identification information can be the frame number corresponding to the original multimedia content within the video.

[0212] Step S1304: The cloud server obtains the original multimedia content based on the identification information, wherein the original multimedia content includes at least one target object to be edited.

[0213] The aforementioned original multimedia content consists of unedited images, including the target object to be edited. Specifically, the original multimedia content is a frame in the identification video that matches the identification information. For example, if the identification information is frame 5, the cloud server can extract frame 5 from the identification video as the aforementioned original multimedia content.

[0214] In one optional embodiment, in the application of an e-commerce shopping platform, the original multimedia content can be the original image of the product. When a user watches a live video of the product on the client, they can select any frame of the live video containing the product and send an edit message based on the edit operation. The cloud server receives the edit message and extracts the frame image from the live video as the original image.

[0215] The target object mentioned above refers to the object in the original multimedia content that needs to be edited. It can be a complete image in the original multimedia content or a part of an image. For example, in the interactive interface of an e-commerce shopping platform's client, the target object can be a complete product image in the original picture or a part of the product image.

[0216] In step S1306, the cloud server responds to the editing command sensed in the interactive interface and generates the edited target image, wherein the editing command is used to edit the content of at least one part of the target object.

[0217] Specifically, the aforementioned editing commands can be triggered by the user through touch operations in the client's interactive interface. For example, the user can click the "Edit" button in the interactive interface to trigger the editing command, or the user can directly perform touch operations on the original multimedia content according to preset gestures to trigger the editing command.

[0218] The target object can be divided into multiple parts, each with different display content, and users can edit the content of different parts. For example, in the interactive interface of an e-commerce shopping platform's client, the original multimedia content can be the original image of the product. If the target object in the original image is a shoe, then the shoe has multiple parts, which may include the heel area, the front surface area of ​​the shoe, the side area of ​​the shoe, etc. Users can modify the content of at least one part.

[0219] It should be noted that the content of the same part of the target object can include multiple modifiable contents. For example, if the target object is a shoe, the modifiable contents in the part of the front surface of the shoe can include the color of the shoe, the pattern on the surface of the shoe, etc. The user can edit multiple display contents in a part at the same time.

[0220] The target image mentioned above is a new image regenerated from the original multimedia content after editing, and can be used for image search. The target image can be the entire image of the modified target object, or it can be only a modified portion of the target object. For example, if the original multimedia content is an image containing shoes, and the user edits the heel, changing a stiletto heel to a chunky heel, the resulting image of a shoe with a chunky heel is the target image.

[0221] In step S1308, the cloud server responds to the search command, performs a retrieval based on the target image, and finds at least one search result that matches the target image from the image library.

[0222] The search command can be triggered by the user through touch operation in the client's interactive interface. After receiving the search command from the client, the cloud server performs a retrieval of the target image. For example, the user clicks the "Search" button in the interactive interface to trigger the search command. The search results can be a list of objects that match the modified object in the target image.

[0223] In one optional embodiment, the image library is an image feature library built based on the vectorized features of the image. After obtaining the target image edited by the user, the vectorized features of the target image can be extracted and matched with the vectorized features in the image feature library to obtain the search results.

[0224] In step S1310, the cloud server sends at least one search result to the client, whereby the client displays at least one search result on the interactive interface.

[0225] Specifically, search results can be displayed in a list format on the interactive interface. For example, in an e-commerce shopping platform application, the interactive interface includes the client interface of the e-commerce platform, where users can select products to purchase from the search results. In an optional embodiment, search results can be sorted and displayed according to their degree of matching with the target image.

[0226] Figure 20 This is a schematic diagram illustrating an optional target object search method according to an embodiment of this application, such as... Figure 20 As shown, client 2001 connects to one or more cloud servers 2002 via a data network connection or electronic connection. The data network connection can be a local area network (LAN) connection, a wide area network (WAN) connection, an internet connection, or other types of data network connection. The aforementioned smart device can perform network services to connect to a cloud server or a group of cloud servers. In an optional embodiment, in an e-commerce shopping platform application, the original multimedia content can be the original image 2003 of a product. The user can select a frame of the recognition video through the interactive interface of client 2001, and click the "Edit" button to trigger an edit message sent to cloud server 2002. The edit message contains the frame number information of the selected image, allowing the cloud server to determine the position of the image in the recognition video. Cloud server 2002 then extracts that frame from the live video as the original image. The user edits the original image in the interactive interface of client 2001, triggering an editing command. The cloud server 2002 generates the edited target image according to the editing command and retrieves the target image according to the search command issued by the user in client 2001. At least one search result is sent to client 2001, and the interactive interface of client 2001 can display the search results to the user.

[0227] In this embodiment, the cloud server receives editing messages from the client, obtains the original multimedia content based on the identification information, responds to the search command, performs a retrieval based on the target image, finds at least one search result matching the target image from the image library, and sends at least one search result back to the client. The client displays at least one search result on the interactive interface, enabling users to edit any frame of the image in the recognition video when searching, and to perform customized image searches based on the edited image. This solves the technical problem of difficulty in achieving customized image searches in the prior art.

[0228] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant description in Embodiment 1, and will not be repeated here.

[0229] Example 6

[0230] According to embodiments of the present invention, an apparatus for implementing the above-described method for searching for a target object is also provided. Figure 14 This is a schematic diagram of a target object search device according to Embodiment 6 of the present invention, as shown below. Figure 14 As shown, the device 1400 includes:

[0231] Display module 1401 is used to display original multimedia content on the interactive interface, wherein the original multimedia content includes at least one target object to be edited; editing module 1402 is used to respond to editing instructions sensed in the interactive interface, determine the editing interaction mode corresponding to the editing instructions, and generate an edited target image according to the determined editing interaction mode, wherein the editing instructions are used to edit the display content in at least one display area of ​​the target object, and the interactive interface provides multiple editing interaction modes; search module 1403 is used to respond to search instructions, perform retrieval based on the target image, and search for at least one search result matching the target image from the image library; display module 1404 is used to display at least one search result on the interactive interface.

[0232] It should be noted that the display module 1401, editing module 1402, search module 1403, and display module 1404 mentioned above correspond to steps S202 to S208 in Embodiment 1. The four modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computing device 10 provided in Embodiment 1.

[0233] As an optional embodiment, the editing module includes: an editing instruction generation submodule, used to generate editing instructions by triggering the editing function and enter the image editing interface on the interactive interface; a mode execution submodule, used to execute at least one editing interaction mode in the image editing interface based on the editing instructions, wherein the editing interaction mode is used to determine the editing method for editing the display content of at least one display area of ​​the target object in the original multimedia content; and an editing result generation submodule, used to edit the display content of at least one display area of ​​the target object in the original multimedia content using the editing interaction mode to generate the edited target image.

[0234] As an optional embodiment, the above-mentioned device further includes: a preview module for triggering the preview function, previewing and displaying the edited target image, and obtaining a preview result; a re-editing module for receiving a re-editing instruction and editing the preview result according to the re-editing instruction to generate a re-edited target image; a re-search module for responding to a re-search instruction, performing a search based on the re-edited target image, and searching the image library for at least one search result that matches the re-edited target image; and a display module for displaying at least one search result on an interactive interface.

[0235] As an optional embodiment, the editing result generation submodule includes: an image attribute acquisition submodule, used to acquire multiple image attributes of the target object, wherein the image attributes represent the display content of the target object in the corresponding display area displayed on the interactive interface; an image attribute determination submodule, used to determine the editable image attributes among the multiple image attributes using an editing interaction mode; and an attribute value editing submodule, used to edit the attribute values ​​of the editable image attributes to generate the edited image attributes, wherein the edited image attributes constitute the editing result.

[0236] As an optional embodiment, the above-mentioned device further includes: an image attribute recognition module, used to recognize the original multimedia content using an image recognition algorithm, and to identify multiple image attributes of the target object in the original multimedia content, wherein the image recognition algorithm is generated by training image samples using a deep learning algorithm.

[0237] As an optional embodiment, the editing result generation submodule includes: a selection submodule, used to determine the selected content to be edited in the image editing interface using an editing interaction mode, wherein the content to be edited is at least one image attribute contained in the target object in the original multimedia content, and the image attribute represents the display content of the target object in the corresponding display area displayed on the interactive interface; a template replacement submodule, used to replace the content to be edited with the selected template content, wherein the template content is displayed in the template library in the image editing interface; and an image attribute generation submodule, used to generate the edited image attributes based on the replacement result, wherein the edited image attributes constitute the editing result.

[0238] As an optional embodiment, the image editing interface displays at least one of the following types of editing areas: a template editing area for displaying at least one template library, wherein different template libraries correspond to different parts of the target object, and the template library includes multiple template contents for the corresponding parts; a template content editing area for displaying multiple attribute parameters of the template content in the template library, wherein the template content is adjusted by editing the attribute parameters; and a historical image area for displaying editing results generated within a predetermined period, wherein the editing results displayed in the historical image area can be directly selected and directly replace the target object in the original multimedia content.

[0239] As an optional embodiment, the editing result generation submodule includes: a line drawing image acquisition submodule, used to acquire a line drawing image of the target object using an editing interaction mode; a drawing board tool invocation submodule, used to display the line drawing image of the target object in the image editing interface and invoke the drawing board tool; and a line drawing image editing submodule, used to edit the line drawing image of the target object using the drawing board tool and generate an editing result.

[0240] As an optional embodiment, the above-mentioned device further includes: a SKETCH detection submodule, used to detect the original multimedia content using the SKETCH detection algorithm, identify the target object, and generate a line drawing image of the target object.

[0241] As an optional embodiment, the editing result generation submodule includes: a key point marking submodule, used to mark editable content from the image editing interface, wherein the content to be edited is key point information of multiple key points of a target object displayed in the original multimedia content, and the multiple key points constitute the outer contour shape of the target object; a dragging submodule, used to respond to the dragging command and obtain the dragging result after executing the dragging command if a dragging command for any one or more key points is detected, wherein the dragging command indicates the target position after the key point that was dragged is moved; and a first editing result generation submodule, used to generate an editing result based on the dragging result.

[0242] As an optional embodiment, the above-mentioned device further includes: a key point detection module, used to detect the original multimedia content using a key point detection algorithm, detect key point information of multiple key points of the target object, and mark each key point displayed in the original multimedia content.

[0243] As an optional embodiment, the editing result generation submodule includes: an image conversion algorithm retrieval submodule, used to retrieve an image conversion algorithm based on the editing interaction mode and in response to interactive editing commands; an image conversion submodule, used to convert the original multimedia content into a three-dimensional model image using the image conversion algorithm, wherein the original multimedia content is a two-dimensional model image, and the target object in the three-dimensional model image is displayed as a stereoscopic image; a mapping submodule, used to map the three-dimensional model image and obtain parameters of multiple editable image attributes in the stereoscopic image, wherein the image attributes characterize the display content of the target object in the corresponding display area on the three-dimensional model image; a parameter display submodule, used to display the parameters of multiple editable image attributes on the image editing interface and display adjustment controls for each parameter; a parameter adjustment submodule, used to obtain the adjustment result after executing the adjustment command if an adjustment command is detected on the adjustment controls of any one or more parameters, wherein the adjustment command is used to drag the progress bar of the parameter adjustment control to the target position; and a second editing result generation submodule, used to generate an editing result based on the adjustment result.

[0244] As an optional embodiment, the above-mentioned device further includes: a preview module, used to trigger the three-dimensional preview function and preview the adjustment result in three-dimensional mode.

[0245] As an optional embodiment, the algorithm invocation submodule includes: a positioning submodule, used to locate the display area that has been edited from the editing result; a region to be rendered submodule, used to obtain the region to be rendered in the original multimedia content based on the positioning result, wherein the region to be rendered corresponds to the display area that has been edited; a rendering result generation submodule, used to input the content of the region to be rendered into an image generation algorithm to generate a rendering result, wherein the image generation algorithm is a training model, and the encoder and decoder included in the training model both adopt convolutional neural network models; and a target image generation submodule, used to generate an edited target image based on the rendering result.

[0246] As an optional embodiment, the positioning submodule includes at least one of the following: a model detection submodule, used to locate the display area where editing has been performed from the editing results using an object detection model; and a selection positioning submodule, used to receive an external selection command to locate the display area where editing has been performed from the editing results.

[0247] As an optional embodiment, the search module is further configured to perform retrieval based on the target image, searching the image library for multiple search results that match the target image. The display module includes at least one of the following: a classification display submodule, configured to classify the multiple search results to obtain classification results, and display the multiple search results on the interactive interface based on the classification results; and a similarity display submodule, configured to obtain the similarity between the multiple search results and the target image, and display the multiple search results on the interactive interface based on the similarity, wherein search results with a similarity lower than a preset value are marked.

[0248] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant description in Embodiment 1, and will not be repeated here.

[0249] Example 7

[0250] According to an embodiment of the present invention, an apparatus for implementing the above-described image search method is also provided. Figure 15 This is a schematic diagram of an image search device according to Embodiment 7 of the present invention, as shown below. Figure 15 As shown, the device 1500 includes:

[0251] The live video display module 1501 is used to display the live video to be identified in the live interface, wherein the live video is used to play product information, the live video consists of multiple frames, and at least one frame displays the product; the frame image extraction module 1502 is used to extract the frame image displaying the product to obtain the original multimedia content, wherein the product displayed in the original multimedia content is the target object to be edited; the editing instruction response module 1503 is used to respond to the editing instruction sensed in the live interface and generate the edited target image, wherein the editing instruction is used to edit the display content in at least one display area of ​​the target object; the search instruction response module 1504 is used to respond to the search instruction, perform retrieval based on the target image, and search for at least one search result matching the target image from the video library; the first search result display module 1505 is used to display at least one search result on the live interface.

[0252] It should be noted that the aforementioned live video display module 1501, frame image extraction module 1502, editing command response module 1503, search command response module 1504, and first search result display module 1505 correspond to steps S1002 to S1010 in Embodiment 2. The five modules and their corresponding steps implement the same examples and application scenarios, but are not limited to the content disclosed in Embodiments 1 and 2. It should also be noted that the aforementioned modules, as part of the device, can run in the computing device 10 provided in Embodiment 1.

[0253] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant description in Embodiment 1, and will not be repeated here.

[0254] Example 8

[0255] According to an embodiment of the present invention, an apparatus for implementing the above-described image search method is also provided. Figure 16 This is a schematic diagram of an image search device according to Embodiment 8 of the present invention, as shown below. Figure 16 As shown, the device includes:

[0256] The original courseware content acquisition module 1601 is used to play teaching videos in the teaching interface and acquire the original courseware content to be edited in the teaching videos, wherein the original courseware content displays different types of teaching content; the target courseware generation module 1602 is used to respond to the editing instructions sensed in the teaching interface and generate the edited target courseware, wherein the editing instructions are used to edit the courseware content in at least one display area in the original courseware content; the target courseware search module 1603 is used to respond to the search instructions, perform a search based on the target courseware, and search for at least one search result matching the target courseware from the courseware library; the second search result display module 1604 is used to display at least one search result on the teaching interface.

[0257] It should be noted that the original courseware content acquisition module 1601, the target courseware search module 1603, and the second search result display module 1604 mentioned above correspond to steps S1102 to S1108 in Embodiment 3. The four modules and their corresponding steps implement the same examples and application scenarios, but are not limited to the content disclosed in Embodiments 1 and 3. It should be noted that the above modules, as part of the device, can run in the computing device 10 provided in Embodiment 1.

[0258] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant description in Embodiment 1, and will not be repeated here.

[0259] Example 9

[0260] According to an embodiment of the present invention, an apparatus for implementing the above-described image search method is also provided. Figure 17 This is a schematic diagram of an image search device according to Embodiment 9 of the present invention, as shown below. Figure 17 As shown, the device includes:

[0261] The system comprises: a physiological video acquisition module 1701 for acquiring physiological video of the area to be examined via medical equipment; a physiological video display module 1702 for displaying physiological video on the examination interface of the medical equipment, wherein the physiological video displays at least one lesion object to be edited; a target pathological image generation module 1703 for generating an edited target pathological image in response to an editing command sensed on the examination interface of the medical equipment, wherein the editing command is used to edit the display content in at least one display area of ​​the lesion object; a pathological image search module 1704 for performing a search based on the target pathological image in response to a search command, and searching for at least one search result matching the target pathological image from an image library; a case record module 1705 for filling the search results into text containing case information to obtain structured case text; and a case text display module 1706 for displaying the structured case text on the examination interface.

[0262] It should be noted that the aforementioned physiological video acquisition module 1701, physiological video display module 1702, target pathological image generation module 1703, pathological image search module 1704, pathological image search module 1705, and case record module 1706 correspond to steps S1202 to S1212 in Embodiment 4. The six modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiments 1 and 4. It should also be noted that the aforementioned modules, as part of the device, can operate in the computing device 10 provided in Embodiment 1.

[0263] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant description in Embodiment 1, and will not be repeated here.

[0264] Example 10

[0265] According to embodiments of the present invention, an apparatus for implementing the above-described method for searching for a target object is also provided. Figure 18 This is a schematic diagram of a target object search device according to Embodiment 10 of the present invention, as shown below. Figure 18 As shown, the device includes:

[0266] The cloud server receives editing messages from the client, the messages carrying identification information representing the identified video. The cloud server acquires original multimedia content based on the identification information, including at least one target object to be edited. The cloud server generates an edited target image in response to an editing command sensed in the interactive interface, the command being used to edit at least one part of the target object. The search result matching module matches the target image in response to a search command, searching the image library for at least one matching search result. The cloud server provides feedback to the client, displaying at least one search result on the interactive interface.

[0267] It should be noted that the above-mentioned editing message receiving module 1801, image acquisition module 1802, target image generation module 1803, search result matching module 1804, and feedback module 1805 correspond to steps S1302 to S1310 in Embodiment 5. The five modules and their corresponding steps implement the same examples and application scenarios, but are not limited to the content disclosed in Embodiments 1 and 5. It should be noted that the above modules, as part of the device, can run in the computing device 10 provided in Embodiment 1.

[0268] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant description in Embodiment 1, and will not be repeated here.

[0269] Example 11

[0270] Embodiments of the present invention also provide a computer-readable storage medium. The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to perform the aforementioned search method for the target object.

[0271] Optionally, in this embodiment, the computer-readable storage medium may be located in any computing device in a group of computing devices in a computer network, or in any mobile terminal in a group of mobile terminals.

[0272] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: displaying original multimedia content on an interactive interface, wherein the original multimedia content includes at least one target object to be edited; responding to an editing instruction sensed in the interactive interface, determining an editing interaction mode corresponding to the editing instruction, and generating an edited target image according to the determined editing interaction mode, wherein the editing instruction is used to edit the display content in at least one display area of ​​the target object, and the interactive interface provides multiple editing interaction modes; responding to a search instruction, performing a retrieval based on the target image, and searching for at least one search result matching the target image from an image library; and displaying at least one search result on the interactive interface.

[0273] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: generating an edited target image in response to an editing instruction sensed in the interactive interface, including: generating an editing instruction by triggering an editing function and entering an image editing interface on the interactive interface; executing at least one editing interaction mode in the image editing interface based on the editing instruction, wherein the editing interaction mode is used to determine an editing method for editing the display content in at least one display area of ​​the target object in the original multimedia content; and editing the display content in at least one display area of ​​the target object in the original multimedia content using the editing interaction mode to generate the edited target image.

[0274] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: after generating the edited target image, triggering a preview function to preview the edited target image and obtain a preview result; receiving a re-editing instruction and editing the preview result according to the re-editing instruction to generate a re-edited target image; responding to a re-search instruction, performing a search based on the re-edited target image to search for at least one search result matching the re-edited target image in the image library; and displaying at least one search result on the interactive interface.

[0275] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: editing the display content of at least one display area of ​​a target object in the original multimedia content using an editing interaction mode, including: acquiring multiple image attributes of the target object, wherein the image attributes characterize the display content of the target object in the corresponding display area displayed on the interactive interface; using the editing interaction mode, determining the editable image attributes among the multiple image attributes; editing the attribute values ​​of the editable image attributes to generate edited image attributes, wherein the edited image attributes constitute the editing result.

[0276] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: before acquiring multiple image attributes of the target object, an image recognition algorithm is used to identify the original multimedia content and identify multiple image attributes of the target image object in the original multimedia content, wherein a deep learning algorithm is used to train image samples to generate the image recognition algorithm.

[0277] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: editing the display content of at least one display area of ​​a target object in the original multimedia content using an editing interaction mode, including: using the editing interaction mode, determining the content to be edited selected in the image editing interface, wherein the content to be edited is at least one image attribute contained in the target object in the original multimedia content, and the image attribute characterizes the display content of the target object in the corresponding display area displayed on the interaction interface; replacing the content to be edited with selected template content, wherein the template content is displayed in the template library in the image editing interface; and generating the edited image attribute based on the replacement result, wherein the edited image attribute constitutes the editing result.

[0278] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: editing the display content of at least one display area of ​​a target object in the original multimedia content using an editing interaction mode, including: marking editable content from the image editing interface, wherein the editable content is key point information of multiple key points of the target object displayed in the original multimedia content, and the multiple key points constitute the outer contour shape of the target object; if a drag command for any one or more key points is detected, responding to the drag command, obtaining the drag result after executing the drag command, wherein the drag command indicates the target position after the key point that was dragged is moved; and generating an editing result based on the drag result.

[0279] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: performing a retrieval based on a target image, searching for multiple search results matching the target image from an image library, and displaying the multiple search results on an interactive interface, including at least one of the following: classifying the multiple search results to obtain classification results, and displaying the multiple search results on the interactive interface based on the classification results; obtaining the similarity between the multiple search results and the target image, and displaying the multiple search results on the interactive interface based on the similarity, wherein search results with a similarity lower than a preset value are marked.

[0280] Example 12

[0281] According to an embodiment of this application, an embodiment of a computer terminal is also provided. This computer terminal can be any one of a group of computer terminals. Optionally, in this embodiment, the aforementioned computer terminal can also be replaced with a mobile terminal or other terminal device.

[0282] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.

[0283] In this embodiment, the computer terminal described above can execute the program code for the following steps in the target object search method of the application: displaying original multimedia content on the interactive interface, wherein the original multimedia content includes at least one target object to be edited; responding to an editing instruction sensed in the interactive interface, determining the editing interaction mode corresponding to the editing instruction, and generating an edited target image according to the determined editing interaction mode, wherein the editing instruction is used to edit the display content in at least one display area of ​​the target object, and the interactive interface provides multiple editing interaction modes; responding to a search instruction, performing a retrieval based on the target image, and searching for at least one search result matching the target image from the image library; and displaying at least one search result on the interactive interface.

[0284] Optionally, Figure 19 This is a structural block diagram of a computer terminal according to Embodiment 12 of this application, as follows: Figure 19 As shown, the computer terminal 1900 may include one or more (only one is shown in the figure) processors 1902, memory 1904, and peripheral interfaces 1906.

[0285] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the image search method and apparatus in this embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned target object search method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the computer terminal 1900 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0286] The processor is used to run programs and can access information and applications stored in memory via a transmission device to perform the following steps: displaying original multimedia content on an interactive interface, wherein the original multimedia content includes at least one target object to be edited; responding to editing instructions sensed in the interactive interface, determining the editing interaction mode corresponding to the editing instructions, and generating an edited target image according to the determined editing interaction mode, wherein the editing instructions are used to edit the display content in at least one display area of ​​the target object, and the interactive interface provides multiple editing interaction modes; responding to search instructions, performing a retrieval based on the target image, and searching for at least one search result matching the target image from an image library; and displaying at least one search result on the interactive interface.

[0287] Those skilled in the art will understand that Figure 19 The structure shown is for illustrative purposes only. The computer terminal can also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a mobile internet device (MID), a PAD, and other terminal devices. Figure 19 This does not limit the structure of the aforementioned electronic device. For example, the computer terminal 1300 may also include components that are more advanced than those described above. Figure 19 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 19 The different configurations shown.

[0288] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0289] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0290] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0291] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0292] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0293] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0294] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0295] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for searching a target object, characterized in that, include: The original multimedia content is displayed on the interactive interface, wherein the original multimedia content includes at least one target object to be edited; In response to an editing command sensed in the interactive interface, the system determines the editing interaction mode corresponding to the editing command and generates an edited target image based on the determined editing interaction mode. The editing command is used to edit the display content in at least one display area of ​​the target object, and the interactive interface provides multiple editing interaction modes. In response to a search command, a retrieval is performed based on the target image, and at least one search result matching the target image is found in the image library; The at least one search result is displayed on the interactive interface; The original multimedia content is a two-dimensional model image. Generating an edited target image based on the determined editing interaction mode includes: converting the two-dimensional model image into a three-dimensional model image based on the editing interaction mode; acquiring multiple editable image attributes from the three-dimensional model image, wherein the image attributes characterize the display content of the target object within the display area on the three-dimensional model image; displaying parameters of the multiple image attributes on the image editing interface of the interactive interface, and displaying adjustment controls for the parameters; in response to detecting an adjustment command executed on the adjustment controls, acquiring the adjustment result after executing the adjustment command, wherein the adjustment command is used to drag the progress bar of the adjustment control to the target position; generating an editing result based on the adjustment result; and rendering the editing result to generate the target image.

2. The method according to claim 1, characterized in that, Responding to editing commands sensed in the interactive interface, determining the editing interaction mode corresponding to the editing command, including: By triggering the editing function, the editing instructions are generated, and the image editing interface is entered; Based on the editing instructions, at least one editing interaction mode is executed in the image editing interface, wherein the editing interaction mode is used to determine the editing method for editing the display content of at least one display area of ​​the target object in the original multimedia content.

3. The method according to claim 2, characterized in that, After generating the edited target image, the method further includes: Trigger the preview function to preview and display the edited target image, and obtain the preview result; Receive a re-edit instruction, and edit the preview result according to the re-edit instruction to generate the re-edited target image; In response to a search command, a search is performed based on the edited target image to find at least one search result that matches the edited target image in the image library; The at least one search result is displayed on the interactive interface.

4. The method according to claim 2, characterized in that, Editing the content displayed in at least one display area of ​​the target object in the original multimedia content using the aforementioned editing interaction mode includes: Obtain multiple image attributes of the target object, wherein the image attributes characterize the display content of the target object in the corresponding display area displayed on the interactive interface; Using the aforementioned editing interaction mode, the editable image attributes among the plurality of image attributes are determined; The attribute values ​​of the editable image attributes are edited to generate edited image attributes, wherein the edited image attributes constitute the editing result.

5. The method according to claim 2, characterized in that, Editing the content displayed in at least one display area of ​​the target object in the original multimedia content using the aforementioned editing interaction mode includes: Using the aforementioned editing interaction mode, the selected content to be edited in the image editing interface is determined, wherein the content to be edited is at least one image attribute contained in the target object in the original multimedia content, and the image attribute characterizes the display content of the target object in the corresponding display area displayed on the interactive interface; The selected template content replaces the content to be edited, wherein the template content is displayed in the template library in the image editing interface; Based on the replacement result, edited image attributes are generated, wherein the edited image attributes constitute the editing result.

6. The method according to claim 2, characterized in that, Editing the content displayed in at least one display area of ​​the target object in the original multimedia content using the aforementioned editing interaction mode includes: The editable content is marked in the image editing interface, wherein the editable content is the key point information of multiple key points of the target object displayed in the original multimedia content, and the multiple key points constitute the outer contour shape of the target object; If a drag command for any one or more key points is detected, respond to the drag command and obtain the drag result after executing the drag command, wherein the drag command indicates the target position after the key point that was dragged is moved; Based on the drag-and-drop results, an edit result is generated.

7. The method according to claim 1, characterized in that, Based on the target image, a search is performed to find multiple search results that match the target image from the image library, and the multiple search results are displayed on the interactive interface, including at least one of the following: The multiple search results are categorized to obtain categorization results, and the multiple search results are displayed on the interactive interface based on the categorization results. The similarity between the multiple search results and the target image is obtained, and the multiple search results are displayed on the interactive interface based on the similarity, wherein search results with a similarity lower than a preset value are marked.

8. A method for searching images, characterized in that, include: The live video to be identified is displayed in the live streaming interface. The live video is used to play product information. The live video consists of multiple frames, and at least one frame displays the product. Extract the frame images that display the product to obtain the original multimedia content, wherein the product displayed in the original multimedia content is the target object to be edited; In response to an editing command sensed in the live streaming interface, an edited target image is generated, wherein the editing command is used to edit the display content in at least one display area of ​​the target object; In response to a search command, a retrieval is performed based on the target image, and at least one search result matching the target image is found in the video library; Display at least one search result on the live streaming interface; The step of generating an edited target image in response to an editing command sensed in the live streaming interface includes: executing at least one editing interaction mode in an image editing interface on the live streaming interface based on the editing command, and generating the target image according to the editing interaction mode; The original multimedia content is a two-dimensional model image. Generating an edited target image based on the editing interaction mode includes: converting the two-dimensional model image into a three-dimensional model image based on the editing interaction mode; obtaining multiple editable image attributes from the three-dimensional model image, wherein the image attributes characterize the display content of the target object within the display area on the three-dimensional model image; displaying parameters of the multiple image attributes on the image editing interface, and displaying adjustment controls for the parameters; in response to detecting an adjustment command executed on the adjustment controls, obtaining the adjustment result after executing the adjustment command, wherein the adjustment command is used to drag the progress bar of the adjustment control to the target position; generating an editing result based on the adjustment result; and rendering the editing result to generate the target image.

9. The method according to claim 8, characterized in that, Responding to editing commands sensed in the live streaming interface, generating an edited target image includes: By triggering the editing function, the editing instructions are generated, and the image editing interface is entered; The editing interaction mode is used to edit the content displayed in at least one display area of ​​the target object in the original multimedia content, and an editing result is generated; An image generation algorithm is invoked to render the edited result, generating the edited target image.

10. A method for searching images, characterized in that, include: Play the teaching video in the teaching interface and obtain the original courseware content to be edited in the teaching video, wherein the original courseware content displays different types of teaching content; In response to the editing instructions sensed in the teaching interface, the edited target courseware is generated, wherein the editing instructions are used to edit the courseware content in at least one display area of ​​the original courseware content; In response to a search command, a search is performed based on the target courseware, and at least one search result matching the target courseware is found in the courseware library; Display at least one search result on the teaching interface; The step of generating edited target courseware in response to editing instructions sensed in the teaching interface includes: executing at least one editing interaction mode in the image editing interface on the teaching interface based on the editing instructions, and generating the target courseware according to the editing interaction mode; The original courseware content is a two-dimensional model image. Generating the target courseware according to the editing interaction mode includes: converting the two-dimensional model image into a three-dimensional model image based on the editing interaction mode; obtaining multiple editable image attributes in the three-dimensional model image, wherein the image attributes are used to characterize the display content of the courseware content within the display area on the three-dimensional model image; displaying parameters of the multiple image attributes on the image editing interface, and displaying adjustment controls for the parameters; in response to detecting an adjustment command executed on the adjustment controls, obtaining the adjustment result after executing the adjustment command, wherein the adjustment command is used to drag the progress bar of the adjustment control to the target position; generating an editing result based on the adjustment result; and rendering the editing result to generate the target courseware.

11. A method for searching images, characterized in that, include: Physiological videos of the area to be tested are collected using medical equipment; The physiological video is displayed on the examination interface of the medical device, wherein the physiological video shows at least one lesion object to be edited; In response to an editing command sensed in the examination interface of the medical device, an edited target pathological image is generated, wherein the editing command is used to edit the display content in at least one display area of ​​the lesion object; In response to a search command, a search is performed based on the target pathological image, and at least one search result matching the target pathological image is found in the image library; The search results are then filled into the text containing the case information to obtain the structured case text. The structured, filled-in case text is displayed on the inspection interface; The method of generating an edited target pathological image in response to an editing command sensed in the examination interface of the medical device includes: executing at least one editing interaction mode in the image editing interface on the examination interface based on the editing command, and generating the target pathological image according to the editing interaction mode, wherein the editing interaction mode is used to determine the editing method for editing the displayed content; The physiological video is a two-dimensional model image. Generating the target pathological image according to the editing interaction mode includes: converting the two-dimensional model image into a three-dimensional model image based on the editing interaction mode; obtaining multiple editable image attributes in the three-dimensional model image, wherein the image attributes are used to characterize the display content of the lesion object within the display area on the three-dimensional model image; displaying parameters of the multiple image attributes on the image editing interface, and displaying adjustment controls for the parameters; in response to detecting an adjustment command executed on the adjustment controls, obtaining the adjustment result after executing the adjustment command, wherein the adjustment command is used to drag the progress bar of the adjustment control to the target position; generating an editing result based on the adjustment result; and rendering the editing result to generate the target pathological image.

12. A method for searching a target object, characterized in that, include: The cloud server receives an editing message from the client, wherein the editing message carries identification information representing the identified video; The cloud server obtains the original multimedia content based on the identification information, wherein the original multimedia content includes at least one target object to be edited; The cloud server responds to the editing instructions sensed in the interactive interface and generates the edited target image, wherein the editing instructions are used to edit the content of at least one part of the target object; The cloud server responds to the search command, performs a retrieval based on the target image, and finds at least one search result that matches the target image from the image library; The cloud server sends the at least one search result back to the client, wherein the client displays the at least one search result on an interactive interface; The cloud server responds to the editing instructions sensed in the interactive interface and generates the edited target image, including: responding to the editing instructions sensed in the interactive interface, determining the editing interaction mode corresponding to the editing instructions, and generating the edited target image according to the determined editing interaction mode, wherein the interactive interface provides multiple editing interaction modes; The original multimedia content is a two-dimensional model image. Generating an edited target image based on the editing interaction mode includes: converting the two-dimensional model image into a three-dimensional model image based on the editing interaction mode; obtaining multiple editable image attributes from the three-dimensional model image, wherein the image attributes characterize the display content of the target object within the display area on the three-dimensional model image; displaying parameters of the multiple image attributes on the image editing interface of the interactive interface, and displaying adjustment controls for the parameters; in response to detecting an adjustment command executed on the adjustment controls, obtaining the adjustment result after executing the adjustment command, wherein the adjustment command is used to drag the progress bar of the adjustment control to the target position; generating an editing result based on the adjustment result; and rendering the editing result to generate the target image.

Citation Information

Patent Citations

  • Garment attribute editable garment image retrieval method

    CN108197180A

  • Picture processing method and device, storage medium and electronic device

    CN110929059A