A search method, apparatus and electronic device

The method of generating fused images through voice input and filtering matching images solves the problem of inefficient searching when users do not save content in time, and enables quick location of the required images or videos.

CN116680421BActive Publication Date: 2026-01-23HISENSE ELECTRONIC TECH (WUHAN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310498237.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-26
Publication Date
2026-01-23
Estimated Expiration
2043-04-26

AI Technical Summary

Technical Problem

When users fail to save images or videos in a timely manner, they need to search through a large amount of historical records to find the corresponding images or videos, resulting in low search efficiency.

Method used

The system generates a fused image containing keyword features by allowing users to describe the desired scenario via voice input. If no modification instruction is received within a preset time, the system filters and displays matching multimedia images, allowing users to select target images or video files.

Benefits of technology

It improves search efficiency for images or videos that were not saved in time, allowing users to automatically find the content they need by simply describing it through voice without having to search through a large amount of historical records.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116680421B_ABST
    Figure CN116680421B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a search method, device and electronic equipment, relates to the technical field of human-computer interaction, and is used to solve the problem of how to improve the search efficiency of a user searching for a picture or a video that has not been reserved in time. The method comprises the following steps: receiving voice information to be recognized; determining at least one first keyword based on the voice information to be recognized; performing a text-to-image operation on the first keyword to generate a first fusion image containing features corresponding to the first keyword, and displaying the first fusion image; in the case where no instruction information for modifying the first fusion image is received within a preset time, screening at least one multimedia image matched with the first fusion image, and displaying the multimedia image; in response to a selection operation on a target image, displaying the target image or a video file with the target image as a starting point of playing; wherein the target image is any one of the multimedia images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of human-computer interaction technology, and in particular to a search method, apparatus and electronic device. Background Technology

[0002] Currently, users can browse pictures or watch videos on electronic devices. However, if users do not save their favorite pictures or videos in time, they will have to search through a large amount of history to find the corresponding images or videos later.

[0003] Therefore, improving the search efficiency for images or videos that were not saved in a timely manner has become an urgent problem to be solved. Summary of the Invention

[0004] To address the aforementioned technical problems, this disclosure provides a search method, apparatus, and electronic device.

[0005] The technical solution disclosed herein is as follows:

[0006] In a first aspect, this disclosure provides an electronic device, comprising: a communicator configured to receive voice information to be recognized; a processor configured to determine at least one first keyword based on the voice information to be recognized received by the communicator; the processor further configured to perform a text-to-image operation on the first keyword to generate a first fused image containing features corresponding to the first keyword, and to display the first fused image; the processor further configured to, if no instruction to modify the first fused image is received within a preset time, filter at least one multimedia image matching the first fused image and display the multimedia image; the processor further configured to, in response to a selection operation on a target image, control a display to display the target image, or control the display to display a video file whose playback starts from the target image; wherein the target image is any one of the multimedia images.

[0007] Secondly, this disclosure provides a search method, comprising: receiving speech information to be recognized; determining at least one first keyword based on the speech information to be recognized; performing a text-to-image operation on the first keyword to generate a first fused image containing features corresponding to the first keyword, and displaying the first fused image; if no instruction to modify the first fused image is received within a preset time, filtering at least one multimedia image that matches the first fused image, and displaying the multimedia image; in response to a selection operation on a target image, displaying the target image, or displaying a video file that starts playing from the target image; wherein the target image is any one of the multimedia images.

[0008] Thirdly, this disclosure provides an electronic device, including: a memory and a processor, wherein the memory is used to store a computer program; and the processor is used to cause the electronic device to implement the search method as provided in any of the second aspects when executing the computer program.

[0009] Fourthly, the present invention provides a computer-readable storage medium comprising: storing a computer program on the computer-readable storage medium, the computer program being executed by a processor using a search method as described in any of the claims provided by the third party.

[0010] Fifthly, the present invention provides a computer program product that, when run on a computer, causes the computer to perform a search method as provided in any of the first aspects.

[0011] It should be noted that the aforementioned computer instructions may be stored, in whole or in part, on the first computer-readable storage medium. The first computer-readable storage medium may be packaged together with the processor of the electronic device, or it may be packaged separately with the processor of the server; this disclosure does not limit this.

[0012] The descriptions of the second, third, fourth, and fifth aspects in this disclosure can be referenced to the detailed description of the first aspect; and the beneficial effects of the descriptions of the second, third, fourth, and fifth aspects can be referenced to the analysis of the beneficial effects of the first aspect, which will not be repeated here.

[0013] In this disclosure, the names of the aforementioned servers do not limit the devices or functional modules themselves. In actual implementation, these devices or functional modules may appear under other names. As long as the functions of each device or functional module are similar to those of this disclosure, they fall within the scope of the claims of this disclosure and their equivalents.

[0014] These or other aspects of this disclosure will become more readily apparent in the following description.

[0015] The technical solution provided in this disclosure has the following advantages compared with the prior art:

[0016] Users can describe the scene they want to see (such as voice information to be recognized) and input it into the electronic device. Upon receiving the voice information, the processor of the electronic device identifies at least one first keyword. The processor then performs a text-to-image conversion on the first keyword, generating a first fused image containing features corresponding to the first keyword, and displays the first fused image. Thus, the electronic device displays the image described by the user's voice, allowing the user to determine if the displayed first fused image is the image they need to search for. If the electronic device does not receive an instruction to modify the first fused image within a preset time, it indicates that the displayed first fused image contains the content of the image the user needs to search for. Since the first fused image generated by the electronic device differs from the actual multimedia image, the electronic device needs to filter for at least one multimedia image that matches the first fused image. The electronic device then displays this multimedia image. The user can then select the desired multimedia image from the displayed multimedia images. In response to the selection of the target image, the electronic device displays the target image or a video file that starts playback from the target image.

[0017] Furthermore, when users use the search method provided in this disclosure to search for images or videos that were not saved in a timely manner, they do not need to review a large amount of historical records. Instead, they can describe the scene they want to see in words, and the electronic device can then automatically search for the images or videos that the user needs. Since the time it takes for users to search for images or videos that were not saved in a timely manner is reduced, the search efficiency for such images or videos can be improved, thus solving the problem of how to improve the search efficiency for such images or videos. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0019] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A schematic diagram of a scenario for the search method provided in an embodiment of this application;

[0021] Figure 2This is one of the structural schematic diagrams of the display device provided in the embodiments of this application;

[0022] Figure 3 This is a second schematic diagram of the structure of the display device provided in the embodiments of this application;

[0023] Figure 4 One of the flowcharts of the search method provided in the embodiments of this application;

[0024] Figure 5 A schematic diagram of the first fused image generation process of the search method provided in the embodiments of this application;

[0025] Figure 6 A schematic diagram of the forward diffusion and backward diffusion processes of the search method provided in the embodiments of this application;

[0026] Figure 7 A schematic diagram of fireworks used in the search method provided in this application embodiment;

[0027] Figure 8 A schematic diagram of the first fused image for the search method provided in the embodiments of this application;

[0028] Figure 9 A schematic diagram of the second fused image for the search method provided in the embodiments of this application;

[0029] Figure 10 A second schematic flowchart illustrating the search method provided in this application embodiment;

[0030] Figure 11 The third flowchart illustrating the search method provided in this application embodiment;

[0031] Figure 12 A fourth schematic flowchart illustrating the search method provided in this application embodiment;

[0032] Figure 13 Fifth flowchart illustrating the search method provided in the embodiments of this application;

[0033] Figure 14 A flowchart illustrating the search method provided in this application embodiment is shown in Figure 6.

[0034] Figure 15 This is a schematic diagram of the structure of a display device provided in an embodiment of this application;

[0035] Figure 16 This is a schematic diagram of a chip system provided in an embodiment of this application. Detailed Implementation

[0036] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0037] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0038] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0039] In this embodiment of the disclosure, Text2Image refers to an image generator used to convert text into an image.

[0040] The display device provided in this application can have various implementation forms, such as a television, a smart television, a laser projection device, a monitor, an electronic bulletin board, an electronic table, etc. Figure 1 and Figure 2 This is one specific embodiment of the display device of this application.

[0041] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device according to one or more embodiments of this application, such as... Figure 1As shown, a user can operate the display device 200 via a mobile terminal 300 and a control device 100. The control device 100 can be a remote control, and communication between the remote control and the display device includes infrared protocol communication, Bluetooth protocol communication, wireless or other wired methods to control the display device 200. The user can input user commands through buttons on the remote control, voice input, control panel input, etc., to control the display device 200. In some embodiments, a mobile terminal, tablet computer, computer, laptop computer, and other smart devices can also be used to control the display device 200.

[0042] In some embodiments, the electronic device provided in this application can be the aforementioned display device 200. When a user needs to search for images or videos that were not saved in time while using the display device 200, the user can describe the image to be searched in voice and input it into the display device 200. For example, the user can wake up the display device 200 using a remote wake-up word, and then input the voice information to be recognized through the audio acquisition device (such as a microphone) of the display device 200; or, the user can input the voice information to be recognized by pressing the voice button on the control device 100 (such as a remote control); or the user can input the voice information to be recognized through an electronic device (such as a mobile phone) that has established a communication connection with the display device 200.

[0043] After receiving the voice information to be recognized, the display device 200 determines at least one first keyword based on the voice information. Then, it performs a text-to-image operation on the first keyword to generate a first fused image containing the features corresponding to the first keyword, and displays the first fused image. In this way, the display device 200 displays the image described by the user via voice, and the user can then determine whether the first fused image displayed by the display device 200 is the image they need to search for. If the display device 200 does not receive an instruction to modify the first fused image within a preset time period, it means that the first fused image displayed by the display device 200 contains the content of the image the user needs to search for. Since the first fused image generated by the display device 200 differs from the actual multimedia image, the display device 200 needs to filter at least one multimedia image that matches the first fused image. Then, the display device 200 displays this multimedia image. The user can then select the desired multimedia image from the multimedia images displayed by the display device 200, such as by selecting the target image via voice control or by pressing the selection button on a remote control. In response to a selection operation of a target image, the display device 200 displays the target image or a video file from which playback begins.

[0044] In some examples, when the target image is a picture file, the display device 200 can display a details page of the target image for the user's convenience. When the target image is a video file starting from the target image, the display device 200 can display a video file starting from the target image.

[0045] In some examples, the electronic device provided in this application embodiment can be the aforementioned server 400. When a user needs to search for images or videos that were not saved in a timely manner while using the display device 200, the user can describe the image to be searched in voice and input it into the display device 200. For example, the user can wake up the display device 200 using a remote wake-up word, and then input the voice information to be recognized through the audio acquisition device (such as a microphone) of the display device 200; or, the user can input the voice information to be recognized by pressing the voice button on the control device 100 (such as a remote control); or the user can input the voice information to be recognized through an electronic device (such as a mobile phone) that has established a communication connection with the display device 200.

[0046] Subsequently, after receiving the voice information to be recognized, the display device 200 sends the voice information to the server 400. Upon receiving the voice information from the display device 200, the server 400 determines at least one first keyword based on the voice information; then, it performs a text-to-image operation on the first keyword to generate a first fused image containing the features corresponding to the first keyword. The server 400 sends the first fused image to the display device 200. Upon receiving the first fused image from the server 400, the display device 200 displays the first fused image. Thus, the display device 200 displays the image described by the user in voice, allowing the user to determine whether the first fused image displayed by the display device 200 is the image they need to search for. If the server 400 does not receive an instruction from the display device 200 to modify the first fused image within a preset time period during which the display device 200 displays the first fused image, it indicates that the first fused image displayed by the display device 200 contains the content of the image the user needs to search for. Since the first fused image generated by the display device 200 differs from the actual multimedia image, the server 400 needs to filter for at least one multimedia image that matches the first fused image. Subsequently, server 400 sends at least one multimedia image to display device 200. Upon receiving the multimedia image from server 400, display device 200 displays the multimedia image. The user can then select the desired multimedia image from the displayed images, such as by voice control or by pressing a selection button on a remote control. In response to the selection, display device 200 displays the selected image or a video file that starts playback from the selected image.

[0047] Figure 2 A hardware configuration block diagram of a display device 200 according to an exemplary embodiment is shown. For example... Figure 2The display device 200 shown includes at least one of the following: a tuner / demodulator 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface 280. The controller includes a central processing unit, a video processor, an audio processor, a graphics processor, RAM, ROM, and a first to nth interface for input / output. The display 260 may be a touch-enabled display, such as a touch screen display. The tuner / demodulator 210 receives broadcast television signals via wired or wireless reception and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals. The detector 230 is used to collect signals from the external environment or signals interacting with the external environment. The controller 250 and the tuner / demodulator 210 may be located in different separate devices; that is, the tuner / demodulator 210 may also be located in an external device of the main device containing the controller 250, such as an external set-top box.

[0048] In some embodiments, the controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the display device 200.

[0049] In some examples, the display device 200 of one or more embodiments is a television 1, and the operating system of the television 1 is the Android system, for example... Figure 3 As shown, TV 1 can be logically divided into an application layer (referred to as "application layer") 21, an application framework layer (referred to as "framework layer") 22, an Android runtime and system library layer (referred to as "system runtime library layer") 23, and a kernel layer 24.

[0050] The application layer 21 includes one or more applications. These applications can be system applications or third-party applications. For example, application layer 21 may include a first application that provides search functionality. The framework layer 22 provides application programming interfaces (APIs) and programming frameworks for the applications in application layer 21. The system runtime library layer 23 supports the upper layer, namely the framework layer 22. When the framework layer 22 is used, the Android operating system runs the C / C++ libraries contained in the system runtime library layer 23 to implement the functions required by the framework layer 22. The kernel layer 24 acts as software middleware between the hardware layer and application layer 21, managing and controlling hardware and software resources.

[0051] In some examples, the first application starts after the TV 1 is turned on. While using the display device 200, if the user needs to search for images or videos that were not saved in time, the user can describe the image to be searched in voice and input it into the display device 200. For example, the user can wake up the display device 200 using a remote wake-up word, and then input the voice information to be recognized through the audio acquisition device (such as a microphone) of the display device 200, or by pressing the voice button on the control device 100 (such as a remote control), or by inputting the voice information to be recognized through an electronic device (such as a mobile phone) that has established a communication connection with the display device 200. After receiving the voice information to be recognized, the receiving unit 201 of the first application determines at least one first keyword based on the voice information received by the receiving unit 201; then, the processing unit 202 performs a text-to-image operation on the first keyword to generate a first fused image containing the features corresponding to the first keyword, and controls the display unit 203 to display the first fused image. If the processing unit 202 determines that the receiving unit 201 has not received any instruction to modify the first fused image within a preset time, it filters at least one multimedia image that matches the first fused image and controls the display unit 203 to display the multimedia image. In response to the selection operation of the target image, the display device 200 displays the target image or a video file whose playback starts from the target image.

[0052] Specifically, the storage unit 203 of the television set 1 is used to store the application of the first application and data such as the operating system of the television set 1.

[0053] In the following embodiments, the television set 1 described above is used as the execution subject for the search method provided in the embodiments of this disclosure to illustrate the method of the present application.

[0054] This application provides a search method, such as... Figure 4 As shown, the search method may include S11-S15.

[0055] S11. Receive the voice information to be recognized.

[0056] In some examples, there may be other irrelevant noise in the speech information to be recognized. Therefore, after receiving the speech information to be recognized, the TV 1 needs to perform preprocessing on the speech information to be recognized, such as echo cancellation, beamforming, noise reduction, gain compensation and silence suppression, to eliminate the interference of noise and improve the accuracy of the search.

[0057] In some examples, when users browse images or videos through different software, there is a problem of asynchronous historical progress across these different software programs. Often, users cannot accurately pinpoint the time point they last viewed, or they may have already viewed a segment on another software and want to start directly from that point, but they don't know the specific timestamp. To address this, the search method provided in this disclosure allows users to describe the scene they want to view using language (such as the voice information to be recognized), such as... Figure 5 As shown, a user describes, "A man wearing a blue-green polo shirt and carrying a military green bag came looking for a curly-haired woman in a white coat, and then the man said, 'I'm a reporter from the ** Evening News…'" Upon receiving the voice information to be recognized, TV 1 determines at least one first keyword based on the voice information, performs a text-to-image conversion operation on the first keyword, generates a first fused image 1 containing the features corresponding to the first keyword, and displays the first fused image 1. Then, the user determines whether the first fused image 1 displayed on TV 1 needs adjustment. For example, if the user needs to adjust the image of the person in the first fused image 1, they can input the instruction "This man has short hair parted at 28". TV 1 then adjusts the first fused image 1 based on the instruction and generates a second fused image. When the user confirms that there are no errors (i.e., no instruction to modify the first fused image is received within a preset time), at least one multimedia image matching the first fused image is selected and displayed.

[0058] S12. Based on the speech information to be recognized, determine at least one first keyword.

[0059] In some examples, the speech information to be recognized can be segmented to obtain at least one theoretical segment. Then, in a pre-defined dictionary, the actual segment that matches the theoretical segment is queried. Based on the actual segment, it is used as the first keyword. Alternatively, the speech information to be recognized can be input into a keyword model to obtain at least one first keyword for the speech information.

[0060] The training process for the keyword model is as follows:

[0061] Obtain training sample data and the labeling results of the training sample data. The training sample data includes at least one historical speech information, and the labeling results include the keywords corresponding to the historical speech information.

[0062] The training sample data is input into the neural network model for learning, and the prediction results of the neural network model on the training sample data are obtained.

[0063] Based on the prediction and labeling results, the network parameters of the neural network model are adjusted until the neural network model converges, thus obtaining the keyword model.

[0064] S13. Perform a text-to-image operation on the first keyword to generate a first fused image containing the features corresponding to the first keyword, and display the first fused image.

[0065] In some examples, the first keyword can be input into text2image for text-to-image conversion, generating a first fused image containing the features corresponding to the first keyword.

[0066] For example, when text2image performs text-to-image conversion, it obtains the desired sample by progressively denoising random Gaussian noise. text2image does this by... Figure 6 The forward and backward diffusion processes shown demonstrate how, when the first keyword (e.g., fireworks) is input, parameters such as... Figure 7 The incredibly realistic images shown ensure a superior user experience.

[0067] The forward expansion is defined as follows: Given a set of data x0~q(x) sampled from the real data distribution, i.e., the original data, Gaussian noise is added to the original data step by step in T steps (note that T is a variable parameter during training), finally obtaining a series of noise-added samples x1, x2, ..., x... T The magnitude of the step number T is affected by β. t constraint:

[0068] / represents the identity matrix.

[0069] From the law of total probability, we can derive:

[0070]

[0071] During this process, the original data x0 gradually loses its unique and distinct characteristics as it iterates through the forward diffusion in steps t. Finally, as T→∞, x... T It is equivalent to a Gaussian distributed noise with isotropic properties.

[0072] Forward diffusion adds noise to the data, while backward diffusion is a noise reduction process. During backward diffusion, we will use Gaussian noise x... T ~N(0, I) as input, from p θ (x t-1 |x t Samples are taken from the sample and the real sample is inferred and reconstructed.

[0073] Here, ForwardDiffusion represents forward diffusion, and Reverse Diffusion represents reverse diffusion.

[0074] In some examples, the features corresponding to the first keyword include any one of people, animals, plants, or objects.

[0075] S14. If no instruction to modify the first fused image is received within a preset time, at least one multimedia image that matches the first fused image is selected and displayed.

[0076] In some examples, because the first blended image generated by TV 1 may differ from the image actually needed by the user, TV 1 will display the first blended image after it is generated. Thus, when the first blended image generated by TV 1 is inconsistent with the image actually needed by the user, such as when the user sends an instruction to TV 1 to adjust the first blended image within a preset time (e.g., 30 seconds) during the time TV 1 displays the first blended image, TV 1 will adjust the first blended image based on the instruction.

[0077] For example, the first fused image generated by television 1 is as follows: Figure 8 As shown. At this point, the user needs to... Figure 8 When caching a Pekingese dog into a Corgi in the image, the user can input instructions into TV 1, such as: "Turn the Pekingese dog next to the girl into a Corgi." After receiving the instructions, TV 1 performs a modification operation, such as: determining that the instructions contain at least one second keyword. Then, based on the second keyword, it determines the image region that needs adjustment, such as: performing image segmentation on the first fused image to determine the image region corresponding to the Pekingese dog. Finally, it replaces the image region corresponding to the Pekingese dog in the first fused image with the Corgi dog, thus generating an image like... Figure 9 The second merged image is shown. Then, television 1 displays this second merged image. If the user needs to adjust the second merged image, television 1 repeatedly performs the modification operation to ensure that the generated merged image meets the user's needs. If the user does not need to adjust the second merged image, and television 1 determines that it has not received an instruction to modify the second merged image within a preset time, it filters at least one multimedia image that matches the second merged image and displays the multimedia image.

[0078] In some examples, the process by which television set 1 determines that the instruction information contains at least one second keyword is the same as the process by which television set 1 determines at least one first keyword based on the voice information to be recognized, and will not be described again here.

[0079] In some examples, since the merged image generated by television 1 is based solely on the user's voice information to be recognized, and the merged image does not represent the actual image, if television 1 does not receive an instruction to modify the first merged image within a preset time, it needs to filter at least one multimedia image that matches the first merged image from a pre-configured multimedia database (containing video files (containing at least two frame images) and image files) and display the multimedia image.

[0080] In some examples, different users have different preferences. Therefore, TV 1 can obtain a user profile of the target account based on the target account's historical viewing history, search history, and registration information, as well as other user-authorized information. Then, when TV 1 filters at least one multimedia image in the multimedia database that matches the first fused image, it can select images or videos with a high probability of user access based on this user profile. For example, if the user likes period dramas, kung fu movies, and celebrity A, then the multimedia database can be filtered to compare images or videos of each of these three categories. Then, from these images or videos, at least one multimedia image that matches the first fused image is selected. This improves the user experience while reducing the computational resource consumption of TV 1.

[0081] Of course, if the target image displayed does not contain the image the user needs, the user can instruct TV 1 to filter. After receiving the re-filtering information, TV 1 will filter at least one multimedia image from the multimedia database, excluding period dramas, kung fu movies, and celebrity A, that matches the first fused image. This improves the search speed while reducing the computational resource consumption of TV 1.

[0082] In some examples, multimedia images include pictures or frame images. Therefore, when selecting at least one multimedia image that matches the first fused image, the similarity between the first fused image and the pictures, and the similarity between the first fused image and the frame images, can be calculated. Then, multimedia images with similarity scores greater than a similarity threshold are selected as at least one multimedia image that matches the first fused image.

[0083] In some examples, when calculating the similarity between the first fused image and the image, or between the first fused image and the frame image, the first fused image can be first converted into a first vector, and the image and frame image can be converted into second vectors respectively. Then, the cosine similarity between the first and second vectors is calculated, and this cosine similarity is used as the similarity between the first fused image and the multimedia image. Alternatively, a target distance (such as Euclidean distance, Manhattan distance, etc.) between the first and second vectors can be calculated, and this target distance is used as the similarity between the first fused image and the multimedia image.

[0084] In some examples, when calculating the similarity between the first fused image and the picture, or the similarity between the first fused image and the frame image, both the first fused image and the multimedia image can be input into the similarity model for calculation to obtain the similarity between the first fused image and the multimedia image.

[0085] The training process of the similarity model is as follows:

[0086] Acquire training sample data and training supervision data. Both training sample data and training supervision data include at least one historical fused image and a historical multimedia image.

[0087] The training sample data is input into the neural network model for learning, resulting in a pre-trained neural network model.

[0088] The neural network model is supervised based on training supervision data to determine the prediction accuracy of the neural network model.

[0089] If the prediction accuracy is less than the preset accuracy, adjust the network parameters of the pre-trained neural network model until the prediction accuracy is greater than or equal to the preset accuracy. Then, determine that the pre-trained neural network model has converged and obtain the similarity model.

[0090] S15. In response to the selection operation of the target image, display the target image, or display a video file that starts playback from the target image; wherein the target image is any one of the multimedia images.

[0091] As can be seen from the above, when users use the search method provided in this disclosure to search for images or videos that were not saved in time, they do not need to review a large amount of historical records. Instead, they can describe the scene they want to see in words, and the electronic device can then automatically search for the images or videos that the user needs. Since the time it takes for users to search for images or videos that were not saved in time is shorter, the search efficiency for images or videos that were not saved in time can be improved.

[0092] In some feasible examples, combining Figure 4 ,like Figure 10As shown, the above S12 can be specifically implemented through the following S120-S123.

[0093] S120. Perform speech-to-text conversion on the speech information to be recognized to obtain the text information corresponding to the speech information to be recognized.

[0094] S121. Segment the text information to obtain at least one theoretical segment.

[0095] S122. In the preset dictionary, query the actual word segmentation that is the same as the theoretical word segmentation.

[0096] S123. Based on the actual word segmentation, the actual word segmentation is used as the first keyword.

[0097] In some feasible examples, combining Figure 4 ,like Figure 11 As shown, the search method provided in this embodiment of the disclosure further includes: S16 and S17.

[0098] S16. If an instruction to modify the first fused image is received within a preset time, at least one second keyword is determined based on the instruction.

[0099] S17. Based on the second keyword, modify the image region in the first fused image corresponding to the second keyword, generate the second fused image, and display the second fused image.

[0100] In some examples, during the display of the Mth merged image, if television 1 determines that no instruction to modify the Mth merged image has been received within a preset time, it filters at least one multimedia image that matches the Mth merged image and displays the multimedia image. If, during the display of the Mth merged image, television 1 determines that an instruction to modify the Mth merged image has been received within a preset time, it determines at least one second keyword based on the instruction. Based on the second keyword, it modifies the image region in the Mth merged image corresponding to the second keyword, generates the (M+1)th merged image, and displays the (M+1)th merged image. Here, M is an integer greater than or equal to 2.

[0101] In some feasible examples, combining Figure 4 ,like Figure 12 As shown, the above S14 can be implemented by the following S140 and S141.

[0102] S140. If no first instruction information for modifying the first fused image is received within a preset time, calculate the similarity between the first fused image and each multimedia image.

[0103] S141. Select multimedia images with a similarity greater than or equal to a similarity threshold as multimedia images that match the first fused image, and display the multimedia images.

[0104] In some feasible examples, combining Figure 12 ,like Figure 13 As shown, the above S140 can be specifically implemented through the following S1400-S1412.

[0105] S1400: If no first instruction to modify the first fused image is received within a preset time, obtain the user profile corresponding to the target account.

[0106] S1401. Based on user profiles, filter multimedia images that match the user profiles.

[0107] S1402. Calculate the similarity between the first fused image and the multimedia image based on the multimedia image that matches the user profile.

[0108] In some feasible examples, combining Figure 12 ,like Figure 14 As shown, the above S140 can be specifically implemented through the following S1403-S1405.

[0109] S1403. Obtain the first vector corresponding to the first fused image and the second vector corresponding to each multimedia image.

[0110] S1404. Based on the first vector and the second vector, calculate the cosine similarity between the first vector and the second vector.

[0111] S1405. Based on cosine similarity, the cosine similarity is used as the similarity between the first fused image and the multimedia image.

[0112] The foregoing mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, it includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0113] This application embodiment can divide the electronic device into functional modules according to the above method example. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing unit. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0114] like Figure 15 As shown in the figure, an embodiment of this application provides a schematic diagram of the structure of a display device 200. It includes a communicator 101, a processor 102, and a display 103.

[0115] The communicator 101 is configured to receive voice information to be recognized; the processor 102 is configured to determine at least one first keyword based on the voice information to be recognized received by the communicator 101; the processor 102 is further configured to perform a text-to-image operation on the first keyword to generate a first fused image containing features corresponding to the first keyword, and to display the first fused image; the processor 102 is further configured to, if no instruction to modify the first fused image is received within a preset time, filter at least one multimedia image that matches the first fused image, and to display the multimedia image; the processor 102 is further configured to, in response to a selection operation on a target image, control the display 103 to display the target image, or control the display 103 to display a video file that starts playing from the target image; wherein, the target image is any one of the multimedia images.

[0116] In some feasible examples, processor 102 is further configured to perform speech-to-text operation on the speech information to be recognized received by communicator 101 to obtain text information corresponding to the speech information to be recognized; processor 102 is further configured to perform word segmentation on the text information to obtain at least one theoretical word segmentation; processor 102 is further configured to query the actual word segmentation that is the same as the theoretical word segmentation in a preset word library; processor 102 is further configured to use the actual word segmentation as the first keyword based on the actual word segmentation.

[0117] In some implementable examples, processor 102 is also configured to, upon receiving an instruction to modify the first fused image within a preset time, determine at least one second keyword based on the instruction; processor 102 is also configured to, based on the second keyword, modify the image region in the first fused image corresponding to the second keyword, generate a second fused image, and control communicator 101 to display the second fused image.

[0118] In some implementable examples, processor 102 is further configured to calculate the similarity between the first fused image and each multimedia image if no first indication information for modifying the first fused image is received within a preset time; processor 102 is further configured to use multimedia images with similarity greater than or equal to a similarity threshold as multimedia images that match the first fused image and to display the multimedia images.

[0119] In some implementable examples, processor 102 is further configured to control communicator 101 to obtain a user profile corresponding to the target account if no first instruction information for modifying the first fused image is received within a preset time; processor 102 is further configured to filter multimedia images that match the user profile obtained by communicator 101; processor 102 is further configured to calculate the similarity between the first fused image and the multimedia image based on the multimedia image that matches the user profile.

[0120] In some implementable examples, the communicator 101 is further configured to acquire a first vector corresponding to the first fused image and a second vector corresponding to each multimedia image; the processor 102 is further configured to calculate a cosine similarity between the first vector and the second vector based on the first vector acquired by the communicator 101 and the second vector acquired by the communicator 101; the processor 102 is further configured to use the cosine similarity as the similarity between the first fused image and the multimedia image.

[0121] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and their functions will not be repeated here.

[0122] Of course, the display device 200 provided in this application embodiment includes, but is not limited to, the modules described above. For example, the display device 200 may also include a memory 103. The memory 103 may be used to store the program code of the display device 200, and may also be used to store data generated by the display device 200 during operation, such as data in write requests.

[0123] As an example, combined Figure 3 The receiving unit 202 in the display device 200 performs the same function as the communicator 101, the processing unit 201 performs the same function as the processor 102, and the storage unit 203 performs the same function as the memory 104.

[0124] like Figure 16As shown, this application embodiment also provides a chip system that can be applied to the display device 200 in the foregoing embodiments. The chip system includes at least one processor 1501 and at least one interface circuit 1502. The processor 1501 may be the processor in the display device 200. The processor 1501 and the interface circuit 1502 are interconnected via a circuit. The processor 1501 can receive and execute computer instructions from the memory of the display device 200 through the interface circuit 1502. When the computer instructions are executed by the processor 1501, the display device 200 can perform the various steps performed by the display device 200 in the foregoing embodiments. Of course, the chip system may also include other discrete devices, which are not specifically limited in this application embodiment.

[0125] This application also provides a computer-readable storage medium for storing computer instructions executed by the display device 200.

[0126] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An electronic device, characterized in that, include: The communicator is configured to receive voice information to be recognized; The processor is configured to determine at least one first keyword based on the voice information to be identified received by the communicator; The processor is further configured to perform a text-to-image operation on the first keyword, generate a first fused image containing features corresponding to the first keyword, and display the first fused image. The processor is further configured to, upon receiving an instruction from a user to modify the first fused image within a preset time, adjust the first fused image based on the instruction to obtain a second fused image and display the second fused image; and if no instruction to modify the second fused image is received within the preset time, filter and display at least one multimedia image that matches the second fused image. The processor is further configured to, if it does not receive an instruction to modify the first fused image within a preset time, filter at least one multimedia image that matches the first fused image and display the multimedia image. The processor is further configured to, in response to a selection operation of a target image, control the display to display the target image, or control the display to display a video file that starts playing from the target image; wherein the target image is any one of the multimedia images.

2. The electronic device according to claim 1, characterized in that, The processor is further configured to perform a speech-to-text operation on the speech information to be recognized received by the communicator to obtain text information corresponding to the speech information to be recognized. The processor is further configured to perform word segmentation on the text information to obtain at least one theoretical word segmentation. The processor is further configured to query the actual word segmentation that is the same as the theoretical word segmentation in the preset word library; The processor is further configured to use the actual word segmentation as the first keyword based on the actual word segmentation.

3. The electronic device according to claim 1, characterized in that, The step of adjusting the first fused image based on the indication information to obtain a second fused image and displaying the second fused image includes: Based on the indicated information, at least one second keyword is determined; Based on the second keyword, the image region in the first fused image corresponding to the second keyword is modified to generate a second fused image, and the communicator is controlled to display the second fused image.

4. The electronic device according to claim 1, characterized in that, The processor is further configured to calculate the similarity between the first fused image and each multimedia image if it does not receive a first indication information to modify the first fused image within a preset time. The processor is further configured to use multimedia images with a similarity greater than or equal to a similarity threshold as multimedia images that match the first fused image, and to display the multimedia images.

5. The electronic device according to claim 4, characterized in that, The user account corresponding to the electronic device is the target account; The processor is further configured to control the communicator to obtain the user profile corresponding to the target account if it does not receive a first indication information for modifying the first fused image within a preset time. The processor is further configured to filter multimedia images that match the user profile obtained by the communicator. The processor is further configured to calculate the similarity between the first fused image and the multimedia image based on a multimedia image that matches the user profile.

6. The electronic device according to claim 4, characterized in that, The communicator is further configured to acquire a first vector corresponding to the first fused image and a second vector corresponding to each multimedia image; The processor is further configured to calculate the cosine similarity between the first vector and the second vector based on the first vector obtained by the communicator and the second vector obtained by the communicator. The processor is further configured to use the cosine similarity as the similarity between the first fused image and the multimedia image.

7. A search method, characterized in that, include: Receive the voice information to be recognized; Based on the speech information to be recognized, at least one first keyword is determined; Perform a text-to-image operation on the first keyword to generate a first fused image containing the features corresponding to the first keyword, and display the first fused image; If a user sends an instruction to modify the first fused image within a preset time, the first fused image is adjusted based on the instruction to obtain a second fused image, and the second fused image is displayed; if no instruction to modify the second fused image is received within the preset time, at least one multimedia image that matches the second fused image is selected and displayed. If no instruction to modify the first fused image is received within a preset time, at least one multimedia image that matches the first fused image is selected and displayed. In response to a selection operation on a target image, the target image is displayed, or a video file that starts playback from the target image is displayed; wherein the target image is any one of the multimedia images.

8. The search method according to claim 7, characterized in that, The step of determining at least one first keyword based on the speech information to be recognized includes: The speech information to be recognized is converted to text to obtain the text information corresponding to the speech information to be recognized. The text information is segmented to obtain at least one theoretical segmentation; In the preset dictionary, query the actual word segmentation that is the same as the theoretical word segmentation; Based on the actual word segmentation, the actual word segmentation is used as the first keyword.

9. The search method according to claim 7, characterized in that, The step of adjusting the first fused image based on the indication information to obtain the second fused image includes: Based on the indicated information, at least one second keyword is determined; Based on the second keyword, the image region corresponding to the second keyword in the first fused image is modified to generate a second fused image.

10. The search method according to claim 7, characterized in that, If no instruction to modify the first fused image is received within a preset time, at least one multimedia image matching the first fused image is selected and displayed, including: If no first instruction to modify the first fused image is received within a preset time, the similarity between the first fused image and each multimedia image is calculated. Multimedia images with a similarity greater than or equal to a similarity threshold are selected as the multimedia images that match the first fused image and are then displayed.

Citation Information

Patent Citations

  • Recommendation information acquisition method, device, system, server and storage medium

    CN108829764A

  • Picture modifying method and electronic equipment

    CN110430356A

  • Image searching method, image searching device and terminal equipment

    CN112926300A

  • Method for carrying out pedestrian search by utilizing text description to generate image

    CN114359132A