Information retrieval method, system and storage medium based on speech recognition

By combining speech recognition and image retrieval technology, user voices are collected and target images are recognized, and the problems of inconvenience and inaccurate search in the prior art are solved, and the effect of simplifying operations and improving the accuracy of information retrieval is achieved.

CN114282090BActive Publication Date: 2025-08-12ZTE CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011044798.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-28
Publication Date
2025-08-12
Estimated Expiration
2040-09-28

AI Technical Summary

Technical Problem

The existing text retrieval, speech recognition and image retrieval methods are inconvenient to search information and the search results are inaccurate, especially when the video playback speed is fast, it is difficult to effectively obtain relevant information.

Method used

Combined with speech recognition and image retrieval technology, by collecting user voice, identifying target text information, obtaining target images, and comparing them in a preset image database to feedback corresponding description information.

Benefits of technology

It simplifies user operations and improves the accuracy of information retrieval. Users can obtain description information of things of interest through voice, improving the user experience of electronic devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114282090B_ABST
    Figure CN114282090B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention disclose a method, system, and medium for information retrieval based on speech recognition. The method comprises: collecting speech information; identifying target text information from the speech information; acquiring a target image based on the target text information; searching a preset image database based on the target image; and, after retrieving descriptive information corresponding to the target image, displaying at least the descriptive information, thereby simplifying user operations and improving retrieval accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of image recognition technology, and in particular to an information retrieval method, system and medium based on speech recognition. Background Art

[0002] With the development of Internet technology, people's access to information has become more diverse. When we encounter something that interests us, the most convenient way is to search online using a mobile phone or computer to find relevant information about it. Currently, common search methods are not limited to text keyword searches, but also include voice search and image search. These search functions require entering text or images in the search box, or entering voice into an electronic device after activating the search box. For example, after seeing a new car, you can view relevant information about it by entering a photo of the new car online, which meets people's diverse needs.

[0003] However, searching for relevant information using the aforementioned text retrieval, voice recognition, or image retrieval methods can present challenges, such as inconvenient user operation and inaccurate search results. For example, if a user is watching a video and sees a new car, but the video plays too fast, they might need to pause the video, take a photo of the car with their phone, and then enter the photo into an online search box to retrieve relevant information about the new car. Therefore, simplifying user operation while improving search accuracy is a pressing issue. Summary of the Invention

[0004] The purpose of one or more embodiments of this specification is to provide an information retrieval method, system, and storage medium based on speech recognition, which can simplify user operations while improving retrieval accuracy.

[0005] To solve the above technical problems, one or more embodiments of this specification are implemented as follows:

[0006] In a first aspect, a voice recognition-based information retrieval method is provided, comprising: collecting user voice emitted by a user; recognizing target text information from the user voice; obtaining a target image based on the target text information; retrieving description information corresponding to the target image in a preset image database; and after retrieving the description information, at least feeding back the description information to the user.

[0007] In the second aspect, a speech recognition-based information retrieval system is proposed, comprising: a speech acquisition module for acquiring speech information; a text recognition module for identifying target text information from the speech information; an image acquisition module for acquiring a target image based on the target text information; a retrieval module for searching in a preset image database based on the target image; and a feedback module for displaying at least the description information corresponding to the target image after the description information is retrieved.

[0008] In a third aspect, a storage medium is proposed for computer-readable storage, wherein the storage medium stores one or more programs, and when the one or more programs can be executed by one or more processors, the steps of the information retrieval method based on speech recognition as described above are implemented.

[0009] As can be seen from the technical solutions provided by one or more embodiments of the present specification, the information retrieval method based on speech recognition provided by the embodiments of the present invention combines speech recognition and image retrieval to improve the accuracy of information retrieval. The information retrieval method is suitable for collecting the user's speech when the user speaks about an object of interest, and then performing speech recognition from the user's speech to obtain target text information related to the object of interest, and then obtaining a target image related to the object of interest based on the target text information. The target image is then retrieved from a preset image database for corresponding description information. The retrieval here can be based on an image comparison between the target image and a standard image in the image database to determine the corresponding description information. After the description information is retrieved, at least the description information is fed back to the user. The feedback can be in various forms, including playing a voice, displaying on a display screen, etc. It can be seen that the information retrieval method based on speech recognition provided by the embodiments of the present invention can feed back the description information of the object of interest to the user after receiving a voice from the user. Due to the combination of speech recognition and image retrieval technology, the accuracy of information retrieval is improved, while the user's search operation is greatly simplified, and the user's experience of using the electronic device is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the description of one or more embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0011] Figure 1 The present invention provides a step diagram of an information retrieval method based on speech recognition.

[0012] Figure 2 This is a schematic diagram of the steps of another information retrieval method based on speech recognition provided by an embodiment of the present invention.

[0013] Figure 3 This is a schematic diagram of the steps of another information retrieval method based on speech recognition provided by an embodiment of the present invention.

[0014] Figure 4 This is a schematic diagram of the steps of another information retrieval method based on speech recognition provided by an embodiment of the present invention.

[0015] Figure 5 This is a schematic diagram of the steps of another information retrieval method based on speech recognition provided by an embodiment of the present invention.

[0016] Figure 6 This is a schematic diagram of the steps of another information retrieval method based on speech recognition provided by an embodiment of the present invention.

[0017] Figure 7 This is a schematic diagram of the steps of another information retrieval method based on speech recognition provided by an embodiment of the present invention.

[0018] Figure 8 This is a schematic diagram of the steps of another information retrieval method based on speech recognition provided by an embodiment of the present invention.

[0019] Figure 9 This is a schematic diagram of the steps of another information retrieval method based on speech recognition provided by an embodiment of the present invention.

[0020] Figure 10 This is a schematic diagram of the steps of another information retrieval method based on speech recognition provided by an embodiment of the present invention.

[0021] Figure 11 This is a schematic diagram of the steps of another information retrieval method based on speech recognition provided by an embodiment of the present invention.

[0022] Figure 12 The present invention provides a structural diagram of an information retrieval system based on speech recognition.

[0023] Figure 13 It is a structural diagram of another information retrieval system based on speech recognition provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0024] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the one or more embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this document.

[0025] An embodiment of the present invention provides a speech recognition-based information retrieval method, which is suitable for providing a user with a paragraph or sentence about an item of interest, and then providing the user with descriptive information related to the item of interest. This simplifies user operations and, based on speech recognition technology, improves the accuracy of information retrieval during the information retrieval process. This method can be applied to electronic devices such as speakers and display devices. The following describes in detail the speech recognition-based information retrieval method and its various steps provided in this specification.

[0026] Example 1

[0027] Reference Figure 1 The figure shows a schematic diagram of the steps of a speech recognition-based information retrieval method provided by an embodiment of the present invention. It can be understood that the speech recognition-based information retrieval method provided by an embodiment of the present invention combines speech recognition and image retrieval to improve the accuracy of information retrieval and simplify the user's information retrieval operation. The speech recognition-based information retrieval method includes:

[0028] Step 10: Collect voice information;

[0029] The information retrieval method based on speech recognition provided by the embodiment of the present invention can be applied to electronic devices used by users. After the information retrieval function is started, the user's speech can be collected, with the aim of capturing the target text information of the things that the user is interested in involved in the user's speech.

[0030] You can set the target text information T , target text information O T It can include: high-rise buildings, cars, watches, whales and other keywords from user voice.

[0031] Step 20: Identify target text information from the voice information;

[0032] After collecting the voice sent by the user, the target text information can be identified from the user voice. At this time, voice recognition is required. For example, if the user voice sent by the user is "high-rise building", the target text information can be identified by voice recognition. TAfter obtaining the target text information, the corresponding image processing is performed based on the target text information, and finally the corresponding description information is fed back to the user.

[0033] See also Figure 12 As shown, the voice acquisition module can collect the user's voice and then perform voice recognition on the user's voice. The subsequent image processing can be performed on the device side or on the server side, which is not limited here. If the subsequent image processing is performed on the server side, after obtaining the target text information, the target text information is sent to the server to realize that the voice acquisition device collects the target information input by the user's voice, and the voice analysis module in the control device analyzes the target (text information). T , and put the target (text information) T Sent to the server device.

[0034] Step 30: Acquire a target image based on the target text information;

[0035] The corresponding target image is obtained based on the target text information. This step can be based on a preset text-image library, and the target image corresponding to the target text information is selected from the text-image library. If there is a display screen displaying an image related to the target text information, the target image corresponding to the target text information can be obtained from the image displayed on the display screen.

[0036] This step is to combine image retrieval with obtaining the corresponding target image based on the target text information. The image can more accurately reflect the target text information, and thus can more accurately feedback the user's user voice, thereby improving the accuracy of information retrieval.

[0037] Step 40: Searching the preset image database according to the target image;

[0038] The pre-set image database contains a large amount of data, and the data storage model is a standard image-description information model. However, the storage format for the description information varies. For example, if the description information involves an animal, the stored description information may include: scientific name, species, distribution area, and protection level. If the description information involves a car, the stored description information may include: make, model, price, etc. If the description information involves a building, the stored description information may include: name, location, characteristics, etc.

[0039] Retrieval based on the target image in a preset image database is to compare the target image with the standard image. Compared with the comparison between text, the comparison between images is closer, and the corresponding description information can be obtained more accurately.

[0040] Step 50: After the description information corresponding to the target image is retrieved, at least the description information is displayed.

[0041] After retrieving the description information corresponding to the target image, the description information is displayed as feedback to the user. This display can be in the form of voice, image, video, or any combination of the three. When sitting in the living room or driving, the user can simply speak to obtain the description information of the object of interest, which simplifies the user's information retrieval process and improves the user's experience of using electronic devices.

[0042] Reference Figure 2 As shown, in some embodiments, the information retrieval method provided by the embodiment of the present invention, step 30: obtaining a target image based on target text information, includes:

[0043] Step 300: determining target image features corresponding to target text information in a preset target feature library;

[0044] The target image features corresponding to the target text information can be determined based on a preset target feature library, and then the target image can be retrieved from the preset target image library based on the target image features. The information retrieval method provided by the present invention in real time is to determine the target image features corresponding to the target text information obtained by speech recognition, and then retrieve relevant standard images and corresponding description information from a subsequent preset image database based on the target image.

[0045] The preset target feature library can be set on the server side, and the target feature library can be regarded as a target feature set. MA , target feature set O MA The target feature library is a pre-trained feature library. The target feature library includes a recognition module and a template image. The recognition module includes recognition features and a recognition algorithm. The recognition features are the recognition features that need to be extracted when recognizing the image corresponding to the target text information, and the recognition algorithm used when recognizing the image corresponding to the target text information. According to the target text information, the corresponding recognition module and template image are found in the target feature library as the target image feature. TM , and then use these target image features O TM Get the corresponding target image.

[0046] It can be seen that the target text information O T and the target image feature O TM There is a corresponding relationship, for example, the target text information is: high-rise building, the corresponding target image feature O is obtained in the preset target feature library TM : Template images, template features, and recognition algorithms of high-rise buildings.

[0047] Step 310: Search in a preset target image library based on the target image features to obtain the target image.

[0048] The preset target image library is set in advance and can collect images related to audio or video previously played by the user's electronic device. This can better meet the user's usage needs. For example, when the user is watching video content, screenshots of the images related to the video content can be collected in real time and saved in the target image library. After the user speaks and needs to obtain descriptive information related to the object of interest, the target image features can be obtained and then searched in the target image library based on the target image features. This search is to compare the target image features with images in the target image library to finally obtain the target image.

[0049] Using the target image feature O TM Perform image retrieval based on image features in the preset target image library C to obtain the target image O R . Get the target image O R Then, if the electronic device used by the user has a display screen, and the electronic device takes a screenshot of the image played by the electronic device, the target image may be displayed on the display screen. R The target image is framed in the image screenshot for the user to view.

[0050] Reference Figure 3 As shown, in some embodiments, in the information retrieval method provided by the embodiments of the present invention, the target feature library includes a recognition module and a template image, and step 300: determining the target image features corresponding to the target text information in the preset target feature library includes:

[0051] Step 301: Determine whether a corresponding recognition module exists in the target feature library based on the target text information;

[0052] The recognition module includes recognition features and recognition algorithms, wherein the recognition features are the recognition features that need to be extracted when identifying the image corresponding to the target text information, and the recognition algorithm used when identifying the image corresponding to the target text information. According to the target text information, the target feature library is searched to see if there is a corresponding recognition module. If there is, the recognition module is used as the target image feature. TM , and then use these target image features O TM The corresponding target image is obtained in the target image library. The target feature library can be continuously updated and trained in order to better represent the target text information and obtain the target image in the target image library more accurately.

[0053] Of course, the finer the division granularity of the recognition module, the better. Ideally, there can be one recognition module for each object, so that the target text information can be expressed more correctly and accurately, and the target image can be captured more accurately.

[0054] Step 302: When a corresponding recognition module exists in the target feature library, determining the corresponding recognition module as the target image feature;

[0055] When there is a corresponding recognition module in the target feature library, the recognition module is used as the target image feature. TM , and then use these target image features O TM Get the corresponding target image. For example, if the target text information is a high-rise building, get the recognition features and recognition algorithm of the high-rise building in the target feature library.

[0056] Step 303: determining whether a corresponding template image exists in the target feature library based on the target text information;

[0057] Next, the target feature library is checked to see if a template image corresponding to the target text exists. This means, for example, if there is an image displaying the target text, such as an image of a high-rise building. If so, the template image is added to the target image features, allowing the target image to be accurately found in the target image library using these features.

[0058] Step 304: Add the corresponding template image to the target image features.

[0059] When a corresponding template image exists in the target feature library, a template matching strategy is executed, that is, the corresponding template image is added to the target image template.

[0060] Reference Figure 4 As shown, in some embodiments, after step 301: determining whether a corresponding recognition module exists in a target feature library based on target text information, the information retrieval method provided by the embodiment of the present invention further includes:

[0061] Step 305: When there is no corresponding recognition module in the target feature library, a universal recognition module is used as the target image feature;

[0062] If there is no recognition module corresponding to the target text information in the target special diagnosis library, a general recognition module can be used as the target image feature. The general recognition module here can be determined according to the division granularity of the recognition module, and the closer to the corresponding recognition module, the better.

[0063] Step 306: determining whether a corresponding template image exists in the target feature library based on the target text information;

[0064] Then determine whether there is a corresponding template image in the target feature library. If there is a corresponding template image, add the corresponding template image to the target image feature. This can more accurately express the target text information and subsequently find the target image more accurately in the target feature library.

[0065] Step 307: Add the corresponding template image to the target image features.

[0066] When there is a corresponding template image in the target feature library, a template matching strategy is adopted to add the corresponding template image to the target image features, so that these target image features can be used to retrieve the target image in the target image library later.

[0067] Reference Figure 9 As shown, first, for the target text information, it is determined whether there is a recognition module in the target feature library. Depending on the situation, the corresponding recognition module is used as the target image feature. In particular, if there is no recognition module in the target feature library, a universal recognition module is selected as the target image feature. In this case, the universal recognition module can be a division of the corresponding recognition module at the next higher level of granularity. This is not an absolute limitation and can be flexibly set. For example, if there is no recognition module corresponding to a high-rise building, the recognition module corresponding to the building can be selected as the universal recognition module for the high-rise building.

[0068] Reference Figure 5 As shown, in some embodiments, after step 10: collecting voice information, the information retrieval method provided by the embodiment of the present invention further includes:

[0069] Step 60: In response to the voice wake-up word recognized from the voice information, the current display interface played by the display screen at the current moment and multiple display interfaces played before and after the current display interface are saved to a preset target image library.

[0070] After the voice acquisition module collects the user's voice, it identifies the voice wake-up word from the user's voice. Here, the voice wake-up word is a keyword in the user's voice, such as a noun or verb excluding function words, auxiliary words, and address words in the user's voice. After identifying the voice wake-up word, the screenshot function of the display screen is activated. The screenshot function can be a screenshot function. The currently captured picture is the current display interface, denoted as P T1 , the screenshot can be saved in the preset target image library C, the time of the screenshot is recorded as T1, and the current display interface P T1 Display, you can display the current display interface in full screen P T1 , you can also change the current display interface P T1 It is only displayed at a fixed position on the display for easy viewing by the user. It can also be hidden. Furthermore, while the user is playing the video content on the display, the images contained in the video content are continuously saved frame by frame in the target image library, forming image data in the target image library.

[0071] Specifically, after the voice wake-up word is recognized, the screenshot function of the display screen is started, and the current display interface and multiple display interfaces played before and after the current display interface can be saved in the target image library, that is, the current display interface P T1 The image set consisting of N display interfaces before and after time T1 is saved in the target image library. For example, the voice wake-up word in the user's voice can be "Xiaobai Xiaobai, take a screenshot". When the voice acquisition module collects the voice wake-up word, the screenshot function is activated on the display screen. The current screenshot time is recorded as T1, and the captured picture is recorded as the current display interface P T1 , and a total of N display interfaces before and after the current screenshot time T1 are combined into an image set, for example, N can be 5.

[0072] Reference Figure 6 As shown, in some embodiments, the information retrieval method provided by the embodiments of the present invention, after executing step 60: saving the current display interface played by the display screen at the current moment and multiple display interfaces played before and after the current display interface to a preset target image library, step 310: searching the preset target image library based on target image features to obtain the target image, includes:

[0073] Step 311: Arrange the multiple display interfaces behind the current display interface in order from closest to farthest from the current display interface, forming a search order for real-time image search;

[0074] Specifically, the target image library C consists of the current display interface P T1 and the current display interface P T1 The total number of screenshots is N. The display interfaces are sorted according to the screenshot time as follows:

[0075]

[0076] Then, multiple display interfaces are sequentially arranged behind the current display interface in the order from near to far from the current display interface, forming a retrieval order for real-time image retrieval:

[0077]

[0078] Then, step 312 is executed: extracting the real-time search image according to the search order;

[0079] According to the retrieval order, the real-time retrieval picture is obtained and the real-time retrieval picture captured at this time is O R .P T1 , the next step is to compare the similarity between the target image features and these search images. Figure 10As shown, each real-time search image can be extracted and compared with the target image features for similarity. If the similarity between the two is not greater than a set threshold, the next real-time search image is extracted until all real-time search images are traversed. Of course, if the similarity between the first real-time search image and the target image features is greater than the set threshold, the real-time search image is used as the target image, and the extraction and similarity comparison of the remaining real-time search images are terminated.

[0080] Step 313: Determine whether the similarity between the real-time search image and the target image feature is greater than a set threshold;

[0081] The target image is obtained by measuring the similarity between the real-time retrieval image and the target image features. It should be noted that the greater the similarity, the closer the real-time retrieval image and the target image features are, and vice versa.

[0082] Step 314: When the similarity is greater than the set threshold, the search is terminated and the real-time searched picture is used as the target image.

[0083] Retrieve the similarity S between the image and the target image features in real time. T1 When the value is greater than the set threshold, the extraction of the real-time retrieval image and the subsequent similarity comparison are terminated, and the real-time retrieval image is used as the target image.

[0084] Reference Figure 7 As shown, in some embodiments, after step 313: determining whether the similarity between the real-time search image and the target image feature is greater than a set threshold, the information retrieval method provided by the embodiment of the present invention further includes:

[0085] Step 315: When the similarity is not greater than the set threshold, the real-time retrieval image corresponding to the one with the largest similarity is used as the target image.

[0086] If the similarity between all real-time retrieval images and the target image features is not greater than the set threshold, then the similarity S is taken. T1 The real-time retrieval image corresponding to the largest one is taken as the target image O R .

[0087] Because O R It is intercepted from the preset target image library C and is related to the target information O T The best matching image, so O R Not necessarily in the screenshot P T1 middle.

[0088] Reference Figure 8As shown, in some embodiments, the information retrieval method provided by the embodiments of the present invention, the preset image database includes standard images and description information of the standard images, step 40: retrieving description information corresponding to the target image in the preset image database, including:

[0089] Step 400: Separate a corresponding image data subset from a preset image database according to target text information;

[0090] The pre-set image database stores different descriptive information for each target text information lock. The image data in the image database can be pre-categorized. If the target text information lock concerns an animal, the stored descriptive information might include scientific name, subject, distribution area, and protection level. If the target text information lock concerns a car, the stored descriptive information might include make, model, and price. If the target text information lock concerns a building, the stored descriptive information might include name, location, and features.

[0091] The storage method of data in the preset image database is standard image and corresponding description information, using target image O R Image retrieval is performed in the preset image database, and the standard image is first found and then the description information corresponding to the standard image is fed back to the user.

[0092] Before retrieving the corresponding description information in the image database based on the target image, the data in the image database can be preliminarily classified according to the target text information to filter out the image data subset related to the target text information. Subsequently, the corresponding standard image can be retrieved based on the target image in the image data subset, thereby reducing the amount of data covered by the retrieval. Figure 11 shown.

[0093] Step 410: performing image retrieval in the image data subset based on the target image to obtain a standard image corresponding to the target image;

[0094] Take the target image O R The purpose of performing image retrieval in an image data subset is to find a corresponding standard image, and existing image retrieval methods can be used.

[0095] like Figure 8 As shown, in some embodiments, in the information retrieval method provided in the embodiments of this specification, step 50: after retrieving the description information, at least feeding back the description information to the user includes:

[0096] Step 500: After acquiring the standard image corresponding to the target image, the standard image and description information are displayed.

[0097] Take the target image O RAfter image retrieval in a subset of image data, standard images and description information can be presented to the user.

[0098] From the above analysis, it can be seen that the information retrieval method based on speech recognition provided by the embodiments of the present invention combines speech recognition and image retrieval to improve the accuracy of information retrieval. This information retrieval method is suitable for collecting the user's speech when the user speaks about an item of interest, then performing speech recognition from the user's speech to obtain target text information related to the item of interest. Based on the target text information, a target image related to the item of interest is then obtained. The target image is then searched for corresponding description information from a preset image database. The search can be based on an image comparison between the target image and a standard image in the image database to determine the corresponding description information. After the description information is retrieved, at least the description information is fed back to the user. The feedback can take various forms, including playing a voice message or displaying it on a display screen. It can be seen that the information retrieval method based on speech recognition provided by the embodiments of the present invention can feed back the description information corresponding to the item of interest to the user after receiving a voice message from the user. Due to the combination of speech recognition and image retrieval technologies, the accuracy of information retrieval is improved, while the user's search operation is greatly simplified, and the user's experience of using the electronic device is enhanced.

[0099] Example 2

[0100] Reference Figure 13 FIG2 is a schematic diagram of the structure of a speech recognition-based information retrieval system 1 provided by an embodiment of the present invention. The information retrieval system combines speech recognition and image retrieval technologies, simplifying the user's retrieval operation while improving the accuracy of information retrieval. The speech recognition-based information retrieval system 1 includes:

[0101] Voice collection module 10, used to collect voice information;

[0102] The information retrieval method based on speech recognition provided by the embodiment of the present invention can be applied to electronic devices used by users. After the information retrieval function is started, the user's speech can be collected, with the aim of capturing the target text information of the things that the user is interested in involved in the user's speech.

[0103] You can set the target text information T , target text information O T It may include keywords from user voice, such as high-rise buildings, cars, watches, whales, etc.

[0104] A text recognition module 20 is used to recognize target text information from voice information;

[0105] After collecting the voice sent by the user, the target text information can be identified from the user voice. At this time, voice recognition is required. For example, if the user voice sent by the user is "high-rise building", the target text information can be identified by voice recognition. T After obtaining the target text information, the corresponding image processing is performed based on the target text information, and finally the corresponding description information is fed back to the user.

[0106] See also Figure 12 As shown, the voice acquisition module can collect the user's voice and then perform voice recognition on the user's voice. The subsequent image processing can be performed on the device side or on the server side, which is not limited here. If the subsequent image processing is performed on the server side, after obtaining the target text information, the target text information is sent to the server to realize that the voice acquisition device collects the target information input by the user's voice, and the voice analysis module in the control device analyzes the target (text information). T , and put the target (text information) T Sent to the server device.

[0107] An image acquisition module 30 is configured to acquire a target image based on target text information;

[0108] The corresponding target image is obtained based on the target text information. This step can be based on a preset text-image library, and the target image corresponding to the target text information is selected from the text-image library. If there is a display screen displaying an image related to the target text information, the target image corresponding to the target text information can be obtained from the image displayed on the display screen.

[0109] This step is to combine image retrieval with obtaining the corresponding target image based on the target text information. The image can more accurately reflect the target text information, and thus can more accurately feedback the user's user voice, thereby improving the accuracy of information retrieval.

[0110] A retrieval module 40 is used to search a preset image database according to a target image;

[0111] The pre-set image database contains a large amount of data, and the data storage model is a standard image-description information model. However, the storage format for the description information varies. For example, if the description information involves an animal, the stored description information may include: scientific name, species, distribution area, and protection level. If the description information involves a car, the stored description information may include: make, model, price, etc. If the description information involves a building, the stored description information may include: name, location, characteristics, etc.

[0112] Retrieval based on the target image in a preset image database is to compare the target image with the standard image. Compared with the comparison between text, the comparison between images is closer, and the corresponding description information can be obtained more accurately.

[0113] The feedback module 50 is configured to at least display the description information after the description information corresponding to the target image is retrieved.

[0114] After retrieving the description information corresponding to the target image, the description information is displayed as feedback to the user. The feedback can be in the form of voice, image, video, or any combination of the three. When sitting in the living room or driving, the user can obtain the description information of the object of interest simply by speaking, which simplifies the user's information retrieval operation and improves the user's experience of using electronic devices.

[0115] In some embodiments, in the information retrieval system based on speech recognition provided by an embodiment of the present invention, the image acquisition module 30 is used to:

[0116] Determining target image features corresponding to target text information in a preset target feature library;

[0117] The target image features corresponding to the target text information can be determined based on a preset target feature library, and then the target image can be retrieved from the preset target image library based on the target image features. The information retrieval method provided by the present invention in real time is to determine the target image features corresponding to the target text information obtained by speech recognition, and then retrieve relevant standard images and corresponding description information from a subsequent preset image database based on the target image.

[0118] The preset target feature library can be set on the server side, and the target feature library can be regarded as a target feature set. MA , target feature set O MA The target feature library is a pre-trained feature library. The target feature library includes a recognition module and a template image. The recognition module includes recognition features and a recognition algorithm. The recognition features are the recognition features that need to be extracted when recognizing the image corresponding to the target text information, and the recognition algorithm used when recognizing the image corresponding to the target text information. According to the target text information, the corresponding recognition module and template image are found in the target feature library as the target image feature. TM , and then use these target image features O TM Get the corresponding target image.

[0119] It can be seen that the target text information O T and the target image feature O TM There is a corresponding relationship, for example, the target text information is: high-rise building, the corresponding target image feature O is obtained in the preset target feature libraryTM : Template images, template features, and recognition algorithms of high-rise buildings.

[0120] Based on the target image features, a search is performed in the preset target image library to obtain the target image.

[0121] The preset target image library is set in advance and can collect images related to audio or video previously played by the user's electronic device. This can better meet the user's usage needs. For example, when the user is watching video content, screenshots of the images related to the video content can be collected in real time and saved in the target image library. After the user speaks and needs to obtain descriptive information related to the object of interest, the target image features can be obtained and then searched in the target image library based on the target image features. This search is to compare the target image features with images in the target image library to finally obtain the target image.

[0122] Using the target image feature O TM Perform image retrieval based on image features in the preset target image library C to obtain the target image O R . Get the target image O R Then, if the electronic device used by the user has a display screen, and the electronic device takes a screenshot of the image played by the electronic device, the target image may be displayed on the display screen. R The target image is framed in the image screenshot for the user to view.

[0123] In some embodiments, in the information retrieval system based on speech recognition provided by the embodiments of the present invention, the target feature library includes a recognition module and a template image, and the image acquisition module 30 is further used to:

[0124] Based on the target text information, it is determined whether there is a corresponding recognition module in the target feature library;

[0125] The recognition module includes recognition features and recognition algorithms, wherein the recognition features are the recognition features that need to be extracted when identifying the image corresponding to the target text information, and the recognition algorithm used when identifying the image corresponding to the target text information. According to the target text information, the target feature library is searched to see if there is a corresponding recognition module. If there is, the recognition module is used as the target image feature. TM , and then use these target image features O TM The corresponding target image is obtained in the target image library. The target feature library can be continuously updated and trained in order to better represent the target text information and obtain the target image in the target image library more accurately.

[0126] Of course, the finer the division granularity of the recognition module, the better. Ideally, there can be one recognition module for each object, so that the target text information can be expressed more correctly and accurately, and the target image can be captured more accurately.

[0127] When a corresponding recognition module exists in the target feature library, the corresponding recognition module is used as the target image feature;

[0128] When there is a corresponding recognition module in the target feature library, the recognition module is used as the target image feature. TM , and then use these target image features O TM Get the corresponding target image. For example, if the target text information is a high-rise building, get the recognition features and recognition algorithm of the high-rise building in the target feature library.

[0129] Based on the target text information, it is determined that a corresponding template image exists in the target feature library;

[0130] Next, the target feature library is checked to see if a template image corresponding to the target text exists. This means, for example, if there is an image displaying the target text, such as an image of a high-rise building. If so, the template image is added to the target image features, allowing the target image to be accurately found in the target image library using these features.

[0131] Add the corresponding template image to the target image features.

[0132] When a corresponding template image exists in the target feature library, a template matching strategy is executed, that is, the corresponding template image is added to the target image template.

[0133] In some embodiments, in the information retrieval system based on speech recognition provided by an embodiment of the present invention, the image acquisition module 30 is further configured to:

[0134] When there is no corresponding recognition module in the target feature library, a universal recognition module is used as the target image feature;

[0135] If there is no recognition module corresponding to the target text information in the target special diagnosis library, a general recognition module can be used as the target image feature. The general recognition module here can be determined according to the division granularity of the recognition module, and the closer to the corresponding recognition module, the better.

[0136] Based on the target text information, it is determined that a corresponding template image exists in the target feature library;

[0137] Then determine whether there is a corresponding template image in the target feature library. If there is a corresponding template image, add the corresponding template image to the target image feature. This can more accurately express the target text information and subsequently find the target image more accurately in the target feature library.

[0138] Add the corresponding template image to the target image features.

[0139] When there is a corresponding template image in the target feature library, a template matching strategy is adopted to add the corresponding template image to the target image features, so that these target image features can be used to retrieve the target image in the target image library later.

[0140] Reference Figure 9 As shown, first, for the target text information, it is determined whether there is a recognition module in the target feature library. Depending on the situation, the corresponding recognition module is used as the target image feature. In particular, if there is no recognition module in the target feature library, a universal recognition module is selected as the target image feature. In this case, the universal recognition module can be a division of the corresponding recognition module at the next higher level of granularity. This is not an absolute limitation and can be flexibly set. For example, if there is no recognition module corresponding to a high-rise building, the recognition module corresponding to the building can be selected as the universal recognition module for the high-rise building.

[0141] In some embodiments, in the information retrieval system based on speech recognition provided by an embodiment of the present invention, in response to recognizing a voice wake-up word from speech information, the image acquisition module 30 is further specifically configured to:

[0142] Arrange the multiple display interfaces behind the current display interface in order from near to far, so as to form a retrieval order for real-time image retrieval;

[0143] Specifically, the target image library C consists of the current display interface P T1 and the current display interface P T1 The total number of screenshots is N. The display interfaces are sorted according to the screenshot time as follows:

[0144]

[0145] Then, multiple display interfaces are sequentially arranged behind the current display interface in the order from near to far from the current display interface, forming a retrieval order for real-time image retrieval:

[0146]

[0147] Extract real-time search images according to the search order;

[0148] According to the retrieval order, the real-time retrieval picture is obtained and the real-time retrieval picture captured at this time is O R .P T1 , the next step is to compare the similarity between the target image features and these search images. Figure 10 As shown, each real-time search image can be extracted and compared with the target image features for similarity. If the similarity between the two is not greater than a set threshold, the next real-time search image is extracted until all real-time search images are traversed. Of course, if the similarity between the first real-time search image and the target image features is greater than the set threshold, the real-time search image is used as the target image, and the extraction and similarity comparison of the remaining real-time search images are terminated.

[0149] Determine whether the similarity between the real-time retrieval image and the target image features is greater than the set threshold;

[0150] The target image is obtained by measuring the similarity between the real-time retrieval image and the target image features. It should be noted that the greater the similarity, the closer the real-time retrieval image and the target image features are, and vice versa.

[0151] When the similarity is greater than the set threshold, the search ends and the real-time search image is used as the target image.

[0152] Retrieve the similarity S between the image and the target image features in real time. T1 When the value is greater than the set threshold, the extraction of the real-time retrieval image and the subsequent similarity comparison are terminated, and the real-time retrieval image is used as the target image.

[0153] From the above analysis, it can be seen that the information retrieval method based on speech recognition provided by the embodiments of the present invention combines speech recognition and image retrieval to improve the accuracy of information retrieval. This information retrieval method is suitable for collecting the user's speech when the user speaks about an item of interest, then performing speech recognition from the user's speech to obtain target text information related to the item of interest. Based on the target text information, a target image related to the item of interest is then obtained. The target image is then searched for corresponding description information from a preset image database. The search can be based on an image comparison between the target image and a standard image in the image database to determine the corresponding description information. After the description information is retrieved, at least the description information is fed back to the user. The feedback can take various forms, including playing a voice message or displaying it on a display screen. It can be seen that the information retrieval method based on speech recognition provided by the embodiments of the present invention can feed back the description information corresponding to the item of interest to the user after receiving a voice message from the user. Due to the combination of speech recognition and image retrieval technologies, the accuracy of information retrieval is improved, while the user's search operation is greatly simplified, and the user's experience of using the electronic device is enhanced.

[0154] Example 3

[0155] An embodiment of the present invention provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the following Figures 1 to 10 The steps of the information retrieval method based on speech recognition are as follows:

[0156] Step 10: Collect the user's voice;

[0157] The information retrieval method based on speech recognition provided by the embodiment of the present invention can be applied to electronic devices used by users. After the information retrieval function is started, the user's speech can be collected, with the aim of capturing the target text information of the things that the user is interested in involved in the user's speech.

[0158] You can set the target text information T , target text information O T It can include: high-rise buildings, cars, watches, whales and other keywords from user voice.

[0159] Step 20: Recognize target text information from the user's voice;

[0160] After collecting the voice sent by the user, the target text information can be identified from the user voice. At this time, voice recognition is required. For example, if the user voice sent by the user is "high-rise building", the target text information can be identified by voice recognition. TAfter obtaining the target text information, the corresponding image processing is performed based on the target text information, and finally the corresponding description information is fed back to the user.

[0161] See also Figure 12 As shown, the voice acquisition module can collect the user's voice and then perform voice recognition on the user's voice. The subsequent image processing can be performed on the device side or on the server side, which is not limited here. If the subsequent image processing is performed on the server side, after obtaining the target text information, the target text information is sent to the server to realize that the voice acquisition device collects the target information input by the user's voice, and the voice analysis module in the control device analyzes the target (text information). T , and put the target (text information) T Sent to the server device.

[0162] Step 30: Acquire a target image based on the target text information;

[0163] The corresponding target image is obtained based on the target text information. This step can be based on a preset text-image library, and the target image corresponding to the target text information is selected from the text-image library. If there is a display screen displaying an image related to the target text information, the target image corresponding to the target text information can be obtained from the image displayed on the display screen.

[0164] This step is to combine image retrieval with obtaining the corresponding target image based on the target text information. The image can more accurately reflect the target text information, and thus can more accurately feedback the user's user voice, thereby improving the accuracy of information retrieval.

[0165] Step 40: Retrieve description information corresponding to the target image from a preset image database;

[0166] The pre-set image database contains a large amount of data, and the data storage model is a standard image-description information model. However, the storage format for the description information varies. For example, if the description information involves an animal, the stored description information may include: scientific name, species, distribution area, and protection level. If the description information involves a car, the stored description information may include: make, model, price, etc. If the description information involves a building, the stored description information may include: name, location, characteristics, etc.

[0167] Retrieval based on the target image in a preset image database is to compare the target image with the standard image. Compared with the comparison between text, the comparison between images is closer, and the corresponding description information can be obtained more accurately.

[0168] Step 50: After the description information is retrieved, at least the description information is fed back to the user.

[0169] After retrieving the description information corresponding to the target image, the description information is fed back to the user in the form of voice, image, video, or any combination of the three. When sitting in the living room or driving, the user can simply speak to obtain the description information of the object of interest, which simplifies the user's information retrieval process and improves the user's experience of using electronic devices.

[0170] From the above analysis, it can be seen that the information retrieval method based on speech recognition provided by the embodiments of the present invention combines speech recognition and image retrieval to improve the accuracy of information retrieval. This information retrieval method is suitable for collecting the user's speech when the user speaks about an item of interest, then performing speech recognition from the user's speech to obtain target text information related to the item of interest. Based on the target text information, a target image related to the item of interest is then obtained. The target image is then searched for corresponding description information from a preset image database. The search can be based on an image comparison between the target image and a standard image in the image database to determine the corresponding description information. After the description information is retrieved, at least the description information is fed back to the user. The feedback can take various forms, including playing a voice message or displaying it on a display screen. It can be seen that the information retrieval method based on speech recognition provided by the embodiments of the present invention can feed back the description information corresponding to the item of interest to the user after receiving a voice message from the user. Due to the combination of speech recognition and image retrieval technologies, the accuracy of information retrieval is improved, while the user's search operation is greatly simplified, and the user's experience of using the electronic device is enhanced.

[0171] In short, the above description is only a preferred embodiment of this specification and is not intended to limit the scope of protection of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this specification shall be included in the scope of protection of this specification.

[0172] The systems, devices, modules, or units described in one or more of the above embodiments may be implemented by a computer chip or entity, or by a product having a certain function. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0173] Computer-readable storage media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0174] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a..." does not preclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0175] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

[0176] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. An information retrieval method based on speech recognition, comprising: Collect voice information; In response to the voice wake-up word recognized from the voice information, saving the current display interface played by the display screen at the current moment and multiple display interfaces played before and after the current display interface to a preset target image library; identifying target text information from the voice information; Determining a target image feature corresponding to the target text information in the preset target feature library; Arrange the multiple display interfaces in order from near to far from the current display interface, behind the current display interface, to form a retrieval order for real-time image retrieval; Extracting the real-time search image according to the search order; Determine whether the similarity between the real-time search image and the target image feature is greater than a set threshold; When the similarity is greater than the set threshold, the search is terminated and the real-time search image is used as the target image; Searching a preset image database according to the target image; After the description information corresponding to the target image is retrieved, at least the description information is displayed.

2. The information retrieval method according to claim 1, wherein the target feature library includes a recognition module and a template image, and determining the target image features corresponding to the target text information in the preset target feature library comprises: Determine whether there is a corresponding recognition module in the target feature library based on the target text information; When a corresponding recognition module exists in the target feature library, determining the corresponding recognition module as the target image feature; Determining, based on the target text information, that a corresponding template image exists in the target feature library; The corresponding template image is added to the target image features.

3. The information retrieval method according to claim 2, further comprising: after determining whether a corresponding recognition module exists in the target feature library based on the target text information; When there is no corresponding recognition module in the target feature library, a universal recognition module is used as the target image feature; Determining, based on the target text information, that a corresponding template image exists in the target feature library; The corresponding template image is added to the target image features.

4. The information retrieval method according to claim 1, further comprising: after determining whether the similarity between the real-time retrieval image and the target image feature is greater than a set threshold; When the similarity is not greater than the set threshold, the real-time retrieval picture corresponding to the one with the largest similarity is used as the target image.

5. The information retrieval method according to claim 1, wherein the preset image database includes standard images and description information of the standard images; Retrieving description information corresponding to the target image in a preset image database includes: Separating a corresponding image data subset from the preset image database according to the target text information; An image search is performed in the image data subset based on the target image to obtain the standard image and the description information corresponding to the target image.

6. The information retrieval method according to claim 5, wherein after the description information is retrieved, at least the description information is displayed, comprising: After the standard image corresponding to the target image is acquired, the standard image and the description information are displayed.

7. A speech recognition-based information retrieval system comprising: Voice collection module, used to collect voice information; The voice acquisition module is further configured to save the current display interface played by the display screen at the current moment and multiple display interfaces played before and after the current display interface to a preset target image library in response to the voice wake-up word recognized from the voice information; A text recognition module, configured to recognize target text information from the voice information; An image acquisition module, configured to determine a target image feature corresponding to the target text information in the preset target feature library; The image acquisition module is further configured to arrange the multiple display interfaces behind the current display interface in order from near to far from the current display interface, so as to form a retrieval order for real-time image retrieval; The image acquisition module is further configured to extract the real-time search image according to the search order; The image acquisition module is further configured to determine whether the similarity between the real-time search image and the target image feature is greater than a set threshold; The image acquisition module is further configured to terminate the search and use the real-time searched image as the target image when the similarity is greater than the set threshold; A retrieval module, configured to search a preset image database based on the target image; The feedback module is configured to at least display the description information after retrieving the description information corresponding to the target image.

8. A storage medium for computer-readable storage, wherein the storage medium stores one or more programs, and when the one or more programs can be executed by one or more processors, the steps of the information retrieval method based on speech recognition as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Information indexing unit and information indexing method

    CN104424257A

  • Video index tag setting method and apparatus, and server

    CN107679227A

  • A voice search method and device based on a webpage video

    CN109697245A