Information pushing method and device based on voice recognition, equipment and medium
Patent Information
- Application Number
- CN202211384659.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-07
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-11-07
AI Technical Summary
[0005]本申请提供一种基于语音识别的信息推送方法、装置、设备和介质,用于解决现有手工输入文本进行信息搜索效率的问题
[0044]本申请实施例提供的基于语音识别的信息推送方法、装置、设备和介质,通过为应用软件新增语音处理功能,获取用户输入至终端设备的语音并提取语音中的关键指示信息,然后可以自动搜索到与关键指示信息匹配的目标信息,并显示在终端设备的显示界面上,如此避免了用户手动输入的过程,提高了信息查询和业务办理的效率。
Smart Images

Figure CN115831112B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial technology, and in particular to an information push method, device, equipment and medium based on voice recognition. Background Technology
[0002] In the field of fintech, various applications are typically developed to facilitate user interaction, allowing users to conduct relevant financial transactions. However, when using these applications, users often find it difficult to quickly find the information they need due to the complex financial information involved, such as dates, times, and card numbers.
[0003] In the prior art, in order to facilitate users to quickly find information, application software usually has a search bar. Users manually enter the target words to be searched in the search bar, and then the application software retrieves related information based on the target words and displays it on the page.
[0004] However, this method requires users to manually input target words, making it inefficient and unable to quickly complete information retrieval in practice. Summary of the Invention
[0005] This application provides a speech recognition-based information push method, apparatus, device, and medium to solve the problem of information search efficiency when manually inputting text.
[0006] In a first aspect, embodiments of this application provide an information push method based on speech recognition, including:
[0007] Acquire the voice input to the terminal device and at least one page contained in the target application on the terminal device;
[0008] Based on the voice, obtain key instruction information;
[0009] Extract at least one keyword from the at least one page and combine them to obtain a semantic set;
[0010] Based on the key indication information and semantic set, target information is determined, and the target information is used to display on the display interface of the terminal device.
[0011] In one possible design of the first aspect, determining the target information based on the key indication information and the semantic set includes:
[0012] Determine whether there are target keywords in the semantic set that match the key indication information;
[0013] If the target keyword exists, then the page containing the target keyword is determined in the at least one page.
[0014] The page containing the target keyword is taken as the target information.
[0015] In another possible design of the first aspect, the method further includes:
[0016] If the target keyword does not exist in the semantic set, obtain the user information stored in the target application;
[0017] Determine whether there is matching information associated with the user information;
[0018] If the matching information exists, then the matching information is used as the target information.
[0019] In another possible design of the first aspect, determining whether there is a target keyword in the semantic set that matches the key indication information includes:
[0020] Obtain the correlation degree between each keyword information in the semantic set and the key indication information, wherein the correlation degree is used to characterize the vector distance between the keyword information and the key indication information;
[0021] If there are keyword information with a relevance greater than a preset threshold, then the keyword information is used as the target keyword.
[0022] In another possible design of the first aspect, before determining whether a target keyword matching the key indication information exists in the semantic set, the method further includes:
[0023] Obtain sample keywords from the page currently displayed on the terminal device and sample indication information associated with the sample keywords. The sample keywords include at least one of date keywords and application operation keywords of the target application.
[0024] The similarity between the sample keywords and sample indication information is obtained, and the preset model is trained and updated based on the similarity to obtain a prediction model. The prediction model is used to predict whether the key indication information matches the keyword information in the semantic set.
[0025] In another possible design of the first aspect, determining the target information based on the key indication information and the semantic set includes:
[0026] Determine whether there are target keywords in the semantic set that match the key indication information;
[0027] If the target keyword exists, then obtain the display information associated with the target keyword;
[0028] The displayed information is used as the target information.
[0029] In another possible design of the first aspect, obtaining key instruction information based on the voice includes:
[0030] The speech is preprocessed, and the preprocessing includes at least one of speech filtering, speech sampling, and speech quantization;
[0031] The preprocessed speech is converted into text information, and semantic analysis is performed on the text information to extract the key indication information.
[0032] In yet another possible design of the first aspect, the method further includes:
[0033] If the matching information does not exist, then the preset matching failure information is obtained as the target information.
[0034] Secondly, embodiments of this application provide an information push device based on voice recognition, comprising:
[0035] The voice acquisition module is used to acquire voice input to the terminal device and at least one page contained in the target application in the terminal device;
[0036] The information acquisition module is used to acquire key instruction information based on the voice.
[0037] The information extraction module is used to extract at least one keyword from the at least one page and combine them to obtain a semantic set;
[0038] The information determination module is used to determine target information based on the key indication information and semantic set, and the target information is used to display on the display interface of the terminal device.
[0039] Thirdly, embodiments of this application provide an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0040] The memory stores computer-executed instructions;
[0041] The processor executes computer execution instructions stored in the memory to implement the above-described method.
[0042] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer instructions that, when executed by a processor, are used to implement the above-described method.
[0043] Fifthly, embodiments of this application provide a computer program product, including computer instructions that, when executed by a processor, implement the above-described method.
[0044] The information push method, apparatus, device, and medium based on voice recognition provided in this application embodiment acquire the voice input by the user to the terminal device and extract key indication information from the voice by adding voice processing function to the application software. Then, it can automatically search for target information that matches the key indication information and display it on the display interface of the terminal device. This avoids the process of manual input by the user and improves the efficiency of information query and business processing. Attached Figure Description
[0045] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application;
[0046] Figure 1 A schematic diagram of the application software provided in the embodiments of this application;
[0047] Figure 2 A flowchart illustrating the information push method based on speech recognition provided in this application embodiment;
[0048] Figure 3 A schematic diagram of the page of the application software provided in another embodiment of this application;
[0049] Figure 4 A schematic diagram of the voice processing flow provided in the embodiments of this application;
[0050] Figure 5 This is a schematic diagram of the structure of an information query system based on speech recognition provided in an embodiment of this application;
[0051] Figure 6 A schematic diagram of the application software provided in yet another embodiment of this application;
[0052] Figure 7 A schematic diagram of the application software provided in another embodiment of the application;
[0053] Figure 8 This is a schematic diagram of the structure of the information push device based on voice recognition provided in the embodiments of this application;
[0054] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0055] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0057] It should be noted that the information push method, apparatus, device, and medium based on voice recognition disclosed herein can be used in the financial field. It can also be used in any field other than finance. The application fields of the information push method, apparatus, device, and medium based on voice recognition disclosed herein are not limited.
[0058] In daily life, many users rely on applications installed on mobile devices (such as smartphones, tablets, and computers) to handle personal business. For example, a user can enter "query personal business history" in the application and click "query." The application will then display the business transactions the user has processed within a specific historical period. Therefore, in the process of handling business, the application requires the user to manually input the query conditions to output the corresponding results. However, in practice, users may be unable to manually input query conditions due to personal factors, such as needing both hands to operate the steering wheel while driving, or having injured or disabled hands that make operation difficult. In such cases, the business process encounters significant obstacles, greatly reducing efficiency, which may be even less efficient than processing at a human window.
[0059] To address the aforementioned issues, this application provides a method, apparatus, device, and medium for information push based on speech recognition. To improve the efficiency of users conducting business using application software, it is necessary to optimize the input function of the application software. Specifically, a speech processing function is added to the application software to acquire the user's speech input to the terminal device and extract key instruction information from the speech. Then, the system can automatically search for target information matching the key instruction information and display it on the terminal device's display interface. This avoids the process of manual input by the user and improves the efficiency of information retrieval and business processing.
[0060] The technical solution of this application will now be described in detail through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0061] Figure 1 The interface diagram of the application software provided in the embodiments of this application is as follows: Figure 1 As shown, taking a mobile phone as an example, the current page displayed on the phone's screen includes a voice button (which the user can long-press to trigger voice capture) and recommended quotes (such as checking history, checking popular events, handling personal business, and handling other business). The user can long-press the voice button and then speak, for example, saying "check history." The terminal device then captures the user's spoken voice, performs voice recognition, finds the user's history records, and displays them on the current page. This eliminates the need for manual text input during the history check process, greatly improving efficiency.
[0062] Example 1
[0063] Figure 2 This is a flowchart illustrating a speech recognition-based information push method provided in an embodiment of this application. This method can be applied to a terminal device or a cloud server connected to the terminal device; that is, the method can be executed on either the terminal device or the cloud server. Taking the cloud server as the executing entity as an example, the method can be implemented through the following steps: Step S201, obtaining the voice input to the terminal device and at least one page contained in the target application on the terminal device.
[0064] In this embodiment, the terminal device carries a microphone. User actions can trigger the microphone to activate and collect sounds from the surrounding environment. After activating the microphone, the user can express their needs via voice, replacing the traditional method of manually inputting text. For example, the target application could be a financial application. Different applications may involve different business functions, and the user may need to adjust the content of their voice input accordingly.
[0065] For example, if the target application is a financial application, the user's voice input could be "check history," while if the target application is a navigation application, the user's voice input might be "navigate home."
[0066] In this embodiment, the target application may include one or more pages, each containing different information. For example, the target application may have a main page with multiple different navigation tabs. Users can navigate to corresponding subpages by clicking these navigation tabs, and the subpages contain detailed tab information. For example, Figure 3 This is a schematic diagram of an application software page provided in another embodiment of this application. Again, taking a mobile phone as the terminal device, after opening the application software on the mobile phone and entering its main page, as shown... Figure 3 As shown, its main page includes multiple navigation tabs, such as "Utility Payments," "Task Center," "Activity Hall," "Local Services," "Concierge Services," "Password Management," "Gold," and "Others." Users can manually click on these navigation tabs to navigate to the corresponding subpages. For example, when a user clicks "Utility Payments," they are taken to the "Utility Payments" subpage, which includes sub-navigation tabs such as "Water Bill Payment," "Electricity Bill Payment," and "Gas Bill Payment." Furthermore, if a user clicks "Water Bill Payment," they are taken to the water bill payment page to pay their water bill.
[0067] Users can also use voice commands instead of manual clicking. For example, if a user says "water bill payment," the page will redirect directly to the water bill payment page.
[0068] Step S202: Obtain key instruction information based on the voice.
[0069] In this embodiment, the voice may contain important information related to the application software. For example, if the user says, "I want to query the detailed records for October 2022," based on the information contained in each page of the application software, the key indication information can be determined as "October 2022" and "detailed records." The key indication information is in text form, while the voice is in audio form. Therefore, when obtaining the key indication information, it is necessary to first recognize the voice, convert it into text, and then extract the key indication information from the converted text.
[0070] In addition, in this embodiment, since the cloud server is used as the execution entity, the terminal device needs to upload the voice to the cloud server, which will then perform voice recognition and extraction of key instruction information. This way, the performance resources of the terminal device are not occupied, and other functions of the terminal device are not affected.
[0071] Step S203: Extract at least one keyword from at least one page and combine them to obtain a semantic set.
[0072] In this embodiment, the application software may include the aforementioned main page, and may also include subpages, with the main page and subpages being related (i.e., one can navigate to any subpage through the main page). Both the main page and subpages contain corresponding keyword information. This keyword information may refer to functional tags (such as the aforementioned navigation tags) present on the main page or subpages, through which corresponding functions can be performed; for example, the aforementioned "utility bill payment" allows for bill payment.
[0073] The semantic set can include multiple keyword information, each of which can originate from a different page. The association between each keyword and its source page can be pre-stored. For example, taking "utility bill payment" as a keyword, which originates from the main page, the association between the keyword information and the page can be pre-built to mark "utility bill payment" as originating from the main page.
[0074] Step S204: Based on the key indication information and semantic set, the target information is determined and displayed on the display interface of the terminal device.
[0075] In this embodiment of the application, the key indication information can be compared one by one with the keyword information in the semantic set. If the semantic set contains the same target keyword as the key indication information, the target information can be determined based on the target keyword and then displayed on the display interface of the terminal device.
[0076] Alternatively, in other implementations, feature vectors for key indication information and keyword information can be calculated. Based on the vector distance between the feature vectors of the key indication information and the feature vectors of each keyword information, the keyword information with the smallest vector distance can be identified as the target keyword.
[0077] In this embodiment, after determining the target keyword, if the target keyword is associated with a subpage, the subpage can be displayed as target information on the terminal device's display interface. Furthermore, in other embodiments, information associated with the target keyword can be pre-set and used as target information.
[0078] This application embodiment acquires the user's voice and extracts key instruction information from the voice. Then, it can automatically search for the target page that matches the key instruction information and display it on the display interface of the terminal device. This avoids the process of manual input by the user and improves efficiency.
[0079] Example 2
[0080] In this embodiment, step S203 can be implemented in the following ways: determine whether there is a target keyword in the semantic set that matches the key indication information; if there is a target keyword, determine the page where the target keyword is located in at least one page; and use the page where the target keyword is located as the target information.
[0081] For example, if a user says "I want to pay my utility bill," the key instruction information obtained through speech recognition and feature extraction is "utility bill payment." In this case, the target keyword matching this in the semantic set is "utility bill payment." (Refer to...) Figure 3 The page containing the target keyword is the main page. In this case, the main page can be used as the target information and displayed on the terminal device's display interface.
[0082] For example, keywords in the semantic set can be categorized according to certain categories. Keywords related to fee changes, such as utility bill payments, electricity bill payments, and water bills, can be grouped into one category. Keywords related to activities, such as task centers and activity halls, can be grouped into another category. This allows for targeted matching and searching based on categories, leading to faster identification of target keywords.
[0083] This application embodiment directly displays the page containing the target keyword on the terminal device's display interface, eliminating the need for manual user operation and improving the interaction efficiency between the user and the application software.
[0084] Example 3
[0085] In this embodiment, when there is no target keyword in the semantic set, the cloud server can obtain the user information stored in the target application, and then check whether there is matching information associated with the user information. If there is matching information, the matching information is used as the target information.
[0086] This can be achieved by pre-storing matching information associated with user information in a database on a cloud server. For example, the matching information could be the user's previous activity records; such as if the user had previously made a utility bill payment, the matching information would be "utility bill payment".
[0087] In other implementations, the matching information associated with user information may also be an account, card number, etc. For different target applications, the user information may be different, and the matching information associated with the user information may also be different. The specific information can be preset according to the actual situation.
[0088] In other embodiments, if no matching information associated with the user information exists, the cloud server directly returns a pre-set matching failure message as the target information and displays it on the terminal device's display interface, prompting the user to enter other voice inputs again.
[0089] This application embodiment sets matching information so that when the target keyword cannot be found in the semantic set, the matching information can be displayed to the user, thereby improving the interaction between the user and the application software.
[0090] Example 4
[0091] Based on the above embodiment 2, in this embodiment, when determining whether there is a target keyword in the semantic set that matches the key indication information, it can be achieved through the following steps: obtain the correlation degree between each keyword information in the semantic set and the key indication information, and the correlation degree is used to characterize the vector distance between the keyword information and the key indication information; if there is keyword information with a correlation degree greater than a preset threshold, then the keyword information is taken as the target keyword.
[0092] In this embodiment, a predetermined association relationship is generated for the key indication information to be determined based on a semantic set. The similarity vector of the predetermined association relationship is used as an update sample set to obtain an association prediction model. The cloud server can call the association prediction model to determine whether there are keyword information matching the key indication information in the semantic set, and the association degree between the keyword information and the key indication information needs to be greater than a preset threshold. Here, vector distance refers to the vector distance between the feature vector of the keyword information and the feature vector of the key indication information.
[0093] When keyword information exists in the semantic set, the server can send a display information identifier, such as "matching information successful", to display the corresponding information content on the terminal device's display interface for the user's convenience.
[0094] In other implementations, the correlation between each keyword information and key indication information in the semantic set can be obtained, and the information can be sorted according to the correlation degree. The keyword information with the highest correlation degree is selected as the target keyword.
[0095] This application embodiment calculates the correlation between keyword information and key indication information, and selects keyword information with a correlation greater than a preset threshold as target keywords, which can improve the accuracy of information push and avoid displaying inaccurate information to users on the display interface of terminal devices.
[0096] Example 5
[0097] Based on Embodiments 2 and 4 above, in some other embodiments, the above method may further include the following steps: obtaining sample keywords and sample indication information associated with the sample keywords in the page currently displayed by the terminal device, wherein the sample keywords include at least one of date keywords and application operation keywords of the target application; obtaining the similarity between the sample keywords and the sample indication information, and training and updating the preset model according to the similarity to obtain a prediction model, wherein the prediction model is used to predict whether the key indication information matches the keyword information in the semantic set.
[0098] In this embodiment, the prediction model can be the correlation prediction model in Embodiment 4 above. The correlation prediction model can be pre-stored in the model library. When the cloud server needs to determine whether the key indication information matches the keyword information in the semantic set, it can directly call the correlation prediction model from the model library and perform measurement estimation.
[0099] For example, taking financial application software on a terminal device as an example, the currently displayed page may include sample keywords such as the user's transfer operation (i.e., application operation keyword), transaction amount, and transaction date (i.e., date keyword). The sample instruction information is determined by the user's voice. For example, if the user's voice is "I want to query the transfer details for January 2022," the corresponding sample instruction information would be "January 2022" and "transfer details." Here, "transfer details" is associated with the sample keyword "transfer operation," and "January 2022" is associated with the sample keyword "transfer date." The feature vectors of the sample keywords and the sample instruction information can be calculated, and the similarity is determined based on the vector distance between the two feature vectors. This similarity is used as training samples to train a prediction model, enabling the trained prediction model to predict the degree of matching between the key instruction information and the keyword information.
[0100] This application embodiment trains a preset model to obtain a prediction model to predict whether key indication information matches keyword information in the semantic set, which can accurately find the corresponding information and display it on the terminal device's display interface, thereby improving the accuracy of information display.
[0101] Example 6
[0102] In this embodiment, determining the target information can be achieved through the following steps: determining whether there is a target keyword in the semantic set that matches the key indication information; if there is a target keyword, obtaining the display information associated with the target keyword; and using the display information as the target information.
[0103] In this embodiment, associated display information can be set for each keyword and stored in the database in advance. When a target keyword that matches the key indication information is found, the display information associated with the target keyword can be retrieved from the database as the target information.
[0104] For example, using the card number "6225001234567890" as the target keyword, you can set the display information associated with the card number, including historical expenditure amount, historical expenditure amount date, historical transfer amount, and historical transfer amount date.
[0105] This application embodiment pre-configures display information associated with each keyword. When there is a target keyword in the semantic set that matches the key indication information, the display information associated with the target keyword can be directly displayed on the display interface of the terminal device, thereby improving the flexibility of information display.
[0106] Example 7
[0107] In this embodiment, the acquisition of key indication information can be achieved through the following steps: preprocessing the speech, including at least one of speech filtering, speech sampling, and speech quantization; converting the preprocessed speech into text information, and performing semantic analysis on the text information to extract the key indication information.
[0108] This embodiment requires the use of speech recognition technology. Speech recognition, also known as Automatic Speech Recognition (ASR), converts the lexical content of human language sentences into computer-readable input. Speech recognition technology mainly includes three aspects: feature extraction, pattern matching rules, and model training. First, feature extraction and pattern matching techniques are used to identify relevant acoustic models of speech. Then, model training forms a language model. During implementation, rapid optimization of the speech is performed to achieve the text-to-speech conversion of the speech signal.
[0109] Speech recognition systems can be categorized based on restrictions imposed on the input speech:
[0110] ① Based on the way the actor speaks, it can be divided into three categories: (1) Connecting word speech recognition system: requires clear pronunciation of each word, and some connected speech phenomena will occur; (2) Isolated word speech recognition system: requires a pause after each word is input; (3) Continuous speech recognition system: requires natural and fluent continuous speech input, and a large number of connected speech and sound changes will occur.
[0111] ② Based on the vocabulary size of the recognition system, it can be divided into three categories: (1) small vocabulary speech recognition system: including speech recognition systems with dozens of words; (2) medium vocabulary speech recognition system: including recognition systems with hundreds to thousands of words; (3) large vocabulary speech recognition system: including speech recognition systems with thousands to tens of thousands of words. With the improvement of computer computing power and recognition system accuracy, the classification of vocabulary-based speech recognition systems is constantly changing.
[0112] ③ Considering the correlation between the actor and the recognition system, it is divided into three categories: (1) Person-independent speech recognition system: the recognized speech is unrelated to the person; (2) Person-specific speech recognition system: considers the recognition of the speech of a specific person; (3) Multi-person speech recognition system: can recognize the speech of a group of people.
[0113] In this embodiment, different speech recognition systems can be selected based on the user's input speech. Simultaneously, during the speech recognition process, the input analog speech signal first undergoes preprocessing, specifically including speech filtering, speech sampling, and speech quantization. After preprocessing, the analog speech signal is converted into a digital signal through analog-to-digital conversion, from which the essential features of the speech are extracted. Specifically, before extracting the essential features of the speech, common short-time analysis techniques can be employed, such as speech signal digitization, endpoint detection, pre-emphasis, and framing, to achieve data compression.
[0114] When extracting the essential features of speech and converting them into text, audio feature extraction technology is involved. Based on this technology, audio feature similarity retrieval is performed. A common method for extracting audio features is using Hidden Markov Models (HMMs). This model is a generative model of dynamic Bayesian networks with the simplest structure. It effectively solves the problem of speech signals being stable in the short term but time-varying in the long term. It can also construct sentence models of continuous speech based on some basic modeling units, and has relatively high modeling accuracy and flexibility. Depending on the size of the collected speech units, HMMs can be divided into whole-word HMMs. The advantage is that it can well describe the characteristics of intra-word phoneme co-pronunciation. Therefore, many small-vocabulary speech recognition systems use whole-word HMMs. For example, if a user says, "I want to check the transaction details of bank card 6225001234567890 from January 1, 2021 to December 31, 2021", the corresponding speech will be converted into text information.
[0115] In this embodiment, after obtaining the text information, key indication information is extracted through semantic analysis. Specifically, a knowledge graph or knowledge base can be used to analyze and obtain key indication information through Natural Language Processing (NLP). For example, in the text "I want to query the transaction details of bank card 6225001234567890 from January 1, 2021 to December 31, 2021", the three key indication information "6225001234567890", "2021-01-01", and "2021-12-31" will be extracted.
[0116] After obtaining the key instructions, since the cloud server uses machine language while the key instructions are in Chinese characters, it is necessary to convert the Chinese key instructions into machine language. The cloud server then uses this machine language to search for matching keywords from a semantic set. For example, the keywords in the semantic set can be stored in a certain index format for easy searching, such as in the form of card number + date.
[0117] This application embodiment, by processing the speech, can obtain accurate key instruction information, improve the effect of voice interaction with users, and avoid speech misrecognition.
[0118] Example 8
[0119] For example, Figure 4 A schematic diagram of the voice processing flow provided in the embodiments of this application, such as Figure 4 As shown, it includes the following steps: Step S401, the front-end sensing device acquires the input speech. Step S402, the front-end server preprocesses the speech. Step S403, the front-end server extracts features from the preprocessed speech. Step S404, metric estimation. Step S405, post-processing. Step S406, the speech processing server obtains the recognition result.
[0120] In this embodiment, the front-end sensing device can be a terminal device. After obtaining the user's voice information, the front-end sensing device sends the user's voice information to the front-end server. After receiving the user's voice information, the front-end server recognizes the user's voice and extracts key indicative information, such as when the user says, "I want to check the transfer details for January 2022." The server obtains the page information currently presented to the user and extracts keyword groups, such as "January 2022" and "transfer," to generate a semantic set. The server performs measurement estimation based on a prediction model to determine whether there is information in the semantic set that matches the user's key indicative information, and whether the correlation between the matching information and the user's indicative information is greater than a set threshold.
[0121] If a matching information exists in the semantic set, the front-end server will send a display information identifier, such as "matching information successful," to the speech processing server. Upon receiving the identifier, the speech processing server will provide the corresponding information to the front-end sensing device, allowing the device to present the speech-recognized information to the user. If no matching information exists in the semantic set, the front-end server will check the database for a matching information associated with the user's information. If a match is found, the above steps will be repeated; otherwise, a specific prompt, such as "matching information failed," will be sent to the speech processing server.
[0122] Example 9
[0123] Furthermore, Figure 5 This is a schematic diagram of the structure of an information query system based on voice recognition provided in an embodiment of this application. This information query system can be applied to the aforementioned information push method; that is, after obtaining the corresponding information through the information query system, it is pushed to the display interface of the terminal device for display. For example... Figure 5 As shown, the information query system includes a voice acquisition module 510, a voice recognition module 520, a semantic analysis module 530, an information query processing module 540, and a database module 550. The voice acquisition module collects voice information, temporarily stores the collected voice information, and sends the stored voice information to the voice recognition module. The voice recognition module performs speech-to-text processing, receiving the voice information, performing voice detection, and sending the converted text information to the semantic analysis module. The semantic analysis module extracts key information from the text information, typically using knowledge graphs or knowledge bases, through natural language analysis, and sends the analysis results to the historical data query processing module. The historical data query processing module converts the received key information into database-recognizable query machine statements, sends them to the database module, extracts data, and performs range queries. The database module stores data information according to a certain index for easy querying.
[0124] For example, Figure 6 A schematic diagram of the application software provided in another embodiment of this application, such as... Figure 6 As shown, users can input voice by long-pressing the virtual microphone button. After the cloud server recognizes and analyzes the voice, the results are displayed on the terminal device's screen for user confirmation. For example, the result after recognition and analysis might be "Help me check the details for July 2022." Once the user confirms that there are no problems, the cloud server then searches for the target information based on the recognition and analysis results.
[0125] For example, Figure 7A schematic diagram of the application software provided in another embodiment of the application, such as... Figure 7 As shown, the user's historical transaction details for July 2022 can be retrieved based on the key information "July 2022" and "details".
[0126] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0127] Figure 8 This is a schematic diagram of the structure of a speech recognition-based information push device 800 provided in an embodiment of this application. The information push device 800 includes a speech acquisition module 810, an information acquisition module 820, an information extraction module 830, and an information determination module 840. The speech acquisition module 810 is used to acquire speech input to the terminal device and at least one page contained in the target application on the terminal device. The information acquisition module 820 is used to acquire key instruction information based on the speech. The information extraction module 830 is used to extract at least one keyword from at least one page and combine them to obtain a semantic set. The information determination module 840 is used to determine the target information based on the key instruction information and the semantic set, and the target information is used to display on the display interface of the terminal device.
[0128] Optionally, the information determination module can be used to: determine whether there is a target keyword in the semantic set that matches the key indication information; if there is a target keyword, determine the page where the target keyword is located in at least one page; and use the page where the target keyword is located as the target information.
[0129] Optionally, it also includes a matching information association module, which is used to obtain user information stored in the target application when the target keyword does not exist in the semantic set; determine whether there is matching information associated with the user information; and if there is matching information, use the matching information as the target information.
[0130] Optionally, the information determination module can be used to: obtain the correlation degree between each keyword information and key indication information in the semantic set, whereby the correlation degree is used to characterize the vector distance between the keyword information and the key indication information; if there is keyword information with a correlation degree greater than a preset threshold, then the keyword information is used as the target keyword.
[0131] Optionally, it also includes a model training module, used to obtain sample keywords and sample indication information associated with the sample keywords in the page currently displayed on the terminal device. The sample keywords include at least one of date keywords and application operation keywords of the target application. The module obtains the similarity between the sample keywords and the sample indication information, and trains and updates the preset model based on the similarity to obtain a prediction model. The prediction model is used to predict whether the key indication information matches the keyword information in the semantic set.
[0132] Optionally, the information determination module can also be used to determine whether there is a target keyword in the semantic set that matches the key indication information; if there is a target keyword, then obtain the display information associated with the target keyword; and use the display information as the target information.
[0133] Optionally, the information acquisition module can be used to: preprocess the speech, including at least one of speech filtering, speech sampling, and speech quantization; convert the preprocessed speech into text information, and perform semantic analysis on the text information to extract key instruction information.
[0134] Optionally, a matching failure module is also included, which is used to obtain preset matching failure information as target information if no matching information exists.
[0135] The apparatus provided in this application embodiment can be used to execute the methods in the above embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.
[0136] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented by processing element calls to software, while others are implemented in hardware. For example, the information determination module can be a separate processing element, or it can be integrated into a chip in the above device. Alternatively, it can be stored as program code in the memory of the above device, and its function can be called and executed by a processing element. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.
[0137] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 9As shown, the electronic device 900 includes at least one processor 910, a memory 920, a bus 930, and a communication interface 940. The processor 910, communication interface 940, and memory 920 communicate with each other via the bus 930. The communication interface is used to communicate with other devices. This communication interface includes a communication interface for data transmission and a display interface or operation interface for human-computer interaction. The processor executes computer instructions stored in the memory, specifically performing the relevant steps in the methods described in the above embodiments. The processor may be a central processing unit, an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. The one or more processors included in the electronic device may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.
[0138] Memory is used to store instructions executed by a computer. Memory may include high-speed RAM, and may also include non-volatile memory, such as at least one disk drive.
[0139] This embodiment also provides a readable storage medium storing computer instructions. When at least one processor of the electronic device executes the computer instructions, the electronic device executes the speech recognition-based information push method provided in the various embodiments described above.
[0140] This embodiment also provides a program product including computer instructions stored in a readable storage medium. At least one processor of the electronic device can read the computer instructions from the readable storage medium, and the processor executes the computer instructions to cause the electronic device to implement the voice recognition-based information push method provided in the various embodiments described above.
[0141] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates an "or" relationship between the preceding and following related objects; in formulas, the character " / " indicates a "division" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0142] It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application. In the embodiments of this application, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. An information push method based on speech recognition, characterized in that, include: Acquire the voice input to the terminal device and at least one page contained in the target application on the terminal device; Based on the voice, obtain key instruction information; Extract at least one keyword from the at least one page and combine them to obtain a semantic set; the keyword information refers to the functional tags present on the page; Based on the key indication information and semantic set, target information is determined, and the target information is used to display on the display interface of the terminal device; The step of determining the target information based on the key indication information and semantic set includes: Determine whether there are target keywords in the semantic set that match the key indication information; If the target keyword exists, then the page containing the target keyword is determined in the at least one page. The page containing the target keyword is taken as the target information; If the target keyword is not found in the semantic set, obtain the user information stored in the target application; Determine whether there is matching information associated with the user information; the matching information is the user's historical operation records; If the matching information exists, then the matching information is used as the target information; The step of determining the target information based on the key indication information and semantic set includes: Determine whether there are target keywords in the semantic set that match the key indication information; If the target keyword exists, then obtain the display information associated with the target keyword, which is pre-set field data stored in the database; The displayed information is used as the target information; Before determining whether a target keyword matching the key indication information exists in the semantic set, the method further includes: Obtain sample keywords from the page currently displayed on the terminal device and sample indication information associated with the sample keywords. The sample keywords include at least one of date keywords and application operation keywords of the target application. The similarity between the sample keywords and sample indication information is obtained, and the preset model is trained and updated based on the similarity to obtain a prediction model. The prediction model is used to predict whether the key indication information matches the keyword information in the semantic set.
2. The method according to claim 1, characterized in that, Determining whether a target keyword matching the key indication information exists in the semantic set includes: Obtain the correlation degree between each keyword information in the semantic set and the key indication information, wherein the correlation degree is used to characterize the vector distance between the keyword information and the key indication information; If there are keyword information with a relevance greater than a preset threshold, then the keyword information is used as the target keyword.
3. The method according to claim 1 or 2, characterized in that, The step of obtaining key instruction information based on the voice includes: The speech is preprocessed, and the preprocessing includes at least one of speech filtering, speech sampling, and speech quantization; The preprocessed speech is converted into text information, and semantic analysis is performed on the text information to extract the key indication information.
4. The method according to claim 1, characterized in that, The method further includes: If the matching information does not exist, then the preset matching failure information is obtained as the target information.
5. An information push device based on speech recognition, characterized in that, include: A voice acquisition module is used to acquire voice input to a terminal device and at least one page contained in a target application on the terminal device; The information acquisition module is used to acquire key instruction information based on the voice. The information extraction module is used to extract at least one keyword from the at least one page and combine them to obtain a semantic set; the keyword information refers to the functional tags present on the page; The information determination module is used to determine target information based on the key indication information and semantic set, and the target information is used to display on the display interface of the terminal device; The information determination module is specifically used to determine whether there is a target keyword matching the key indication information in the semantic set; if the target keyword exists, the page where the target keyword is located is determined in the at least one page; and the page where the target keyword is located is taken as the target information. If the target keyword is not found in the semantic set, obtain the user information stored in the target application; Determine whether there is matching information associated with the user information; if the matching information exists, then use the matching information as the target information; the matching information is the user's historical operation record; The information determination module is further configured to determine whether there is a target keyword in the semantic set that matches the key indication information; if the target keyword exists, then the display information associated with the target keyword is obtained, wherein the display information is field data that is pre-set and stored in the database; and the display information is used as the target information. The model training module is used to obtain sample keywords and sample indication information associated with the sample keywords in the page currently displayed by the terminal device. The sample keywords include at least one of date keywords and application operation keywords of the target application. The module obtains the similarity between the sample keywords and the sample indication information, and trains and updates the preset model based on the similarity to obtain a prediction model. The prediction model is used to predict whether the key indication information matches the keyword information in the semantic set.
6. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, are used to implement the method as described in any one of claims 1-4.
8. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the method described in any one of claims 1-4.
Citation Information
Patent Citations
Page switching method and device of application program, computer equipment and storage medium
CN110147216A