Voice search method and apparatus, storage medium, and electronic device
By displaying search terms and recommended text sets during voice search, and allowing users to select and merge text, the problem of inaccurate voice search results is solved, resulting in higher search accuracy and better reflection of intent.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2022-08-04
- Publication Date
- 2026-07-24
Smart Images

Figure CN117555982B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and more specifically, to a voice search method and apparatus, storage medium, and electronic device. Background Technology
[0002] Currently, in voice search scenarios, existing voice search products only convert the user's voice content into text. When a user initiates a voice-interactive search, the system performs text recognition based on the user's voice and initiates a search based on the recognized text, similar to text search. This approach cannot fully express the user's search intent. Because it only converts the user's voice content into text, the search results are not accurate enough and do not match the user's true search intent.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This application provides a voice search method and apparatus, storage medium and electronic device to at least solve the technical problem that voice search results cannot truly reflect the user's search intent, resulting in low accuracy of search results.
[0005] According to one aspect of the embodiments of this application, a voice search method is provided, comprising: in response to a first voice acquisition operation, displaying a first search term and a first recommended text set, wherein the first search term is a word identified in first voice acquired by the first voice acquisition operation, and the first recommended text set includes text determined based on the first search term; if a first recommended text selection operation is detected when displaying the first recommended text set, in response to the first recommended text selection operation, selecting a first recommended text in the first recommended text set; displaying a first search text obtained by merging the first search term and the first recommended text, and displaying a second recommended text set, wherein the second recommended text set includes text determined based on the first search text; if a search operation is detected when displaying the first search text, in response to the search operation, displaying a first search result, wherein the first search result includes results obtained by searching based on the first search text.
[0006] According to another aspect of the embodiments of this application, a voice search device is also provided, comprising: a first display module, configured to display a first search term and a first recommended text set in response to a first voice acquisition operation, wherein the first search term is a word identified in first voice acquired by the first voice acquisition operation, and the first recommended text set includes text determined based on the first search term; a selection module, configured to select a first recommended text in the first recommended text set in response to a first recommended text selection operation when a first recommended text selection operation is detected while the first recommended text set is being displayed; a second display module, configured to display a first search text obtained by merging the first search term and the first recommended text, and to display a second recommended text set, wherein the second recommended text set includes text determined based on the first search text; and a third display module, configured to display a first search result in response to a search operation initiated when a search operation is detected while the first search text is being displayed, wherein the first search result includes results obtained by searching based on the first search text.
[0007] Optionally, the device is further configured to: when a second recommended text selection operation is detected while displaying the second recommended text set, in response to the second recommended text selection operation, select the second recommended text in the second recommended text set; display the second search text obtained by merging the first search text and the second recommended text, and display a third recommended text set, wherein the third recommended text set includes text determined based on the second search text; when the search initiation operation is detected while displaying the second search text, in response to the search initiation operation, display the second search result, wherein the second search result includes results obtained by searching based on the second search text.
[0008] Optionally, the device is further configured to: when displaying the first search term and the first set of recommended texts, and in response to a second voice acquisition operation, display a third search text, wherein the third search text is text obtained by merging the first search term and the second search term, and the second search term is a word identified in the second voice acquired by the second voice acquisition operation; and when the search operation is detected while displaying the third search text, in response to the search operation, display a third search result, wherein the third search result includes results obtained by searching based on the third search text.
[0009] Optionally, the device is further configured to: when displaying the first search text and the second recommended text set, in response to a third voice acquisition operation, display a fourth search text and a fourth recommended text set, wherein the fourth search text is text obtained by merging the first search text and the third search term, the third search term is a word identified in the third voice acquired by the third voice acquisition operation, and the fourth recommended text set includes text determined based on the fourth search text; and when the search operation is detected while displaying the fourth search text, in response to the search operation, display a fourth search result, wherein the fourth search result includes results obtained by searching based on the fourth search text.
[0010] Optionally, the device is configured to select a first recommended text in the first recommended text set in response to a first recommended text selection operation when a first recommended text selection operation is detected while the first recommended text set is being displayed, by: acquiring a touch interaction operation performed on the first recommended text while the first recommended text set is being displayed, wherein the touch interaction operation is used to select the first recommended text from the first recommended text set; and moving the first recommended text to the position displayed by the first search term in response to the touch interaction operation, so as to select the first recommended text.
[0011] Optionally, the device is configured to display the first search text obtained by merging the first search term and the first recommended text, and to display the second set of recommended texts in the following manner: when the first recommended text moves to the position where the first search term is displayed, the first search term and the first recommended text are merged to obtain the first search text; the displayed first search term is replaced with the first search text, and the displayed first set of recommended texts is replaced with the second set of recommended texts.
[0012] The device is characterized in that, after displaying the first search text obtained by merging the first search term and the first recommended text, and displaying the second set of recommended texts, if a deletion operation is detected when displaying the first search text and an associated virtual button is executed, the first recommended text in the first search text is deleted, and the first search term and the first set of recommended texts are displayed, wherein the virtual button is used to trigger the deletion of the associated recommended text; or if an editing operation is detected when displaying the first search text and an editing interface for editing the first search text is detected, the first recommended text is deleted in the editing interface, and the first search term and the first set of recommended texts are displayed.
[0013] Optionally, the device is configured to display a first search text obtained by merging the first search term and the first recommended text, and to display a second set of recommended texts in the following manner: when the first recommended text is selected and a fourth voice acquisition operation is obtained, in response to the fourth voice acquisition operation, the first search text is displayed, and the second set of recommended texts is displayed, wherein the first search text includes the first search term, the first recommended text, and a fourth search term, the fourth search term being a word identified in the fourth voice acquired by the fourth voice acquisition operation, and the fourth search term being located after the first search term and the first recommended text in the first search text.
[0014] Optionally, the device is configured to respond to a first voice acquisition operation by displaying a first search term and a first set of recommended texts in the following manner: In response to the first voice acquisition operation, a first string and a first data structure are sent from a target client to a target server, wherein the first string includes the text content of the first search term, and the first data structure is used to record first recommended text parameters associated with the first search term; the target client receives the first string and a second data structure returned by the target server, wherein the second data structure includes the first set of recommended texts and a first insertion parameter for each recommended text in the first set of recommended texts, the first set of recommended texts including recommended texts determined based on a first group of words in the first search term, and the first insertion parameter indicating the position of each recommended text in the first set of recommended texts inserted into the first group of words; the target client displays the first search term and the first set of recommended texts based on the first string and the second data structure.
[0015] Optionally, the device is configured to, in response to the first recommended text selection operation, select the first recommended text in the first recommended text set when a first recommended text selection operation is detected while displaying the first recommended text set: in response to the first recommended text selection operation, modify the second data structure to a third data structure on the target client, and send the first string and the third data structure to the target server, wherein the third data structure includes a target field indicating that the first recommended text has been selected; the device is configured to display the first search text obtained by merging the first search term and the first recommended text, and display the second recommended text set: receive the second string returned by the target server and the fourth data structure updated by the target server according to the third data structure, wherein the second string is used to represent the first search text, and the fourth data structure includes the position of the first recommended text inserted into the first group of words determined by the target server according to the target field and the second recommended text set; display the first search text and the second recommended text set on the target client according to the second string and the fourth data structure.
[0016] Optionally, the apparatus is further configured to: when the third data structure also records a second recommended text parameter associated with the first search text, receive the second string and the fourth data structure returned by the target server on the target client, wherein the fourth data structure further includes a second recommended text set and a second insertion parameter for each recommended text in the second recommended text set, the second recommended text set including recommended text determined by the target server based on a second group of words in the first search text excluding the first group of words, and the second insertion parameter indicating the position of each recommended text in the second recommended text set inserted into the second group of words; and draw a target search interface on the target client based on the second string and the fourth data structure, wherein the target search interface displays the first search text and the second recommended text set.
[0017] Optionally, the device is configured to display the first search text obtained by merging the first search term and the first recommended text, and to display a second set of recommended texts, in the following manner: if the string stored on the target client is the first string, and the first string is the same as the first string returned by the target server, the first string returned by the target server is used as the string stored on the target client, and the data structure stored on the target client is updated from the first data structure to the second data structure; if the string stored on the target client is different from the first string returned by the target server, and the string stored on the target client is updated to a third string during the process of receiving the first string and the second data structure returned by the target server, the third string stored on the target client and the first string returned by the target server are merged, the merged fourth string is updated to the string stored on the target client, and the first data structure and the second data structure stored on the target client are merged, and the merged fifth data structure is updated to the data structure stored on the target client; a target search interface is drawn on the target client based on the fourth string and the fifth data structure, wherein the target search interface displays the first search text and the second set of recommended texts.
[0018] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, and the computer program is configured to execute the above-described voice search method when it is run.
[0019] According to another aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the voice search method described above.
[0020] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-described voice search method through the computer program.
[0021] In this embodiment, in response to a first voice acquisition operation, a first search term and a first recommended text set are displayed. The first search term is a word identified from the first voice acquired during the first voice acquisition operation. The first recommended text set includes text determined based on the first search term. When a first recommended text selection operation is detected while displaying the first recommended text set, in response to the first recommended text selection operation, the first recommended text is selected in the first recommended text set, and the first search text obtained by merging the first search term and the first recommended text is displayed. A second recommended text set is also displayed, where the second recommended text set includes text determined based on the first search text. When a search operation is detected while displaying the first search text, in response to the search operation, the first search result is displayed. The first search result includes results obtained based on the first search text. By intelligently displaying recommended text based on real-time speech recognition when the user performs a voice search, the user can select the recommended text to form the search text for further searching. This combines voice input capabilities with natural language learning, solving the problem of unscientific omissions that easily occur when users input voice data. By simplifying the user's input content according to the user's voice habits, the content of the user's voice input is optimized, making it more expressive of the speaker's intent. This improves the accuracy of search results, making the search results more realistically reflect the user's search intent. This solves the technical problem that voice search results cannot accurately reflect the user's search intent, resulting in low accuracy.
[0022] In addition, users can continue to use voice commands to conduct further searches, or supplement the generated search text with voice, which improves the interactive efficiency of voice search and further optimizes the accuracy of voice search in reflecting the user's true search intent.
[0023] Furthermore, by combining touch interaction with recommended text with real-time voice search, natural language sentences (i.e., the first search text) can be generated for searching, better helping users express their search needs in voice scenarios. Attached Figure Description
[0024] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0025] Figure 1 This is a schematic diagram of an application environment for an optional voice search method according to an embodiment of this application;
[0026] Figure 2This is a flowchart illustrating an optional voice search method according to an embodiment of this application;
[0027] Figure 3 This is a schematic diagram of an optional voice search method according to an embodiment of this application;
[0028] Figure 4 This is a schematic diagram of another optional voice search method according to an embodiment of this application;
[0029] Figure 5 This is a schematic diagram of another optional voice search method according to an embodiment of this application;
[0030] Figure 6 This is a schematic diagram of another optional voice search method according to an embodiment of this application;
[0031] Figure 7 This is a schematic diagram of another optional voice search method according to an embodiment of this application;
[0032] Figure 8 This is a schematic diagram of another optional voice search method according to an embodiment of this application;
[0033] Figure 9 This is a schematic diagram of another optional voice search method according to an embodiment of this application;
[0034] Figure 10 This is a schematic diagram of another optional voice search method according to an embodiment of this application;
[0035] Figure 11 This is a schematic diagram of another optional voice search method according to an embodiment of this application;
[0036] Figure 12 This is a schematic diagram of another optional voice search method according to an embodiment of this application;
[0037] Figure 13 This is a schematic diagram of another optional voice search method according to an embodiment of this application;
[0038] Figure 14 This is a schematic diagram of another optional voice search method according to an embodiment of this application;
[0039] Figure 15 This is a schematic diagram of another optional voice search method according to an embodiment of this application;
[0040] Figure 16 This is a schematic diagram of another optional voice search method according to an embodiment of this application;
[0041] Figure 17This is a schematic diagram of the structure of an optional voice search device according to an embodiment of this application;
[0042] Figure 18 This is a schematic diagram of the structure of an optional voice search product according to an embodiment of this application;
[0043] Figure 19 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of this application. Detailed Implementation
[0044] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0045] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0046] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:
[0047] The text field: The text content entered by the user via voice.
[0048] The insertContent field has a default text data structure, which internally consists of arrays, each containing the following:
[0049] index: The number of characters inserted after the first character in the text content indicated by the text field above;
[0050] insertCount: The number of times the recommended text has been inserted.
[0051] allowInsert: Whether to allow recommendations. The default is true, which allows recommendations. If changed to false, no recommended text will be displayed in the corresponding position.
[0052] wordList: The actual recommended text, with the selected field indicating that the user has already selected the currently recommended text.
[0053] The present application will be described below with reference to embodiments:
[0054] According to one aspect of the embodiments of this application, a voice search method is provided. Optionally, in this embodiment, the above-described voice search method can be applied to, for example... Figure 1 The hardware environment shown consists of server 101 and terminal device 103. For example... Figure 1 As shown, server 101 is connected to terminal 103 via a network and can be used to provide services for terminal devices or applications installed on terminal devices. Applications can be video applications, instant messaging applications, browser applications, educational applications, game applications, etc. Database 105 can be set up on the server or independently of the server to provide data storage services for server 101, such as a search data storage server. The network can include, but is not limited to, wired networks and wireless networks. The wired network includes local area networks (LANs), metropolitan area networks (MANs), and wide area networks (WANs). The wireless network includes Bluetooth, Wi-Fi, and other networks that enable wireless communication. Terminal device 103 can be a terminal configured with applications and can include, but is not limited to, at least one of the following: mobile phones (such as Android phones, iOS phones, etc.), laptops, tablets, PDAs, MIDs (Mobile Internet Devices), PADs, desktop computers, smart TVs, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, and other computer devices. The server can be a single server, a server cluster consisting of multiple servers, or a cloud server. The application 107 using the above-mentioned voice search method is displayed through terminal device 103 or other connected display devices.
[0055] Combination Figure 1 As shown, the above voice search method can be implemented in terminal device 103 through the following steps:
[0056] S1, in response to the first voice acquisition operation, the terminal device 103 displays the first search term and the first recommended text set, wherein the first search term is a word identified in the first voice acquired by the first voice acquisition operation, and the first recommended text set includes text determined based on the first search term;
[0057] S2, if a first recommended text selection operation is detected when the first recommended text set is displayed, the first recommended text is selected in the first recommended text set on the terminal device 103 in response to the first recommended text selection operation;
[0058] S3, the terminal device 103 displays the first search text obtained by merging the first search term and the first recommended text, and displays the second recommended text set, wherein the second recommended text set includes text determined based on the first search text;
[0059] S4, if a search operation is detected when the first search text is displayed, the terminal device 103 displays the first search result in response to the search operation, wherein the first search result includes the result obtained by searching based on the first search text.
[0060] Optionally, in this embodiment, the above-described voice search method can also be implemented via a server, for example, Figure 1 It is implemented in server 101 shown; or it is implemented jointly by the terminal device and the server.
[0061] The above is merely an example, and this embodiment does not impose any specific limitations.
[0062] Alternatively, as an alternative implementation method, such as Figure 2 As shown, the above voice search method includes:
[0063] S202, in response to the first voice acquisition operation, a first search term and a first recommended text set are displayed, wherein the first search term is a word identified in the first voice acquired by the first voice acquisition operation, and the first recommended text set includes text determined based on the first search term;
[0064] Optionally, in the embodiments of this application, the first voice acquisition operation may include, but is not limited to, a voice acquisition operation triggered by the user through voice interaction, a voice acquisition operation triggered after performing an interactive operation on a virtual button used to start the voice acquisition operation, or a voice acquisition operation triggered by a voice trigger command.
[0065] For example, Figure 3 This is a schematic diagram of an optional voice search method according to an embodiment of this application, such as... Figure 3 As shown, virtual button 302 is a pre-configured virtual button in the search interface. Virtual button 302 is used to display the voice search interface after a trigger operation is obtained. Virtual button 304 is pre-configured on the voice search interface. Virtual button 304 is used to start collecting the first voice after a trigger operation is obtained.
[0066] The above is merely an example, and this application does not impose any specific limitations.
[0067] It should be noted that the first voice acquisition operation described above is a real-time voice acquisition operation. Therefore, the first search term generated by the real-time recognition and conversion of the first voice will be displayed on the voice search interface.
[0068] Optionally, in this embodiment of the application, the first speech is speech collected in real time by a speech acquisition device, which may include, but is not limited to, a microphone, a recording device, etc. The first search term is a search term obtained in real time based on the first speech, which may include, but is not limited to, artificial intelligence-based speech recognition technologies such as automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition technology.
[0069] Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or computers-controlled machines to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.
[0070] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0071] Key technologies in speech technology include Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Voiceprint Recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech emerging as one of the most promising methods.
[0072] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.
[0073] In this embodiment, the first search term can be obtained by recognizing the first speech using speech technology, and then the first recommended text set can be obtained based on the first search term using natural language processing technology.
[0074] It should be noted that the aforementioned first recommended text set is obtained by segmenting the first search term, acquiring corresponding recommended texts for each segment, and then filtering each recommended text based on semantics. The number of recommended texts in the aforementioned first recommended text set is pre-configured by the system and can be determined based on the number of segments of the first search term. For example, when the first search term is allowed to be segmented into 1 segment, the number of recommended texts in the aforementioned first recommended text set is 3; when the first search term is allowed to be segmented into 2 segments, the number of recommended texts in the aforementioned first recommended text set is 4; and the more segments, the more recommended texts in the first recommended text set.
[0075] For example, Figure 4 This is a schematic diagram of another optional voice search method according to an embodiment of this application, such as... Figure 4 As shown, by performing a trigger operation on the virtual button 402, the first voice is collected. During the collection of the first voice, the first voice is recognized in real time, and the first search term "Xi'an" generated based on the first voice is displayed. When the first search term "Xi'an" is displayed, the first set of recommended texts "snacks", "weather", "attractions" and "universities" obtained based on the first search term "Xi'an" are also displayed.
[0076] The above is merely an example, and this application does not impose any specific limitations.
[0077] S204, if a first recommended text selection operation is detected when the first recommended text set is displayed, in response to the first recommended text selection operation, the first recommended text is selected in the first recommended text set;
[0078] Optionally, in the embodiments of this application, the first recommended text selection operation may include, but is not limited to, a combination of one or more selection operation methods such as clicking, long pressing, dragging, double-clicking, gestures, and voice.
[0079] For example, Figure 5 This is a schematic diagram of another optional voice search method according to an embodiment of this application, such as... Figure 5 As shown, by dragging the virtual button "snacks" to select and move the recommended text "snacks", when the virtual button "snacks" is dragged to the area where the first search term "Xi'an" is located, the first recommended text "snacks" is selected in the first recommended text set.
[0080] The above is merely an example, and this application does not impose any specific limitations.
[0081] S206, display the first search text obtained by merging the first search term and the first recommended text, and display the second recommended text set, wherein the second recommended text set includes text determined based on the first search text;
[0082] Optionally, in this embodiment of the application, the first search text is the text obtained by merging the first search term and the first recommended text. During the merging process, the first recommended text is inserted into the first search term according to the word segmentation position of the first recommended text after word segmentation of the first search term. Alternatively, a natural language processing model can be used to generate a natural sentence as the first search text based on the first search term and the first recommended text.
[0083] Optionally, in the embodiments of this application, the second recommended text set is a set of recommended texts obtained by re-segmenting the first search text and recommending it based on the segmented texts that have not generated recommended texts. That is, each time the recommended text set is displayed, it is a set of recommended texts obtained based on the text content added in the previous time.
[0084] For example, Figure 6 This is a schematic diagram of another optional voice search method according to an embodiment of this application, such as... Figure 6 As shown, when the virtual button corresponding to the recommended text "snacks" in the first recommended text set is dragged to the area where the first search term "Xi'an" is located, the first search text "Xi'an snacks" is displayed on the search interface, and the first search text is segmented into two words: "Xi'an" and "snacks". Since the word "Xi'an" has already generated recommended text, the second recommended text set mentioned above is displayed based on the word "snacks", including: "where to eat", "types", and "historical origin".
[0085] The above is merely an example, and this application does not impose any specific limitations.
[0086] S208, if a search operation is detected when the first search text is displayed, in response to the search operation, the first search result is displayed, wherein the first search result includes the result obtained by searching based on the first search text.
[0087] Optionally, in this embodiment of the application, the above-mentioned search initiation operation may include, but is not limited to, a combination of one or more selection operation methods such as clicking, long pressing, dragging, double-clicking, gestures, and voice. A trigger operation can be performed on the virtual button used to initiate the search operation to start a search based on the first search text. When the search operation is detected, the first search result is displayed.
[0088] Optionally, in this embodiment of the application, the first search result is the result obtained by searching based on the first search text. The first search result may include, but is not limited to, search results of the same or different search dimensions such as web pages, images, social media, videos, maps, music, news, and academic materials.
[0089] For example, Figure 7 This is a schematic diagram of another optional voice search method according to an embodiment of this application, such as... Figure 7 As shown, by clicking the virtual "Search" button to initiate a search, the system searches for the first search text "Xi'an snacks". Finally, the display interface switches from the search interface to the search results interface, where the first search result obtained based on the first search text is displayed.
[0090] The above is merely an example, and this application does not impose any specific limitations.
[0091] In this embodiment, in response to a first voice acquisition operation, a first search term and a first recommended text set are displayed. The first search term is a word identified from the first voice acquired during the first voice acquisition operation. The first recommended text set includes text determined based on the first search term. When a first recommended text selection operation is detected while displaying the first recommended text set, in response to the first recommended text selection operation, the first recommended text is selected in the first recommended text set, and the first search text, obtained by merging the first search term and the first recommended text, is displayed. A second recommended text set is also displayed, containing text determined based on the first search text. When a search operation is detected while displaying the first search text, in response to the search operation, a first search result is displayed. The first search result includes results obtained based on the first search text. By intelligently displaying recommended text based on real-time speech recognition when the user performs a voice search, the user can select the recommended text to form the search text for further searching. This combines voice input capabilities with natural language learning, solving the problem of unscientific omissions that easily occur when users input voice data. By simplifying the user's input content according to the user's voice habits, the content of the user's voice input is optimized, making it more expressive of the speaker's intent. This improves the accuracy of search results, making the search results more realistically reflect the user's search intent. This solves the technical problem that voice search results cannot accurately reflect the user's search intent, resulting in low accuracy.
[0092] As an optional approach, the above method also includes:
[0093] If a second recommended text selection operation is detected when the second recommended text set is displayed, the second recommended text is selected in response to the second recommended text selection operation.
[0094] The system displays the second search text, which is obtained by merging the first search text and the second recommended text, and displays the third recommended text set, which includes the text determined based on the second search text.
[0095] If a search operation is detected while the second search text is being displayed, the second search result is displayed in response to the search operation, wherein the second search result includes the results obtained from the search based on the second search text.
[0096] Optionally, in the embodiments of this application, the second recommended text selection operation may be the same as or different from the first recommended text selection operation, and may include, but is not limited to, a combination of one or more selection operation methods such as clicking, long pressing, dragging, double-clicking, gestures, and voice.
[0097] Optionally, in this embodiment of the application, the text obtained by merging the first search text and the second recommended text can be a natural sentence obtained based on a natural language processing model. The second search text can also be inserted into the first search text during the merging process according to the word segmentation position of the second recommended text after word segmentation of the first search text.
[0098] Optionally, in the embodiments of this application, the aforementioned third recommended text set is a recommended text set obtained by re-segmenting the second search text and recommending it based on the segmentation of words that have not generated recommended text. That is, each time the recommended text set is displayed, it is a recommended text set obtained based on the previously added text content.
[0099] For example, Figure 8 This is a schematic diagram of another optional voice search method according to an embodiment of this application, such as... Figure 8 As shown, by dragging the virtual button "Where to Eat", the recommended text "Where to Eat" is selected and moved. When the virtual button "Where to Eat" is dragged to the area where the first search text "Xi'an Snacks" is located, the second recommended text "Where to Eat" is selected in the second recommended text set. When the virtual button corresponding to the recommended text "Where to Eat" in the second recommended text set is dragged to the area where the first search text "Xi'an Snacks" is located, the second search text "Where to Eat Xi'an Snacks" is displayed on the search interface. The second search text is then segmented into three words: "Xi'an", "snacks", and "Where to Eat". Since the words "Xi'an" and "snacks" have already generated recommended text, the third recommended text set, including "How to Eat Most Conveniently", "Map", and "Established Brands", is displayed based on the word "Where to Eat". By initiating a search operation by clicking the "Search" virtual button, a search is performed based on the second search text "Where to Eat Xi'an Snacks". Finally, the display interface switches from the search interface to the search results interface, and the second search result obtained based on the second search text is displayed on the search results interface.
[0100] The above is merely an example, and this application does not impose any specific limitations.
[0101] As an alternative approach, the method also includes:
[0102] When displaying the first search term and the first set of recommended texts, in response to the second voice acquisition operation, the third search text is displayed, wherein the third search text is the text obtained by merging the first search term and the second search term, and the second search term is the word identified in the second voice acquired by the second voice acquisition operation;
[0103] If a search operation is detected while the third search text is being displayed, the third search result is displayed in response to the search operation, wherein the third search result includes the results obtained from the search based on the third search text.
[0104] Optionally, in the embodiments of this application, the above-mentioned second voice acquisition operation may include, but is not limited to, a voice acquisition operation triggered by the user through voice interaction, a voice acquisition operation triggered after performing an interactive operation on a virtual button used to start the voice acquisition operation, or a voice acquisition operation triggered by a voice trigger command.
[0105] It should be noted that the second voice acquisition operation is a real-time voice acquisition operation. Therefore, the third search text will be displayed on the voice search interface. The third search text is the search text obtained by merging the second search term generated by the real-time voice recognition and conversion with the first search term.
[0106] For example, Figure 9 This is a schematic diagram of another optional voice search method according to an embodiment of this application, such as... Figure 9 As shown, by performing a first voice acquisition operation on the virtual button, the first voice is acquired and recognized to obtain the first search term. Then, by performing a second voice acquisition operation on the virtual button, the second voice is acquired and recognized to obtain the second search term. The first and second search terms are merged to display the merged third search text. A new set of recommended texts, including "history," "modern," "literature," and "novel," is also displayed based on the second search term. By performing a search operation on the "search" virtual button, a search is initiated based on the third search text "Xi'an people." Finally, the display interface switches from the search interface to the search results interface, displaying the third search result obtained based on the third search text.
[0107] The above is merely an example, and this application does not impose any specific limitations.
[0108] Through this embodiment, users can continue to send voice commands for further searches, improving the interactive efficiency of voice search and further optimizing the accuracy of voice search in reflecting the user's true search intent.
[0109] As an alternative approach, the method also includes:
[0110] When displaying the first search text and the second recommended text set, in response to the third voice acquisition operation, the fourth search text and the fourth recommended text set are displayed, wherein the fourth search text is the text obtained by merging the first search text and the third search term, the third search term is the word identified in the third voice acquired by the third voice acquisition operation, and the fourth recommended text set includes the text determined based on the fourth search text.
[0111] If a search operation is detected while the fourth search text is being displayed, the fourth search result is displayed in response to the search operation, wherein the fourth search result includes the results obtained from the search based on the fourth search text.
[0112] Optionally, in the embodiments of this application, the aforementioned third voice acquisition operation may include, but is not limited to, a voice acquisition operation triggered by the user through voice interaction, a voice acquisition operation triggered after performing an interactive operation on a virtual button used to start the voice acquisition operation, or a voice acquisition operation triggered by a voice trigger command.
[0113] It should be noted that the third voice acquisition operation mentioned above is a real-time voice acquisition operation. Therefore, the fourth search text mentioned above will be displayed on the voice search interface. The fourth search text is the search text obtained by merging the third search term generated by the real-time voice recognition and conversion with the first search text.
[0114] For example, Figure 10 This is a schematic diagram of another optional voice search method according to an embodiment of this application, such as... Figure 10 As shown, by selecting the first recommended text on the virtual button "Snacks," the first search text and a second set of recommended texts are displayed. Then, by selecting the third voice recording on the virtual button, the third voice recording is captured and recognized to obtain the third search term. The first search text and the third search term are merged to display the merged fourth search text. A new set of recommended texts, including "Map," "Transportation," and "Region," is displayed based on the third search term. By selecting the "Search" virtual button to initiate a search based on the fourth search text "How to choose snacks in Xi'an?", the display switches from a search interface to a search results interface, displaying the fourth search result obtained based on the fourth search text.
[0115] The above is merely an example, and this application does not impose any specific limitations.
[0116] Through this embodiment, users can supplement the generated search text with their voice, which improves the interactive efficiency of voice search and further optimizes the accuracy of voice search in reflecting the user's true search intent.
[0117] As an optional approach, if a selection operation of the first recommended text is detected when displaying the first recommended text set, in response to the selection operation, the first recommended text is selected from the first recommended text set, including:
[0118] When displaying the first set of recommended texts, obtain the touch interaction operation performed on the first recommended texts, wherein the touch interaction operation is used to select the first recommended texts from the first set of recommended texts;
[0119] In response to a touch interaction, the first recommended text is moved to the position where the first search term is displayed to select the first recommended text.
[0120] Optionally, in this embodiment of the application, the above-mentioned touch interaction operation may include, but is not limited to, a combination of one or more selection operation methods such as clicking, long pressing, dragging, double-clicking, gestures, and voice. The above-mentioned moving the first recommended text to the position displayed by the first search term may include, but is not limited to, configuring corresponding UI (User Interface) objects for the first recommended text and the first search term respectively, performing touch interaction operation on the UI object corresponding to the first recommended text to select the first recommended text, and determining that the first recommended text has been selected when the UI object corresponding to the first recommended text overlaps with the UI object corresponding to the first search term.
[0121] by Figure 5 For example, the UI object "snacks" corresponds to the aforementioned first recommended text, and the UI object "Xi'an" corresponds to the aforementioned first search term. When the UI object "snacks" is moved to the UI object "Xi'an" through touch interaction, causing the UI object "snacks" and the UI object "Xi'an" to overlap, it is determined that the recommended text corresponding to the selected UI object "snacks" is the aforementioned first recommended text.
[0122] As an optional approach, display the first search text obtained by merging the first search term and the first recommended text, and display the second set of recommended texts, including:
[0123] When the first recommended text moves to the position where the first search term is displayed, the first search term and the first recommended text are merged to obtain the first search text;
[0124] The first search term will be replaced with the first search text, and the first set of recommended texts will be replaced with the second set of recommended texts.
[0125] Optionally, in this embodiment of the application, when the first recommended text moves to the position where the first search term is displayed, the first recommended text and the first search term can be merged to obtain the first search text. At this time, a replacement operation can be used to replace the displayed first search term with the first search text and replace the displayed first recommended text set with the second recommended text set, so as to display the merged first search text in real time and display the second recommended text set corresponding to the first search text.
[0126] by Figure 6 For example, the UI object "snacks" corresponds to the aforementioned first recommended text, and the UI object "Xi'an" corresponds to the aforementioned first search term. When the UI object "snacks" is moved to the position displayed by the UI object "Xi'an" through touch interaction, the UI objects "snacks" and "Xi'an" are merged to obtain the UI object "Xi'an snacks", which is used as the aforementioned first search text. In the subsequent display of the search interface, the first recommended text set containing the UI object "snacks" is replaced with the second recommended text set, and the UI object "Xi'an" is replaced with the UI object "Xi'an snacks" as the aforementioned first search text.
[0127] As an optional solution, after displaying the first search text obtained by merging the first search term and the first recommended text, and displaying the second set of recommended texts, the method further includes: if a deletion operation is detected when a virtual button associated with the first recommended text is executed while displaying the first search text, deleting the first recommended text from the first search text, and displaying the first search term and the first set of recommended texts, wherein the virtual button is used to trigger the deletion of the associated recommended text; or if an editing operation is detected when an editing operation is executed on the first search text while displaying the first search text, displaying an editing interface for editing the first search text; deleting the first recommended text in the editing interface, and displaying the first search term and the first set of recommended texts.
[0128] Optionally, in this embodiment of the application, if a deletion operation is detected when the first search text is displayed, the first recommended text in the first search text can be deleted, and the first search term and the first recommended text set can be displayed.
[0129] The aforementioned detection of a deletion operation when displaying the first search text may include, but is not limited to, detecting a deletion operation performed on a virtual button associated with the first recommended text. The virtual button is associated with the added first recommended text. By performing a touch interaction operation on the virtual button, the virtual button is triggered to delete the first recommended text in the first search text, so as to redisplay the first search term and the first recommended text set.
[0130] It should be noted that the first set of recommended texts that is redisplayed may include the first set of recommended texts that have been deleted, or it may be the set of recommended texts in the first set of recommended texts with the deleted first set of recommended texts removed.
[0131] The aforementioned detection of a deletion operation when displaying the first search text may also include, but is not limited to, detecting an editing operation on the first search text. The editing operation is used to trigger an editing interface for displaying the first search text. By deleting the first recommended text in the first search text in the editing interface, the first search term and the first recommended text set are redisplayed.
[0132] As an optional approach, display the first search text obtained by merging the first search term and the first recommended text, and display the second set of recommended texts, including:
[0133] When the first recommended text has been selected and the fourth voice acquisition operation has been obtained, in response to the fourth voice acquisition operation, the first search text is displayed and the second set of recommended texts is displayed. The first search text includes the first search term, the first recommended text, and the fourth search term. The fourth search term is a word identified in the fourth voice acquired by the fourth voice acquisition operation. The fourth search term is located after the first search term and the first recommended text in the first search text.
[0134] Optionally, in the embodiments of this application, the above-mentioned fourth voice acquisition operation may include, but is not limited to, a voice acquisition operation triggered by the user through voice interaction, a voice acquisition operation triggered after performing an interactive operation on a virtual button used to start the voice acquisition operation, or a voice acquisition operation triggered by a voice trigger command.
[0135] It should be noted that the above-mentioned fourth voice acquisition operation is a real-time voice acquisition operation. Therefore, the above-mentioned first search text will be displayed on the above-mentioned voice search interface. The first search text is the search text obtained by merging the fourth search term generated by the real-time recognition and conversion of the fourth voice with the first search term and the first recommended text.
[0136] For example, Figure 11 This is a schematic diagram of another optional voice search method according to an embodiment of this application, such as... Figure 11As shown, the search interface is displayed, including the first search term and the first recommended text set. When the first recommended text is selected and the fourth voice acquisition operation is obtained, the first search term, the first recommended text and the fourth search term are merged and the merged first search text is displayed. A new recommended text set is also displayed based on the first recommended text and the third search term, including: "map", "transportation", and "area".
[0137] The above is merely an example, and this application does not impose any specific limitations.
[0138] This embodiment combines touch interaction with recommended text with real-time voice search to generate natural language sentences for searching (i.e., the first search text), which better helps users express their search needs in voice scenarios.
[0139] As an optional approach, in response to the first voice acquisition operation, the first search term and the first set of recommended texts are displayed, including:
[0140] In response to the first voice acquisition operation, the first string and the first data structure are sent from the target client to the target server, wherein the first string includes the text content of the first search term, and the first data structure is used to record the first recommended text parameters associated with the first search term;
[0141] The target client receives a first string and a second data structure returned by the target server. The second data structure includes a first set of recommended texts and a first insertion parameter for each recommended text in the first set of recommended texts. The first set of recommended texts includes recommended texts determined based on a first group of words in the first search term. The first insertion parameter indicates the position of each recommended text in the first set of recommended texts inserted into the first group of words.
[0142] On the target client, display the first search term and the first set of recommended texts based on the first string and the second data structure.
[0143] Optionally, in this embodiment, the first string may include, but is not limited to, a Text field for recording the text content of the first search term, and the first data structure may include, but is not limited to, an insertContent field for recording the first recommended text parameters associated with the first search term. When the first search term is the first search term identified, the insertContent field is set to an empty array so that the data structure can be modified subsequently to achieve the synchronization and uniqueness of the search text related data.
[0144] It should be noted that the above-mentioned first recommended text parameters may include, but are not limited to, the set of positions where recommended text is allowed to be inserted after the first search term is segmented, whether the first search term allows the insertion of recommended text, the maximum number of times recommended text is allowed to be inserted, the display color and size of the recommended text that is allowed to be inserted, and sensitive words and interference words used to filter recommended text.
[0145] Optionally, in this embodiment of the application, the second data structure includes a first set of recommended texts determined by the server based on the first search term, and the position of each recommended text in the first set of recommended texts when inserted into the first search term, which is the first insertion parameter.
[0146] Optionally, in this embodiment of the application, the above-mentioned display of the first search term and the first recommended text set on the target client based on the first string and the second data structure can be understood as the client merging the first string and the second data structure with the locally stored string and data structure, and finally drawing the UI object corresponding to the search interface, so as to facilitate the subsequent triggering of the recommended text selection operation or the initiation of the search operation.
[0147] Optionally, in this embodiment of the application, the first group of words is a group of words obtained after performing word segmentation on the first search term, and the first insertion parameter indicates the position of each recommended text when it is inserted into the first group of words.
[0148] As an alternative solution,
[0149] When a first recommended text selection operation is detected while displaying the first recommended text set, in response to the first recommended text selection operation, the first recommended text is selected in the first recommended text set, including: in response to the first recommended text selection operation, the second data structure is modified to a third data structure on the target client, and the first string and the third data structure are sent to the target server, wherein the third data structure includes a target field indicating that the first recommended text has been selected;
[0150] The system displays the first search text obtained by merging the first search term and the first recommended text, and displays the second recommended text set, including: receiving a second string returned by the target server and a fourth data structure updated by the target server according to a third data structure, wherein the second string is used to represent the first search text, and the fourth data structure includes the position of the first recommended text inserted into the first group of words determined by the target server according to the target field and the second recommended text set; and displays the first search text and the second recommended text set on the target client according to the second string and the fourth data structure.
[0151] Optionally, in this embodiment, when the first recommended text selection operation is obtained, the second data structure can be modified to a third data structure on the client. The modification involves adding the aforementioned target field to the second data structure to indicate that the recommended text including the target field is the recommended text selected by the user. Then, by sending the third data structure including the target field to the target server, the target server can continue to generate a new set of related recommended texts based on the selected recommended text and update the third data structure to a fourth data structure. The fourth data structure includes, relative to the third data structure, the new set of recommended texts generated based on the recommended text selected by the user. At the same time, the server will also modify the first string to the second string according to the recommended text selected by the user indicated in the third data structure, so that the client can modify the displayed search text content.
[0152] As an optional approach, the above method also includes:
[0153] In the case that the third data structure also records the second recommended text parameters associated with the first search text, the target client receives the second string and the fourth data structure returned by the target server. The fourth data structure further includes a second recommended text set and a second insertion parameter for each recommended text in the second recommended text set. The second recommended text set includes recommended texts determined by the target server based on the second group of words in the first search text, excluding the first group of words. The second insertion parameter indicates the position of each recommended text in the second recommended text set inserted into the second group of words.
[0154] On the target client, a target search interface is drawn based on the second string and the fourth data structure, where the target search interface displays a first search text and a second set of recommended texts.
[0155] Optionally, in the embodiments of this application, the second recommended text parameters may include, but are not limited to, the set of positions where recommended text is allowed to be inserted after the first search text is segmented, whether the first search text allows the insertion of recommended text, the historical number / remaining number / maximum number of times recommended text can be inserted, the display color and size of the recommended text that can be inserted, etc., used to filter sensitive words, interference words, etc. in the recommended text.
[0156] Optionally, in this embodiment of the application, the second group of words is a group of words obtained after performing word segmentation on the first search text, and the second insertion parameter indicates the position of each recommended text when it is inserted into the second group of words.
[0157] As an optional approach, display the first search text obtained by merging the first search term and the first recommended text, and display the second set of recommended texts, including:
[0158] If the string stored on the target client is the first string, and the first string is the same as the first string returned by the target server, then the first string returned by the target server is used as the string stored on the target client, and the data structure stored on the target client is updated from the first data structure to the second data structure.
[0159] If the string stored on the target client is different from the first string returned by the target server, and if the string stored on the target client is updated to the third string during the process of receiving the first string and the second data structure returned by the target server, the third string stored on the target client and the first string returned by the target server are merged, the merged fourth string is updated to the string stored on the target client, the first data structure and the second data structure stored on the target client are merged, and the merged fifth data structure is updated to the data structure stored on the target client.
[0160] On the target client, a target search interface is drawn based on the fourth string and the fifth data structure, where the target search interface displays a first search text and a second set of recommended texts.
[0161] Optionally, in this embodiment, the string stored on the target client is the first string, and the first string being the same as the first string returned by the target server can be understood as follows: during the process of obtaining the first search text from the target server, the user did not input any new voice or the user did not select any new recommended text. Therefore, the first string stored on the client is the same as the first string returned by the target server. The string stored on the target client being different from the first string returned by the target server can be understood as follows: during the process of obtaining the first search text from the target server, the user inputs new voice or the user selects new recommended text. Therefore, the string stored on the client is the updated third string, which is different from the first string returned by the target server. Therefore, it is necessary to merge the first string and the third string into a fourth string and display them accordingly.
[0162] Optionally, in this embodiment of the application, updating the data structure stored on the target client from the first data structure to the second data structure can be understood as merging data structures. That is, using the second data structure returned by the server as the data structure stored locally on the client. Merging the first and second data structures stored on the target client and updating the merged fifth data structure to the data structure stored on the target client can be understood as merging the data structure stored on the client with the data structure returned by the server.
[0163] The following specific examples will further explain this application:
[0164] one, Figure 12 This is a schematic diagram of another optional voice search method according to an embodiment of this application, such as... Figure 12 As shown, it includes the following steps:
[0165] 1. When the user enables voice input, the text is set to an empty string; and the native solution is used to perform voice-to-text conversion on the client locally before saving the text to the cache.
[0166] 2. Submit the converted text and empty data structure to the server. Figure 13 This is a schematic diagram of another optional voice search method according to an embodiment of this application, such as... Figure 13 As shown;
[0167] 3. The server processes data based on whether the data structure is empty:
[0168] 3.1 If the data structure is not empty, modify the text content according to the content recorded in the data structure, calculate the new data structure and save it;
[0169] 3.2 If the data structure is empty, skip this step and proceed directly to step 4;
[0170] 4. Based on the instructions of the text and data structure, perform text content splitting and intelligent content addition;
[0171] 5. Return the processed text and data structure to the client;
[0172] 6. The client renders the UI (User Interface) and refreshes the page based on the data structure returned by the server;
[0173] 7. The user performs an action, selecting words recommended by AI (Artificial Intelligence) and interacting with the system;
[0174] 8. Modify the local text and data structure, save it, upload it to the server, and then execute step 3 and step 9 simultaneously.
[0175] 9. Continue voice input and continue local STT (sound to text) conversion.
[0176] 10. Modify the local text, continue to merge the data structure and save it. After uploading it to the server, execute step 3 and simultaneously perform step 9.
[0177] It's important to note that after continuous voice input, the system performs voice recognition. After processing the recognized text and prepared data locally, it uploads it to the server. However, during voice recognition, the user is also allowed to select AI-recommended words and phrases simultaneously. This step modifies the local text and prepared data. Subsequently, after the AI-recommended words and phrases request is returned, the local text and data structure are modified again. Since the same data has three sources for modification, the synchronization and uniqueness of this data need to be ensured through the following methods:
[0178] Since the entire data update has only two sources: local STT data and server-returned data, two local member variables are used to handle the uniqueness issue. Variable 1: local string, used to store the local audio converted to text data and the text data returned by the server using the text field; Variable 2: data structure, insertContent field, the default field is an empty array.
[0179] It should also be noted that the aforementioned voice search service differs from general AI intelligent recommendation, which simply adds intelligent recommendation text at the end of the text. Instead, the entire text uploaded for the first time is first segmented into sentences, and AI-recommended phrases are added after each reasonable complete word. For subsequent uploaded text, only the parts that have not undergone AI intelligent recommendation are segmented into sentences and AI-recommended phrases are added. Furthermore, based on the user's possible actions, the overall data structure and text are processed. Therefore, by converting the data structure back to a string for data caching, the length of data for intelligent word segmentation and recall in the above processing flow can be reduced, thereby reducing time consumption. At the same time, using parallel requests for AI automatic word recommendation can further reduce time consumption.
[0180] The main interaction flow between the client and server includes the following: Figure 14 This is a schematic diagram of another optional voice search method according to an embodiment of this application. The data structure can be as follows: Figure 14 As shown:
[0181] S1, the client determines whether it is a new pronunciation. If it is a new pronunciation, the text field is set to an empty string and the insertcontent field is set to an empty array.
[0182] S2, the client uses stt to convert the user's pronunciation into text and saves it in the text field, and submits both the text field and the insertcontent field to the server, such as... Figure 13 As shown;
[0183] S3. After receiving the data, the server first checks whether there is a wordlist field in insertContent. If there is, it continues to check the index field data X in wordlist, performs word segmentation and AI-generated word association requests from the Xth character of the text string, and adds new recommended text data to wordlist.
[0184] S4, the server checks if the wordlist contains the select field; if not, it returns directly.
[0185] S5: After receiving the data, the client updates the UI and gives the user a choice. If the user makes a choice, the client sends the data to the server.
[0186] S6, if the server detects the select field, it needs to insert the word selected by the user at the position containing the select field, adjust the index field of this item and all subsequent items, add the selected text at the specified position in the text field, and return the modified text and insertContent fields to the client.
[0187] II. Overall design at both ends:
[0188] Figure 15 This is a schematic diagram of another optional voice search method according to an embodiment of this application, such as... Figure 15 As shown, the client includes, but is not limited to, the following processes:
[0189] S1, determine whether to initiate a new voice input;
[0190] S2 converts voice input into text (stt);
[0191] S3, client data upload;
[0192] S4, client-side UI rendering;
[0193] S5, client interaction response and implementation request;
[0194] S6 temporarily saves and merges client data.
[0195] Figure 15 This is a schematic diagram of another optional voice search method according to an embodiment of this application, such as... Figure 15 As shown, the server-side process includes, but is not limited to, the following:
[0196] S1, performs word segmentation, recall, and blacklisting on the text field uploaded by the client;
[0197] S2, AI-recommended word processing is performed on the data processed by S1;
[0198] S3, process the text and index fields based on the insertcontent field uploaded by the client;
[0199] S4 merges the data generated by S2 and S3 and sends it to the client.
[0200] Through innovative interaction between the client and server, this application combines voice input with natural language learning to more conveniently and quickly correct text omitted by users during speech, and intelligently predicts the content the user wants to say based on their input habits. First, this application directly converts the user's voice input into text, simplifying the input process. Furthermore, by incorporating AI predictive capabilities, users can experience automatic prediction functions previously only felt when inputting text, making input more convenient and intelligent. This application solves the problem of unscientific omissions that easily occur during voice input and simplifies the user's input based on their speech habits, making it more effective in conveying the sender's meaning.
[0201] Furthermore, speech recognition can be performed online via a server, while word segmentation and the retrieval of sensitive and intervention words can be handled on the client side.
[0202] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0203] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0204] According to another aspect of the embodiments of this application, a voice search apparatus for implementing the above-described voice search method is also provided. For example... Figure 17 As shown, the device includes:
[0205] The first display module 1702 is configured to respond to a first voice acquisition operation and display a first search term and a first recommended text set, wherein the first search term is a word identified in the first voice acquired by the first voice acquisition operation, and the first recommended text set includes text determined based on the first search term.
[0206] The selection module 1704 is used to select the first recommended text in the first recommended text set in response to the first recommended text selection operation when a first recommended text selection operation is detected while the first recommended text set is being displayed.
[0207] The second display module 1706 is used to display the first search text obtained by merging the first search term and the first recommended text, and to display a second set of recommended texts, wherein the second set of recommended texts includes text determined based on the first search text;
[0208] The third display module 1708 is configured to, in response to the initiated search operation, display a first search result when a search operation is detected while displaying the first search text, wherein the first search result includes results obtained from searching based on the first search text.
[0209] As an optional solution, the device is also used for:
[0210] If a second recommended text selection operation is detected when the second recommended text set is displayed, the second recommended text is selected in the second recommended text set in response to the second recommended text selection operation;
[0211] The system displays a second search text obtained by merging the first search text and the second recommended text, and displays a third set of recommended texts, wherein the third set of recommended texts includes texts determined based on the second search text.
[0212] If the search operation is detected while the second search text is being displayed, in response to the search operation, a second search result is displayed, wherein the second search result includes the result obtained from the search based on the second search text.
[0213] As an optional solution, the device is also used for:
[0214] When displaying the first search term and the first set of recommended texts, in response to the second voice acquisition operation, a third search text is displayed, wherein the third search text is a text obtained by merging the first search term and the second search term, and the second search term is a word identified in the second voice acquired by the second voice acquisition operation;
[0215] If the search operation is detected while the third search text is being displayed, a third search result is displayed in response to the search operation, wherein the third search result includes results obtained from searching based on the third search text.
[0216] As an optional solution, the device is also used for:
[0217] When the first search text and the second recommended text set are displayed, in response to the third voice acquisition operation, a fourth search text and a fourth recommended text set are displayed, wherein the fourth search text is the text obtained by merging the first search text and the third search term, the third search term is the word identified in the third voice acquired by the third voice acquisition operation, and the fourth recommended text set includes the text determined based on the fourth search text;
[0218] If the search operation is detected while the fourth search text is being displayed, in response to the search operation, a fourth search result is displayed, wherein the fourth search result includes results obtained from searching based on the fourth search text.
[0219] As an optional solution, the device is configured to select a first recommended text in the first recommended text set in response to a first recommended text selection operation when a first recommended text selection operation is detected while the first recommended text set is being displayed:
[0220] When displaying the first set of recommended texts, obtain the touch interaction operation performed on the first recommended text, wherein the touch interaction operation is used to select the first recommended text from the first set of recommended texts;
[0221] In response to the touch interaction, the first recommended text is moved to the position where the first search term is displayed to select the first recommended text.
[0222] As an optional solution, the device is used to display the first search text obtained by merging the first search term and the first recommended text, and to display the second set of recommended texts in the following manner:
[0223] When the first recommended text moves to the position where the first search term is displayed, the first search term and the first recommended text are merged to obtain the first search text;
[0224] The first search term displayed will be replaced with the first search text, and the first set of recommended texts displayed will be replaced with the second set of recommended texts.
[0225] As an optional solution, the device is further configured to: after displaying the first search text obtained by merging the first search term and the first recommended text, and displaying the second set of recommended texts, if a deletion operation is detected when displaying the first search text on a virtual button associated with the first recommended text, delete the first recommended text from the first search text, and display the first search term and the first set of recommended texts, wherein the virtual button is used to trigger the deletion of the associated recommended text; or if an editing operation is detected when displaying the first search text, display an editing interface for editing the first search text; delete the first recommended text in the editing interface, and display the first search term and the first set of recommended texts.
[0226] As an optional solution, the device is used to display the first search text obtained by merging the first search term and the first recommended text, and to display the second set of recommended texts in the following manner:
[0227] When the first recommended text has been selected and a fourth voice acquisition operation has been performed, in response to the fourth voice acquisition operation, the first search text is displayed and the second set of recommended texts is displayed. The first search text includes the first search term, the first recommended text, and the fourth search term. The fourth search term is a word identified in the fourth voice acquired by the fourth voice acquisition operation. The fourth search term is located after the first search term and the first recommended text in the first search text.
[0228] As an alternative, the device is configured to respond to a first voice acquisition operation by displaying a first search term and a first set of recommended text in the following manner:
[0229] In response to the first voice acquisition operation, the first string and the first data structure are sent from the target client to the target server, wherein the first string includes the text content of the first search term, and the first data structure is used to record the first recommended text parameters associated with the first search term;
[0230] The target client receives the first string and the second data structure returned by the target server, wherein the second data structure includes the first recommended text set and the first insertion parameter of each recommended text in the first recommended text set, the first recommended text set includes recommended text determined according to the first group of words in the first search term, and the first insertion parameter indicates the position of each recommended text in the first recommended text set inserted into the first group of words;
[0231] The first search term and the first set of recommended texts are displayed on the target client based on the first string and the second data structure.
[0232] As an alternative solution,
[0233] The device is configured to, in response to the first recommended text selection operation, select a first recommended text in the first recommended text set when a first recommended text selection operation is detected while displaying the first recommended text set: in response to the first recommended text selection operation, modify the second data structure to a third data structure on the target client, and send the first string and the third data structure to the target server, wherein the third data structure includes a target field indicating that the first recommended text has been selected;
[0234] The device is configured to display a first search text obtained by merging the first search term and the first recommended text, and to display a second set of recommended texts, in the following manner: receiving a second string returned by the target server and a fourth data structure updated by the target server according to the third data structure, wherein the second string represents the first search text, and the fourth data structure includes the position of the first recommended text inserted into the first group of words as determined by the target server according to the target field, and the second set of recommended texts; and displaying the first search text and the second set of recommended texts on the target client according to the second string and the fourth data structure.
[0235] As an optional solution, the device is also used for:
[0236] In the case where the third data structure also records the second recommended text parameters associated with the first search text, the target client receives the second string returned by the target server and the fourth data structure, wherein the fourth data structure further includes the second recommended text set and the second insertion parameters of each recommended text in the second recommended text set, the second recommended text set including the recommended text determined by the target server based on the second group of words in the first search text excluding the first group of words, and the second insertion parameters indicating the position of each recommended text in the second recommended text set inserted into the second group of words;
[0237] A target search interface is drawn on the target client based on the second string and the fourth data structure, wherein the target search interface displays the first search text and the second set of recommended texts.
[0238] As an optional solution, the device is used to display the first search text obtained by merging the first search term and the first recommended text, and to display the second set of recommended texts in the following manner:
[0239] If the string stored on the target client is the first string, and the first string is the same as the first string returned by the target server, the first string returned by the target server is used as the string stored on the target client, and the data structure stored on the target client is updated from the first data structure to the second data structure.
[0240] If the string stored on the target client is different from the first string returned by the target server, and during the process of receiving the first string and the second data structure returned by the target server, the string stored on the target client is updated to a third string, the third string stored on the target client and the first string returned by the target server are merged, the merged fourth string is updated to the string stored on the target client, the first data structure and the second data structure stored on the target client are merged, and the merged fifth data structure is updated to the data structure stored on the target client.
[0241] A target search interface is drawn on the target client based on the fourth string and the fifth data structure, wherein the target search interface displays the first search text and the second set of recommended texts.
[0242] According to one aspect of this application, a computer program product is provided, comprising a computer program / instructions containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via communication section 1809, and / or installed from removable media 1811. When the computer program is executed by central processing unit 1801, it performs various functions provided in embodiments of this application.
[0243] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0244] Figure 18 A schematic block diagram of a computer system architecture for implementing an electronic device according to embodiments of the present application is shown.
[0245] It should be noted that, Figure 18 The computer system 1800 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0246] like Figure 18 As shown, the computer system 1800 includes a central processing unit (CPU) 1801, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 1802 or programs loaded from storage section 1808 into random access memory (RAM). The RAM 1803 also stores various programs and data required for system operation. The CPU 1801, ROM 1802, and RAM 1803 are interconnected via a bus 1804. An input / output interface 1805 (I / O interface) is also connected to the bus 1804.
[0247] The following components are connected to the input / output interface 1805: an input section 1806 including a keyboard, mouse, etc.; an output section 1807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1808 including a hard disk, etc.; and a communication section 1809 including a network interface card such as a local area network card, modem, etc. The communication section 1809 performs communication processing via a network such as the Internet. A drive 1180 is also connected to the input / output interface 1805 as needed. Removable media 1811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on the drive 1180 as needed so that computer programs read from them can be installed into the storage section 1808 as needed.
[0248] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1809, and / or installed from removable medium 1811. When the computer program is executed by central processing unit 1801, it performs various functions defined in the system of this application.
[0249] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described voice search method is also provided. This electronic device may be... Figure 1 The terminal device or server shown. This embodiment uses this electronic device as an example for illustration. Figure 19As shown, the electronic device includes a memory 1902 and a processor 1904. The memory 1902 stores a computer program, and the processor 1904 is configured to execute the steps of any of the above method embodiments via the computer program.
[0250] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0251] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0252] S1, in response to the first voice acquisition operation, display the first search term and the first recommended text set, wherein the first search term is a word identified in the first voice acquired by the first voice acquisition operation, and the first recommended text set includes text determined based on the first search term;
[0253] S2, if a first recommended text selection operation is detected when the first recommended text set is displayed, in response to the first recommended text selection operation, the first recommended text is selected in the first recommended text set;
[0254] S3, display the first search text obtained by merging the first search term and the first recommended text, and display the second recommended text set, wherein the second recommended text set includes the text determined based on the first search text;
[0255] S4, if a search operation is detected when the first search text is displayed, in response to the search operation, the first search result is displayed, wherein the first search result includes the result obtained by searching based on the first search text.
[0256] Alternatively, as those skilled in the art will understand, Figure 19 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones (such as Android phones, iOS phones, etc.), tablets, PDAs, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 19 This does not limit the structure of the aforementioned electronic devices or electronic equipment. For example, electronic devices or electronic equipment may also include components that are more... Figure 19 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 19 The different configurations shown.
[0257] The memory 1902 can be used to store software programs and modules, such as the program instructions / modules corresponding to the voice search method and device in this embodiment. The processor 1904 executes various functional applications and data processing by running the software programs and modules stored in the memory 1902, thereby realizing the aforementioned voice search method. The memory 1902 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1902 may further include memory remotely located relative to the processor 1904, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 1902 may be used, but is not limited to, to store information such as search terms and recommended text. As an example, such as... Figure 19 As shown, the memory 1902 may include, but is not limited to, the first display module 1702, the selection module 1704, the second display module 1706, and the third display module 1708 of the voice search device. Furthermore, it may include, but is not limited to, other module units of the voice search device, which will not be elaborated upon in this example.
[0258] Optionally, the aforementioned transmission device 1906 is used to receive or send data via a network. Specific examples of the network may include wired and wireless networks. In one example, the transmission device 1906 includes a Network Interface Controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 1906 is a Radio Frequency (RF) module used for wireless communication with the Internet.
[0259] In addition, the aforementioned electronic device also includes: a display 1908 for displaying the aforementioned search results; and a connection bus 1910 for connecting the various module components in the aforementioned electronic device.
[0260] In other embodiments, the aforementioned terminal device or server can be a node in a distributed system, wherein the distributed system can be a blockchain system, which is a distributed system formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer (P2P) network, and any form of computing device, such as a server, terminal, or other electronic device, can become a node in the blockchain system by joining this peer-to-peer network.
[0261] According to one aspect of this application, a computer-readable storage medium is provided, from which a processor of a computer device reads computer instructions, and the processor executes the computer instructions, causing the computer device to perform the voice search method provided in the various alternative implementations of the above-described voice search aspect.
[0262] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0263] S1, in response to the first voice acquisition operation, display the first search term and the first recommended text set, wherein the first search term is a word identified in the first voice acquired by the first voice acquisition operation, and the first recommended text set includes text determined based on the first search term;
[0264] S2, if a first recommended text selection operation is detected when the first recommended text set is displayed, in response to the first recommended text selection operation, the first recommended text is selected in the first recommended text set;
[0265] S3, display the first search text obtained by merging the first search term and the first recommended text, and display the second recommended text set, wherein the second recommended text set includes the text determined based on the first search text;
[0266] S4, if a search operation is detected when the first search text is displayed, in response to the search operation, the first search result is displayed, wherein the first search result includes the result obtained by searching based on the first search text.
[0267] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0268] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0269] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0270] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0271] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.
[0272] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0273] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0274] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A voice search method, characterized in that, include: In response to a first voice acquisition operation, a first search term and a first recommended text set are displayed, wherein the first search term is a word identified in the first voice acquired by the first voice acquisition operation, and the first recommended text set includes text determined based on the first search term; If a first recommended text selection operation is detected when the first recommended text set is displayed, the first recommended text is selected in the first recommended text set in response to the first recommended text selection operation; The system displays the first search text obtained by merging the first search term and the first recommended text, and displays the second set of recommended texts, wherein the second set of recommended texts includes text determined based on the first search text; If a search operation is detected while the first search text is being displayed, in response to the search operation, a first search result is displayed, wherein the first search result includes results obtained from searching based on the first search text; The step of displaying the first search text obtained by merging the first search term and the first recommended text, and displaying the second set of recommended texts, includes: when the first recommended text is selected and a fourth voice acquisition operation is obtained, in response to the fourth voice acquisition operation, displaying the first search text and displaying the second set of recommended texts, wherein the first search text includes the first search term, the first recommended text, and a fourth search term, the fourth search term being a word identified in the fourth voice acquired by the fourth voice acquisition operation, the fourth search term being located after the first search term and the first recommended text in the first search text, the first search text being the fourth search term generated by real-time recognition and conversion of the fourth voice, and the first search term being the search text obtained by merging the first search term and the first recommended text.
2. The method according to claim 1, characterized in that, The method further includes: If a second recommended text selection operation is detected when the second recommended text set is displayed, the second recommended text is selected in the second recommended text set in response to the second recommended text selection operation; The system displays a second search text obtained by merging the first search text and the second recommended text, and displays a third set of recommended texts, wherein the third set of recommended texts includes texts determined based on the second search text. If the search operation is detected while the second search text is being displayed, in response to the search operation, a second search result is displayed, wherein the second search result includes the result obtained from the search based on the second search text.
3. The method according to claim 1, characterized in that, The method further includes: When displaying the first search term and the first set of recommended texts, in response to the second voice acquisition operation, a third search text is displayed, wherein the third search text is a text obtained by merging the first search term and the second search term, and the second search term is a word identified in the second voice acquired by the second voice acquisition operation; If the search operation is detected while the third search text is being displayed, a third search result is displayed in response to the search operation, wherein the third search result includes results obtained from searching based on the third search text.
4. The method according to claim 1, characterized in that, The method further includes: When the first search text and the second recommended text set are displayed, in response to the third voice acquisition operation, a fourth search text and a fourth recommended text set are displayed, wherein the fourth search text is the text obtained by merging the first search text and the third search term, the third search term is the word identified in the third voice acquired by the third voice acquisition operation, and the fourth recommended text set includes the text determined based on the fourth search text; If the search operation is detected while the fourth search text is being displayed, in response to the search operation, a fourth search result is displayed, wherein the fourth search result includes results obtained from searching based on the fourth search text.
5. The method according to claim 1, characterized in that, When a first recommended text selection operation is detected while displaying the first recommended text set, in response to the first recommended text selection operation, selecting the first recommended text from the first recommended text set includes: When displaying the first set of recommended texts, obtain the touch interaction operation performed on the first recommended text, wherein the touch interaction operation is used to select the first recommended text from the first set of recommended texts; In response to the touch interaction, the first recommended text is moved to the position where the first search term is displayed to select the first recommended text.
6. The method according to claim 5, characterized in that, The process of displaying the first search text obtained by merging the first search term and the first recommended text, and displaying the second set of recommended texts, includes: When the first recommended text moves to the position where the first search term is displayed, the first search term and the first recommended text are merged to obtain the first search text; The first search term displayed will be replaced with the first search text, and the first set of recommended texts displayed will be replaced with the second set of recommended texts.
7. The method according to claim 1, characterized in that, After displaying the first search text obtained by merging the first search term and the first recommended text, and displaying the second set of recommended texts, the method further includes: If a deletion operation is detected on the virtual button associated with the first recommended text when the first search text is displayed, the first recommended text in the first search text is deleted, and the first search term and the first set of recommended texts are displayed, wherein the virtual button is used to trigger the deletion of the associated recommended text; or If an editing operation is detected on the first search text while it is being displayed, an editing interface for editing the first search text is displayed; the first recommended text is deleted in the editing interface, and the first search term and the first set of recommended texts are displayed.
8. The method according to any one of claims 1 to 7, characterized in that, The response to the first voice acquisition operation, displaying the first search term and the first recommended text set, includes: In response to the first voice acquisition operation, the first string and the first data structure are sent from the target client to the target server, wherein the first string includes the text content of the first search term, and the first data structure is used to record the first recommended text parameters associated with the first search term; The target client receives the first string and the second data structure returned by the target server, wherein the second data structure includes the first recommended text set and the first insertion parameter of each recommended text in the first recommended text set, the first recommended text set includes recommended text determined according to the first group of words in the first search term, and the first insertion parameter indicates the position of each recommended text in the first recommended text set inserted into the first group of words; The first search term and the first set of recommended texts are displayed on the target client based on the first string and the second data structure.
9. The method according to claim 8, characterized in that, When a first recommended text selection operation is detected while displaying the first recommended text set, in response to the first recommended text selection operation, the first recommended text is selected in the first recommended text set, which includes: in response to the first recommended text selection operation, the second data structure is modified to a third data structure on the target client, and the first string and the third data structure are sent to the target server, wherein the third data structure includes a target field indicating that the first recommended text has been selected; The step of displaying the first search text obtained by merging the first search term and the first recommended text, and displaying the second recommended text set, includes: receiving a second string returned by the target server and a fourth data structure updated by the target server according to the third data structure, wherein the second string is used to represent the first search text, and the fourth data structure includes the position of the first recommended text inserted into the first group of words determined by the target server according to the target field and the second recommended text set; and displaying the first search text and the second recommended text set on the target client according to the second string and the fourth data structure.
10. The method according to claim 9, characterized in that, The method further includes: In the case where the third data structure also records the second recommended text parameters associated with the first search text, the target client receives the second string returned by the target server and the fourth data structure, wherein the fourth data structure further includes the second recommended text set and the second insertion parameters of each recommended text in the second recommended text set, the second recommended text set including the recommended text determined by the target server based on the second group of words in the first search text excluding the first group of words, and the second insertion parameters indicating the position of each recommended text in the second recommended text set inserted into the second group of words; A target search interface is drawn on the target client based on the second string and the fourth data structure, wherein the target search interface displays the first search text and the second set of recommended texts.
11. The method according to claim 9, characterized in that, The process of displaying the first search text obtained by merging the first search term and the first recommended text, and displaying the second set of recommended texts, includes: If the string stored on the target client is the first string, and the first string is the same as the first string returned by the target server, the first string returned by the target server is used as the string stored on the target client, and the data structure stored on the target client is updated from the first data structure to the second data structure. If the string stored on the target client is different from the first string returned by the target server, and during the process of receiving the first string and the second data structure returned by the target server, the string stored on the target client is updated to a third string, the third string stored on the target client and the first string returned by the target server are merged, the merged fourth string is updated to the string stored on the target client, the first data structure and the second data structure stored on the target client are merged, and the merged fifth data structure is updated to the data structure stored on the target client. A target search interface is drawn on the target client based on the fourth string and the fifth data structure, wherein the target search interface displays the first search text and the second set of recommended texts.
12. A voice search device, characterized in that, include: A first display module is configured to respond to a first voice acquisition operation by displaying a first search term and a first recommended text set, wherein the first search term is a word identified in the first voice acquired by the first voice acquisition operation, and the first recommended text set includes text determined based on the first search term; The selection module is used to select the first recommended text in the first recommended text set in response to the first recommended text selection operation when a first recommended text selection operation is detected while the first recommended text set is being displayed. The second display module is used to display the first search text obtained by merging the first search term and the first recommended text, and to display the second recommended text set, wherein the second recommended text set includes text determined based on the first search text; The third display module is configured to, in response to the initiated search operation, display a first search result when a search operation is detected while displaying the first search text, wherein the first search result includes results obtained by searching based on the first search text; The device is configured to display a first search text obtained by merging the first search term and the first recommended text, and to display a second set of recommended texts in the following manner: when the first recommended text is selected and a fourth voice acquisition operation is obtained, in response to the fourth voice acquisition operation, the first search text is displayed, and the second set of recommended texts is displayed, wherein the first search text includes the first search term, the first recommended text, and a fourth search term, the fourth search term being a word identified in the fourth voice acquired by the fourth voice acquisition operation, the fourth search term being located after the first search term and the first recommended text in the first search text, the first search text being the fourth search term generated by real-time recognition and conversion of the fourth voice, and the search text obtained by merging the first search term and the first recommended text.
13. The apparatus according to claim 12, characterized in that, The device is also used for: If a second recommended text selection operation is detected when the second recommended text set is displayed, the second recommended text is selected in the second recommended text set in response to the second recommended text selection operation; The system displays a second search text obtained by merging the first search text and the second recommended text, and displays a third set of recommended texts, wherein the third set of recommended texts includes texts determined based on the second search text. If the search operation is detected while the second search text is being displayed, in response to the search operation, a second search result is displayed, wherein the second search result includes the result obtained from the search based on the second search text.
14. The apparatus according to claim 12, characterized in that, The device is also used for: When displaying the first search term and the first set of recommended texts, in response to the second voice acquisition operation, a third search text is displayed, wherein the third search text is a text obtained by merging the first search term and the second search term, and the second search term is a word identified in the second voice acquired by the second voice acquisition operation; If the search operation is detected while the third search text is being displayed, a third search result is displayed in response to the search operation, wherein the third search result includes results obtained from searching based on the third search text.
15. The apparatus according to claim 12, characterized in that, The device is also used for: When the first search text and the second recommended text set are displayed, in response to the third voice acquisition operation, a fourth search text and a fourth recommended text set are displayed, wherein the fourth search text is the text obtained by merging the first search text and the third search term, the third search term is the word identified in the third voice acquired by the third voice acquisition operation, and the fourth recommended text set includes the text determined based on the fourth search text; If the search operation is detected while the fourth search text is being displayed, in response to the search operation, a fourth search result is displayed, wherein the fourth search result includes results obtained from searching based on the fourth search text.
16. The apparatus according to claim 12, characterized in that, The device is configured to select a first recommended text in the first recommended text set in response to a first recommended text selection operation when a first recommended text selection operation is detected while the first recommended text set is being displayed, by: When displaying the first set of recommended texts, obtain the touch interaction operation performed on the first recommended text, wherein the touch interaction operation is used to select the first recommended text from the first set of recommended texts; In response to the touch interaction, the first recommended text is moved to the position where the first search term is displayed to select the first recommended text.
17. The apparatus according to claim 16, characterized in that, The device is used to display the first search text, which is obtained by merging the first search term and the first recommended text, and to display the second set of recommended texts in the following manner: When the first recommended text moves to the position where the first search term is displayed, the first search term and the first recommended text are merged to obtain the first search text; The first search term displayed will be replaced with the first search text, and the first set of recommended texts displayed will be replaced with the second set of recommended texts.
18. The apparatus according to claim 12, characterized in that, After displaying the first search text obtained by merging the first search term and the first recommended text, and displaying the second set of recommended texts, the device is further configured to: If a deletion operation is detected on the virtual button associated with the first recommended text when the first search text is displayed, the first recommended text in the first search text is deleted, and the first search term and the first set of recommended texts are displayed, wherein the virtual button is used to trigger the deletion of the associated recommended text; or If an editing operation is detected on the first search text while it is being displayed, an editing interface for editing the first search text is displayed; the first recommended text is deleted in the editing interface, and the first search term and the first set of recommended texts are displayed.
19. The apparatus according to any one of claims 12 to 18, characterized in that, The device is configured to respond to a first voice acquisition operation by displaying a first search term and a first set of recommended texts in the following manner: In response to the first voice acquisition operation, the first string and the first data structure are sent from the target client to the target server, wherein the first string includes the text content of the first search term, and the first data structure is used to record the first recommended text parameters associated with the first search term; The target client receives the first string and the second data structure returned by the target server, wherein the second data structure includes the first recommended text set and the first insertion parameter of each recommended text in the first recommended text set, the first recommended text set includes recommended text determined according to the first group of words in the first search term, and the first insertion parameter indicates the position of each recommended text in the first recommended text set inserted into the first group of words; The first search term and the first set of recommended texts are displayed on the target client based on the first string and the second data structure.
20. The apparatus according to claim 19, characterized in that, The device is configured to, in response to the first recommended text selection operation, select a first recommended text in the first recommended text set when a first recommended text selection operation is detected while displaying the first recommended text set: in response to the first recommended text selection operation, modify the second data structure to a third data structure on the target client, and send the first string and the third data structure to the target server, wherein the third data structure includes a target field indicating that the first recommended text has been selected; The device is configured to display a first search text obtained by merging the first search term and the first recommended text, and to display a second set of recommended texts, in the following manner: receiving a second string returned by the target server and a fourth data structure updated by the target server according to the third data structure, wherein the second string represents the first search text, and the fourth data structure includes the position of the first recommended text inserted into the first group of words as determined by the target server according to the target field, and the second set of recommended texts; and displaying the first search text and the second set of recommended texts on the target client according to the second string and the fourth data structure.
21. The apparatus according to claim 20, characterized in that, The device is also used for: In the case where the third data structure also records the second recommended text parameters associated with the first search text, the target client receives the second string returned by the target server and the fourth data structure, wherein the fourth data structure further includes the second recommended text set and the second insertion parameters of each recommended text in the second recommended text set, the second recommended text set including the recommended text determined by the target server based on the second group of words in the first search text excluding the first group of words, and the second insertion parameters indicating the position of each recommended text in the second recommended text set inserted into the second group of words; A target search interface is drawn on the target client based on the second string and the fourth data structure, wherein the target search interface displays the first search text and the second set of recommended texts.
22. The apparatus according to claim 20, characterized in that, The device is used to display the first search text, which is obtained by merging the first search term and the first recommended text, and to display the second set of recommended texts in the following manner: If the string stored on the target client is the first string, and the first string is the same as the first string returned by the target server, the first string returned by the target server is used as the string stored on the target client, and the data structure stored on the target client is updated from the first data structure to the second data structure. If the string stored on the target client is different from the first string returned by the target server, and during the process of receiving the first string and the second data structure returned by the target server, the string stored on the target client is updated to a third string, the third string stored on the target client and the first string returned by the target server are merged, the merged fourth string is updated to the string stored on the target client, the first data structure and the second data structure stored on the target client are merged, and the merged fifth data structure is updated to the data structure stored on the target client. A target search interface is drawn on the target client based on the fourth string and the fifth data structure, wherein the target search interface displays the first search text and the second set of recommended texts.
23. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program can be executed by a terminal device or computer at runtime as described in any one of claims 1 to 11.
24. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 11.
25. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 11 through the computer program.