A method, device, electronic device and storage medium for processing warning text
By recognizing the user's voice information in real time and displaying the input progress in the business interface, the problem of users not reading the warning text is solved, which improves business processing efficiency and saves resources.
Patent Information
- Application Number
- CN202110476641.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-29
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2041-04-29
AI Technical Summary
Existing online commitment guarantee schemes cannot guarantee that users actually read the warning text, resulting in low efficiency and waste of network resources.
By displaying warning text in the business interface, it receives and recognizes the user's voice information in real time, displays the voice input progress, and saves the voice information when the similarity threshold is met.
Ensuring that users actually read the warning text improves business processing efficiency and saves time and network resources.
Smart Images

Figure CN113761120B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a method, device, electronic device, and storage medium for processing warning text. Background Art
[0002] Users who violate platform usage guidelines on messaging, social, or payment apps or their websites will be penalized by the platform. Depending on the severity of the violation, the platform may offer an opportunity to waive the penalty. For more serious offenses, the offending user will be educated on the rules and required to submit a pledge. Other apps, such as those for stocks, wealth management, insurance, and online lending, display a letter of commitment or other reminder text to warn users of risks when opening an account or making a payment.
[0003] Existing online commitment guarantee solutions are generally as follows Figure 1 As shown, methods such as checking the read checkbox and signing online cannot guarantee that the user has actually read the content. Therefore, if the user only checks the box perfunctorily without carefully reading the text content, and needs to reread it later, it is inefficient, delays time, and wastes network resources. Summary of the Invention
[0004] The embodiments of the present application provide a method, device, electronic device, and storage medium for processing warning texts to solve the problem that online warning solutions are inefficient and waste network resources.
[0005] This embodiment of the present application provides a method for processing a warning text, including:
[0006] In response to an execution operation triggered for a target business in the business interface, displaying a warning text corresponding to the target business;
[0007] receiving in real time a voice message based on the warning text input;
[0008] Performing real-time speech recognition on the speech information to obtain real-time speech text;
[0009] Displaying the voice input progress in the warning text based on the real-time voice text;
[0010] When the voice input progress indicates that the voice information input is completed, the voice information is saved in association with the target service.
[0011] On the other hand, an embodiment of the present application further provides a device for processing a warning text, comprising:
[0012] A display unit, configured to display a warning text corresponding to a target business in response to an execution operation triggered for the target business in the business interface;
[0013] An input unit, configured to receive in real time a voice message input based on the warning text;
[0014] A recognition unit, configured to perform real-time speech recognition on the speech information to obtain a real-time speech text;
[0015] The display unit is further configured to display the voice input progress in the warning text based on the real-time voice text;
[0016] A storage unit is configured to store the voice information in association with the target service when the voice input progress indicates that the voice information input is completed.
[0017] Optionally, the recognition unit is further configured to determine the similarity between the real-time voice text and the warning text;
[0018] The display unit is further configured to prompt the user to re-enter the voice when the similarity is less than a similarity threshold; and to execute a step of displaying the voice input progress in the warning text based on the real-time voice text when the similarity is greater than or equal to the similarity threshold.
[0019] Optionally, the display unit is specifically used to:
[0020] Matching the real-time voice text with the warning text, and locating the text position of the real-time voice text in the warning text;
[0021] The voice real-time text is displayed in the warning text based on the text position to show the voice input progress.
[0022] Optionally, the display unit is specifically used to:
[0023] Matching the real-time voice text with the warning text to determine a voice comparison text in the warning text that corresponds to the real-time voice text;
[0024] The voice control text is displayed on the page where the warning text is located.
[0025] Optionally, the storage unit is further configured to:
[0026] Determining the speech characteristics of the target service based on the speech information;
[0027] The voice information and the speech spectrum features are stored in association with the target service.
[0028] Optionally, the display unit is further used to:
[0029] In response to a release operation triggered for a target service, displaying a release rejection interface for the target service; the release rejection interface includes a warning text;
[0030] Play the voice information saved in the target service association.
[0031] Optionally, the display unit is further used to:
[0032] In response to an input operation for the warning text, starting a timer;
[0033] If the voice information input is not completed based on the voice input progress and the timing duration exceeds the set time limit, the reception of the voice information is stopped and an input timeout prompt interface is displayed.
[0034] On the other hand, an embodiment of the present application further provides an electronic device, including:
[0035] a memory for storing executable instructions;
[0036] The processor is configured to read and execute the executable instructions stored in the memory to implement the method described above.
[0037] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, which, when instructions in the computer-readable storage medium are executed by a processor, enables the processor to execute the method described above.
[0038] In an embodiment of the present application, the user operates in the business interface of the terminal, and the terminal responds to the execution operation for the target business and presents the corresponding warning text to the user, so that the user can read aloud according to the content in the warning text. The terminal receives the voice information input based on the warning text, and performs real-time voice recognition on the voice information to obtain the real-time voice text. While receiving the voice information, the terminal displays the voice input progress in the warning text based on the real-time voice text, so that the user can read the warning text aloud. When the voice input progress indicates that the voice information input is completed, the voice information is saved in association with the corresponding target business. Thus, by comparing the voice information entered by the user, it is ensured that the user has actually read the warning text. In this way, the problem of the user only checking the box perfunctorily without reading the text content and needing to review it again later is avoided. This improves the user's efficiency in handling business and saves time and network resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 This is a schematic diagram of a warning page of a terminal in the related art;
[0040] Figure 2 This is a schematic diagram of the application architecture of the warning text processing method in the embodiment of the present application;
[0041] Figure 3 A detailed flowchart of a method for processing a warning text in an embodiment of the present application is shown;
[0042] Figures 4A to 4I This is a schematic diagram of an interface for processing warning texts by a terminal in an embodiment of the present application;
[0043] Figures 5A to 5F This is a schematic diagram of an interface for processing a warning text by a terminal in another embodiment of the present application;
[0044] Figure 6 A schematic diagram of the implementation process of the warning text processing method provided in a specific embodiment of the present application;
[0045] Figure 7 A schematic diagram of the structure of the warning text processing device provided in an embodiment of the present application;
[0046] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] To make the objectives, technical solutions, and advantages of this application more clear, this application will be further described in detail below with reference to the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0048] To facilitate understanding of the embodiments of this application, several concepts are briefly introduced below:
[0049] ASR (Automatic Speech Recognition) is a technology that converts human speech into text. Speech recognition is a multidisciplinary field, closely linked to acoustics, phonetics, linguistics, digital signal processing theory, information theory, computer science, and other disciplines. Its goal is to convert the lexical content of human speech into computer-readable input, such as keystrokes, binary codes, or character sequences.
[0050] Text similarity matching: refers to a technology that compares the similarity of text content. In the embodiment of this application, it refers to whether the converted text obtained by converting the user's voice input is consistent with the content of the provided warning text. There are three methods for measuring text similarity: one is the traditional method based on keyword matching, such as N-gram (N-gram model) similarity; the second is to map the text to a vector space and then use methods such as cosine similarity; the third is deep learning methods, such as deep learning semantic matching models (Deep Structured Semantic Models, DSSM) based on user click data, convolutional neural networks (CNN), and long short-term memory networks (Long Short-Term Memory, LSTM) and other methods.
[0051] A spectrogram is a graph of the speech spectrum. It is typically generated by processing the received time-domain signal. It displays a time-dependent Fourier analysis, showing how the speech signal spectrum changes over time. The abscissa of a spectrogram is time, the ordinate is frequency, and the value at each coordinate represents the energy of the speech data. Because it uses a two-dimensional plane to represent three-dimensional information, the energy value is represented by color: darker colors indicate stronger speech energy at that point.
[0052] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0053] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0054] Natural language processing (NLP) is a key area of research in computer science and artificial intelligence. It studies the theories and methods that enable effective communication between humans and computers using natural language. Natural language processing (NLP) integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language we use in everyday life—and is closely linked to the study of linguistics. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robotic question answering, and knowledge graphs.
[0055] Machine Learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning.
[0056] Cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to access computing power, storage space, and information services as needed. The network that provides these resources is called the "cloud." To users, these resources appear infinitely scalable and can be accessed at any time, used on demand, expanded at any time, and paid for on a pay-per-use basis.
[0057] As a provider of cloud computing infrastructure, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an Infrastructure as a Service (IaaS) platform) is established. Various types of virtual resources are deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.
[0058] Based on logical functional divisions, the Platform as a Service (PaaS) layer can be deployed on top of the IaaS layer, and the Software as a Service (SaaS) layer can be deployed on top of the PaaS layer. SaaS can also be deployed directly on top of IaaS. PaaS is a platform for software execution, such as databases and web containers. SaaS is a variety of business software, such as web portals and text messaging apps. Generally speaking, SaaS and PaaS are layers above IaaS.
[0059] Users of communication, social, and payment apps or websites who violate platform usage guidelines will be penalized accordingly. Depending on the severity of the violation, the platform may offer users the opportunity to waive the penalty. For more serious violations, the platform may require users to receive guidance on the rules and submit a commitment guarantee. Furthermore, in other apps, such as those for stocks, wealth management, insurance, and online lending, users may be presented with text messages providing commitments or risk warnings during the transaction process.
[0060] Related technical solutions are generally as follows Figure 1 As shown, the user is provided with a reading option on the warning page. After reading, the user can check the box to proceed with the subsequent business. Alternatively, the warning text is displayed to the user on the warning page, and the user signs on the display after reading. However, none of the above solutions can guarantee that the user has read the text.
[0061] In view of this, in an embodiment of the present application, the user operates in the business interface of the terminal, and the terminal responds to the execution operation for the target business and presents the corresponding warning text to the user, so that the user can read aloud according to the content in the warning text. The terminal receives the voice information input based on the warning text, and performs real-time voice recognition on the voice information to obtain the real-time voice text. While receiving the voice information, the terminal displays the voice input progress in the warning text based on the real-time voice text, so that the user can read the warning text aloud. When the voice input progress indicates that the voice information input is completed, the voice information is saved in association with the corresponding target business. Thus, by comparing the voice information entered by the user, it is ensured that the user has actually read the warning text. In this way, the problem of the user only checking the text perfunctorily without reading the text content is avoided, and then having to review it again. This improves the user's efficiency in handling business and saves time and network resources.
[0062] The preferred embodiments of the present application are further described in detail below with reference to the accompanying drawings.
[0063] In specific implementation, the process of processing warning text can be applied to various application scenarios. Figure 2, which is a schematic diagram of the application architecture of the warning text processing method in an embodiment of the present application, includes a server 100 and a terminal device 200.
[0064] The terminal device 200 can be a mobile or fixed electronic device. For example, it can be a mobile phone, tablet computer, laptop computer, desktop computer, various wearable devices, smart TV, in-vehicle device, or other electronic device capable of performing the above functions. The terminal device 200 can display articles, short messages, and other content used for warnings to the user, receive voice information read by the user through a voice input device such as a microphone, send the voice information to the server 100, and receive text comparison results sent by the server 100 and display them to the user.
[0065] The terminal device 200 and the server 100 can be connected via the Internet to enable communication between them. Optionally, the above-mentioned Internet uses standard communication technologies and / or protocols. The Internet is typically the Internet, but it can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or any combination of a virtual private network. In some embodiments, technologies and / or formats including Hypertext Markup Language (HTML) and Extensible Markup Language (XML) are used to represent data exchanged over the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.
[0066] The server 100 can provide various network services for the terminal device 200, and the server 100 can use cloud computing technology for information processing. Among them, the server 100 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected through wired or wireless communication, and this application does not limit this.
[0067] Specifically, the server 100 may include a processor 110 (Center Processing Unit, CPU), a memory 120, an input device 130 and an output device 140, etc. The input device 130 may include a keyboard, a mouse, a touch screen, etc., and the output device 140 may include a display device, such as a liquid crystal display (LCD), a cathode ray tube (CRT), etc.
[0068] The memory 120 may include a read-only memory (ROM) and a random access memory (RAM), and provides program instructions and data stored in the memory 120 to the processor 110. In an embodiment of the present invention, the memory 120 may be used to store the program of the warning text processing method in an embodiment of the present invention.
[0069] The processor 110 calls the program instructions stored in the memory 120, and the processor 110 is used to execute the steps of any one of the warning text processing methods in the embodiments of the present invention according to the obtained program instructions.
[0070] Based on the above design concept, refer to Figure 3 As shown, in the embodiment of the present application, the detailed process of implementing the processing of the warning text is as follows:
[0071] Step 301: In response to an execution operation triggered for a target service in a service interface, the terminal displays a warning text corresponding to the target service.
[0072] Specifically, the terminal can display the corresponding service interface during the execution of the target service and present warning text in the service interface in response to the user's execution operation for the target service. In specific implementations, the warning text can be generated by the terminal sending an execution request to the server based on the execution operation. The server determines the corresponding warning text and sends it to the terminal, which then displays the warning text. Alternatively, the terminal can pre-retrieve the corresponding warning text from the server based on the target service and store it. When the user triggers the execution operation, the terminal directly retrieves and displays the warning text from memory.
[0073] For example, see Figure 4A In response to the user's login operation, the terminal sends a login request to the server. Based on the account information in the login request, the server determines that the account has been complained by multiple people and there are violations, and then sends a login restriction response to the terminal. The login restriction response contains a warning text. After receiving the login restriction response, the terminal presents the following to the user: Figure 4B When the user operates in the prompt interface, such as clicking "Next", the terminal will display the warning text, such as Figure 4C shown.
[0074] For example, when a user applies for a loan in an APP, he needs to enter the identity information of the loan party. Figure 4D After the user clicks "Continue Input", the terminal responds to the click operation and displays a warning text to the user, such as Figure 4E shown.
[0075] Step 302: The terminal receives voice information based on the warning text input in real time.
[0076] In the specific implementation process, the terminal device can receive voice information through a voice input device such as a microphone and convert the voice information into text information. In addition, the terminal device can also use a camera to shoot video and receive voice information during the shooting process.
[0077] Generally speaking, the interface that displays the warning text also displays a button that prompts the user to input a voice message, such as Figure 4C or Figure 4E As shown, a button of "click to start" is displayed in the interface. In this way, after the user clicks the corresponding button, the terminal starts to receive voice information in response to the user's click operation.
[0078] Step 303: The terminal performs real-time speech recognition on the voice information to obtain real-time speech text.
[0079] Here, the terminal device may perform voice recognition on the received voice information to obtain real-time voice text; or the terminal device may send the voice information to a server, and the server may convert the voice information into real-time voice text.
[0080] In the embodiment of the present application, ASR is used to recognize voice information and obtain corresponding real-time text of the voice. Specifically, the received voice information can be converted into text unit by unit, and the context is combined with the conversion to increase the accuracy of the conversion.
[0081] A unit here can contain one or more words. In order to facilitate subsequent voice input and text comparison, the warning text is divided into multiple units. Here, the warning text is split according to the rules to obtain multiple units, and the number of words contained in each unit is not limited. For example, the warning text can be divided according to paragraphs to obtain multiple units, and a paragraph in a warning text is a unit; or the warning text can be divided according to sentences to obtain multiple units, and a sentence in a warning text is a unit; or the warning text can be divided according to words to obtain multiple units, and a word in a warning text is a unit.
[0082] The speech recognition process can be mainly divided into four steps: "input-encoding-decoding-output".
[0083] Specifically, the terminal receives the sound produced by the user reading the warning text through a microphone. The sound itself is an audio.
[0084] After the audio signal is processed, it is split into frames (millisecond level), and the split small waveforms are converted into multi-dimensional vector information according to the characteristics of the human ear.
[0085] These frame information are identified as states, where the state can be understood as an intermediate process, a process smaller than a phoneme.
[0086] The states are then combined to form phonemes, usually three states are combined into one phoneme.
[0087] Finally, the phonemes are combined into words and strung together into sentences, thus realizing the conversion from speech to text.
[0088] Furthermore, after the terminal performs real-time speech recognition on the voice information to obtain real-time speech text, the method further includes:
[0089] Determine the similarity between the real-time text of the speech and the warning text;
[0090] When the similarity is less than the similarity threshold, prompt the user to re-enter the voice;
[0091] When the similarity is greater than or equal to the similarity threshold, step 304 is executed.
[0092] Specifically, in the embodiment of the present application, text similarity matching is performed between the real-time speech text and the warning text to determine the similarity between the two. The text similarity matching method is not limited and can be performed using cosine similarity, term frequency-inverse document frequency (TF-IDF) method, DSSM, etc.
[0093] In an embodiment of the present application, the warning text is divided into multiple units for text similarity matching. The units here can be divided according to paragraphs, that is, one paragraph is one unit; or they can be divided according to sentences to obtain multiple units, one sentence is one unit; or they can be divided according to words, one word is one unit. After the voice information is subjected to voice recognition to obtain the corresponding real-time voice text, each time a unit text is obtained, the unit text is matched with the corresponding content in the warning text for text similarity to obtain the corresponding similarity value. In this way, multiple similarity values can be obtained for the total warning text.
[0094] Afterwards, based on preset conditions, the multiple similarity values obtained are used to determine whether the voice information received by the terminal meets the requirements.
[0095] Here, each similarity must be greater than or equal to a preset similarity threshold, that is, the similarity value between each unit of voice information obtained by speech recognition and the corresponding warning text must be greater than or equal to the similarity threshold. This ensures that the user is fully aware of the entire warning text.
[0096] There is no restriction on the similarity threshold in the specific implementation process. For example, all similarity thresholds are the same, or multiple similarity thresholds can be set.
[0097] If the similarity does not meet the preset conditions, the voice input is determined to have failed. Specifically, each time a voice message is received, a determination is made as to whether the similarity between the real-time voice text and the corresponding warning text is less than a similarity threshold. Thus, if it is determined that the similarity between the currently received real-time voice text and the corresponding warning text is less than the preset similarity threshold, the voice message reception is stopped and a recording failure prompt interface is displayed.
[0098] During the specific implementation process, after each voice message is received, the currently received voice message is subjected to voice recognition to determine the real-time voice text corresponding to the voice message. The similarity between the real-time voice text and the corresponding warning text is calculated, and the relationship between the similarity and the preset similarity threshold is determined. If the similarity is greater than or equal to the similarity threshold, the next voice message is continued to be received and compared until all voice messages are received. If the similarity is less than the similarity threshold, it is determined that the voice message does not match the warning text, and the voice input is determined to have failed, and subsequent voice messages are stopped from being received. An input failure prompt interface is displayed, prompting the user to re-enter the voice message, such as Figure 4F In this case, the user can continue the operation on the input failure prompt interface and try again to input the voice; or, the user can click "Cancel Input" to terminate the voice input.
[0099] Step 304: The terminal displays the voice input progress in the warning text based on the real-time voice text.
[0100] During the specific implementation, when the similarity meets a preset condition, that is, the similarity is greater than or equal to the similarity threshold, the terminal performs the step of displaying the voice input progress in the alert text based on the real-time voice text. To facilitate the user's reading, the voice input progress is displayed in the alert text, thereby indicating that the user has read the alert text.
[0101] There is no restriction on the way of displaying the voice input progress in the warning text. The voice input progress can be displayed directly in the warning text or outside the warning text.
[0102] In a preferred embodiment, the terminal displays the voice input progress in the warning text based on the real-time voice text, including:
[0103] Matching the real-time voice text with the warning text, and locating the text position of the real-time voice text in the warning text;
[0104] Based on the text position, the real-time text of the voice input is displayed separately in the alert text to show the progress of the voice input.
[0105] During the specific implementation process, after the terminal receives the voice information and recognizes the voice information as real-time voice text, it matches the real-time voice text with the warning text for similarity, and displays the real-time voice text whose similarity with the warning text is greater than or equal to the similarity threshold according to its corresponding position in the warning text.
[0106] The specific display method of the present application is not limited in the embodiment, for example, it can be to add underline, highlight text, bold text, increase text background color, etc. Figure 4GAs shown, the terminal uses bold text to mark the part of the warning text that the user has read aloud. Figure 4H As shown, the terminal uses underlines to mark the part of the warning text that the user has read aloud.
[0107] In another preferred embodiment, the terminal displays the voice input progress in the warning text based on the real-time voice text, including:
[0108] Matching the real-time voice text with the warning text to determine the voice comparison text in the warning text that corresponds to the real-time voice text;
[0109] The voice-to-text is displayed on the page where the warning text is located.
[0110] In addition to directly displaying the real-time voice text in the warning text, the embodiment of the present application can also display the real-time voice text in an area outside the warning text. Specifically, based on the result of the similarity matching, the voice comparison text corresponding to the real-time voice text in the warning text is determined, and then the voice comparison text is displayed on the page. For the convenience of comparison, it is generally necessary to display the voice comparison text and the warning text on the same page. Figure 4I As shown, the black box displays the warning text that the user has read aloud, that is, the voice comparison text, and the black box and the warning text are located on the same page.
[0111] Step 305: When the voice input progress indicates that the voice information input is completed, the terminal associates and saves the voice information with the target service.
[0112] The terminal saves all received voice messages based on the warning text. The saving here can be that the terminal saves the voice messages in the terminal's memory; or the terminal can send all voice messages to the server, which will associate and save them with the target service.
[0113] In an embodiment of the present application, the user operates in the business interface of the terminal, and the terminal responds to the execution operation for the target business and presents the corresponding warning text to the user, so that the user can read aloud according to the content in the warning text. The terminal receives the voice information input based on the warning text, and performs real-time voice recognition on the voice information to obtain the real-time voice text. While receiving the voice information, the terminal displays the voice input progress in the warning text based on the real-time voice text, so that the user can read the warning text aloud. When the voice input progress indicates that the voice information input is completed, the voice information is saved in association with the corresponding target business. Thus, by comparing the voice information entered by the user, it is ensured that the user has actually read the warning text. In this way, the problem of the user only checking the box perfunctorily without reading the text content and needing to review it again later is avoided. This improves the user's efficiency in handling business and saves time and network resources.
[0114] Furthermore, in order to save network resources and avoid wasting time, the embodiment of the present application limits the user's recording time, that is, the user needs to read the entire warning text aloud within a limited time. In a preferred embodiment, in response to an execution operation triggered for a target service in the service interface, after displaying the warning text corresponding to the target service, and before receiving a voice message input based on the warning text in real time, the following is also included:
[0115] In response to an input operation for the warning text, starting a timer;
[0116] After receiving voice messages based on alert text input in real time, it also includes:
[0117] If the voice input progress indicates that the voice message input is not completed and the timing exceeds the set time limit, the voice message reception will be stopped and the input timeout prompt interface will be displayed.
[0118] In the specific implementation process, during the voice recording process, a warning text can be displayed to the user in the business interface, such as Figure 5A As shown, the user needs to read the recording within a limited time. Specifically, the user can click Figure 5A Click the "Start" button in the Start menu to start recording. At this time, the terminal starts timing.
[0119] After receiving the voice message, the terminal sends the voice message to the platform server. The platform server converts the voice into text, obtains the corresponding converted sub-text, locates the corresponding source sub-text content in the warning text, and sends the location information back to the terminal. In this way, the terminal can mark and display the text that the user has read aloud based on the received location information. The corresponding display interface of the terminal is to underline the part of the text that has been read according to the speed of the user's reading, such as Figure 5B shown.
[0120] During recording, the user can pause and resume recording at any time by operating in the display interface. Figure 5B If you click “Pause Recording” in the menu, the terminal will stop receiving voice messages and display the following message: Figure 5C If the user needs to continue recording, click Figure 5C Click "Continue Recording" in the menu, and the terminal will continue to receive voice information through the microphone.
[0121] When the user finishes reading, the terminal can automatically stop recording; or the terminal can allow the user to operate in the display interface, such as clicking Figure 5D Click "End Recording" in the dialog box, and the terminal stops recording in response to this operation.
[0122] When the terminal stops recording, the user can listen to the recording again. For example, click Figure 5E In the “Click to Listen”, the terminal plays the voice that the user has just recorded. In addition, if the user is not satisfied with the current recording, he can click “Re-record”, and the terminal will delete the voice that has just been received and receive the voice information again.
[0123] During the recording process, the user's voice information can be converted into a conversion sub-text in real time, and then the conversion sub-text and the corresponding source sub-text in the warning text are matched to ensure that the user is reading and recording according to the displayed warning text content.
[0124] After the user completes voice input, he can click Figure 5E Click "Finish" to end the voice information input. At the same time, the terminal or platform server will archive the voice information file.
[0125] In a preferred embodiment, when the voice input progress indicates that the voice information input is completed, after the voice information is associated and saved with the target service, the method further includes:
[0126] In response to a release operation triggered for a target service, displaying a release rejection interface for the target service; the release rejection interface includes a warning text;
[0127] Play the voice information saved in the target service association.
[0128] In the specific implementation process, the platform server can save the voice information read aloud by the user. The platform retains the user's voice guarantee. If the user violates the rules again, the platform can present the last recorded voice information for questioning and use it as a basis to refuse to lift the penalty again.
[0129] For example, the user follows Figure 5E The warning text shown in is read aloud, and after the terminal receives the voice message read aloud by the user, it sends the corresponding voice message to the platform server for storage. If the user violates the rules again later, the platform will cancel the user's account according to the rules, and blacklist the user, prohibiting the user from registering. In this case, when the user re-registers an account on the platform, the terminal responds to the user's registration operation and sends a registration request to the server. Based on the account information in the login request, the server determines that the account still has violations after entering the warning text, and then sends the warning text and the corresponding voice message that the user has read aloud to the terminal. The terminal presents the following to the user. Figure 5F The rejection cancellation interface shown indicates that the user's registration request is rejected, and the warning text is displayed in the rejection cancellation interface, and the voice message previously recorded by the user is played.
[0130] In another preferred embodiment, when the voice input progress indicates that the voice information input is completed, the method further includes:
[0131] Determine the speech characteristics of the target business based on voice information;
[0132] Save voice information associated with the target service, including:
[0133] The voice information and spectral features are stored in association with the target service.
[0134] During the specific implementation process, after the warning text is recorded, the spectral features of the entire recording need to be converted and saved based on the received voice information. The spectral features here correspond to the specific user and the warning text, for example, it can be a spectrogram. The spectrogram is a graphical representation of the sound signal. Its horizontal axis represents time, and its vertical axis represents frequency. The amplitude of the voice at each frequency point is distinguished by color. The terminal mainly extracts parameters such as the fundamental frequency spectrum and envelope of the speaker's voice, the energy of the fundamental frequency frame, the frequency of occurrence of the fundamental frequency resonance peak and its trajectory. The fundamental frequency and harmonic frequency of the user's voice are displayed as bright lines on the spectrogram. The feature spectrum will be recorded and stored in conjunction with the user's identity information and the corresponding warning text. Since the spectrogram is unique, it can be used as an identifier for the user and the target service.
[0135] In addition, the spectrogram can be based on partial or full voice information. Since voiceprint identification generally takes more than 20 seconds, the length of the user's voice recording should be taken into consideration when generating the spectrogram. It should not be too long, as this will affect the user's completion rate; nor should it be too short, as this will affect the effectiveness of subsequent voiceprint identification.
[0136] The following is a specific example to illustrate the implementation process of the warning text processing method provided in the embodiment of the present application. Figure 6 shown.
[0137] The interface first explains the requirements for recording the voice guarantee and displays the content of the guarantee that needs to be read. After the user clicks "Start Recording", the recording process begins. The system uses the microphone function of the front-end device (computer, mobile phone, etc.), and the ASR unit, NLP unit, and progress control module start working simultaneously.
[0138] After the user enters their voice, the terminal sends the voice information to the server, which recognizes and converts the user's voice input into real-time text. The real-time text here is updated in real time, and the real-time text of the voice conversion is updated as the user progresses.
[0139] The server performs a similarity match between the real-time text converted from the user's voice and the warning text. Taking into account the user's accent, the similarity threshold set for text similarity matching is 0.5 (this is a reference value, and a perfect match score of 1).
[0140] After comparing the real-time voice text with the warning text, the server performs two operations: one is to output the similarity value to determine whether the user's voice meets the preset conditions; the other is to locate the user's reading progress.
[0141] 1) If the similarity value is less than 0.5, the user's voice information is determined to be inconsistent with the requirements, the terminal interrupts the recording and prompts to re-record;
[0142] 2) If the similarity value is greater than or equal to 0.5, the user's voice information is determined to be consistent with the requirements. The server determines the progress of the user's recording based on the positioning result and notifies the terminal to display the user's recorded content with a mark;
[0143] 3) If the user's voice message completely matches the last sentence of the warning text and the user does not enter any new voice message at this time, the recording process is completed. The terminal prompts the user that the recording has been completed and can be listened to or re-recorded;
[0144] 4) According to the system recording time requirement, if the server determines that the recording time has reached the upper limit, it will actively interrupt the recording and prompt a timeout, requiring the user to re-record.
[0145] After the recording is completed, the server converts and saves the characteristic spectrogram of the entire recording, and binds it with the user's identity information on the platform and the recorded voice information for storage.
[0146] The application scenario of the specific embodiment 2 is that when a user is handling banking business online, the bank APP displays a warning text through the interface to remind the user to pay attention to property safety. The specific processing process is as follows:
[0147] The terminal responds to a click operation triggered by the user in the banking service interface by displaying a corresponding warning text and a corresponding operation method.
[0148] The terminal responds to the input operation triggered by the user in the interface, receives the voice information in real time, and starts timing.
[0149] The terminal performs voice recognition on the voice information received in real time to obtain real-time voice text.
[0150] The terminal compares the real-time voice text with the warning text to determine the similarity between the real-time voice text and the warning text:
[0151] If the similarity between the two is less than the similarity threshold, the user stops receiving the voice message and is prompted to re-enter the message.
[0152] If the similarity between the two is greater than or equal to the similarity threshold, determining the text position of the real-time voice text in the warning text;
[0153] A horizontal line is placed under the text position in the warning text corresponding to the real-time voice text to show the progress of the received voice information.
[0154] During the process of receiving a voice message, if it is determined that the timing duration exceeds the set time limit, the reception of the voice message will be stopped and a prompt interface indicating that the input timeout has occurred will be displayed to the user.
[0155] When the marked horizontal line in the warning text reaches the end of the warning text, it indicates that the voice information is received completely, then the voice input completion interface is displayed to the user, and the voice information is sent to the server.
[0156] After receiving the voice information, the server determines the spectral features of the voice information.
[0157] The server associates the voice information and the determined spectral features with the user's account and saves them.
[0158] The application scenario of the specific embodiment 3 is that when a user logs in to a social APP, the social APP determines that the user has previously violated the rules based on the user's account information, and then displays a warning text to the user through the terminal to warn the user. The specific processing process is as follows:
[0159] The terminal responds to the login operation triggered by the user in the login interface and obtains the corresponding account and password information.
[0160] The terminal sends a login request to the server, which contains the user's account number, password and other information.
[0161] The server performs a security check on the user's account based on the received account number and password, and determines that the user has violated the rules and has been restricted from logging in.
[0162] The server sends a login response to the terminal, which includes the user's login restriction information and warning text.
[0163] The terminal displays a login restriction prompt interface to the user based on the received login restriction information.
[0164] In response to the user's confirmation operation on the login restriction prompt interface, the terminal displays a warning text to the user and displays an input instruction of the warning text.
[0165] In response to the user's input operation, the terminal starts to receive the voice information in real time and starts timing.
[0166] The terminal performs voice recognition on the voice information received in real time to obtain real-time voice text.
[0167] The terminal compares the real-time voice text with the warning text to determine the similarity between the real-time voice text and the warning text:
[0168] If the similarity between the two is less than the similarity threshold, the user stops receiving the voice message and is prompted to re-enter the message.
[0169] If the similarity between the two is greater than or equal to the similarity threshold, determining the text position of the real-time voice text in the warning text;
[0170] In the display interface of the warning text, the real-time text of the recorded voice is also displayed to show the progress of the received voice information.
[0171] During the process of receiving a voice message, if it is determined that the timing duration exceeds the set time limit, the reception of the voice message will be stopped and a prompt interface indicating that the input timeout has occurred will be displayed to the user.
[0172] When the displayed real-time voice text reaches the end of the warning text, it indicates that the voice information reception is completed, and the voice input completion interface is displayed to the user.
[0173] The terminal determines the spectral features of the voice information and sends the voice information and the spectral features to the server.
[0174] The server associates the received voice information and spectral features with the user's account and saves them.
[0175] After a period of time, the terminal again responds to the login operation triggered by the user in the login interface and obtains the corresponding account and password information.
[0176] The terminal sends a login request to the server, which contains the user's account number, password and other information.
[0177] The server performs a security check on the user's account based on the received account number and password. If it is determined that the user still has violations after entering the voice information, the server obtains the user's voice information and spectral features.
[0178] The server sends a login rejection response to the terminal, which includes the voice information and spectral features entered by the user last time.
[0179] The terminal displays a prompt interface for refusing login to the user, and displays the warning text corresponding to the voice message previously entered by the user.
[0180] In response to the user's play operation, the terminal plays the voice information previously recorded by the user to the user.
[0181] Corresponding to the above method embodiment, the embodiment of the present application also provides a warning text processing device. Figure 7 This is a schematic diagram of the structure of the warning text processing device provided in the embodiment of the present application; Figure 7 As shown, the warning text processing device includes:
[0182] The display unit 701 is configured to display a warning text corresponding to a target service in response to an execution operation triggered for the target service in the service interface;
[0183] An input unit 702 is configured to receive voice information input based on the warning text in real time;
[0184] The recognition unit 703 is used to perform real-time speech recognition on the speech information to obtain real-time speech text;
[0185] The display unit 701 is further configured to display the voice input progress in the warning text based on the real-time voice text;
[0186] The saving unit 704 is configured to save the voice information in association with the target service when the voice input progress indicates that the voice information input is completed.
[0187] Optionally, the recognition unit 703 is further configured to determine the similarity between the real-time voice text and the warning text;
[0188] The display unit 701 is further configured to prompt the user to re-enter the voice when the similarity is less than a similarity threshold; and to execute a step of displaying the voice input progress in the warning text based on the real-time voice text when the similarity is greater than or equal to the similarity threshold.
[0189] Optionally, the display unit 701 is specifically configured to:
[0190] Matching the real-time voice text with the warning text, and locating the text position of the real-time voice text in the warning text;
[0191] The voice real-time text is displayed in the warning text based on the text position to show the voice input progress.
[0192] Optionally, the display unit 701 is specifically configured to:
[0193] Matching the real-time voice text with the warning text to determine a voice comparison text in the warning text that corresponds to the real-time voice text;
[0194] The voice control text is displayed on the page where the warning text is located.
[0195] Optionally, the storage unit 704 is further configured to:
[0196] Determining the speech characteristics of the target service based on the speech information;
[0197] The voice information and the speech spectrum features are stored in association with the target service.
[0198] Optionally, the display unit 701 is further configured to:
[0199] In response to a release operation triggered for a target service, displaying a release rejection interface for the target service; the release rejection interface includes a warning text;
[0200] Play the voice information saved in the target service association.
[0201] Optionally, the display unit 701 is further configured to:
[0202] In response to an input operation for the warning text, starting a timer;
[0203] If the voice input progress indicates that the voice information input is not completed and the timer duration exceeds the set time limit, the voice information is stopped from being received and an input timeout prompt interface is displayed. Corresponding to the above method embodiment, the embodiment of the present application also provides an electronic device.
[0204] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application; Figure 8 As shown, in the embodiment of the present application, the electronic device 80 includes: a processor 81, a display 82, a memory 83, an input device 86, a bus 85 and a communication device 84; the processor 81, the memory 83, the input device 86, the display 82 and the communication device 84 are all connected through the bus 85, and the bus 85 is used to transmit data between the processor 81, the memory 83, the display 82, the communication device 84 and the input device 86.
[0205] Among them, the memory 83 can be used to store software programs and modules, such as the program instructions / modules corresponding to the warning text processing method in the embodiment of the present application. The processor 81 executes various functional applications and data processing of the electronic device 80 by running the software programs and modules stored in the memory 83, such as the warning text processing method provided in the embodiment of the present application. The memory 83 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application application, etc.; the data storage area can store data created according to the use of the electronic device 80 (such as training samples, feature extraction networks, etc.). In addition, the memory 83 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0206] The processor 81 is the control center of the electronic device 80. It connects the various components of the electronic device 80 using a bus 85 and various interfaces and lines. It executes or runs software programs and / or modules stored in the memory 83 and calls data stored in the memory 83 to perform various functions of the electronic device 80 and process data. Optionally, the processor 81 may include one or more processing units, such as a CPU, a GPU (Graphics Processing Unit), a digital processing unit, etc.
[0207] In the embodiment of the present application, the processor 81 displays the segmented image to the user through the display 82.
[0208] The input device 86 is primarily used to obtain user input operations. Depending on the electronic device, the input device 86 may also vary. For example, if the electronic device is a computer, the input device 86 may be a mouse, keyboard, or other input device; if the electronic device is a portable device such as a smartphone or tablet, the input device 86 may be a touch screen.
[0209] An embodiment of the present application further provides a computer storage medium, in which computer executable instructions are stored. The computer executable instructions are used to implement the warning text processing method described in any embodiment of the present application.
[0210] In some possible implementations, various aspects of the warning text processing method provided in the present application may also be implemented in the form of a program product, which includes program code. When the program product is run on a computer device, the program code is used to enable the computer device to perform the steps of the warning text processing method according to various exemplary embodiments of the present application described above in this specification. For example, the computer device may perform the following steps: Figure 2 The warning text processing flow in steps S201 to S204 is shown.
[0211] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0212] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0213] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0214] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0215] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.
[0216] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A method for processing warning text, characterized in that: include: In response to an execution operation triggered for a target business in the business interface, displaying a warning text corresponding to the target business; receiving in real time at least one voice message based on the warning text input; After receiving each voice message, performing real-time voice recognition on the voice message to obtain corresponding real-time voice text; Determining the similarity between the real-time text of the voice and the warning text; when the similarity is less than a similarity threshold, prompting the user to re-enter the voice; When the similarity is greater than or equal to a similarity threshold, displaying the voice input progress in the warning text based on the real-time voice text; When the voice input progress indicates that the input of the at least one voice message is completed, determining the speech spectrum feature of the target service based on the at least one voice message; The at least one voice message and the speech spectrum feature are stored in association with the target service, and the speech spectrum feature is used as an identifier of the object and the target service, and is bound and stored with the identity information of the object and the warning text; In response to a release operation triggered for a target service, displaying a release rejection interface for the target service; the release rejection interface includes a warning text; Play at least one voice message stored in association with the target service.
2. The method according to claim 1, characterized in that The displaying of the voice input progress in the warning text based on the real-time voice text includes: Matching the real-time voice text with the warning text, and locating the text position of the real-time voice text in the warning text; The voice real-time text is displayed in the warning text based on the text position to show the voice input progress.
3. The method according to claim 1, characterized in that The displaying of the voice input progress in the warning text based on the real-time voice text includes: Matching the real-time voice text with the warning text to determine a voice comparison text in the warning text that corresponds to the real-time voice text; The voice control text is displayed on the page where the warning text is located.
4. The method according to any one of claims 1 to 3, characterized in that After displaying the warning text corresponding to the target service in response to the execution operation triggered for the target service in the service interface and before receiving in real time at least one voice message input based on the warning text, the method further includes: In response to an input operation for the warning text, starting a timer; After receiving at least one voice message based on the warning text input in real time, the method further includes: If the voice input progress indicates that the input of at least one voice message is not completed and the timing duration exceeds the set time limit, the reception of the voice message is stopped and an input timeout prompt interface is displayed.
5. A warning text processing device, characterized in that: include: A display unit, configured to display a warning text corresponding to a target business in response to an execution operation triggered for the target business in the business interface; An input unit, configured to receive in real time at least one voice message input based on the warning text; The recognition unit is configured to perform real-time speech recognition on each voice message after receiving the voice message to obtain a corresponding real-time speech text; The display unit is further configured to determine the similarity between the real-time voice text and the warning text; when the similarity is less than a similarity threshold, prompt the user to re-enter the voice; and when the similarity is greater than or equal to the similarity threshold, display the voice input progress in the warning text based on the real-time voice text. a storage unit, configured to determine, when the voice input progress indicates that the input of the at least one voice message is completed, a speech spectrum feature of the target service based on the at least one voice message; The at least one voice message and the speech spectrum feature are stored in association with the target service, and the speech spectrum feature is used as an identifier of the object and the target service, and is bound and stored with the identity information of the object and the warning text; The display unit is further configured to: in response to a release operation triggered for a target service, display a release rejection interface for the target service; the release rejection interface includes a warning text; and play at least one voice message associated with and saved in association with the target service.
6. An electronic device, characterized in that: include: a memory for storing executable instructions; A processor, configured to read and execute the executable instructions stored in the memory to implement the method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor, the processor is enabled to perform the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Recitation-assistant document display method and recitation-assistant document display system
CN101593438A
Smart terminal teller machine and application thereof
CN109859410A
Illegal account processing method and device, terminal, server and storage medium
CN112231666A