Text generation method, device, electronic device and storage medium
By generating replacement information, the problem of inaccurate display of digital formats in voice input is solved, user-friendly digital format modification is achieved, and the accuracy and user experience of input text are improved.
Patent Information
- Application Number
- CN202210148464.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-17
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-02-17
AI Technical Summary
During the voice input process, the format display of digital information is difficult to meet the user's personalized preferences, resulting in inaccurate identification and display, and the need for correct information transmission is not met.
Provide a text generation method, by generating replacement information to replace digital information in voice input with digital information in other formats, users can choose a suitable replacement method to realize quick modification of digital format.
It improves the accuracy of displaying digital information in voice input scenarios, meets users' personalized needs, reduces user modification costs, and improves the accuracy of input text.
Smart Images

Figure CN114518805B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, in particular to fields such as voice technology and intelligent search, and specifically to a text generation method, device, electronic device, and storage medium. Background Art
[0002] An input method refers to a coding method used to enter various symbols into a computer or other electronic device. The coding method for Chinese character input associates the sound, shape, and meaning of characters with specific keys, and then combines them according to the character to complete the input. With the development of computer technology, voice input has also evolved into a method for Chinese character input. Voice input is a simple and easy-to-use input method that uses the operator's speech to convert the computer into Chinese characters. Summary of the Invention
[0003] The present disclosure provides a text generation method, device, electronic device, and storage medium.
[0004] According to one aspect of the present disclosure, a text generation method is provided, comprising: in response to receiving voice information, generating initial text information corresponding to the voice information; in response to detecting that the initial text information includes first digital information, generating at least one replacement information, the at least one replacement information being used to represent that the first digital information is replaced with second digital information in another format; and in response to receiving a selection operation for target replacement information in the at least one replacement information, replacing the first digital information in the initial text information with target digital information in the second digital information corresponding to the target replacement information, to obtain target text information corresponding to the voice information.
[0005] According to another aspect of the present disclosure, a text generation device is provided, comprising: a first generation module for generating initial text information corresponding to the voice information in response to receiving voice information; a second generation module for generating at least one replacement information in response to detecting that the initial text information includes first digital information, the at least one replacement information being used to represent that the first digital information is replaced with second digital information in another format; and a first replacement module for replacing the first digital information in the initial text information with target digital information corresponding to the target replacement information in the second digital information in response to receiving a selection operation for target replacement information in the at least one replacement information, thereby obtaining target text information corresponding to the voice information.
[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the text generation method of the present disclosure.
[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the text generation method of the present disclosure.
[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the text generation method of the present disclosure when executed by a processor.
[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0011] Figure 1 Schematically illustrates an exemplary system architecture to which the text generation method and apparatus according to an embodiment of the present disclosure can be applied;
[0012] Figure 2 The flowchart of the text generation method according to the embodiment of the present disclosure is schematically shown;
[0013] Figure 3 A schematic diagram schematically illustrates a text generation method according to an embodiment of the present disclosure;
[0014] Figure 4 Schematically illustrates a schematic diagram of generating at least one replacement information according to an embodiment of the present disclosure;
[0015] Figure 5 A schematic diagram schematically illustrates a text generation method according to another embodiment of the present disclosure;
[0016] Figure 6 A block diagram schematically illustrates a text generation device according to an embodiment of the present disclosure; and
[0017] Figure 7 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0018] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0019] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.
[0020] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0021] With the development of voice technology and the improvement of voice input accuracy, more and more users are choosing voice input when entering text, especially when two hands are not convenient. Voice input can provide great convenience for users. When voice input includes numerical information, each input method mainly determines the display format of Arabic numerals, Chinese characters, or a specific format after recognizing the number by establishing various rules. For example, when a long string of numbers is detected, it will be adjusted to Arabic numerals for display.
[0022] In the process of realizing the concept of the present disclosure, the inventors discovered that displaying numbers in a suitable format is an unavoidable difficulty in the voice input process. The difficulty includes, for example: the combination of numbers and scenarios are changeable, the display strategy is difficult to cover all situations, and when the same digital information is processed with multiple rules, unreasonable display results may be produced. Users may have personalized preferences, and there is no absolute standard for the display format of numbers. As the scenario changes, users may prefer different display formats for the same digital information. During the voice input process, users have different habits of oral input, which makes it more difficult to correctly identify and display numbers. Some scenarios cannot be judged based on voice. Therefore, the display results cannot always meet the needs of users, and numbers are often key information, and their correct expression is crucial to the correct communication of information.
[0023] The present disclosure provides a text generation method, apparatus, electronic device, and storage medium. The text generation method includes: in response to receiving voice information, generating initial text information corresponding to the voice information; in response to detecting that the initial text information includes first digital information, generating at least one replacement information, the at least one replacement information being used to represent the replacement of the first digital information with second digital information in another format; and in response to receiving a selection operation for target replacement information in the at least one replacement information, replacing the first digital information in the initial text information with target digital information in the second digital information corresponding to the target replacement information, thereby obtaining target text information corresponding to the voice information.
[0024] Figure 1 An exemplary system architecture to which the text generation method and apparatus according to an embodiment of the present disclosure can be applied is schematically shown.
[0025] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure. This does not mean that the embodiments of the present disclosure cannot be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the text generation method and apparatus may be applied may include a terminal device, but the terminal device may implement the text generation method and apparatus provided by the embodiments of the present disclosure without interacting with a server.
[0026] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0027] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).
[0028] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0029] The server 105 may be a server that provides various services, such as a background management server that provides support for the content browsed by users using the terminal devices 101, 102, and 103 (for example only). The background management server may analyze and process the received user requests and other data, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server for a distributed system, or a server combined with a blockchain.
[0030] It should be noted that the text generation method provided in the embodiment of the present disclosure can generally be executed by the terminal device 101, 102, or 103. Accordingly, the text generation apparatus provided in the embodiment of the present disclosure can also be provided in the terminal device 101, 102, or 103.
[0031] Alternatively, the text generation method provided in the embodiment of the present disclosure may also be generally executed by the server 105. Accordingly, the text generation apparatus provided in the embodiment of the present disclosure may generally be provided in the server 105. The text generation method provided in the embodiment of the present disclosure may also be performed by a server or server cluster that is different from the server 105 and that is capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the text generation apparatus provided in the embodiment of the present disclosure may also be provided in a server or server cluster that is different from the server 105 and that is capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.
[0032] For example, when a text message needs to be generated, the terminal devices 101, 102, and 103 can obtain voice information and then send the obtained voice information to the server 105. In response to receiving the voice information, the server 105 generates initial text information corresponding to the voice information. In response to detecting that the initial text information includes first digital information, the server 105 generates at least one replacement information, the at least one replacement information being used to represent the replacement of the first digital information with second digital information in another format. In response to receiving a selection operation for target replacement information in the at least one replacement information, the server replaces the first digital information in the initial text information with target digital information corresponding to the target replacement information in the second digital information, thereby obtaining target text information corresponding to the voice information. Alternatively, a server or server cluster capable of communicating with the terminal devices 101, 102, 103 and / or the server 105 analyzes the voice information and obtains the target text information corresponding to the voice information.
[0033] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0034] Figure 2 The flowchart of the text generation method according to the embodiment of the present disclosure is schematically shown.
[0035] like Figure 2 As shown, the method includes operations S210 to S230.
[0036] In operation S210, in response to receiving voice information, initial text information corresponding to the voice information is generated.
[0037] In operation S220, in response to detecting that the initial text information includes the first digital information, at least one replacement information is generated, where the at least one replacement information is used to represent that the first digital information is replaced with second digital information in another format.
[0038] In operation S230, in response to receiving a selection operation for target replacement information in at least one replacement information, the first digital information in the initial text information is replaced with target digital information corresponding to the target replacement information in the second digital information to obtain target text information corresponding to the voice information.
[0039] According to an embodiment of the present disclosure, the first digital information can represent various digital-related information in the initial text information. Other formats may include at least one of Chinese character format, Arabic numeral format, and Roman numeral format. The replacement information may include at least one of replacement information representing the target digital information to be replaced with the Chinese character format, replacement information representing the target digital information to be replaced with the Arabic numeral format, and replacement information representing the target digital information to be replaced with the Roman numeral format. The second digital information can represent digital information in other formats that are different from the digital format of the first digital information, and the number of the second digital information can be the same as the number of the replacement information, that is, each replacement information corresponds to a second digital information that is the same as the voice information of the first digital information. The above-mentioned target text information can be determined based on the initial text information after the digital format is replaced.
[0040] Figure 3 The following schematically shows a text generation method according to an embodiment of the present disclosure.
[0041] like Figure 3As shown, clicking the microphone 310 can trigger the voice input page 300 and start detecting voice information. The detected voice information can be converted into corresponding initial text information in real time and displayed in the text input box 320. When it is detected that the initial text information includes first digital information, replacement information 330 and 340 can be generated, but it is not limited to this. The replacement information 330 can, for example, represent that the first digital information is replaced with numbers in Chinese format for display. The replacement information 340 can, for example, represent that the first digital information is replaced with numbers in Arabic numeral format for display.
[0042] According to the embodiments of the present disclosure, see Figure 3 As shown, for example, based on the voice information, initial text message 321 "I teach Class 2, Grade 3, with approximately 10 to 20 students" can be obtained, and initial text message 321 can be determined to include first numerical information such as "32" and "10 to 20." By clicking on replacement message 330, the user can replace the first numerical information in initial text message 321 with Chinese characters, resulting in target text message 321, for example, "I teach Class 2, Grade 3, with approximately 10 to 20 students." By clicking on replacement message 340, the user can replace the first numerical information in initial text message 321 with Arabic numerals, resulting in target text message 321, for example, "I teach Class 32, with approximately 10 to 20 students."
[0043] Through the above-described embodiments of the present disclosure, a solution has been implemented for recognizing and displaying numbers in voice input scenarios, allowing users to quickly modify the format of numbers. Based on this solution, when number-related content is recognized, at least one replacement message is generated, providing a quick replacement option, reducing the user's cost of modifying the number and providing text information that better meets actual application needs.
[0044] In conjunction with specific embodiments, Figure 2 The method shown is further explained.
[0045] According to an embodiment of the present disclosure, in response to detecting that the initial text information includes the first numeric information, generating at least one replacement information may include: generating intermediate information in response to detecting that the initial text information includes the first numeric information; and generating at least one replacement information in response to receiving a selection operation on the intermediate information.
[0046] According to an embodiment of the present disclosure, the intermediate information may include, for example, a pop-up window indicating that the digital format can be replaced. Selecting the pop-up window may indicate that the first digital information in the initial text message is allowed to be replaced later.
[0047] Figure 4 The diagram schematically shows a schematic diagram of generating at least one replacement information according to an embodiment of the present disclosure.
[0048] like Figure 4 As shown, clicking microphone 410 can trigger voice input page 400, start detecting voice information, and convert the detected voice information into corresponding initial text information in real time, which is displayed in text input box 420, such as initial text information 421. When it is detected that the initial text information 421 includes the first digital information, a pop-up window 430 for obtaining replacement information 431, 432, etc. can be first generated. Pop-up window 430 can indicate that a format replacement operation will be performed on the first digital information in initial text information 421. When the user clicks on pop-up window 430, replacement information 431, 432 for implementing the format replacement can be generated, for example, but is not limited to this.
[0049] Through the above-mentioned embodiments of the present disclosure, a method of obtaining at least one replacement information based on the intermediate information can add a step of confirming the operation to be replaced before selecting the replacement information, which can effectively reduce the situation where the first digital information is incorrectly replaced due to the incorrect selection of a certain replacement information, thereby improving the user experience.
[0050] According to an embodiment of the present disclosure, the text generation method may further include: in response to detecting that a cursor has moved to a position corresponding to third numerical information in the initial text information, generating at least one presentation information; the voice information corresponding to each presentation information in the at least one presentation information is the same as the voice information corresponding to the third numerical information; and in response to receiving a selection operation for a target presentation information, replacing the third numerical information with the target presentation information.
[0051] According to an embodiment of the present disclosure, the third digital information may represent at least one of each number-related information in the initial text information and other information in a predefined format. For example, the other information in the predefined format may include information in a format such as X time X moment. The position corresponding to the third digital information may include at least one of: a position before the first number in the third digital information where a cursor can be placed; a position between two adjacent numbers in the third digital information where a cursor can be placed; and a position after the last number in the third digital information where a cursor can be placed, but is not limited to these.
[0052] According to an embodiment of the present disclosure, display information may include at least one of time-based display information, mass-based display information, and formula-based display information. For example, time-based display information may include at least one of the following formats: XX:XX, X hour, and X minute. For example, mass-based display information may include at least one of the following formats: XX grams, XX kg. For example, formula-based display information may include various calculation formats, such as X×Y=Z.
[0053] According to an embodiment of the present disclosure, third digital information and display information with the same voice information can be pre-stored in a database as matching pairs. Each matching item in the matching pair can be digital information with a preset structure, such as XX:XX, X hour X minute, XX grams, etc., to match the numbers in actual applications. For example, there is no phonetic difference between "working 30 days" and "working 31 days", or "three quarter past three" and "3.1 grams".
[0054] According to an embodiment of the present disclosure, the display information may also be display information with the same semantics as the semantic information corresponding to the third digital information. The third digital information and the display information with the same semantics may be pre-stored in a database as matching pairs. Each matching item in the matching pair may be digital information with a preset structure, such as XX:XX, X o'clock X o'clock, X.30 o'clock, etc., to match the numbers in actual applications. For example, "8:30", "8:30", and "8:30" have the same semantics.
[0055] According to an embodiment of the present disclosure, the cursor can locate the third digital information of the corresponding structure. When the user moves the cursor near a certain third digital information, at least one display information corresponding to the third digital information can be generated based on the matching items related to the third digital information stored in the database.
[0056] For example, when a user inputs "liu yi" by voice, they may mean the "Six One" in Children's Day, or they may actually mean "61" due to their habit of expression. If the initial text information displays "Six One", the display information of "61" can be generated. Based on the positioning of the cursor, it is also possible to replace Arabic and Chinese characters, such as replacing "20" in the initial text information with "Twenty", and to replace ambiguous numerical information, such as replacing "3:15" with "3.1 grams".
[0057] Figure 5 The following schematically shows a text generation method according to another embodiment of the present disclosure.
[0058] like Figure 5 As shown, clicking the microphone 510 can trigger the voice input page 500, start detecting voice information, and convert the detected voice information into corresponding initial text information in real time, which is displayed in the text input box 520. When it is detected that the initial text information includes the first digital information, replacement information 531 and 532 can be generated.
[0059] According to the embodiments of the present disclosure, see Figure 5As shown, for example, based on the voice information, initial text message 521 "3.1 grams in ten minutes." When the user moves cursor 540 between "grams" and "read" in initial text message 521, the third numeric information "3.1 grams" can be located, and display information 541 and 542 can be generated accordingly. For example, the user can select display information 541 to replace "3.1 grams" in initial text message 521 with "three quarters." The resulting target text information can be, for example, "three quarters in ten minutes."
[0060] Through the above-mentioned embodiments of the present disclosure, specific digital information in the initial text information can be located, and corresponding display information can be generated, providing a quick replacement option to realize the modification of specific digital information, reducing the problem of inaccurate replacement results that may be caused by one-click replacement, and improving the user experience.
[0061] Figure 6 The block diagram of the text generation device according to an embodiment of the present disclosure is schematically shown.
[0062] like Figure 6 As shown, the text generating apparatus 600 includes a first generating module 610 , a second generating module 620 and a replacing module 630 .
[0063] The first generating module 610 is configured to generate initial text information corresponding to the voice information in response to receiving the voice information.
[0064] The second generating module 620 is configured to generate at least one replacement information in response to detecting that the initial text information includes the first digital information. The at least one replacement information is used to represent that the first digital information is replaced with second digital information in another format.
[0065] The first replacement module 630 is used to replace the first digital information in the initial text information with the target digital information corresponding to the target replacement information in the second digital information in response to receiving a selection operation for the target replacement information in at least one replacement information, so as to obtain the target text information corresponding to the voice information.
[0066] According to an embodiment of the present disclosure, the text generating apparatus further includes a third generating module and a second replacing module.
[0067] The third generating module is configured to generate at least one presentation information in response to detecting that the cursor moves to a position corresponding to the third numeric information in the initial text information. The voice information corresponding to each presentation information in the at least one presentation information is the same as the voice information corresponding to the third numeric information.
[0068] The second replacing module is configured to replace the third digital information with the target display information in response to receiving a selection operation on the target display information.
[0069] According to an embodiment of the present disclosure, the second generation module includes a first generation unit and a second generation unit.
[0070] The first generating unit is configured to generate intermediate information in response to detecting that the initial text information includes the first digital information.
[0071] The second generating unit is configured to generate at least one replacement information in response to receiving a selection operation on the intermediate information.
[0072] According to an embodiment of the present disclosure, the replacement information may include at least one of the following: replacement information representing that the first digital information is replaced with target digital information in Chinese character format, replacement information representing that the first digital information is replaced with target digital information in Arabic numeral format, and replacement information representing that the first digital information is replaced with target digital information in Roman numeral format.
[0073] According to an embodiment of the present disclosure, the presentation information may include at least one of the following: time-based presentation information, quality-based presentation information, and formula-based presentation information.
[0074] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0075] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the text generation method of the present disclosure.
[0076] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the text generation method of the present disclosure.
[0077] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the text generation method of the present disclosure is implemented.
[0078] Figure 7A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0079] like Figure 7 As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0080] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0081] The computing unit 701 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the text generation method. For example, in some embodiments, the text generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the text generation method described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the text generation method by any other suitable means (e.g., via firmware).
[0082] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0083] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0084] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0085] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0086] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0087] A computer system may include clients and servers. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0088] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0089] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A text generation method, comprising: In response to receiving the voice information, generating initial text information corresponding to the voice information; In response to detecting that the initial text information includes first digital information, generating at least one replacement information, the at least one replacement information being used to represent that the first digital information is replaced with second digital information in another format; as well as In response to receiving a selection operation for target replacement information in the at least one replacement information, replacing the first digital information in the initial text information with target digital information corresponding to the target replacement information in the second digital information, to obtain target text information corresponding to the voice information; The method further comprises: In response to detecting that the cursor moves to a position corresponding to third digital information in the initial text information, generating at least one display information corresponding to the third digital information based on matching items related to the third digital information stored in a database, the display information including at least one of the following: time-based display information, quality-based display information, and formula-based display information; The third digital information and the display information with the same voice information are pre-stored in the database in the form of matching pairs, and each matching item representing the display information in the matching pair includes digital information of a preset structure of a different type.
2. The method according to claim 1, further comprising: In response to receiving a selection operation on target presentation information, the third digital information is replaced with the target presentation information.
3. The method according to claim 1, wherein In response to detecting that the initial text information includes the first digital information, generating at least one replacement information comprises: generating intermediate information in response to detecting that the initial text information includes first digital information; and In response to receiving a selection operation on the intermediate information, the at least one replacement information is generated.
4. The method according to any one of claims 1 to 3, wherein: The replacement information includes at least one of the following: replacement information representing that the first digital information is replaced with target digital information in Chinese character format, replacement information indicating that the first digital information is replaced with target digital information in Arabic numeral format; as well as The replacement information represents that the first digital information is replaced with target digital information in Roman numeral format.
5. A text generation device comprising: A first generating module, configured to generate initial text information corresponding to the voice information in response to receiving the voice information; a second generating module, configured to generate at least one replacement information in response to detecting that the initial text information includes the first digital information, wherein the at least one replacement information is used to represent that the first digital information is replaced with second digital information in another format; as well as a first replacing module configured to, in response to receiving a selection operation for target replacement information in the at least one replacement information, replace the first digital information in the initial text information with target digital information corresponding to the target replacement information in the second digital information, thereby obtaining target text information corresponding to the voice information; The device further comprises: a third generating module configured to, in response to detecting that the cursor moves to a position corresponding to the third digital information in the initial text information, generate at least one display information corresponding to the third digital information based on matching items related to the third digital information stored in a database, the display information including at least one of the following: time-based display information, quality-based display information, and formula-based display information; The third digital information and the display information with the same voice information are pre-stored in the database in the form of matching pairs, and each matching item representing the display information in the matching pair includes digital information of a preset structure of a different type.
6. The apparatus according to claim 5, further comprising: The second replacing module is configured to replace the third digital information with the target display information in response to receiving a selection operation on the target display information.
7. The device according to claim 5, wherein The second generation module includes: a first generating unit, configured to generate intermediate information in response to detecting that the initial text information includes first digital information; and The second generating unit is configured to generate the at least one replacement information in response to receiving a selection operation on the intermediate information.
8. The device according to any one of claims 5 to 7, wherein: The replacement information includes at least one of the following: replacement information representing that the first digital information is replaced with target digital information in Chinese character format, replacement information indicating that the first digital information is replaced with target digital information in Arabic numeral format; as well as The replacement information represents that the first digital information is replaced with target digital information in Roman numeral format.
9. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 4.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.
11. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Character input error correction method and device
CN103473003A
Speech conversion error correction method and device
CN107679032A