Document operation method, device, equipment and storage medium
By extracting text content in the document and performing voice playback and contactless command recognition, the problem of obtaining document content when manual operation is inconvenient, and the convenience and freedom of document operation are improved.
Patent Information
- Application Number
- CN201911158402.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-22
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2039-11-22
AI Technical Summary
When driving and other things are inconvenient to manual operation, users cannot obtain document content in time, resulting in poor convenience in document processing.
By extracting the text content of the document and temporarily saving it to a preset position, voice play and monitor contactless instructions, identifying and performing corresponding operations, including stopping, deleting, updating, annotating and rolling back, etc.
It realizes that document content can be obtained in a timely manner when manual operation cannot be operated, enriches document operation scenarios, and improves the freedom and convenience of document operation.
Smart Images

Figure CN112836548B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer application technology, and in particular to a document operation method, apparatus, device and storage medium. Background Art
[0002] Document processing is an essential part of today's work and life. Documents are often sent as attachments, requiring manual opening to access the document content. However, this manual operation significantly limits document processing scenarios. For example, when users are driving, or in situations where manual operation is inconvenient, they are unable to freely process documents, resulting in delayed access to document content and a lack of convenience. Summary of the Invention
[0003] The present application provides a method, device, system and storage medium for document operation.
[0004] An embodiment of the present application provides a document operation method, including: extracting text content from a document to be processed and temporarily storing the text content in a preset storage location; playing the text content by voice and monitoring contactless instructions for the text content; identifying the contactless instructions and performing corresponding operations on the text content corresponding to the contactless instructions.
[0005] An embodiment of the present application provides a document operation device, including: a temporary storage module, used to extract text content from a document to be processed and temporarily store the text content in a preset storage location; an instruction recognition module, used to play the text content in voice and monitor contactless instructions for the text content; a document operation module, used to recognize the contactless instructions and perform corresponding operations on the text content corresponding to the contactless instructions.
[0006] An embodiment of the present application provides a device, which includes: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the document operation method as described in any one of the embodiments of the present application.
[0007] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method for document operation as described in any one of the embodiments of the present application is implemented.
[0008] The embodiments of the present application provide a document operation method, apparatus, device, and storage medium. By temporarily storing the text content in the document to be processed, playing the temporarily stored text content by voice and obtaining corresponding contactless instructions, and operating the text content according to the contactless instructions, users who are unable to manually process and operate documents can obtain document content, thereby improving the freedom and convenience of document operations.
[0009] With respect to the above embodiments and other aspects of the present application and their implementation, further description is provided in the accompanying drawings, detailed description and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 A flowchart of the document operation method in the embodiment of the present application;
[0011] Figure 2 Another flow chart of the document operation method in the embodiment of the present application;
[0012] Figure 3 This is a structural diagram of a document operation device in an embodiment of the present application;
[0013] Figure 4 This is a structural diagram of the document operation device in an embodiment of the present application. DETAILED DESCRIPTION
[0014] To make the purpose, technical solutions and advantages of this application more clear, the embodiments of this application will be described in detail below with reference to the accompanying drawings. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of this application can be combined with each other in any way.
[0015] Figure 1 This is a flow chart of the document operation method in the embodiment of the present application, see Figure 1 The embodiments of the present application may be applicable to situations where manual document operation is inconvenient, such as while driving. The method may be performed by a document operation device in the embodiments of the present application. The device may be implemented in software and / or hardware and may generally be integrated into a smart terminal or an in-vehicle terminal. The method in the embodiments of the present application includes:
[0016] Step 101: extract text content from a document to be processed and temporarily store the text content in a preset storage location.
[0017] The document to be processed may be a document that needs to be operated on, and may include a text document obtained through email, instant messaging software, and text messages. The document to be processed may be located in the body of an email or in an attachment to an email. The document to be processed may be a text document in a format such as PDF, WORD, EXCEL, PPT, or TXT. Furthermore, the document to be processed may be a document in an image format or a compressed file format. The text content may be a sentence, paragraph, or fixed-length text in the text content of the document to be processed. The text content may be obtained by recognizing the document to be processed, and the recognition method may include image recognition, text recognition, and optical character recognition. The preset storage location may be a pre-set storage space for storing text content, and may specifically include memory, cache, auxiliary storage, and hard disk storage.
[0018] Specifically, the text content of the document to be processed can be read through text recognition or image recognition technology, and the text content can be stored in a pre-set memory. It is understood that the text content can be stored in the pre-set storage location in the form of storing the entire text content or first dividing the text content into multiple parts and storing them in the pre-set storage location. For example, in the pre-set storage location, each character of the text content is stored in a different memory address, or a segment of the text content can be stored in a content address.
[0019] Step 102: Play the text content in voice and monitor contactless instructions for the text content.
[0020] Among them, voice playback can be the output of text content in a preset temporary storage location in the form of voice; contactless instructions can be instructions for operating text content, and contactless instructions can mean that the user does not directly contact the smart terminal or vehicle-mounted terminal that executes the method of the embodiment of the present application. For example, contactless instructions can include voice operation instructions, gesture operation instructions, and action operation instructions, etc.
[0021] In an embodiment of the present application, the stored text content can be played by voice, and the text content can be played in the order of the memory addresses in the preset storage location. The text content with the smaller memory address value can be played by voice first. While the voice is playing, the contactless instructions can be monitored to obtain the contactless instructions for operating the text content. Specifically, the user can be monitored through a microphone or camera to obtain the user's words or facial expressions, and the user's words or gestures of stopping can be used as contactless instructions.
[0022] Step 103: Identify the contactless instruction and perform a corresponding operation on the text content corresponding to the contactless instruction.
[0023] Among them, recognition can be an operation of recognizing contactless instructions. The operations can specifically include stop, delete, update, annotation and rollback operations. Recognition can include image recognition or semantic recognition, and the corresponding operation can be determined according to the contactless instructions.
[0024] Specifically, the acquired contactless instructions can be recognized, and the operation content of the contactless instructions can be determined through voice recognition or image recognition. The acquired contactless instructions may include multiple ones, and the corresponding operation content can be identified one by one in the order in which the contactless instructions were acquired. After the operation content is acquired, the corresponding text content can be operated. The contactless instructions acquired during text content playback can be contactless instructions for operating the text content being played by voice. The text content can be stopped, deleted, updated, and rewound based on the operation content of the contactless instructions. For example, if a user says a contactless command such as "go back to the previous paragraph" or "go back two paragraphs", the user can ask a question in a question-and-answer format, such as "Do you want to go back two paragraphs?" to assist in the recognition of the contactless instruction. The user can recognize numbers in the voice instruction, such as keywords such as "one", "two", and "two", and can control the pointer in the preset storage location to retract a corresponding distance to find the text content in the corresponding position in the preset storage location, and then retract to the corresponding position for editing or voice playback of the text content.
[0025] The technical solution of the embodiment of the present application extracts text content from the document to be processed, temporarily stores the text content in a preset storage location, plays the text content by voice and obtains contactless instructions corresponding to the text content, identifies the contactless instructions and performs corresponding operations, thereby realizing contactless document operations when manual operations are inconvenient, obtaining document content in a timely manner, enriching the document operation scenarios, and improving the freedom and convenience of document operations.
[0026] Figure 2 This is another flow chart of the document operation method in the embodiment of the present application; the embodiment of the present application extracts sentence segments based on text symbols, see Figure 2 , the document operation method of the embodiment of the present application includes:
[0027] Step 201: Identify the document format of the document to be processed.
[0028] The document to be processed may be a document received by a user, may be in the form of an email attachment, and may be in a file format corresponding to the document to be processed, including formats such as TXT, WORD, EXCEL, and PPT.
[0029] Specifically, when monitoring document files in an intelligent terminal or vehicle-mounted terminal that implements the document operation method of the present application, when the document file is acquired by the monitored intelligent terminal or vehicle-mounted terminal, the document file can be used as a document to be processed. For example, the document file can be scanned in the memory or interface at regular intervals. The method of acquiring the document file can include wireless network transmission and Bluetooth transmission. After acquiring the document to be processed, the document format of the document to be processed can be acquired. For example, the storage format can be acquired by using regular matching in the attribute information of the document to be processed. Furthermore, when the target document is stored in compression, the document to be processed can be decompressed before acquiring the document format.
[0030] Step 202: Determine that the document format is not a preset format in a preset list, and delete the document to be processed.
[0031] Among them, the preset list can store a list of document formats that can perform document operations. When the document format of the target document exists in the preset list, the document operation method of this application can be implemented on the target document. It can be understood that the preset format can be the document format stored in the preset list, specifically common document formats, such as PDF, TXT, WORD and EXCEL.
[0032] Specifically, the document format of the document to be processed can be searched in the preset list. If the document format is a preset format in the preset list, it can be determined that the document to be processed can be processed. If the document format is not a preset format in the preset list, it can be determined that the document to be processed cannot be processed or the document to be processed is a protected document. The document operation method in the embodiment of the present application cannot be implemented on the document to be processed, and the document to be processed can be deleted. Exemplarily, a whitelist and a blacklist can be set in an intelligent terminal or a vehicle-mounted terminal that implements the embodiment of the present application. The user can choose to store the document formats that allow document operations in the whitelist and the document formats that prohibit document operations in the blacklist. When a target document of a confidential type is received, it will not be played by voice, thereby ensuring the confidentiality of the document operation.
[0033] Step 203: Segment the extracted text content into at least one stored text.
[0034] The stored text may be a minimum storage unit stored in a preset storage location, and the stored text may be in bytes or in text segments of natural text.
[0035] Specifically, the acquired text content may be divided by special words, characters or letters, and the text content may be divided into storage texts in bytes or text segments.
[0036] In one embodiment, the extracted text content is segmented into at least one storage text, including: segmenting the text content according to text symbols in the text content; and using the text content between two adjacent text symbols as the storage text.
[0037] The text symbols may include punctuation marks such as commas, semicolons, and periods in natural languages, and may also include special symbols used for text segmentation, such as semicolons, commas, and spaces in EXCEL files.
[0038] In an embodiment of the present application, text symbols can be used to divide text content into stored texts, text symbols in text content can be identified, and the text content between every two adjacent symbols can be used as a stored text. It can be understood that when there are no text symbols in the text content, the text content can be divided using a fixed step size or special characters to generate at least one stored text, which is convenient for subsequent processes to determine the operation object of the contactless instruction.
[0039] Step 204: Store each stored text in a preset storage location.
[0040] In an embodiment of the present application, an identification number can be determined for each stored text, and the identification number can be the memory address of the storage location of each stored text. It can be understood that the storage address of the stored text can be used as the memory address, or the address in the preset storage location after the stored text is stored can be used as the memory address. The memory address can be a manually set logical address, or it can be the actual physical address of the storage medium.
[0041] Step 205: Acquire the text content temporarily stored in the preset storage location according to the memory address, and play the text content in voice.
[0042] Among them, the memory address can be the storage address of the text content in a preset storage location, and the text content can correspond to at least one memory address. For example, each character in the text content can correspond to a different memory address, or different sentence segments in the text content can correspond to different memory addresses. It can be understood that when the position of the same character in the text content is different, the corresponding memory address may also be different.
[0043] Specifically, the text content is obtained from a preset storage location according to the memory address, and the obtained text content can be played as a voice.
[0044] Step 206: Monitor contactless instructions for the text content, wherein the contactless instructions include at least one of voice instructions, gesture instructions, and action instructions.
[0045] Specifically, contactless instructions may refer to instructions that can perform document operations without contacting the smart terminal or the vehicle-mounted terminal, and may include voice instructions, gesture instructions, and action instructions. In an embodiment of the present application, contactless instructions may be one or more of voice instructions, gesture instructions, and action instructions at the same time. Among them, voice instructions may be the voice spoken by the user to control the document, such as stop, return to the previous section, or continue playing, and the monitoring method of voice instructions may be achieved through a microphone or a voice sensor; gesture instructions may be the hand movements of the user to control the document, which may include waving or sliding in different directions, etc., and may be monitored and obtained through an image sensor; action instructions may be the body movements of the user to control the document, which may include head expressions and limb movements, etc., and may be obtained through an image sensor.
[0046] Step 207: Identify the contactless instruction to determine at least one operation among a stop operation, a delete operation, an update operation, an annotation operation, and a rollback operation.
[0047] In an embodiment of the present application, the contactless instruction can be identified, and different contactless instructions can be identified in different ways. If the contactless instruction is a voice instruction, the instruction operation can be determined by semantic recognition. If the contactless instruction is a gesture instruction or an action instruction, the instruction operation can be determined by image recognition. Different operations can operate on text content in different ways. Operations may include pause, delete, rollback, update, annotation and play, etc. The text content being operated can be the text content played by voice when the contactless instruction is obtained. It is understandable that multiple operations can be recognized at the same time, for example, a stop operation and a delete operation can be recognized at the same time, and a rollback operation and an annotation operation can also be recognized at the same time.
[0048] Step 2071: When it is determined that the operation is a stop operation, stop the voice playback of the text content.
[0049] Specifically, if the operation is determined to be a stop operation, the text content can be stopped from being retrieved from the preset storage location according to the memory address, and the voice playback of the text content can be stopped. For example, if a user voice instruction such as "temporarily" or "wait a moment" is detected during the voice broadcast process, the text content at the pause point can be identified and playback can be suspended; at this time, the pause pointer can point to the data stored in the memory address of the text content, that is, the pause pointer now points to the document content at the pause point.
[0050] Step 2072: When it is determined that the operation is a deletion operation, confirm the deletion object corresponding to the deletion operation, and delete the deletion object in the text content.
[0051] Specifically, when it is determined that the operation is a deletion operation, the deletion object can be confirmed according to the text content played by voice. For example, when playing the document content "server system components", the user can be asked whether to delete "server". When the user says "yes" or "right", it is confirmed that the deletion object corresponding to the deletion operation is "server", and the "server" stored at the storage address pointed to by the pointer can be deleted.
[0052] Step 2073: When it is determined that the operation is an update operation, confirm the update object and update data corresponding to the update operation, and use the update data to replace the update object in the text content.
[0053] In the embodiment of the present application, when it is determined that the operation is an update operation, the update object and update data corresponding to the update operation can be confirmed. The update object and update data of the user can be confirmed through multiple rounds of conversations. For example, an inquiry mode such as "Is the part you need to modify the component part" can be issued, and "component part" can be used as the update object of the update operation. The inquiry method "What is your replacement content" can be used to ask the user to input the update data. After determining the update object, the pointer can be pointed to the storage address of the update object, and the update data can be used to replace the data corresponding to the storage address.
[0054] Step 2074: When it is determined that the operation is a comment operation, confirm the comment object and comment data corresponding to the comment operation, and use the comment data to make a comment at the comment object in the text content.
[0055] Specifically, when it is determined that the operation is a comment operation, the comment object and comment data corresponding to the comment operation can be confirmed. The comment object and comment data of the user can be confirmed through multiple rounds of conversations. For example, an inquiry mode such as "Is the patent application the one you need to comment on" can be issued, and "patent application" can be used as the comment object of the comment operation. The inquiry method "What is your comment content" can be used to ask the user to input the comment data. After determining the comment object, the pointer can be pointed to the storage address of the comment object, and the comment data can be stored at the storage address as a comment on the text content. Further, the comment data can be associated and stored with a comment identification number for identifying the comment data.
[0056] Step 2075: When it is determined that the operation is a rollback operation, confirm the rollback reference object and rollback distance corresponding to the rollback operation, search for the rollback object at the rollback distance with the rollback reference object as the reference in the text content, and start voice playback or operation from the rollback object in the text content.
[0057] Specifically, to determine whether the operation is a backoff operation, the backoff reference object and backoff distance can be identified through semantic recognition or image recognition. For example, when the text content voice plays "server system component", the memory address of "server system component" can be used as the backoff reference object, and the user voice such as "backoff to the previous paragraph" or "backoff to the previous two paragraphs" can be monitored. The backoff operation is recognized in the form of a repeated question and answer "Do you want to backoff to the previous two paragraphs?", and then the numbers in the user voice command, such as keywords such as "one", "two", and "two" are read as the backoff distance. The pointer in the memory is controlled to backoff the corresponding length of the backoff distance based on the reference backoff object, find the corresponding storage location in the memory, and then backoff to the corresponding location for editing. Further, the backoff distance can be automatically generated. For example, it can be generated based on text symbols in the text content according to words such as "one", "two", or "two". The memory address distance of the text content before the reference backoff object can be grouped as a backoff distance. It is understandable that the backoff distance can vary according to the length of the text content, and the backoff distance of different sentence segments can be different.
[0058] The technical solution of the embodiment of the present application identifies the document format of the document to be processed, and when it is determined that the document format is not a preset format in a preset list, the document to be processed is deleted, the text content is divided into at least one storage text and stored in a preset storage location, the text content stored in the preset storage location is obtained according to the memory address and played, the contactless instructions for operating the text content are monitored, the operation of the contactless instructions is identified, and the text content is operated to realize contactless document operation, and the document is operated when manual operation is not possible, which enriches the document operation scenarios, improves the freedom of document operation, and enhances the convenience of user document operation.
[0059] In one embodiment, confirming the deletion object corresponding to the deletion operation and deleting the deletion object in the text content includes: obtaining the deletion object input by the user through at least one method of semantic recognition, image recognition and gesture recognition; obtaining the storage location of the deletion object in the preset storage location according to the memory address; and deleting the deletion object in the storage location.
[0060] In an embodiment of the present application, the user's deletion object can be obtained through semantic recognition, image recognition, or gesture recognition. For example, if the user says "delete previous sentence", the previous sentence of the text content currently being played can be obtained as the deletion object. The stored data can be searched in a preset storage location based on the memory address of the deletion object and deleted. Furthermore, after the deletion object is deleted, the storage address of the text content following the deletion object can be decremented.
[0061] Furthermore, the text content temporarily stored in a preset storage location can be converted into a common document format, such as Word or PDF, and the document name can be set by voice, and then stored persistently. If the document needs to be sent, an application such as email can be called to send the stored document.
[0062] Figure 3 This is a schematic diagram of the structure of a document operation device in an embodiment of the present application. The document operation device provided in this embodiment of the present application can execute the document operation method provided in any embodiment of the present application, and has the corresponding modules and beneficial effects of executing the method. The device can be implemented by software and / or hardware, and specifically includes: a temporary storage module 301, an instruction recognition module 302, and a document operation module 303.
[0063] The temporary storage module 301 is used to extract text content from the document to be processed and temporarily store the text content in a preset storage location.
[0064] The instruction recognition module 302 is used to play the text content by voice and monitor contactless instructions for the text content.
[0065] The document operation module 303 is configured to identify the contactless instruction and perform corresponding operations on the text content corresponding to the contactless instruction.
[0066] The technical solution of the embodiment of the present application is to extract text content from the document to be processed through the temporary storage module 301 and temporarily store the text content in a preset storage location. The instruction recognition module 302 plays the text content in voice and obtains contactless instructions corresponding to the text content. The document operation module 303 recognizes the contactless instructions and performs corresponding operations, thereby realizing contactless document operations when manual operations are inconvenient, obtaining document content in a timely manner, enriching the document operation scenarios, and improving the freedom and convenience of document operations.
[0067] In one embodiment, the document operation device may further include a pre-processing module specifically configured to: identify the document format of the document to be processed; and if it is determined that the document format is not a preset format in a preset list, delete the document to be processed.
[0068] In one embodiment, the temporary storage module 301 includes:
[0069] The content extraction unit is used to divide the extracted text content into at least one stored text.
[0070] The storage unit is used to store each of the stored texts in a preset storage location.
[0071] In one embodiment, the content extraction unit includes:
[0072] The segmentation subunit is used to segment the text content according to the text symbols in the text content.
[0073] The extraction subunit is used to store the text content between two adjacent text symbols as text.
[0074] In one embodiment, the instruction recognition module 302 includes:
[0075] The voice playing unit is used to obtain the text content temporarily stored in the preset storage location according to the memory address, and play the text content in voice.
[0076] An instruction monitoring unit is used to monitor non-contact instructions for the text content, wherein the non-contact instructions include at least one of voice instructions, gesture instructions and action instructions.
[0077] In one embodiment, the document operation module 303 includes:
[0078] The operation recognition unit is used to recognize the contactless instruction to determine at least one operation among a stop operation, a delete operation, an update operation, an annotation operation and a rollback operation.
[0079] An operation execution unit, configured to execute at least one of the following corresponding operations on the text content corresponding to the contactless instruction:
[0080] When it is determined that the operation is a stop operation, stopping the voice playback of the text content;
[0081] When it is determined that the operation is a deletion operation, confirming the deletion object corresponding to the deletion operation, and deleting the deletion object in the text content;
[0082] When it is determined that the operation is an update operation, confirming the update object and update data corresponding to the update operation, and replacing the update object in the text content with the update data;
[0083] When it is determined that the operation is an annotation operation, confirming the annotation object and annotation data corresponding to the annotation operation, and annotating the annotation object in the text content using the annotation data;
[0084] When it is determined that the operation is a backoff operation, the backoff reference object and backoff distance corresponding to the backoff operation are confirmed, and the backoff object at the backoff distance is searched for within the text content based on the backoff reference object, and voice playback or operation is started from the backoff object of the text content.
[0085] In one embodiment, the operation execution unit is specifically used to obtain the deletion object input by the user through at least one method of semantic recognition, image recognition and gesture recognition; obtain the storage location of the deletion object in the preset storage location according to the memory address; and delete the deletion object in the storage location.
[0086] Figure 4 This is a schematic diagram of the structure of the document operation device in the embodiment of the present application. Figure 4 As shown, the device includes a processor 40, a memory 41, an input device 42 and an output device 43; the number of processors 40 in the device can be one or more. Figure 4 In the figure, a processor 40 is used as an example; the device processor 40, memory 41, input device 42 and output device 43 can be connected by a bus or other means. Figure 4 The bus connection is taken as an example.
[0087] The memory 41, as a computer-readable storage medium, can be used to store software programs, computer executable programs, and modules, such as the modules corresponding to the document operation device in the embodiment of the present application (temporary storage module 301, instruction recognition module 302, and document operation module 303). The processor 40 executes the software programs, instructions, and modules stored in the memory 41 to execute various functional applications and data processing of the device, thereby implementing the above-mentioned document operation method.
[0088] The memory 41 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created based on the use of the terminal, etc. In addition, the memory 41 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 41 may further include a memory remotely located relative to the processor 40, and these remote memories may be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0089] The input device 42 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The output device 43 may include a display device such as a display screen.
[0090] Exemplarily, the device in the embodiment of the present application may include a display, a processing device (such as a central processing unit, etc.), a read-only memory, an input / output (I / O) interface, etc., wherein the input / output interface may include a touch screen, a touchpad, a camera, a speaker, and a microphone. The device in the embodiment of the present application also includes the following modules: a document monitoring module, a voice monitoring module, a control module, a text recognition module, a text editing module, a dump module, a voice recognition module, a voice-to-text conversion module, a voice playback module, a document output module, etc. The device in the embodiment of the present application can implement a document operation method, and the specific steps can be as follows:
[0091] (1) The document monitoring module monitors the terminal device and can set an interface on the wireless network side to monitor and identify the received data. It can also set a timer to scan the memory to obtain newly added document data; the document source can be file information obtained through wireless network transmission, Bluetooth transmission, etc.
[0092] (2) After detecting that the terminal device has received a new document, the document monitoring module sends a request to the control module, and the control module sends a first instruction to the text recognition module.
[0093] (3) After receiving the instruction, the text recognition module identifies the document format. If it is a compressed file, it will first be decompressed and then recognized. If it is not a compressed file, the received document will be directly opened with the terminal's built-in reader. The format or attribute of the received document can be judged, and a whitelist can be set. If the document format is determined to be a common format such as PDF, Word or TXT, it will be added to the whitelist and the third-party software in the terminal will be called to open the document. If the format is not recognized as a common document format or the document format cannot be recognized, the received document will be added to the blacklist and deleted.
[0094] (4) After the text recognition module starts to start the document content, it immediately sends a start identifier to the control module, and the control module sends a second instruction to the dump module, calling the dump module to store the recognized text content; the dump module converts the recognized text content into characters and then transfers it to the local memory. During the storage process, the document content is marked, specifically: the punctuation marks in the text are recognized, and the content between each two adjacent symbols is stored as a section and marked accordingly. Different marks are marked for different contents, and all marks in the document are not repeated; the mark serves as an index identifier for this section of content; the transfer process can be set to store synchronously with the text recognition, so as to operate in parallel. The processing speed can be improved; the storage can also be performed after all recognition is completed. The above steps can be provided to the user for setting options in the operation interface of the terminal device; there is no restriction on the form or format of storage, and it can be stored according to the preset format and storage location; after the storage is completed, the storage end mark is sent to the control module, and the control module sends the third instruction to the voice-to-text conversion module, and the voice-to-text conversion module voice broadcasts the recognized text content; at the same time, the control module synchronously starts the voice monitoring module. If you do not choose to pause with a command during the voice playback process, and do not interact with the smart device, then after the playback is completed, you can choose to request to replay or delete directly with a command, and the process ends. If the voice monitoring module detects the user's voice during the voice broadcast process, such as "tentative" or "wait a moment", the pause request is sent to the control module. After receiving the pause request, the control module sends a fourth instruction to the voice playback module to pause, and identifies the content at the pause point, matches it with the marked sections of content in the memory, and locates it to the corresponding position in the memory; the positioning step is specifically as follows: if the user's voice is detected during the voice broadcast process, such as "tentative" or "wait a moment", the content at the pause point is identified, and the pointer points to the address where the index identifier of this section of content is stored, that is, the pointer points to the document content at the pause point; after receiving instructions such as "modify" and "delete", the content of the corresponding storage position identified by the index identifier pointed to by the pointer is modified, overwritten or directly deleted, and then the pointer points to the storage address of the index identifier of the next section.
[0095] (5) If the voice monitoring module detects voice commands such as "modify", "add", or "delete", it converts the relevant commands into requests and sends them to the control module.
[0096] (6) After receiving the request, the control module sends a fifth instruction to the voice-to-text conversion module, and issues a confirmation question to the user in the form of voice playback.
[0097] (7) During the confirmation process of multiple rounds of inquiries, if the playback content needs to be rolled back, the voice monitoring module detects the content of "roll back to ***" and returns the monitored content to the control module. After the control module recognizes the text content, it locates the corresponding marked position in the memory.
[0098] (8) After multiple rounds of inquiries are finally confirmed, an inquiry end mark is sent to the control module. After receiving the inquiry end mark, the control module sends a sixth instruction to the text editing module. The text editing module performs corresponding editing operations such as adding, deleting or modifying the text content at the corresponding position in the memory according to the mark, and sends an editing completion mark to the control module after the modification is completed.
[0099] (9) After receiving the editing completion mark, the control module sends the seventh instruction to the voice playback module. After receiving the instruction, the voice playback module plays the message "Editing completed" to the user.
[0100] (10) After the playback is completed to the user, the control module sends an eighth instruction to the voice playback module. The instruction carries the marked position of the edited text. The voice playback module continues to play from this marked position and repeats the above steps (4)-(9) until the document playback is completed.
[0101] (11) After the playback is completed, a playback completion mark is sent to the control module. After receiving the above mark, the control module sends an instruction to the voice playback module, sends a prompt message to the user that "the entire document has been played" and asks the user whether to "save" or "send".
[0102] (12) The voice monitoring module sends the monitored user instructions to the control module. After receiving the instructions, the control module performs the corresponding operation. If saving is required, the edited document is saved locally and converted into a common document format, such as word or pdf. A certain document name is set by voice and then saved in a fixed location. If the document needs to be sent, the control module sends the ninth instruction to the document output module. After receiving the instruction, the document output module calls applications such as email to output the document.
[0103] (13) After the above steps are completed, the voice monitoring module is turned off to prevent erroneous reception of information, and the process ends.
[0104] An embodiment of the present application further provides a storage medium containing computer-executable instructions. When the computer-executable instructions are executed by a computer processor, the computer-executable instructions are used to perform a document operation method, the method comprising:
[0105] Extract the text content from the document to be processed and temporarily store the text content in a preset storage location; play the text content by voice and listen for contactless instructions for the text content; identify the contactless instructions and perform corresponding operations on the text content corresponding to the contactless instructions. Of course, the computer-executable instructions of the storage medium containing computer-executable instructions provided in the embodiment of the present application are not limited to the operations of the method described above, but can also be used to execute the document operation method provided in any embodiment of the present application.
[0106] Through the above description of the implementation methods, those skilled in the art can clearly understand that the present application can be implemented with the help of software and necessary general-purpose hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as a computer's floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0107] It is worth noting that in the embodiment of the above-mentioned document operation device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application.
[0108] The above description is merely an exemplary embodiment of the present application and is not intended to limit the scope of protection of the present application.
[0109] It will be appreciated by those skilled in the art that the term user terminal covers any suitable type of wireless user equipment, such as a mobile phone, a portable data processing device, a portable web browser or a vehicle-mounted mobile station.
[0110] In general, various embodiments of the present application may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device, although the present application is not limited thereto.
[0111] Embodiments of the present application may be implemented by executing computer program instructions by a data processor of a mobile device, for example, in a processor entity, or by hardware, or by a combination of software and hardware. The computer program instructions may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages.
[0112] The block diagram of any logical flow in the accompanying drawings of the present application can represent program steps, or can represent interconnected logical circuits, modules and functions, or can represent a combination of program steps and logical circuits, modules and functions. The computer program can be stored on a memory. The memory can have any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as but not limited to read-only memory (ROM), random access memory (RAM), optical memory device and system (digital versatile disc DVD or CD optical disc) etc. Computer-readable media can include non-transient storage media. The data processor can be any type suitable for the local technical environment, such as but not limited to a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a programmable logic device (FPGA) and a processor based on a multi-core processor architecture.
[0113] The above description of exemplary embodiments of the present application has been provided by way of exemplary and non-limiting examples. However, various modifications and adaptations to the above embodiments will be apparent to those skilled in the art, when considered in conjunction with the accompanying drawings and claims, without departing from the scope of the present invention. Therefore, the proper scope of the present invention will be determined by reference to the claims.
Claims
1. A document operation method, characterized in that: include: Extracting text content from a document to be processed and temporarily storing the text content in a preset storage location; Playing the text content by voice and monitoring contactless instructions for the text content; Identifying the contactless instruction and performing a corresponding operation on the text content corresponding to the contactless instruction; The identifying the contactless instruction and performing a corresponding operation on the text content corresponding to the contactless instruction includes: Identifying the contactless instruction to determine at least one of a stop operation, a delete operation, and an update operation; Perform at least one of the following corresponding operations on the text content corresponding to the contactless instruction: When it is determined that the operation is a stop operation, stopping the voice playback of the text content; When it is determined that the operation is a deletion operation, confirming the deletion object corresponding to the deletion operation, and deleting the deletion object in the text content; When it is determined that the operation is an update operation, the update object and update data corresponding to the update operation are confirmed, and the update object in the text content is replaced with the update data.
2. The method according to claim 1, characterized in that Before extracting the text content in the document to be processed, it also includes: Identify the document format of the document to be processed; It is determined that the document format is not a preset format in the preset list, and the document to be processed is deleted.
3. The method according to claim 1, characterized in that The step of extracting text content from a document to be processed and temporarily storing the text content in a preset storage location includes: Segmenting the extracted text content into at least one stored text; Each of the stored texts is stored in a preset storage location.
4. The method according to claim 3, characterized in that The step of dividing the extracted text content into at least one stored text comprises: Segmenting the text content according to text symbols in the text content; The text content between two adjacent text symbols is used as stored text.
5. The method according to claim 1, characterized in that The voice plays the text content and monitors the contactless instructions for the text content, including: Acquire the text content temporarily stored in the preset storage location according to the memory address, and play the text content in voice; Monitoring contactless instructions for the text content, wherein the contactless instructions include at least one of voice instructions, gesture instructions, and action instructions.
6. The method according to claim 1, characterized in that The identifying the contactless instruction and performing a corresponding operation on the text content corresponding to the contactless instruction further includes: Recognizing the contactless instruction to determine at least one of an annotation operation and a backoff operation; Perform at least one of the following corresponding operations on the text content corresponding to the contactless instruction: When it is determined that the operation is an annotation operation, confirming the annotation object and annotation data corresponding to the annotation operation, and annotating the annotation object in the text content using the annotation data; When it is determined that the operation is a backoff operation, the backoff reference object and backoff distance corresponding to the backoff operation are confirmed, and the backoff object at the backoff distance is searched for within the text content based on the backoff reference object, and voice playback or operation is started from the backoff object of the text content.
7. The method according to claim 6, characterized in that The confirming the deletion object corresponding to the deletion operation and deleting the deletion object in the text content includes: Acquire the deletion object input by the user through at least one method of semantic recognition, image recognition, and gesture recognition; Acquire the storage location of the deleted object in the preset storage location according to the memory address; Delete the deletion object in the storage location.
8. A document operation device, characterized in that: include: A temporary storage module is used to extract text content from a document to be processed and temporarily store the text content in a preset storage location; A command recognition module, configured to play the text content by voice and monitor contactless commands for the text content; a document operation module, configured to identify the contactless instruction and perform a corresponding operation on the text content corresponding to the contactless instruction; The document operation module is further used to: Identifying the contactless instruction to determine at least one of a stop operation, a delete operation, and an update operation; Perform at least one of the following corresponding operations on the text content corresponding to the contactless instruction: When it is determined that the operation is a stop operation, stopping the voice playback of the text content; When it is determined that the operation is a deletion operation, confirming the deletion object corresponding to the deletion operation, and deleting the deletion object in the text content; When it is determined that the operation is an update operation, the update object and update data corresponding to the update operation are confirmed, and the update object in the text content is replaced with the update data.
9. A device, characterized in that The device comprises: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the document operation method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the document operation method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
E-book annotation adding method, electronic equipment and computer storage medium
CN109634501A