Working note generation method and device
By collecting and analyzing the operation information of the smart terminal and using OCR and NLP technologies to generate work notes, the problem of low efficiency in identifying and classifying operation behaviors in the existing technology is solved, and valuable work notes are automatically generated, and the efficiency of asset accumulation is improved.
Patent Information
- Application Number
- CN202110290101.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-18
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-03-18
AI Technical Summary
The existing technology is difficult to effectively identify and classify the operation behavior of smart terminals, resulting in low efficiency in generating work notes and ineffective summary of work assets.
By collecting operation information and operation time of the smart terminal, using OCR and NLP technology to analyze and identify screenshot image information and video playback information, generate text image information and text video information, and use pre-established recognition models to integrate this information to generate valuable work notes.
It realizes accurate identification and classification of smart terminal operation behaviors, automatically generates valuable work notes, reduces labor costs, and improves asset accumulation efficiency.
Smart Images

Figure CN113033536B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information processing technology, and in particular to a work note generating method and device. Background Art
[0002] At present, employees of enterprises record work items and summarize work content every day, which is convenient for understanding their work status in the short term and can form work assets in the long term. However, few people can keep recording work notes. The main reasons are: 1. Daily work is already very busy, and recording work notes requires extra time, and some details are difficult to recall afterwards; 2. Some work items have a large time span, and simple diary notes cannot reflect the complete situation of the items, and it is necessary to integrate records of multiple days; 3. Traditional notes are highly personalized and casual, which is not conducive to sharing with other team members. After recording, they will not be read again, and the overall value is low. Nowadays, almost everyone's work is carried out on smart terminals, such as computers, and many valuable assets are hidden in the computer operation behavior records. Smart terminal operation behavior refers to the interaction between people and smart terminals, including keyboard and mouse events, browsing web pages, sending emails, file editing, file transfer, file printing, copying and pasting, etc. Recording personal smart terminal operation behavior records, mining the value in them, automatically producing personal work notes and summarizing assets has become an urgent need for modern workers.
[0003] Therefore, how to intelligently identify and classify smart terminal operation behaviors and use artificial intelligence technology to automatically generate work notes is a technical problem that needs to be solved urgently. Summary of the invention
[0004] In view of the problems existing in the prior art, the main purpose of the embodiments of the present invention is to provide a work note generation method and device, so as to realize the recognition and classification of smart terminal operation behaviors, automatically generate valuable work notes, and thus help individuals and teams to easily summarize assets.
[0005] In order to achieve the above object, an embodiment of the present invention provides a work note generation method, the method comprising:
[0006] Collect multiple pieces of operation information of the intelligent terminal and the operation time corresponding to each piece of operation information; wherein the operation information includes screenshot image information, text processing information and video playback information;
[0007] Analyze and identify the screenshot image information and the video playback information to obtain text-type image information and text-type video information corresponding to the screenshot image information and the video playback information respectively;
[0008] Work notes are generated based on the text processing information, the text-type image information and the text-type video information, as well as the operation time corresponding to the text processing information, the text-type image information and the text-type video information, using a pre-established recognition model.
[0009] Optionally, in one embodiment of the present invention, the text processing information includes user input information, page browsing information, text copying information and file processing information.
[0010] Optionally, in an embodiment of the present invention, the parsing and identifying the screenshot image information and the video playback information to obtain text-type image information and text-type video information corresponding to the screenshot image information and the video playback information respectively includes:
[0011] Performing image recognition on the screenshot image information to obtain text-type image information corresponding to the screenshot image information;
[0012] Perform voice recognition on the video playback information to obtain text-type video information corresponding to the video playback information.
[0013] Optionally, in an embodiment of the present invention, generating work notes using a pre-established recognition model according to the text processing information, the text-based image information, the text-based video information, and the operation time corresponding to the text processing information, the text-based image information, and the text-based video information includes:
[0014] Inputting the text processing information, text image information and text video information into a pre-established recognition model to determine labels corresponding to the text processing information, text image information and text video information respectively;
[0015] Work notes are generated according to the labels and operation times respectively corresponding to the text processing information, the text image information and the text video information.
[0016] Optionally, in one embodiment of the present invention, the recognition model is pre-established in the following manner:
[0017] Acquire multiple text-based historical operation information of the smart terminal and labels corresponding to each piece of text-based historical operation information, and pre-process the text-based historical operation information data;
[0018] The textual historical operation information and corresponding labels that have undergone data preprocessing are used as training sample data to train a preset initial natural language processing model to obtain the recognition model.
[0019] An embodiment of the present invention further provides a work note generating device, the device comprising:
[0020] An information collection module, used to collect multiple pieces of operation information of the intelligent terminal and the operation time corresponding to each piece of operation information; wherein the operation information includes screenshot image information, text processing information and video playback information;
[0021] An information identification module, used to parse and identify the screenshot image information and the video playback information, and obtain text-type image information and text-type video information corresponding to the screenshot image information and the video playback information respectively;
[0022] The note generation module is used to generate work notes based on the text processing information, the text-type image information and the text-type video information, as well as the operation time corresponding to the text processing information, the text-type image information and the text-type video information, using a pre-established recognition model.
[0023] Optionally, in one embodiment of the present invention, the text processing information includes user input information, page browsing information, text copying information and file processing information.
[0024] Optionally, in an embodiment of the present invention, the information identification module includes:
[0025] An image recognition unit, used to perform image recognition on the screenshot image information to obtain text image information corresponding to the screenshot image information;
[0026] The speech recognition unit is used to perform speech recognition on the video playback information to obtain text-type video information corresponding to the video playback information.
[0027] Optionally, in one embodiment of the present invention, the note generation module includes:
[0028] A label generation unit, used for inputting the text processing information, text image information and text video information into a pre-established recognition model to determine labels corresponding to the text processing information, text image information and text video information respectively;
[0029] The note generating unit is used to generate work notes according to the labels and operation times respectively corresponding to the text processing information, the text image information and the text video information.
[0030] Optionally, in one embodiment of the present invention, the device also includes a model building module, which is used to obtain multiple text-based historical operation information of the smart terminal and labels corresponding to each piece of text-based historical operation information, and pre-process the text-based historical operation information data; the pre-processed text-based historical operation information and the corresponding labels are used as training sample data to train a preset initial natural language processing model to obtain the recognition model.
[0031] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above method when executing the program.
[0032] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program for executing the above method.
[0033] The present invention identifies and classifies the operation behaviors of smart terminals, accurately records various matters, integrates various matters, automatically generates valuable work notes, reduces labor costs, and accurately counts time, thereby helping individuals and teams to easily summarize assets. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0035] Figure 1 A flowchart of a method for generating work notes according to an embodiment of the present invention;
[0036] Figure 2 A flowchart of information identification in an embodiment of the present invention;
[0037] Figure 3 A flowchart of work note generation in an embodiment of the present invention;
[0038] Figure 4 A flow chart of establishing a recognition model in an embodiment of the present invention;
[0039] Figure 5 This is a schematic diagram of the structure of a system for applying a work note generation method in an embodiment of the present invention;
[0040] Figure 6 A flowchart of a system for applying a work note generation method in an embodiment of the present invention;
[0041] Figure 7 This is a schematic diagram of the structure of a work note generating device according to an embodiment of the present invention;
[0042] Figure 8 This is a schematic diagram of the structure of an information identification module in an embodiment of the present invention;
[0043] Fig. 9 This is a schematic diagram of the structure of a note generation module in an embodiment of the present invention;
[0044] Fig.10 It is a structural schematic diagram of a work note generating device in a specific embodiment of the present invention;
[0045] Fig.11 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0046] The embodiments of the present invention provide a work note generation method and device, which can be used in the financial field or other fields. It should be noted that the work note generation method and device of the present invention can be used in the financial field, and can also be used in any field other than the financial field. The application field of the work note generation method and device of the present invention is not limited.
[0047] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0048] like Figure 1 The flowchart of a method for generating work notes according to an embodiment of the present invention is shown. The execution subject of the method for generating work notes according to an embodiment of the present invention includes but is not limited to a computer. The method shown in the figure includes:
[0049] Step S1, collecting multiple pieces of operation information of the smart terminal and the operation time corresponding to each piece of operation information; wherein the operation information includes screenshot image information, text processing information and video playback information.
[0050] Among them, the smart terminal can be a computer, and multiple operation information of the smart terminal and the corresponding operation time are collected. The operation information is a record of the operation behavior of the smart terminal, specifically including screenshot image information, text processing information and video playback information. Specifically, the screenshot image information can be collected by timed screenshots, and the text processing information includes user input information, page browsing information, text copy information and file processing information. The video playback information can be recorded when the smart terminal plays a video, thereby collecting the video playback information. In addition, when collecting each operation information, the operation time corresponding to the operation information is also collected. The operation time can be a moment or a time period, such as the screenshot operation time when the screenshot image information is collected, or the playback time period corresponding to the video playback time.
[0051] Furthermore, the operation information can be obtained in the following ways.
[0052] 1. When entering text (keyboard, etc.), obtain the user input information;
[0053] 2. When you click a file (with a mouse, etc.), get the file name information, that is, collect file processing information. You can use the existing tool gtName.exe to get the file name to the clipboard when you click a file with a mouse;
[0054] 3. When copying and pasting, get the text or file content in the clipboard, that is, collect text copy information. You can use the existing python tool win32clipboard to read and save the clipboard text every preset time, such as 0.2 seconds, and compare the clipboard content with the previous content. If there is any change, save it;
[0055] 4. When browsing web pages, use a web crawler to obtain the content of web page browsing, that is, collect page browsing information. For example, use the existing tool Diffbot web crawler, each time you enter a new URL or click a network connection, automatically obtain the URL in the IE address bar, and automatically pass the URL to the Diffbot crawler tool, you can get the text content of the current web page and save it.
[0056] 5. When processing files (saving, transferring or printing files), obtain the file content, i.e., collect file processing information;
[0057] 6. Timing screenshots and image saving, that is, collecting screenshot image information, and saving the time of the above events and related software at the same time. The existing windows function GetForegroundWindow can be used to obtain the window currently operated by the user, and the software currently used by the smart terminal can be obtained.
[0058] Specifically, the collected and counted intelligent terminal operation behavior records are shown in Table 1. The operation time includes the start time point, time consumption and / or end time point of the operation behavior.
[0059] Table 1
[0060]
[0061]
[0062] Step S2, analyzing and identifying the screenshot image information and the video playback information to obtain text-type image information and text-type video information corresponding to the screenshot image information and the video playback information respectively.
[0063] Among them, OCR recognition is performed on the screenshot image information, and OCR (optical character recognition) text recognition uses the tesseract-OCR engine to identify the text information of the screenshot image information of the smart terminal, and obtain the text image information corresponding to the screenshot image information. In addition, voice recognition is performed on the video playback information to obtain the text video information corresponding to the video playback information. The obtained text information is summarized and saved in the local record to form a complete text operation information record for subsequent operation behavior recognition and classification.
[0064] Specifically, the text image information obtained by recognizing the screenshot image information is shown in Table 2.
[0065] Table 2
[0066]
[0067] Step S3, generating work notes using a pre-established recognition model according to the text processing information, the text-type image information and the text-type video information, as well as the operation time corresponding to the text processing information, the text-type image information and the text-type video information.
[0068] Among them, the recognition model can be pre-established using NLP natural language processing technology. Input text processing information, text image information and text video information into the recognition model, and output the labels corresponding to each information. Labeling of text processing information, text image information and text video information is implemented. Specifically, the labels include meeting, training, email, etc. In this way, each piece of operation information is labeled using the recognition model to determine which type of operation each piece of operation information belongs to. Specifically, the labels are stored in a preset label expert library, as shown in Table 3, which is a preset label expert library.
[0069] Table 3
[0070] Tag Expert Database Meeting Training Document editing programming mail ......
[0071] Furthermore, the labels of each piece of operation information and the corresponding operation time are integrated to generate a work note. Specifically, each piece of operation information in the work note can be sorted according to a preset sort, operation time proportion or label. The work note can be shown in Table 4, where the status of the matter can be filled in by the user.
[0072] Table 4
[0073]
[0074] As an embodiment of the present invention, the word processing information includes user input information, page browsing information, word copying information and file processing information.
[0075] Among them, user input information can be obtained by inputting text (keyboard, etc.), and page browsing information can be obtained by using a web crawler to obtain the content of the web page when browsing the web page. In addition, text copy information can be obtained by obtaining the text or file content in the clipboard when copying and pasting, and file processing information can be obtained by obtaining file content when processing files (file saving, file transfer or file printing).
[0076] As an embodiment of the present invention, Figure 2 As shown, the screenshot image information and the video playback information are analyzed and identified, and the text image information and text video information corresponding to the screenshot image information and the video playback information are obtained, including:
[0077] Step S21, performing image recognition on the screenshot image information to obtain text image information corresponding to the screenshot image information.
[0078] Step S22, performing voice recognition on the video playback information to obtain text-type video information corresponding to the video playback information.
[0079] Among them, OCR recognition is performed on the screenshot image information, and OCR (optical character recognition) text recognition uses the tesseract-OCR engine to identify the text information of the screenshot image information of the smart terminal, and obtain the text image information corresponding to the screenshot image information. In addition, voice recognition is performed on the video playback information to obtain the text video information corresponding to the video playback information. The obtained text information is summarized and saved in the local record to form a complete text operation information record for subsequent operation behavior recognition and classification.
[0080] As an embodiment of the present invention, Figure 3 As shown, according to the text processing information, the text image information and the text video information, and the operation time corresponding to the text processing information, the text image information and the text video information, using a pre-established recognition model, generating a work note includes:
[0081] Step S31, inputting the text processing information, text image information and text video information into a pre-established recognition model to determine the labels corresponding to the text processing information, text image information and text video information respectively.
[0082] Among them, the recognition model can be pre-established using NLP natural language processing technology. The text processing information, text image information and text video information are input into the recognition model, and the output is the label corresponding to each information. The text processing information, text image information and text video information are labeled. Specifically, the labels include meeting, training, email, etc. In this way, each piece of operation information is labeled using the recognition model to determine which operation type each piece of operation information belongs to.
[0083] Step S32, generating work notes according to the labels and operation times corresponding to the text processing information, the text image information and the text video information respectively.
[0084] The labels of each piece of operation information and the corresponding operation time are integrated to generate a work note. Specifically, each piece of operation information in the work note can be sorted according to a preset sort, operation time proportion or label.
[0085] As an embodiment of the present invention, Figure 4 As shown, the recognition model is pre-established in the following way:
[0086] Step S41, obtaining multiple text-based historical operation information of the smart terminal and labels corresponding to each piece of text-based historical operation information, and pre-processing the text-based historical operation information data.
[0087] The text-based operation behavior record data of the smart terminal is obtained from the local smart terminal or the database, that is, multiple text-based historical operation information, corresponding historical operation time and corresponding labels are obtained. In addition, after obtaining the text-based historical operation information, the text-based historical operation information can be manually labeled to obtain the label corresponding to the text-based historical operation information. The multiple text-based historical operation information, the corresponding historical operation time and the corresponding label are used as training sample data.
[0088] Furthermore, the text-based historical operation information data is preprocessed. Specifically, the sample training data is corrected, the characteristic values of the sample data are extracted, and feature dimension reduction, feature null value processing, and target value conversion processing are performed.
[0089] Step S42, using the pre-processed textual historical operation information and corresponding labels as training sample data, training the preset initial natural language processing model to obtain the recognition model.
[0090] Among them, natural language processing technology is used to make a semantic analysis model, specifically, the FastText model. Among them, FastText is the existing mainstream natural language processing training model. Specifically, the model training process includes inputting the training sample data into the initial FastText model, and according to the training results, outputting the probability of each text-type historical operation information belonging to different labels, thereby completing the pre-establishment of the recognition model.
[0091] In a specific embodiment of the present invention, Figure 5 FIG. 1 is a schematic diagram of the structure of a system for applying a work note generation method in an embodiment of the present invention. The work flow diagram of the system is as follows: Figure 6 As shown, specifically including:
[0092] Step 1: The system collects the operation behaviors of smart terminals in the background, mainly including text records, screenshots, etc., and also records the event time and the software used at the time.
[0093] Step 2: The system uses OCR image recognition technology to convert the screenshot image into text content.
[0094] Step 3: The system uses NLP natural language processing technology to record text operations, and after model training, summarizes the work item element information based on the model recognition results.
[0095] Step 4: The system selects the top priority items each day and generates work notes based on historical items.
[0096] Step 5: Users evaluate and score the generated work notes for use in iterative model training and tuning.
[0097] In this embodiment, if Figure 5 The system shown specifically includes: intelligent terminal operation behavior recording module 1, screenshot image processing module 2, operation behavior recognition and classification module 3, and work note generation module 4. The system is located between the user's intelligent terminal and the team asset repository. Without the need for human intervention, it analyzes and processes the intelligent terminal operation records through OCR image recognition, NLP natural language processing and other technologies, and then achieves the purpose of automatically generating work notes.
[0098] Operation behavior recording module 1: used to record the user's smart terminal operation behavior, which can be obtained in the following ways:
[0099] When inputting text (keyboard, etc.), obtain user input information; when clicking a file (mouse, etc.), obtain file name information, that is, collect file processing information. With the help of the existing tool gtName.exe, when the mouse clicks a file, obtain the file name to the clipboard; when copying and pasting, obtain the text or file content in the clipboard, that is, collect text copy information. The existing python tool win32clipboard can be used to read and save the clipboard text every preset time, such as 0.2 seconds, and compare the clipboard content with the previous content. If there is a change, save it; when browsing the web page, use a web crawler to obtain the content of the web page browsing, that is, collect page browsing information. For example, using the existing tool Diffbot web crawler, each time a new URL is entered or a network connection is clicked, the URL in the IE address bar is automatically obtained, and the URL is automatically passed to the Diffbot crawler tool, so that the text content of the current web page can be obtained and saved. When processing files (file saving, file transfer or file printing), obtain the file content, that is, collect file processing information; regularly capture and save images, that is, collect screenshot image information, and simultaneously save the time and related software of the above events. The window currently being operated by the user can be obtained through the existing Windows function GetForegroundWindow, and the software currently used by the smart terminal can be obtained.
[0100] Screenshot image processing module 2: For the OCR recognition of smart terminal screenshot images, the OCR text recognition module uses the tesseract-OCR engine to recognize the text information of the screenshot image and summarize it into the saved record of the operation behavior recording module 1 to form a complete text-based operation behavior record for subsequent smart terminal operation behavior recognition and classification.
[0101] Operation behavior recognition and classification module 3: For text-based operation behavior records, NLP natural language processing technology is mainly used. After model training, the recognized operation information is obtained according to the model recognition results. The text-based operation behavior records are used as model training sample data, which specifically includes the following steps:
[0102] (a) Sample data: extract text-based operation behavior record data;
[0103] (b) Data preprocessing: process and correct sample data, extract sample data feature values, perform feature dimension reduction, feature null value processing, and target value conversion processing. At the same time, label the sample data. For example: January 29th, attend distributed technology training in the second conference room, split into the time January 29th, the location of the second conference room, and the event of attending distributed technology training; match the event content with the label expert database, and the label is "Event type: training, other keywords: distributed technology";
[0104] (c) Model selection and training: For model training, the present invention uses natural language processing technology for semantic analysis. FastText is the mainstream natural language processing training model, so the FastText model is selected. A training sample data is input into the FastText model, and the probability of different labels belonging to the training sample data is output according to the training results.
[0105] (d) Model evaluation: Use the trained model to predict samples in the validation set or test set to determine whether the trained model meets the requirements. If not, perform iterative optimization.
[0106] Generate work notes module 4: For the items that have been processed by natural language, the same items are combined and described, and the frequency and time span of the same items are counted. If the frequency is high and the time span is long, it will be set as a priority item. The system selects items with a preset number of priorities, such as TOP10, generates work notes according to a fixed format template, and summarizes them in the team asset repository.
[0107] Since employees are busy with their daily work, it is difficult to record work notes in a timely and easy manner. Based on computer operation records, the present invention uses OCR image recognition and NLP natural language processing technology to automatically generate work notes and summarize assets. Its advantages are as follows: users can get a work note without manual operation, reducing labor costs and increasing asset accumulation. The operation records are relatively complete, important matters will not be missed, and time statistics are accurate. The note format is standard, which is conducive to team integration and sharing.
[0108] like Figure 7 FIG. 1 is a schematic diagram of a structure of a work note generating device according to an embodiment of the present invention. The device shown in the figure includes:
[0109] The information collection module 10 is used to collect multiple pieces of operation information of the intelligent terminal and the operation time corresponding to each piece of operation information; wherein the operation information includes screenshot image information, text processing information and video playback information.
[0110] Among them, the smart terminal can be a computer, and multiple operation information of the smart terminal and the corresponding operation time are collected. The operation information is a record of the operation behavior of the smart terminal, specifically including screenshot image information, text processing information and video playback information. Specifically, the screenshot image information can be collected by timed screenshots, and the text processing information includes user input information, page browsing information, text copy information and file processing information. The video playback information can be recorded when the smart terminal plays a video, thereby collecting the video playback information. In addition, when collecting each operation information, the operation time corresponding to the operation information is also collected. The operation time can be a moment or a time period, such as the screenshot operation time when the screenshot image information is collected, or the playback time period corresponding to the video playback time.
[0111] The information identification module 20 is used to analyze and identify the screenshot image information and the video playback information to obtain text-type image information and text-type video information corresponding to the screenshot image information and the video playback information respectively.
[0112] Among them, OCR recognition is performed on the screenshot image information, and OCR (optical character recognition) text recognition uses the tesseract-OCR engine to identify the text information of the screenshot image information of the smart terminal, and obtain the text image information corresponding to the screenshot image information. In addition, voice recognition is performed on the video playback information to obtain the text video information corresponding to the video playback information. The obtained text information is summarized and saved in the local record to form a complete text operation information record for subsequent operation behavior recognition and classification.
[0113] The note generation module 30 is used to generate work notes based on the text processing information, the text-type image information and the text-type video information, as well as the operation time corresponding to the text processing information, the text-type image information and the text-type video information, using a pre-established recognition model.
[0114] Among them, the recognition model can be pre-established using NLP natural language processing technology. The text processing information, text image information and text video information are input into the recognition model, and the output is the label corresponding to each information. The text processing information, text image information and text video information are labeled. Specifically, the labels include meeting, training, email, etc. In this way, each piece of operation information is labeled using the recognition model to determine which operation type each piece of operation information belongs to.
[0115] Furthermore, the labels of each piece of operation information and the corresponding operation time are integrated to generate a work note. Specifically, each piece of operation information in the work note can be sorted according to a preset sort, operation time proportion or label.
[0116] As an embodiment of the present invention, the word processing information includes user input information, page browsing information, word copying information and file processing information.
[0117] As an embodiment of the present invention, Figure 8 As shown, the information identification module includes:
[0118] An image recognition unit 21 is used to perform image recognition on the screenshot image information to obtain text image information corresponding to the screenshot image information;
[0119] The voice recognition unit 22 is used to perform voice recognition on the video playback information to obtain text-type video information corresponding to the video playback information.
[0120] As an embodiment of the present invention, Fig. 9 As shown, the note generation module includes:
[0121] The label generation unit 31 is used to input the text processing information, the text image information and the text video information into a pre-established recognition model to determine labels corresponding to the text processing information, the text image information and the text video information respectively;
[0122] The note generating unit 32 is used to generate work notes according to the labels and operation times respectively corresponding to the text processing information, the text image information and the text video information.
[0123] As an embodiment of the present invention, Fig.10 As shown, the device also includes a model building module 40, which is used to obtain multiple text-based historical operation information of the smart terminal and the labels corresponding to each text-based historical operation information, and pre-process the text-based historical operation information data; the text-based historical operation information and the corresponding labels that have been pre-processed are used as training sample data to train the preset initial natural language processing model to obtain the recognition model.
[0124] Based on the same application concept as the above-mentioned work note generation method, the present invention also provides the above-mentioned work note generation device. Since the principle of solving the problem by the work note generation device is similar to that of the work note generation method, the implementation of the work note generation device can refer to the implementation of the work note generation method, and the repeated parts will not be repeated.
[0125] The present invention identifies and classifies the operation behaviors of smart terminals, accurately records various matters, integrates various matters, automatically generates valuable work notes, reduces labor costs, and accurately counts time, thereby helping individuals and teams to easily summarize assets.
[0126] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above method when executing the program.
[0127] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program for executing the above method.
[0128] like Fig.11As shown, the electronic device 600 may further include: a communication module 110, an input unit 120, an audio processing unit 130, a display 160, and a power supply 170. It is worth noting that the electronic device 600 does not necessarily have to include Fig.11 In addition, the electronic device 600 may also include Fig.11 For components not shown, reference may be made to the prior art.
[0129] like Fig.11 As shown, the central processor 100 is sometimes also referred to as a controller or an operation control, and may include a microprocessor or other processor devices and / or logic devices. The central processor 100 receives inputs and controls the operations of various components of the electronic device 600.
[0130] The memory 140 may be, for example, one or more of a cache, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory or other suitable devices. The above-mentioned information related to the failure may be stored, and a program for executing the relevant information may also be stored. The CPU 100 may execute the program stored in the memory 140 to implement information storage or processing.
[0131] The input unit 120 provides input to the CPU 100. The input unit 120 is, for example, a key or a touch input device. The power supply 170 is used to provide power to the electronic device 600. The display 160 is used to display display objects such as images and text. The display may be, for example, an LCD display, but is not limited thereto.
[0132] The memory 140 may be a solid-state memory, such as a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It may also be a memory that saves information even when the power is off, can be selectively erased, and is provided with more data, examples of which are sometimes referred to as EPROMs, etc. The memory 140 may also be some other type of device. The memory 140 includes a buffer memory 141 (sometimes referred to as a buffer). The memory 140 may include an application / function storage unit 142, which is used to store application programs and function programs or processes for executing the operation of the electronic device 600 through the central processor 100.
[0133] The memory 140 may also include a data storage unit 143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 144 of the memory 140 may include various drivers for communication functions of the electronic device and / or for executing other functions of the electronic device (such as messaging applications, address book applications, etc.).
[0134] The communication module 110 is a transmitter / receiver 110 that transmits and receives signals via an antenna 111. The communication module (transmitter / receiver) 110 is coupled to the central processor 100 to provide input signals and receive output signals, which may be the same as the case of a conventional mobile communication terminal.
[0135] Based on different communication technologies, multiple communication modules 110 may be provided in the same electronic device, such as a cellular network module, a Bluetooth module and / or a wireless LAN module. The communication module (transmitter / receiver) 110 is also coupled to a speaker 131 and a microphone 132 via an audio processor 130 to provide an audio output via the speaker 131 and receive an audio input from the microphone 132, thereby realizing a common telecommunication function. The audio processor 130 may include any suitable buffer, decoder, amplifier, etc. In addition, the audio processor 130 is also coupled to the central processor 100, so that the sound can be recorded on the local machine through the microphone 132, and the sound stored on the local machine can be played through the speaker 131.
[0136] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0137] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0138] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0139] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0140] The present invention uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A method for generating work notes, characterized in that: The method comprises: Collect multiple pieces of operation information of the intelligent terminal and the operation time corresponding to each piece of operation information; wherein the operation information includes screenshot image information, text processing information and video playback information; Analyze and identify the screenshot image information and the video playback information to obtain text-type image information and text-type video information corresponding to the screenshot image information and the video playback information respectively; Generate work notes using a pre-established recognition model according to the text processing information, the text image information, the text video information, and the operation time corresponding to the text processing information, the text image information, and the text video information; The text processing information includes user input information, page browsing information, text copy information and file processing information, wherein the user input information is obtained by inputting text on the keyboard, the page browsing information is obtained by using a web crawler to obtain the content of the web page when browsing the web page, the text copy information is obtained by obtaining the text or file content in the clipboard when copying and pasting, and the file processing information is obtained by obtaining the file content when saving, transferring or printing the file; The analyzing and identifying the screenshot image information and the video playback information to obtain text-type image information and text-type video information corresponding to the screenshot image information and the video playback information respectively includes: Performing image recognition on the screenshot image information to obtain text-type image information corresponding to the screenshot image information; Perform voice recognition on the video playback information to obtain text-type video information corresponding to the video playback information.
2. The method according to claim 1, characterized in that The generating of the work notes according to the text processing information, the text image information, the text video information, and the operation time corresponding to the text processing information, the text image information, and the text video information by using a pre-established recognition model includes: Inputting the text processing information, text image information and text video information into a pre-established recognition model to determine labels corresponding to the text processing information, text image information and text video information respectively; Work notes are generated according to the labels and operation times respectively corresponding to the text processing information, the text image information and the text video information.
3. The method according to claim 1, characterized in that The recognition model is pre-established in the following manner: Acquire multiple text-based historical operation information of the smart terminal and labels corresponding to each piece of text-based historical operation information, and pre-process the text-based historical operation information data; The textual historical operation information and corresponding labels that have undergone data preprocessing are used as training sample data to train a preset initial natural language processing model to obtain the recognition model.
4. A work note generating device, characterized in that: The device comprises: An information collection module, used to collect multiple pieces of operation information of the intelligent terminal and the operation time corresponding to each piece of operation information; wherein the operation information includes screenshot image information, text processing information and video playback information; An information identification module, used to parse and identify the screenshot image information and the video playback information, and obtain text-type image information and text-type video information corresponding to the screenshot image information and the video playback information respectively; a note generation module, for generating work notes using a pre-established recognition model according to the text processing information, the text image information, the text video information, and the operation time corresponding to the text processing information, the text image information, and the text video information; The text processing information includes user input information, page browsing information, text copy information and file processing information, wherein the user input information is obtained by inputting text on the keyboard, the page browsing information is obtained by using a web crawler to obtain the content of the web page when browsing the web page, the text copy information is obtained by obtaining the text or file content in the clipboard when copying and pasting, and the file processing information is obtained by obtaining the file content when saving, transferring or printing the file; The information identification module comprises: An image recognition unit, used to perform image recognition on the screenshot image information to obtain text image information corresponding to the screenshot image information; The speech recognition unit is used to perform speech recognition on the video playback information to obtain text-type video information corresponding to the video playback information.
5. The device according to claim 4, characterized in that The note generation module includes: A label generation unit, used for inputting the text processing information, text image information and text video information into a pre-established recognition model to determine labels corresponding to the text processing information, text image information and text video information respectively; The note generating unit is used to generate work notes according to the labels and operation times respectively corresponding to the text processing information, the text image information and the text video information.
6. The device according to claim 4, characterized in that The device also includes a model building module, which is used to obtain multiple text-based historical operation information of the smart terminal and labels corresponding to each piece of text-based historical operation information, and pre-process the text-based historical operation information data; the pre-processed text-based historical operation information and the corresponding labels are used as training sample data to train a preset initial natural language processing model to obtain the recognition model.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 3 is implemented.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for executing the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Automated medical note generation system utilizing text, audio and video data
US20190057760A1