A text recognition method, device, equipment and storage medium in a task mining scenario

CN117079287BActive Publication Date: 2025-09-09SHANGHAI YISAIQI SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311104798.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-30
Publication Date
2025-09-09
Estimated Expiration
2043-08-30

AI Technical Summary

Technical Problem

该种方法,可以提升压缩后UI图像ocr识别准确率的上升,但是这种上升还是无法达到针对原始图像进行ocr识别的准确率,因为该种处理方法,无法利用图像的上下文信息

Benefits of technology

[0052] The present invention proposes a text recognition method in a task mining scenario. On the premise of compressing the collected user operation data to save storage space, it neither affects the recognition accuracy of the OCR model nor requires the technical problem of re-labeling the training data, thereby realizing effective text recognition in the task mining scenario.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117079287B_ABST
    Figure CN117079287B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, equipment and storage medium for text recognition in a task mining scenario, belonging to the field of text recognition technology. The method comprises: obtaining the target user's operation behavior data on the computer desktop, the operation behavior data including a video stream; performing a segmentation operation on the video stream to obtain multiple frames of UI images; determining keyframe UI images and non-keyframe UI images from the multiple frames of UI images; performing compression of a first target compression quality on the keyframe UI images, and performing compression of a second target compression quality on the non-keyframe UI images to obtain a compressed UI image; in response to receiving an instruction to perform task mining, restoring the compressed UI image based on a pre-trained super-resolution model to obtain a restored UI image; inputting the restored UI image into a pre-trained text recognition model for recognition to obtain a text recognition result. Effective text recognition in a task mining scenario is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a text recognition method, device, equipment and storage medium in a task mining scenario, and belongs to the technical field of text recognition. Background Art

[0002] In task mining products, it is often necessary to record user operations, that is, to record every step of the user's operation through screenshots. In the long run, the saved image data will inevitably affect storage space. Therefore, from the perspective of saving storage space and reducing costs, it is necessary to compress the collected UI images. In existing technologies, in order to achieve a good compression effect, the display quality of the image is often compromised and the image resolution is reduced. Reducing the resolution of the collected UI images will affect the accuracy of subsequent OCR recognition; in actual applications, after deep compression, the accuracy rate may drop from 90% to 10%.

[0003] How can we reduce the resolution to save storage space without affecting the accuracy of OCR recognition?

[0004] One feasible solution is to retrain an OCR model for the reduced resolution. Specifically, you can collect compressed UI images, re-annotate the compressed UI images, and then train a new optical character recognition OCR model. However, this method requires re-annotating the images, which will increase the cost. The purpose of reducing storage space is to reduce costs. Therefore, this method is not applicable. Moreover, the compressed UI images will lose a lot of image information. Even if the corresponding OCR model is retrained after annotation, its recognition accuracy is usually not as good as the OCR model trained on the original UI images.

[0005] Another possible solution is to apply a CV image processing algorithm to the compressed image before OCR recognition. For example, you can sharpen the compressed image before performing OCR recognition on the sharpened image. This method can improve the OCR recognition accuracy of compressed UI images, but it still cannot reach the accuracy of OCR recognition of the original image because it cannot utilize the image's contextual information.

[0006] Therefore, how to compress the collected UI images to save storage space without affecting the subsequent OCR recognition accuracy and without the need to re-label the UI images is a technical problem that needs to be solved. Summary of the Invention

[0007] Purpose: In view of at least one of the above technical problems, the present invention provides a text recognition method, device, equipment and storage medium in a task mining scenario, which is used to solve the technical problem of compressing the collected user operation data to save storage space without affecting the recognition accuracy of the OCR model and without the need to re-label the training data.

[0008] Technical solution: To solve the above technical problems, the technical solution adopted by the present invention is:

[0009] In a first aspect, the present invention provides a text recognition method in a task mining scenario, the method comprising:

[0010] Acquire the target user's operating behavior data on the computer desktop; wherein the operating behavior data includes a video stream;

[0011] Slicing the video stream to obtain multiple frames of UI images;

[0012] Determine a key frame UI image and a non-key frame UI image from the multiple frames of UI image;

[0013] Performing compression at a first target compression quality on the keyframe UI image and performing compression at a second target compression quality on the non-keyframe UI image to obtain compressed keyframe UI images and non-keyframe UI images; wherein the target compression quality is used to represent the storage space occupied by the compressed UI image, and the first target compression quality is higher than the second target compression quality;

[0014] In response to receiving an instruction to perform task mining, restoring the compressed keyframe UI image and the non-keyframe UI image based on a pre-trained super-resolution model to obtain a restored UI image;

[0015] The restored UI image is input into the pre-trained text recognition model for recognition to obtain the text recognition result.

[0016] In some embodiments, the multiple UI image frames carry a time identifier for identifying the time sequence of each UI image; and determining the key frame UI image and the non-key frame UI image from the multiple UI image frames includes:

[0017] Generate image vectors for each frame of UI image;

[0018] In a time sequence consisting of time stamps of each frame of UI image, similarity judgment is performed on the image vectors of the UI image in the next frame and the UI image in the previous frame;

[0019] If the image vector of the UI image in the subsequent frame is similar to the image vector of the UI image in the previous frame, the UI image in the subsequent frame is a non-keyframe UI image; if not similar, the UI image in the subsequent frame is a keyframe UI image.

[0020] Furthermore, in some embodiments, determining the key frame UI image and the non-key frame UI image from the multiple frames of UI images specifically includes:

[0021] Performing a segmentation process on multiple frames of UI images to obtain image blocks corresponding to each UI image, wherein the image blocks also carry a time stamp;

[0022] Generate image vectors of each image block corresponding to each UI image;

[0023] In a time sequence formed by the time stamps of each frame of UI image, similarity judgment is performed on each image block in the UI image of the next frame and the corresponding image block in the UI image of the previous frame;

[0024] If the image block of the UI image in the subsequent frame is similar to the corresponding image block of the UI image in the previous frame, the image block of the UI image in the subsequent frame is a non-key image block; if not similar, the image block of the UI image in the subsequent frame is a key image block;

[0025] A UI image including key image blocks whose number exceeds a threshold is determined as a key frame UI image; and a UI image including key image blocks whose number does not exceed the threshold is determined as a non-key frame UI image.

[0026] Furthermore, in some embodiments, performing compression with a first target compression quality on the key frame UI image and performing compression with a second target compression quality on the non-key frame UI image to obtain compressed key frame UI images and non-key frame UI images includes:

[0027] Performing compression with a first target compression quality on key image blocks in each UI image, and performing compression with a second target compression quality on non-key image blocks in each UI image, to obtain compressed key image blocks and non-key image blocks;

[0028] The image blocks belonging to each frame of UI image are stitched together to obtain compressed key-frame UI images and non-key-frame UI images.

[0029] In some embodiments, the super-resolution model training method includes:

[0030] The training dataset is composed of the original UI images and the corresponding compressed UI images;

[0031] Based on the pre-constructed loss function, the super-resolution model is trained using the training dataset until a preset condition is met, thereby obtaining a trained super-resolution model.

[0032] Furthermore, the loss function is constructed based on the loss between the original UI image and the corresponding compressed UI image.

[0033] In some embodiments, the generating of the image vectors of the image blocks corresponding to the UI images adopts one of the following methods:

[0034] Method 1: Use a convolutional neural network to extract features from each image block and use the extracted features as the image vector of the image block;

[0035] Method 2: Construct a corresponding image pyramid for each image block. Perform multiple downsampling on a certain image block to obtain images of different resolutions to construct an image pyramid. Perform feature extraction on the image of each pyramid, and use the features extracted from each pyramid as the image vector of the corresponding image block.

[0036] Method 3: Input each image block into the pre-trained Vision Transformer model and use the output of the Vision Transformer model as the image vector of each image block.

[0037] In a second aspect, the present invention provides a text recognition device in a task mining scenario, the device comprising:

[0038] An acquisition module, which is used to acquire the target user's operating behavior data on the computer desktop; wherein the operating behavior data includes a video stream;

[0039] A segmentation module, configured to segment the video stream to obtain multiple frames of UI images;

[0040] A key frame determining module, configured to determine a key frame UI image and a non-key frame UI image from the multiple frames of UI images;

[0041] a compression module configured to compress keyframe UI images at a first target compression quality and non-keyframe UI images at a second target compression quality, thereby obtaining compressed keyframe UI images and non-keyframe UI images; wherein the target compression quality is used to represent the storage space occupied by the compressed UI images, and the first target compression quality is higher than the second target compression quality;

[0042] A restoration module, configured to, in response to receiving an instruction to perform task mining, restore the compressed keyframe UI image and the non-keyframe UI image based on a pre-trained super-resolution model to obtain a restored UI image;

[0043] The text recognition module is used to input the restored UI image into a pre-trained text recognition model for recognition to obtain text recognition results.

[0044] In a third aspect, the present invention provides a device comprising:

[0045] Memory;

[0046] processor;

[0047] as well as

[0048] computer programs;

[0049] The computer program is stored in the memory and is configured to be executed by the processor to implement the method described in the first aspect above.

[0050] In a fourth aspect, the present invention provides a storage medium having a computer program stored thereon, wherein the computer program implements the method described in the first aspect when executed by a processor.

[0051] Beneficial effects: The text recognition method, device, equipment, and storage medium provided by the present invention in the task mining scenario have the following advantages:

[0052] The present invention proposes a text recognition method in a task mining scenario. On the premise of compressing the collected user operation data to save storage space, it neither affects the recognition accuracy of the OCR model nor requires the technical problem of re-labeling the training data, thereby realizing effective text recognition in the task mining scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A schematic diagram of an application scenario of a text recognition method in a task mining scenario according to an embodiment of the present invention;

[0054] Figure 2 A flowchart of a text recognition method in a task mining scenario according to an embodiment of the present invention;

[0055] Figure 3 A schematic diagram of a text recognition device according to an embodiment of the present invention;

[0056] Figure 4 A schematic diagram of an electronic device provided according to an embodiment of the present invention. DETAILED DESCRIPTION

[0057] The present invention will be further described below in conjunction with the accompanying drawings and examples. The following examples are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.

[0058] like Figure 1 As shown, before describing the embodiments of the present invention in detail, a specific scenario example is used to reveal the scenario in which the technical solution of the present invention can be used. The method of the present invention can be configured in Figure 1 In the server, the server can communicate with multiple user terminals.

[0059] Multiple user terminals, such as Figure 1 As shown in 110-130, the user terminal is used to collect specific computer desktop operations performed by the user at the user terminal, thereby generating operation behavior data, and transmitting the collected operation behavior data to the server through the network.

[0060] After receiving user operation behavior data (video streams) transmitted by multiple terminals, the server 140 segments the video streams to obtain multiple frames of UI images, and then determines keyframe UI images and non-keyframe UI images from the multiple frames of UI images. A lower degree of compression (lower compression ratio) is then performed on the keyframe UI images, and a higher degree of compression (higher compression ratio) is performed on the non-keyframe UI images. The compressed UI images are then stored to save storage space. For example, the compressed UI images may be stored in the database 150.

[0061] When the server 140 receives an instruction to perform task mining, the instruction may be issued by one of the user terminals 110-130 mentioned above. The task mining instruction may carry a user ID and a time ID. The user ID indicates the specific user's operation behavior data for which the task mining is to be performed, and the time ID indicates the specific time interval's operation behavior data for which the task mining is to be performed.

[0062] The server 140 can extract the compressed UI image corresponding to the instruction from the database 150 based on the instruction for performing task mining, then restore the compressed UI image, and finally perform text recognition on the restored UI image. The recognized text can be used for task mining.

[0063] The aforementioned server 140 can be an electronic device with certain computing and processing capabilities. For example, the server can be a server in a distributed system, or a system with multiple processors, memories, network communication modules, etc. operating in coordination. The server can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host equipped with artificial intelligence technology. The server can also be a server cluster formed by multiple servers. Alternatively, with the advancement of science and technology, the server can also be a new technical means capable of implementing the corresponding functions of the embodiments of this specification. For example, it can be a new form of "server" based on quantum computing.

[0064] The user terminal can be an electronic device with network access capability. Specifically, for example, the terminal can be a desktop computer, tablet computer, laptop computer, smart phone, etc. Alternatively, the terminal can also be software that can run on the electronic device.

[0065] The aforementioned networks may be any type of network that can support data communications using any of a variety of available protocols, including but not limited to TCP / IP, SNA, IPX, etc. The one or more networks may be a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0066] The database 150 described above can be located locally on the server, or can be remote from the server and can communicate with the server via a network-based or dedicated connection. The database can be of different types. In some embodiments, the database used by the server can be a relational database. One or more of these databases can store, update, and retrieve data to and from the database in response to commands.

[0067] Example 1

[0068] First, as Figure 2 As shown, this embodiment provides a text recognition method in a task mining scenario, including:

[0069] Step S100: Acquire the target user's operating behavior data on the computer desktop; wherein the operating behavior data includes a video stream.

[0070] In this embodiment, the above-mentioned operation behavior data can be directly recorded by recording the user's computer desktop operation.

[0071] In this embodiment, the above-mentioned operation behavior data may be a video stream.

[0072] Step S200: Segment the video stream to obtain multiple frames of UI images.

[0073] In this embodiment, the multi-frame UI images obtained by segmentation carry a time stamp. The time stamp is used to identify the time sequence of each UI image. The time stamp can be real-time data collected when the user operates on the computer desktop. For example, the time interval of a video stream is 2023 / 8 / 18 / 9:57-2023 / 8 / 18 / 12:03, then the time stamp pair of each UI image is obtained by segmenting 2023 / 8 / 18 / 9:57-2023 / 8 / 18 / 12:03. For example, the time stamp of a frame of UI image can be: 2023 / 8 / 18 / 10:01.

[0074] Accordingly, in step S200, when performing segmentation, the time information of the real-time occurrence of the user operation may not be considered, that is, only the order of the UI images in the time sequence is considered. In other words, the time identifier can also be used only to represent the time sequence of the UI image frames, without carrying the real-time time information of the user operation. For example, the time identifier of a UI image frame may be 00023, and the time identifier of a subsequent UI image frame may be 00078.

[0075] Step S300: Determine a key-frame UI image and a non-key-frame UI image from the multiple frames of UI images.

[0076] In this embodiment, a UI image with a changed image content can be defined as a keyframe image. This change in image content is indicated by a change in the image content of the subsequent frame relative to the previous frame exceeding a threshold. A change in the image content of the subsequent UI frame relative to the previous frame exceeding the threshold indicates that the user has performed a specific operation on the computer desktop, such as a business system, which has caused the content of the UI image to change. Furthermore, this operation can be a critical operation. Specifically, dragging the mouse pointer on the computer desktop will also cause the image content of the subsequent UI frame to change relative to the previous UI frame, but this change is often relatively small. Accordingly, such desktop operations are often not critical operations for task mining. Conversely, if a user clicks a button in a business system to cause a new window to pop up, this operation will cause the image content of the subsequent UI frame to change significantly relative to the previous UI frame. Accordingly, this operation can also be a critical operation.

[0077] As described above, step 300 may include:

[0078] Generate image vectors for each frame of UI image;

[0079] In a time sequence consisting of time stamps of each frame of UI image, similarity judgment is performed on the image vectors of the UI image in the next frame and the UI image in the previous frame;

[0080] If the image vector of the UI image in the subsequent frame is similar to the image vector of the UI image in the previous frame, the UI image in the subsequent frame is a non-keyframe UI image; if not similar, the UI image in the subsequent frame is a keyframe UI image.

[0081] Step S400: Perform compression of a first target compression quality on the key frame UI image, and perform compression of a second target compression quality on the non-key frame UI image to obtain compressed key frame UI images and non-key frame UI images; wherein the target compression quality is used to characterize the storage space occupied by the compressed UI image, and the first target compression quality is higher than the second target compression quality.

[0082] In this embodiment, the above-mentioned compression operation is represented by a process of executing an image compression algorithm on the key-frame UI image and the non-key-frame UI image.

[0083] In this embodiment, the keyframe UI images and non-keyframe UI images are compressed to within the range of target compression quality. Specifically, the first target compression quality can be set to 30% of the original UI image storage volume, and the second target compression quality can be set to 10% of the original UI image storage volume. The purpose of this step is to compress the collected images to save storage space. As long as the above-mentioned purpose can be achieved, any image compression algorithm can be used. It should be noted that after the compression is performed, the original UI image needs to be discarded, and then the compressed keyframe UI image and non-keyframe UI image are stored to save storage space.

[0084] Step S500: In response to receiving an instruction to perform task mining, the compressed key-frame UI image and the non-key-frame UI image are restored based on a pre-trained super-resolution model to obtain a restored UI image.

[0085] In this embodiment, only when receiving a text recognition instruction, more specifically, when the server receives a task mining instruction transmitted by the client, which is issued by the user through the user terminal, the server restores the compressed UI image.

[0086] The instruction may carry a user ID and a time ID. The user ID indicates the specific user's UI image is to be restored, and the time ID indicates the specific time period within which the UI image is to be restored.

[0087] In this embodiment, the super-resolution model can adopt models such as Unet, FSRCNN, SRCNN, and LapSRN. The training input of the super-resolution model is the original UI image and the corresponding compressed UI image. In this embodiment, a loss function can be constructed based on the loss between the original UI image and the corresponding compressed UI image to control the training process.

[0088] It should be noted that this super-resolution model is pre-trained, and this solution is for actual application scenarios, not training scenarios.

[0089] Step S600: Input the restored UI image into a pre-trained text recognition model for recognition to obtain a text recognition result.

[0090] In this embodiment, the above-mentioned text recognition model has fixed model parameters and is an existing text recognition model. This solution does not re-annotate new UI images to train the text recognition model, does not increase new annotation costs, and does not reduce the accuracy of subsequent OCR recognition while compressing the collected UI images, and does not increase annotation costs.

[0091] It should be noted that UI images are artificially designed and computer-generated images, which often have certain rules. When users perform specific tasks, they often perform specific tasks in a certain window, and the content outside the window may not be very necessary for task mining. In other words, for a frame of UI image, different areas have different importance. For example, the importance of the window title area is higher than that of other areas, because the window title can be extracted from the window title area. Furthermore, when users perform specific business operations, it is often one or several areas that change, and there are also multiple areas that will not change. For example, in a certain business system, the user keeps clicking on each button to pop up the corresponding window, and performs specific operations in each pop-up box. The area outside the corresponding pop-up box will not change. Accordingly, in the operation of performing compression in step S400, a smaller degree of compression can be performed on the key area, and a larger degree of compression can be performed on the non-key area.

[0092] In summary, the present invention proposes a more fine-grained method for determining key frame UI images and a more fine-grained compression method.

[0093] The above-mentioned more fine-grained method for determining keyframe UI images is described as follows:

[0094] Step 300, determining a key frame UI image and a non-key frame UI image from the multiple frames of UI images, may specifically include:

[0095] Step 301: Perform segmentation processing on multiple frames of UI images to obtain image blocks corresponding to each UI image, wherein the image blocks also carry time stamps;

[0096] There are several ways to slice a UI image. The simplest method is to directly slice the UI image into uniform 16*16 or 8*8 rectangular image blocks. Regardless of the method used, as long as the UI image is sliced ​​into multiple image blocks, it is important to note that all UI images should be sliced ​​using the same method. This is because only when the same cutting method is used can subsequent image block comparison operations be meaningful.

[0097] Step 302: Generate image vectors for each image block corresponding to each UI image;

[0098] There are many ways to generate image vectors for each image block corresponding to each UI image, which can be:

[0099] Method 1: Use a convolutional neural network to extract features from each image block and use the extracted features as the vector representation of the image block.

[0100] Method 2: Build an image pyramid for each image block. Specifically, downsample the image block multiple times to obtain images of different resolutions to construct an image pyramid. Feature extraction is then performed on each pyramid image, for example using a convolutional neural network. The extracted features from each pyramid are then used as a vector representation of the corresponding image block. This method captures multi-scale image information and provides a richer image representation.

[0101] Method 3: Directly input each image patch into a Vision Transformer model, such as ViT-B / 16 or ViT-L / 16, which has been pre-trained on large-scale image datasets. The output of the Vision Transformer model is used as the vector representation of each image patch.

[0102] Step 303: In a time sequence consisting of the time stamps of each frame of UI image, each image block in the UI image of the subsequent frame is judged for similarity with the corresponding image block in the UI image of the previous frame. If the image block in the UI image of the subsequent frame is similar to the corresponding image block in the UI image of the previous frame, the image block in the UI image of the subsequent frame is a non-key image block; if not, the image block in the UI image of the subsequent frame is a key image block. The specific similarity can be controlled by a threshold.

[0103] In this embodiment, the corresponding image blocks are represented as image blocks with the same position information in the two preceding and following frames of images. It is easy to understand that it is meaningless to perform similarity comparison on image blocks at different positions.

[0104] Step 304: Determine a UI image whose number of key image blocks exceeds a threshold as a key-frame UI image; and determine a UI image whose number of key image blocks does not exceed the threshold as a non-key-frame UI image.

[0105] For example, the threshold may be 3 or 5. The specific setting may correspond to the segmentation method in step 301. If there are many segmented image blocks, the threshold may be higher. If there are few segmented image blocks, the threshold may be lower.

[0106] On this basis, the above-mentioned more fine-grained compression method is described as follows:

[0107] Step 400, performing compression with a first target compression quality on the key frame UI image and performing compression with a second target compression quality on the non-key frame UI image to obtain compressed key frame UI images and non-key frame UI images, may include:

[0108] Step 401: Compression is performed at a first target compression quality on key image blocks in each UI image, and compression is performed at a second target compression quality on non-key image blocks in each UI image, to obtain compressed key image blocks and non-key image blocks;

[0109] Step 402: Combine the image blocks belonging to each frame of UI image to obtain compressed key-frame UI images and non-key-frame UI images.

[0110] In some embodiments, in addition to the stitching operation as in step 402, other methods can also be used. Specifically, when cutting in 301, the position information of each image block in the corresponding UI image can be recorded, such as the coordinates of the upper left corner and the lower left corner; in 401, compression of corresponding compression parameters can be performed on different areas in the UI image based on the position information.

[0111] In some cases, the application software running on the user terminal often changes, and the existing application software will also undergo version changes, and the UI interface will not remain unchanged. Therefore, when the user terminal adds a new application software, the data collector configured in the user terminal collects the UI image of the new application software, and the pre-trained super-resolution model has not been trained for the new application software in advance. Therefore, the restoration effect may be poor, and accordingly, the subsequent OCR recognition result will also be poor. To address this situation, in some implementations, the following method can be used:

[0112] Receive an instruction to train a super-resolution model, and store an original target UI image and a corresponding compressed target UI image, wherein the target UI image may be a UI image of a new application or a UI image of an updated version; it should be noted that the original UI image is only saved when the super-resolution model is needed for training;

[0113] The super-resolution model is trained using the stored original target UI image and the corresponding compressed target UI image.

[0114] Example 2

[0115] In the second aspect, based on Example 1, Figure 3 As shown, this embodiment provides a text recognition device in a task mining scenario, including:

[0116] An acquisition module, which is used to acquire the target user's operating behavior data on the computer desktop; wherein the operating behavior data includes a video stream;

[0117] A segmentation module, configured to segment the video stream to obtain multiple frames of UI images;

[0118] A key frame determining module, configured to determine a key frame UI image and a non-key frame UI image from the multiple frames of UI images;

[0119] a compression module configured to compress keyframe UI images at a first target compression quality and non-keyframe UI images at a second target compression quality, thereby obtaining compressed keyframe UI images and non-keyframe UI images; wherein the target compression quality is used to represent the storage space occupied by the compressed UI images, and the first target compression quality is higher than the second target compression quality;

[0120] A restoration module, configured to, in response to receiving an instruction to perform task mining, restore the compressed keyframe UI image and the non-keyframe UI image based on a pre-trained super-resolution model to obtain a restored UI image;

[0121] The text recognition module is used to input the restored UI image into a pre-trained text recognition model for recognition to obtain text recognition results.

[0122] Example 3

[0123] In the third aspect, based on Example 1, Figure 4 As shown, this embodiment provides a device, including:

[0124] Memory;

[0125] processor;

[0126] as well as

[0127] computer programs;

[0128] The computer program is stored in the memory and is configured to be executed by the processor to implement the method described in embodiment 1.

[0129] Example 4

[0130] In a fourth aspect, based on Example 1, this embodiment provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the method described in Example 1 is implemented.

[0131] Example 5

[0132] In a fifth aspect, based on Example 1, this embodiment proposes a hardware system that applies the text recognition method in a task mining scenario in Example 1. This method can be applied to both the terminal side and the server side.

[0133] A terminal can be any electronic device with network access capabilities, such as a desktop computer, tablet computer, laptop computer, smartphone, digital assistant, shopping guide terminal, or television.

[0134] A server can be an electronic device with certain computing and processing capabilities. For example, a server can be a distributed system server, a system with multiple processors, memories, network communication modules, and the like operating in coordination. A server can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host equipped with artificial intelligence technology. A server can also be a server cluster consisting of multiple servers. Alternatively, with the advancement of science and technology, a server can also be a new technological means capable of implementing the functions described in the embodiments of this specification. For example, it can be a new form of "server" based on quantum computing.

[0135] Example 6:

[0136] In a sixth aspect, based on Example 1, this embodiment further provides a computer program product comprising instructions, which, when executed by a computer, enables the computer to execute a text recognition method in a task mining scenario in Example 1.

[0137] It should be understood that the specific examples herein are only intended to help those skilled in the art better understand the embodiments of this specification, rather than to limit the scope of the present invention.

[0138] It can be understood that in the various implementations of this specification, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the implementation methods of this specification.

[0139] It can be understood that the various embodiments described in this specification can be implemented individually or in combination, and the embodiments in this specification are not limited to this.

[0140] Unless otherwise indicated, all technical and scientific terms used in the embodiments of this specification have the same meaning as those commonly understood by those skilled in the art in the technical field of this specification. The terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the scope of this specification. The term "and / or" used in this specification includes any and all combinations of one or more related listed items. The singular forms "a", "above", and "the" used in the embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.

[0141] It is understood that the processor in the embodiments of this specification can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiment can be completed by hardware integrated logic circuits in the processor or software instructions. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this specification can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this specification can be directly implemented as being executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0142] It will be understood that the memory in the embodiments of this specification may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0143] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this specification.

[0144] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, devices and units can refer to the corresponding processes in the aforementioned method implementation methods and will not be repeated here.

[0145] In the several embodiments provided in this specification, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0146] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of this embodiment.

[0147] In addition, each functional unit in each embodiment of this specification may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0148] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this specification, or the part that contributes to the prior art, or the part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this specification. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0149] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A text recognition method in a task mining scenario, characterized in that: The method comprises: Acquire the target user's operating behavior data on the computer desktop; wherein the operating behavior data includes a video stream; Slicing the video stream to obtain multiple frames of UI images; Determine a key frame UI image and a non-key frame UI image from the multiple frames of UI image; Performing compression at a first target compression quality on the keyframe UI image and performing compression at a second target compression quality on the non-keyframe UI image to obtain compressed keyframe UI images and non-keyframe UI images; wherein the target compression quality is used to represent the storage space occupied by the compressed UI image, and the first target compression quality is higher than the second target compression quality; In response to receiving an instruction to perform task mining, restoring the compressed keyframe UI image and the non-keyframe UI image based on a pre-trained super-resolution model to obtain a restored UI image; Input the restored UI image into the pre-trained text recognition model for recognition to obtain the text recognition result; Among them, the multi-frame UI images carry a time identifier for identifying the time sequence of each UI image; determining the key frame UI image and the non-key frame UI image from the multi-frame UI images specifically includes: performing a cutting process on the multi-frame UI images to obtain an image block corresponding to each UI image, wherein the image block also carries a time identifier; generating an image vector of each image block corresponding to each UI image; in a time sequence composed of the time identifiers of each frame UI image, performing a similarity judgment on each image block in the UI image of the subsequent frame and the corresponding image block in the UI image of the previous frame; if the image block of the UI image of the subsequent frame is similar to the image block corresponding to the UI image of the previous frame, the image block of the UI image of the subsequent frame is a non-key image block; if not similar, the image block of the UI image of the subsequent frame is a key image block; determining a UI image containing a number of key image blocks exceeding a threshold as a key frame UI image; and determining a UI image containing a number of key image blocks not exceeding a threshold as a non-key frame UI image.

2. The text recognition method in the task mining scenario according to claim 1 is characterized in that: Performing compression with a first target compression quality on the key frame UI image and performing compression with a second target compression quality on the non-key frame UI image to obtain compressed key frame UI images and non-key frame UI images, including: Performing compression of a first target compression quality on key image blocks in each UI image, and performing compression of a second target compression quality on non-key image blocks in each UI image, to obtain compressed key image blocks and non-key image blocks; The image blocks belonging to each frame of UI image are stitched together to obtain compressed key-frame UI images and non-key-frame UI images.

3. The text recognition method in the task mining scenario according to claim 1 is characterized in that: The training method of the super-resolution model includes: The training dataset is composed of the original UI images and the corresponding compressed UI images; Based on the pre-constructed loss function, the super-resolution model is trained using the training dataset until a preset condition is met, thereby obtaining a trained super-resolution model.

4. The text recognition method in the task mining scenario according to claim 3 is characterized in that: The loss function is constructed based on the loss between the original UI image and the corresponding compressed UI image.

5. The text recognition method in the task mining scenario according to claim 1 is characterized in that: The image vectors of the image blocks corresponding to the UI images are generated by using one of the following methods: Method 1: Use a convolutional neural network to extract features from each image block and use the extracted features as the image vector of the image block; Method 2: Construct a corresponding image pyramid for each image block. Perform multiple downsampling on a certain image block to obtain images of different resolutions to construct an image pyramid. Perform feature extraction on the image of each pyramid, and use the features extracted from each pyramid as the image vector of the corresponding image block. Method 3: Input each image block into the pre-trained Vision Transformer model and use the output of the Vision Transformer model as the image vector of each image block.

6. A text recognition device in a task mining scenario, characterized in that: The device comprises: An acquisition module, which is used to acquire the target user's operating behavior data on the computer desktop; wherein the operating behavior data includes a video stream; A segmentation module, configured to segment the video stream to obtain multiple frames of UI images; a key frame determination module, which is used to determine key frame UI images and non-key frame UI images from the multiple frames of UI images; the multiple frames of UI images carry time identifiers for identifying the time sequence of each UI image; determining key frame UI images and non-key frame UI images from the multiple frames of UI images, specifically comprising: performing a segmentation process on the multiple frames of UI images to obtain image blocks corresponding to each UI image, wherein the image blocks also carry time identifiers; generating image vectors of each image block corresponding to each UI image; performing similarity judgment on each image block in the UI image of the subsequent frame with the corresponding image block in the UI image of the previous frame in a time sequence composed of the time identifiers of each frame of UI images; if the image block of the UI image of the subsequent frame is similar to the image block corresponding to the UI image of the previous frame, the image block of the UI image of the subsequent frame is a non-key image block; if not similar, the image block of the UI image of the subsequent frame is a key image block; determining a UI image containing a number of key image blocks exceeding a threshold as a key frame UI image; and determining a UI image containing a number of key image blocks not exceeding a threshold as a non-key frame UI image; a compression module configured to compress keyframe UI images at a first target compression quality and non-keyframe UI images at a second target compression quality, thereby obtaining compressed keyframe UI images and non-keyframe UI images; wherein the target compression quality is used to represent the storage space occupied by the compressed UI images, and the first target compression quality is higher than the second target compression quality; A restoration module, configured to, in response to receiving an instruction to perform task mining, restore the compressed keyframe UI image and the non-keyframe UI image based on a pre-trained super-resolution model to obtain a restored UI image; The text recognition module is used to input the restored UI image into a pre-trained text recognition model for recognition to obtain text recognition results.

7. A device, characterized in that include: Memory; processor; as well as computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method according to any one of claims 1 to 5.

8. A storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Text recognition model training method and system, electronic equipment and storage medium

    CN116343230A

  • Identifying and redressing shadows in connection with digital watermarking and fingerprinting; and dynamic signal detection adapted based on lighting information

    WO2011163378A1