Device interface data collection method, device, equipment and storage medium
By pre-training the main interface discriminator and intelligent segmentation technology, the device interface is automatically identified and presented, the occlusion sub-interface is restored and intelligent segmentation and identification is performed, which solves the problems of low efficiency and poor accuracy of equipment data acquisition, and realizes efficient and accurate data acquisition and real-time monitoring.
Patent Information
- Application Number
- CN202211110854.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-09-13
AI Technical Summary
During the digital transformation process, there are difficulties in obtaining multi-source heterogeneous multimodal data, especially in the problems of device data interface blocking and third-party software system maintenance, resulting in low data acquisition efficiency and poor accuracy, affecting business decisions.
The pre-trained main interface discriminator automatically recognizes and presents the required main interface, determines and restores the target sub-interface that is blocked or exceeds the screen boundary, and uses intelligent segmentation and character classification discriminator for data recognition to realize automated data acquisition.
It realizes data acquisition for hidden or unenabled main interfaces, and automatically restores sub-interfaces that are blocked or exceeded boundaries, improving the accuracy and efficiency of data acquisition, reducing manual error rates, and supporting real-time data monitoring.
Smart Images

Figure CN115629831B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data acquisition technology, and in particular to a method, device, equipment and storage medium for data acquisition of a device interface. Background Art
[0002] Under the influence of various factors such as technological progress, changes in business models, consumption upgrades, and rising labor costs, the landscape of many industries in the era of large-scale industry has changed. For all walks of life, digital transformation is no longer an option but a must. In the process of digital transformation, the acquisition of multi-source heterogeneous (structured, unstructured, etc.) multimodal data is the source. For example, there are a large number of multi-source heterogeneous multimodal (text, pictures, videos, voice, etc.) data in industrial production equipment, government affairs systems of government departments, equipment and instruments in medical institutions, market analysis software in financial institutions, and shopping systems in the retail industry. However, the equipment, instruments or software systems in these industries are difficult to obtain data. For example, in industrial production, a large number of production equipment comes from domestic or international procurement, and there are problems such as the procurement equipment not opening data interfaces to enterprises. At the same time, the supporting software system may be developed by a third party, and due to the operation problems of the third party, it is common that the software is no longer updated and maintained. The data acquisition of equipment and software systems is heavily dependent on external parties. These problems have greatly delayed the digitalization process of various industries.
[0003] In many industries, due to the lack of core equipment operation data, some units still obtain equipment and system data by visual observation, hand copying, and mental memorization. However, these methods have the following disadvantages: labor-intensive, requiring a large amount of manual labor to copy data, which brings labor costs; complex processes, and the complexity of the equipment interface bring difficulties in collection, such as interface classification problems, interface overlap problems, interface displacement problems, etc.; error-prone, relying on manual labor, unstable efficiency, easy to make mistakes in input, resulting in inaccurate results and affecting business; slow response. It is impossible to obtain equipment and system data in real time, and decision-making responses are not timely. Summary of the invention
[0004] The embodiments of the present application provide a data collection method, apparatus, device and storage medium for a device interface. In order to have a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary is not a general review, nor is it intended to identify key / important components or describe the scope of protection of these embodiments. Its only purpose is to present some concepts in a simple form as a preface to the detailed description that follows.
[0005] In a first aspect, an embodiment of the present application provides a method for collecting data from a device interface, comprising:
[0006] Identify the required main interface according to the pre-trained main interface discriminator, and automatically present the required main interface on the device to be collected;
[0007] Determine the image template to be collected, locate the target sub-interface similar to the image template on the required main interface, and automatically restore the blocked part of the target sub-interface or the part beyond the screen boundary;
[0008] The target sub-interface is intelligently segmented to obtain a single character image, and the single character image is input into a pre-trained character classification discriminator to obtain the recognized data to be collected.
[0009] In some optional embodiments, identifying the required main interface according to the pre-trained main interface discriminator and automatically presenting the required main interface on the device to be collected includes:
[0010] Capture the main interface of the device to be collected;
[0011] Input the intercepted main interface into a pre-trained main interface discriminator to determine whether the intercepted main interface is the desired main interface;
[0012] When the captured main interface is the required main interface, the required main interface is automatically presented;
[0013] When the captured main interface is the desktop interface, the desired main interface is automatically searched and presented through the keyboard and mouse controller;
[0014] When the captured main interface is the interface of other software, the interface is switched to the desktop interface through the keyboard and mouse controller, and then the keyboard and mouse controller is used to automatically search and present the required main interface.
[0015] In some optional embodiments, before identifying the required main interface according to the pre-trained main interface discriminator, the method further includes:
[0016] Intercepting a preset number of main interfaces, the preset number of main interfaces including the required main interface, desktop interface and other software interfaces;
[0017] Divide a preset number of main interfaces into training sets, test sets, and validation sets;
[0018] The main interface discriminator is trained according to the training set, the test set, the validation set and the classification neural network model to obtain a trained main interface discriminator.
[0019] In some optional embodiments, determining an image template to be captured, locating a target sub-interface similar to the image template on a desired main interface, and automatically restoring a blocked portion of the target sub-interface or a portion beyond a screen boundary, includes:
[0020] Intercept the image template to be collected on the required main interface, and obtain the length and width of the image template;
[0021] According to the required main interface, the image template and the correlation coefficient matching algorithm, the similarity between each position in the required main interface and the image template is obtained;
[0022] The position coordinates with the greatest similarity are taken as the starting coordinates, and a sub-image with the same length and width as the image template is intercepted on the required main interface with the starting coordinates as the starting point;
[0023] Calculate the HSV matching confidence between the sub-image and the image template. If the HSV matching confidence is greater than or equal to the preset threshold, the sub-image is the target sub-interface.
[0024] If the HSV matching confidence is less than the preset threshold, the sub-image is occluded or exceeds the screen boundary, and the occluded part of the sub-image or the part exceeding the screen boundary is automatically restored. If the HSV matching confidence of the restored sub-image is greater than or equal to the preset threshold, the restored sub-image is the target sub-interface.
[0025] In some optional embodiments, according to the required main interface, the image template and the correlation coefficient matching algorithm, the similarity between each position in the required main interface and the image template is obtained, including:
[0026] Grayscale the required main interface and image template to obtain grayscale matrices respectively;
[0027] The starting point coordinates of the required main interface, the starting point coordinates of the image template, the grayscale matrix of the required main interface and the grayscale matrix of the image template are input into the correlation coefficient matching algorithm to obtain the similarity between each position in the required main interface and the image template.
[0028] In some optional embodiments, automatically restoring the blocked portion of the sub-image or the portion beyond the screen boundary, if the HSV matching confidence of the restored sub-image is greater than or equal to a preset threshold, then the restored sub-image is the target sub-interface, including:
[0029] Click the unobstructed part of the sub-image by using the keyboard and mouse controller to obtain the restored sub-image;
[0030] Calculate the HSV matching confidence between the restored sub-image and the image template. If the HSV matching confidence is greater than or equal to the preset threshold, the restored sub-image is the target sub-interface.
[0031] If the HSV matching confidence is less than the preset threshold, the required main interface is evenly divided into four quadrants with the center of the required main interface as the origin, and the sub-image is dragged by the keyboard and mouse controller to move the preset distance in the quadrant direction symmetrical to the center of the quadrant where it is located to obtain the restored sub-image;
[0032] The HSV matching confidence between the restored sub-image and the image template is calculated. If the HSV matching confidence is greater than or equal to the preset threshold, the restored sub-image is the target sub-interface.
[0033] In some optional embodiments, intelligent segmentation is performed on the target sub-interface to obtain a single character image, including:
[0034] Convert the target sub-interface into a grayscale image to obtain a grayscale matrix corresponding to the target sub-interface;
[0035] According to the grayscale matrix corresponding to the target sub-interface, the grayscale image is converted into a binary image;
[0036] The pixel values of each column and each row in the binary image are accumulated, and the intersection of the mutation point of the accumulated pixel value of each column and the mutation point of the pixel value of each row is used as the segmentation point to segment a single character image.
[0037] In a second aspect, an embodiment of the present application provides a data acquisition device for a device interface, comprising:
[0038] The automatic interface recognition and presentation module is used to recognize the required main interface according to the pre-trained main interface discriminator, and automatically present the required main interface on the device to be collected;
[0039] A target sub-interface determination module is used to determine the image template to be collected, locate the target sub-interface similar to the image template on the required main interface, and automatically restore the blocked part of the target sub-interface or the part beyond the screen boundary;
[0040] The interface data acquisition module is used to intelligently segment the target sub-interface to obtain a single character image, and input the single character image into a pre-trained character classification discriminator to obtain the recognized data to be collected.
[0041] In a third aspect, an embodiment of the present application provides a data acquisition device for a device interface, including a processor and a memory storing program instructions, wherein the processor is configured to execute the data acquisition method for the device interface provided by the above embodiment when executing the program instructions.
[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable medium having computer-readable instructions stored thereon, and the computer-readable instructions are executed by a processor to implement a data collection method for a device interface provided in the above embodiment.
[0043] The technical solution provided by the embodiments of the present application may have the following beneficial effects:
[0044] The data collection method for the device interface provided in the embodiment of the present application can automatically identify and present the required main interface. It is different from the usual data collection, which requires manual click to place the visual interface in the front, and cannot handle the situation where the sub-interface in the interface is moved manually. The present application creatively uses the main interface matching and template matching technology to intelligently present the main interface to be collected at the front of the screen, and can adjust the partially obscured sub-interface to become fully visible. And the present application creatively uses the data intelligent segmentation technology, which can effectively locate all characters, and combines the data intelligent recognition technology to identify the character category, achieving 100% accuracy, greatly improving work efficiency, and significantly reducing the error rate of handwriting. The method does not affect the normal use of the equipment, allows changes and displacements of the visual interface, can monitor data changes in real time, automatically pause when the operator uses the equipment, and automatically restart after use, without terminating the process.
[0045] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0047] Figure 1 is a flow chart of a method for collecting data from a device interface according to an exemplary embodiment;
[0048] Figure 2 is a framework diagram of a data collection method for a device interface according to an exemplary embodiment;
[0049] Figure 3 is a flow chart of a method for collecting data from a device interface according to an exemplary embodiment;
[0050] Figure 4 is a schematic diagram showing intelligent identification and presentation of a main interface according to an exemplary embodiment;
[0051] Figure 5 is a schematic diagram showing intelligent template matching and restoration according to an exemplary embodiment;
[0052] Figure 6 is a schematic diagram of intelligent data segmentation and recognition according to an exemplary embodiment;
[0053] Figure 7 This is a sample display of required collection sub-interface data according to an exemplary embodiment;
[0054] Figure 8is a schematic diagram of a data acquisition device for a device interface according to an exemplary embodiment;
[0055] Fig. 9 is a schematic structural diagram of an electronic device according to an exemplary embodiment;
[0056] Fig.10 It is a schematic diagram of a computer storage medium according to an exemplary embodiment. DETAILED DESCRIPTION
[0057] The following description and the drawings sufficiently illustrate specific embodiments of the invention to enable those skilled in the art to practice them.
[0058] It should be clear that the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0059] When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are only examples of systems and methods consistent with some aspects of the present invention as detailed in the attached claims.
[0060] At present, whether it is production equipment, government affairs systems or other types of hardware and software facilities, the data involved are the core basic data assets of various industries and units, and are the foundation of digital transformation. It is necessary to break through the blockade of equipment data interfaces, get rid of the constraints of third-party software system developers, and grasp the core data in hand. In response to some of the above problems, there are already some technical solutions on the market, which are as follows: One is based on OCR optical character recognition technology to identify data information in the specified area of the screen and write it into an Excel file or database. The second is data collection based on robotic process automation (RPA), which solidifies the data collection process and performs scheduled operations to replace the tedious and boring manual work. The process robot can work 24 hours a day. However, all of the above traditional technical solutions have some problems.
[0061] Solution 1 is based on OCR optical character recognition technology. If the application interface on the software system screen is not fixed, such as when the resolution of the main interface changes or the position of the sub-interface changes, the solution cannot accurately locate the target interface, and mis-positioning may occur, resulting in a significant decrease in acquisition accuracy. The effect is poor for low-pixel level characters (such as single text, numbers, etc.), which are prone to mis-positioning and mis-classification, resulting in insufficient recognition accuracy, requiring manual re-proofreading, which is time-consuming and labor-intensive.
[0062] Solution 2 is based on robot collection, which usually solidifies the collection process. For a large number of complex interfaces in the software system, the target collection interface needs to be manually specified, and the ability to intelligently determine the target collection interface is lacking. It is even more difficult to handle complex situations such as interface overlap and target interface coverage in the system, resulting in data collection stagnation. For example, in industrial production equipment, interface data will cause overall pixel offset and size changes due to system reasons (such as resolution changes, system freezes, etc.), and the process cannot be reversed, which will affect the RPA fixed process. Therefore, this solution is not suitable for scenarios where pixel coordinates or sizes will change.
[0063] To sum up, traditional data collection solutions have the following problems: it is impossible to intelligently determine whether the required target interface exists, nor is it possible to open the required target interface; the system interface may be moved or blocked, making it difficult to locate, resulting in a significant decrease in collection accuracy; the existing character recognition model has poor recognition effect on low-pixel characters, and the collection results need to be re-calibrated.
[0064] Based on this, the embodiment of the present application provides a data collection method for a device interface, which can collect (hidden) visual interface data on any operating system in real time. At the same time, a solution for resetting the sub-interface is proposed, and combined with data intelligent segmentation and recognition technology, the following problems are solved: uniformly applicable to data collection scenarios of hidden and unopened main interfaces, the main interface can be presented to the front of the screen; uniformly applicable to scenarios where the sub-interface is blocked or partially outside the boundary of the main interface, the sub-interface can be reset; solve the problem of low positioning and recognition accuracy of existing OCR models. Adapt to changing processes and monitor data in real time. Manually operated equipment will pause the collection process and restart intelligently.
[0065] The following will describe in detail the data collection method for the device interface provided by the embodiment of the present application in conjunction with the accompanying drawings. Figure 1 , the method specifically includes the following steps.
[0066] S101 identifies the required main interface according to the pre-trained main interface discriminator, and automatically presents the required main interface on the device to be collected.
[0067] Among them, the equipment to be collected may be production equipment in industry, government affairs system equipment in government departments, equipment and instruments in medical institutions, market analysis software equipment in financial institutions, shopping system equipment in the retail industry, etc., and no specific limitation is made without applying for an embodiment.
[0068] To collect data on the interface of the device to be collected, first, the number of collections and the collection cycle can be set, such as single collection, continuous collection or other collection methods. Then, the pre-trained main interface discriminator identifies the required main interface and automatically presents the required main interface on the device to be collected.
[0069] Specifically, the main interface of the device to be collected can be captured by using a screenshot device to capture the current main interface of the device to be collected, and the captured main interface is input into a pre-trained main interface discriminator to determine whether the captured main interface is the required main interface. The required main interface is the software or web page interface required for collecting data. When the captured main interface is the required main interface, the required main interface is automatically presented. When the captured main interface is a desktop interface, the required main interface is automatically searched and presented through a keyboard and mouse controller. For example, if the discriminator identifies the interface as a Windows desktop interface, the keyboard and mouse controller is used to automatically click the system software icon at the specified location and automatically enter the account password, or automatically click the browser icon and automatically enter the required URL, and the required main interface is presented.
[0070] When the captured main interface is another software interface, the interface is switched to the desktop interface through the keyboard and mouse controller, and then the desired main interface is automatically searched and presented through the keyboard and mouse controller. For example, the interface is switched to the Windows desktop interface using the keyboard and mouse controller, and then the keyboard and mouse controller is used to automatically click the system software icon at the specified location and automatically enter the account password, or automatically click the browser icon and automatically enter the required URL to present the desired main interface.
[0071] In some optional embodiments, before the required main interface is identified by the pre-trained main interface discriminator, the method further includes: capturing a preset number of main interfaces through a screenshot device, the preset number of main interfaces including the required main interface, desktop interface and other software interfaces, labeling the captured main interfaces, marking the main interface required for collecting data as the required main interface, marking the captured windows desktop interface as the desktop interface, and marking other captured main interfaces not related to the collected data as other software interfaces. The labeled images are divided into a training set, a test set and a validation set, and an image classification neural network model is constructed. The image classification model in the prior art can be used to train the main interface discriminator according to the training set, the test set, the validation set and the classification neural network model. On the basis that the accuracy rate on the training set and the validation set reaches 100%, the test set is tested. If the accuracy rate on the test set also reaches 100%, the model parameters in the trainer are fixed, and the trainer is transformed into a main interface discriminator to obtain a trained main interface discriminator.
[0072] Figure 4 is a schematic diagram showing intelligent recognition and presentation of a main interface according to an exemplary embodiment. Figure 4As shown, first, a large number of main interfaces are captured through a screenshot device, and then the captured main interfaces are classified and annotated, and a training set, a test set, and a validation set are constructed. The main interface discriminator is trained according to the training set, the test set, the validation set, and the classification neural network model to obtain a main interface discriminator with an accuracy of 100%. The main interface of the device currently captured is identified according to the main interface discriminator to determine whether the captured main interface is the required main interface. When the captured main interface is the required main interface, the required main interface is automatically presented. When the captured main interface is the desktop interface, the keyboard and mouse controller is used to automatically search and present the required main interface. When the captured main interface is the interface of other software, the keyboard and mouse controller is used to switch the interface to the desktop interface, and then the keyboard and mouse controller is used to automatically search and present the required main interface.
[0073] According to this step, the required main interface can be automatically identified and actively presented in front of the screen of the device to be collected. 1. It is uniformly applicable to data collection scenarios of hidden and unopened main interfaces, and the main interface can be presented to the front of the screen.
[0074] S102 determines the image template to be collected, locates a target sub-interface similar to the image template on the required main interface, and automatically restores the blocked part of the target sub-interface or the part beyond the screen boundary.
[0075] In one possible implementation, the interface to be captured may be a part of the required main interface. First, this part is captured as the image template to be captured. Then, the area of the main interface to be captured may change. During the continuous change, the target sub-interface similar to the image template is located on the required main interface.
[0076] Specifically, the image template to be collected is intercepted on the desired main interface, and the length and width of the image template are obtained. A screenshot device can be used to intercept the desired sub-interface in the desired main interface as the image template to be collected, and the length and width of the template pixels are output.
[0077] Furthermore, according to the required main interface, the image template and the correlation coefficient matching algorithm, the similarity between each position in the required main interface and the image template is obtained. First, the required main interface and the image template are grayed to obtain gray matrices respectively; then the starting point coordinates of the required main interface, the starting point coordinates of the image template, the gray matrix of the required main interface and the gray matrix of the image template are input into the correlation coefficient matching algorithm to obtain the similarity between each position in the required main interface and the image template.
[0078] In one possible implementation, the correlation coefficient matching algorithm is as follows:
[0079]
[0080] Among them, (x, y) represents the coordinates of the starting point of the required main interface, (x', y') represents the coordinates of the starting point of the image template, T represents the grayscale matrix of the image template, and I represents the grayscale matrix of the required main interface. R(x, y) represents the similarity between each position in the required main interface and the image template. If the pixel size of the required main interface is 1920*1080, then R(x, y) is the 1920*1080 similarity matrix. Each value ranges from 0 to 1, 1 represents the best match, and 0 represents no correlation.
[0081] The position coordinates with the greatest similarity are taken as the starting coordinates, and a sub-image with the same length and width as the image template is intercepted on the required main interface with the starting coordinates as the starting point, wherein the starting point is the point at the upper left corner of the intercepted image.
[0082] Furthermore, the HSV matching confidence between the sub-image and the image template is calculated. If the HSV matching confidence is greater than or equal to a preset threshold, the sub-image is the target sub-interface.
[0083] Specifically, the sub-image and the image template are converted from RGB three-channel color images to HSV images respectively. The specific conversion method is: in the RGB image, R: the brightness value in the red channel; G: the brightness value in the green channel; B: the brightness value in the blue channel.
[0084] R'=R / 255,G'=G / 255,B'=B / 255. Find HSV:
[0085] C max =max(R',G',B'),C min =min(R',G',B'),Δ=C max -C min ;
[0086] In HSV image:
[0087] If C max =R', H = (G'-B') / (C max -C min )*60;
[0088] If C max =G',H=120+(B'-R') / (C max -C min )*60;
[0089] If C max =B',H=240+(R'-G') / (C max -C min )*60;
[0090] S=(C max -Cmin ) / C max ;
[0091] V=C max ;
[0092] After calculating the HSV values of the sub-image and the image template, they are converted into vectors (H1, S1, V1) and (H2, S2, V2) respectively, and then the template discriminator with a built-in cosine similarity formula is used to calculate the HSV matching confidence of the two. If the HSV matching confidence is greater than or equal to the preset threshold, the sub-image is the target sub-interface. If the HSV matching confidence is less than the preset threshold, the sub-image is obscured or exceeds the screen boundary, and the obscured part of the sub-image or the part that exceeds the screen boundary is automatically restored. If the HSV matching confidence of the restored sub-image is greater than or equal to the preset threshold, the restored sub-image is the target sub-interface. Among them, the preset threshold can be set according to the actual situation, and the embodiments of the present disclosure do not impose specific restrictions.
[0093] Specifically, the unobstructed part of the sub-image is clicked by the keyboard and mouse controller to obtain the restored sub-image, and the HSV matching confidence of the restored sub-image and the image template is calculated. If the HSV matching confidence is greater than or equal to a preset threshold, the restored sub-image is the target sub-interface.
[0094] If the HSV matching confidence is less than the preset threshold, the required main interface is evenly divided into four quadrants with the center of the required main interface as the origin, and the sub-image is dragged by the keyboard and mouse controller to move the preset distance in the quadrant direction symmetrical to the center of the quadrant where it is located, where the preset dragging distance can be set according to the actual situation, and the restored sub-image is obtained. The HSV matching confidence of the restored sub-image and the image template is calculated. If the HSV matching confidence is greater than or equal to the preset threshold, the restored sub-image is the target sub-interface. If the matching confidence of the restored sub-image is still less than the preset threshold, the correlation coefficient matching algorithm is re-performed to confirm the sub-image with high correlation again.
[0095] After obtaining the target sub-interface, use a screenshot device to capture the target sub-interface.
[0096] Figure 5 is a schematic diagram showing a template intelligent matching and restoration according to an exemplary embodiment. Figure 5 As shown, first, determine the image template to be collected, then obtain a sub-image with high similarity to the image template through the correlation coefficient matching algorithm, convert the sub-image and the image template from RGB three-channel color images to HSV images, and calculate the HSV matching confidence of the sub-image and the image template. If the HSV matching confidence is greater than or equal to the preset threshold, the sub-image is the target sub-interface.
[0097] If the HSV matching confidence is less than the preset threshold, the sub-image is blocked or exceeds the screen boundary, and the blocked part of the sub-image or the part that exceeds the screen boundary is automatically restored. First, click on the unblocked part of the image and mark it to get the restored sub-image displayed. If the HSV matching confidence of the restored sub-image is greater than or equal to the preset threshold, the restored sub-image is the target sub-interface.
[0098] If it is less than the preset threshold and has been marked (the unobstructed part has been clicked), the sub-image may exceed the screen boundary. Use the keyboard and mouse controller to drag the sub-image to the center to obtain the restored sub-image. Calculate the HSV matching confidence of the restored sub-image and the image template. If the HSV matching confidence is greater than or equal to the preset threshold, the restored sub-image is the target sub-interface. If the matching confidence of the restored sub-image is still less than the preset threshold, re-perform the correlation coefficient matching algorithm to confirm the sub-image with high correlation again.
[0099] According to this step, the target sub-interface to be collected can be intelligently acquired, and a sub-interface that is partially blocked or exceeds the screen boundary can be adjusted to become fully visible.
[0100] S103 intelligently segments the target sub-interface to obtain a single character image, and inputs the single character image into a pre-trained character classification discriminator to obtain recognized data to be collected.
[0101] Specifically, the target sub-interface is converted into a grayscale image to obtain a grayscale matrix corresponding to the target sub-interface. According to the grayscale matrix corresponding to the target sub-interface, the grayscale image is converted into a binary image. That is, a threshold is preset, and elements greater than the threshold in the grayscale matrix become 255, and elements less than the threshold become 0.
[0102] Accumulate the pixel values of each column and each row in the binary image, find the column where the accumulated pixel values in each column change suddenly, find the row where the accumulated pixel values in each row change suddenly, use the intersection of the mutation point of each column and the mutation point of each row as the segmentation point, and segment out a single character image. According to this step, all non-adhesive (single) character images can be segmented.
[0103] Input a single character image into the pre-trained character classification discriminator to obtain the recognized data to be collected. Figure 7 As shown, it is a sample display of the data required for the collection sub-interface. Through intelligent segmentation technology, all characters can be effectively located and the tables in the interface can be recognized.
[0104] Before using the character classification discriminator for recognition, it also includes training the character classification discriminator. Specifically, a large number of single-character binary image samples are obtained and labeled by a labeler, and then divided into training sets and test sets according to the ratio of 8:2. A multi-classification trainer is established, which is a convolutional neural network model. The model contains two convolutional layers, a pooling layer and a fully connected layer. And configure the training parameters, including the total number of training steps, batch size, learning rate and optimization algorithm. Use the test set to test the accuracy of the model. After the accuracy reaches 100%, save the model to obtain the trained character classification discriminator.
[0105] Figure 6 is a schematic diagram of intelligent data segmentation and recognition according to an exemplary embodiment. Figure 6 As shown, first, the character classification discriminator is trained, and when the accuracy reaches 100%, the model is saved to obtain the trained character classification discriminator. The segmented character image is input into the character classification discriminator to obtain the recognized data to be collected.
[0106] Furthermore, the data is intelligently segmented and the identified data is stored. The specific storage method is to store the data through a memory, which can store the data in a database or a local folder, including mysql, oracle, sql server, db2, postgresql, xlsx, csv, etc. According to the parameters set by the timing, the above steps are repeated regularly and output to the preset database or local folder, including mysql, oracle, sql server, db2, postgresql, xlsx, csv, etc. The database or local file can be updated in real time.
[0107] In order to facilitate understanding of the data collection method for the device interface provided in the embodiment of the present application, the following is a Figure 2 and 3 Further details are given.
[0108] like Figure 2 As shown, the data collection method provided by the embodiment of the present application includes: first setting the collection mode, including setting the number of collections (continuous collection, single collection, interval collection) and setting the collection cycle. Further, the main interface is identified and automatically presented, first the main interface is obtained through the screenshot device, then the required main interface is identified according to the pre-trained main interface discriminator, and further the required main interface is automatically presented through the keyboard and mouse controller.
[0109] Furthermore, through intelligent template matching and positioning, the target sub-interface to be captured is obtained, including determining the image template to be captured through a screenshot device, and then obtaining a sub-image with high similarity to the image template through a correlation coefficient matching algorithm, and converting the sub-image and the image template from RGB three-channel color images to HSV images through an image processor, and calculating the HSV matching confidence of the sub-image and the image template. If the HSV matching confidence is greater than or equal to a preset threshold, the sub-image is the target sub-interface. If the HSV matching confidence is less than the preset threshold, the sub-image is blocked or exceeds the screen boundary, and the blocked part of the sub-image or the part exceeding the screen boundary is automatically restored. First, click on the unobstructed part of the image through the keyboard and mouse controller and mark it to obtain the restored sub-image displayed. If the HSV matching confidence of the restored sub-image is greater than or equal to the preset threshold, the restored sub-image is the target sub-interface.
[0110] If it is less than the preset threshold and has been marked (the unobstructed part has been clicked), the sub-image may exceed the screen boundary. Use the keyboard and mouse controller to drag the sub-image to the center to obtain the restored sub-image. Calculate the HSV matching confidence of the restored sub-image and the image template. If the HSV matching confidence is greater than or equal to the preset threshold, the restored sub-image is the target sub-interface.
[0111] Furthermore, intelligent data segmentation and recognition are performed to obtain the identified data to be collected. The target sub-interface is converted into a grayscale image to obtain the grayscale matrix corresponding to the target sub-interface. According to the grayscale matrix corresponding to the target sub-interface, the grayscale image is converted into a binary image. The pixel values of each column and each row in the binary image are accumulated by the image segmentor, and the columns with the accumulated pixel values mutated in each column and the rows with the accumulated pixel values mutated in each row are found. The intersection of the mutation point of each column and the mutation point of each row is used as the segmentation point to segment a single character image.
[0112] Train the character classification discriminator. Specifically, obtain a large number of single-character binary image samples and use a labeler to mark them, and divide them into training sets and test sets according to the ratio of 8:2. Establish a multi-classification trainer, which is a convolutional neural network model. The model contains two convolutional layers, a pooling layer, and a fully connected layer. And configure the training parameters, including the total number of training steps, batch size, learning rate, and optimization algorithm. Use the test set to test the accuracy of the model. Save the model after the accuracy reaches 100% to obtain the trained character classification discriminator.
[0113] The segmented character images are recognized by the trained character classification discriminator.
[0114] Finally, the identified data is stored. The data can be stored in a database or local folder, including mysql, oracle, sql server, db2, postgresql, xlsx, csv, etc.
[0115] Figure 3 It is a flow chart of intelligent interface recognition and data collection method, such as Figure 3 As shown, the device interface recognition and data collection method provided in the embodiment of the present application includes five steps: timing setting, intelligent recognition and presentation of the main interface, intelligent matching positioning and restoration of templates, intelligent data segmentation and recognition, and data storage.
[0116] The data collection method provided in the embodiment of the present application is different from the usual data collection method, which requires manual click to place the visual interface in the foreground, and cannot handle the situation where the sub-interface in the interface is manually moved. The present invention creatively uses the main interface matching and template matching technology to intelligently present the main interface to be collected in the foreground of the screen, and can adjust the partially blocked sub-interface to be fully visible.
[0117] Different from the existing OCR optical character recognition which has the problem of low recognition accuracy for low-pixel-level data, the present invention creatively uses data intelligent segmentation technology, which can effectively locate all characters, and combines data intelligent recognition technology to identify character categories, achieving 100% accuracy, greatly improving work efficiency and significantly reducing the error rate of handwriting.
[0118] Compared with traditional RPA automated robot processes, this method does not affect the normal use of the equipment, allows changes and displacements of the visual interface, can monitor data changes in real time, automatically pause when the operator uses the equipment, and automatically restart after use without terminating the process.
[0119] The present application also provides a device interface data acquisition device, which is used to execute the device interface data acquisition method of the above embodiment, such as Figure 8 As shown, the device comprises:
[0120] The automatic interface recognition and presentation module 801 is used to recognize the required main interface according to the pre-trained main interface discriminator, and automatically present the required main interface on the device to be collected;
[0121] The target sub-interface determination module 802 is used to determine the image template to be collected, locate the target sub-interface similar to the image template on the required main interface, and automatically restore the blocked part of the target sub-interface or the part beyond the screen boundary;
[0122] The interface data acquisition module 803 is used to intelligently segment the target sub-interface to obtain a single character image, and input the single character image into a pre-trained character classification discriminator to obtain recognized data to be collected.
[0123] It should be noted that the data acquisition device for the device interface provided in the above embodiment only uses the division of the above functional modules as an example when executing the data acquisition method for the device interface. In actual applications, the above functional allocation can be completed by different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the data acquisition device for the device interface provided in the above embodiment and the data acquisition method embodiment for the device interface belong to the same concept, and the implementation process thereof is detailed in the method embodiment, which will not be repeated here.
[0124] The embodiment of the present application also provides an electronic device corresponding to the device interface data collection method provided in the above embodiment, so as to execute the above device interface data collection method.
[0125] Please refer to Fig. 9 , which shows a schematic diagram of an electronic device provided by some embodiments of the present application. Fig. 9 As shown, the electronic device includes: a processor 900, a memory 901, a bus 902 and a communication interface 903, and the processor 900, the communication interface 903 and the memory 901 are connected via the bus 902; the memory 901 stores a computer program that can be run on the processor 900, and when the processor 900 runs the computer program, it executes the data acquisition method of the device interface provided in any of the aforementioned embodiments of the present application.
[0126] The memory 901 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 903 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used.
[0127] The bus 902 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory 901 is used to store programs, and the processor 900 executes the program after receiving the execution instruction. The data acquisition method of the device interface disclosed in any implementation of the embodiment of the present application may be applied to the processor 900, or implemented by the processor 900.
[0128] The processor 900 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 900. The above processor 900 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a readily available programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor to be executed, or the hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 901, and the processor 900 reads the information in the memory 901 and completes the steps of the above method in combination with its hardware.
[0129] The electronic device provided in the embodiment of the present application and the data collection method of the device interface provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, operated or implemented therein.
[0130] The present application also provides a computer-readable storage medium corresponding to the device interface data collection method provided in the above embodiment. Fig.10 The computer-readable storage medium shown is a CD 1000 on which a computer program (ie, a program product) is stored. When the computer program is run by a processor, the device interface data acquisition method provided in any of the aforementioned embodiments will be executed.
[0131] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.
[0132] The computer-readable storage medium provided in the above-mentioned embodiments of the present application and the data acquisition method of the device interface provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0133] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0134] The above embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the attached claims.
Claims
1. A data collection method for a device interface, characterized in that: include: Identify the required main interface according to the pre-trained main interface discriminator, and automatically present the required main interface on the device to be collected; Determine the image template to be collected, locate the target sub-interface similar to the image template on the required main interface, and automatically restore the blocked part of the target sub-interface or the part beyond the screen boundary; including: intercepting the image template to be collected on the required main interface and obtaining the length and width of the image template; obtaining the similarity between each position in the required main interface and the image template according to the required main interface, the image template and the correlation coefficient matching algorithm; taking the position coordinates with the greatest similarity as the starting coordinates, and taking the starting coordinates as the starting point, intercepting a sub-image with the same length and width as the image template on the required main interface; calculating the HSV matching confidence of the sub-image and the image template, if the HSV matching confidence is greater than or equal to a preset threshold, the sub-image is the target sub-interface; if the HSV matching confidence is less than the preset threshold, the sub-image is blocked or exceeds the screen boundary, and the blocked part of the sub-image or the part beyond the screen boundary is automatically restored, and if the HSV matching confidence of the restored sub-image is greater than or equal to the preset threshold, the restored sub-image is the target sub-interface; The target sub-interface is intelligently segmented to obtain a single character image, and the single character image is input into a pre-trained character classification discriminator to obtain recognized data to be collected.
2. The method according to claim 1, characterized in that Identify the required main interface according to the pre-trained main interface discriminator, and automatically present the required main interface on the device to be collected, including: Intercept the main interface of the device to be collected; Input the intercepted main interface into a pre-trained main interface discriminator to determine whether the intercepted main interface is the desired main interface; When the captured main interface is the desired main interface, the desired main interface is automatically presented; When the captured main interface is a desktop interface, automatically searching and presenting the required main interface through a keyboard and mouse controller; When the captured main interface is an interface of other software, the interface is switched to the desktop interface through the keyboard and mouse controller, and then the keyboard and mouse controller is used to automatically search and present the required main interface.
3. The method according to claim 1, characterized in that Before identifying the required main interface according to the pre-trained main interface discriminator, it also includes: Intercepting a preset number of main interfaces, wherein the preset number of main interfaces include a required main interface, a desktop interface, and other software interfaces; Dividing the preset number of main interfaces into a training set, a test set, and a validation set; The main interface discriminator is trained according to the training set, the test set, the validation set and the classification neural network model to obtain a trained main interface discriminator.
4. The method according to claim 1, characterized in that: According to the required main interface, the image template and the correlation coefficient matching algorithm, the similarity between each position in the required main interface and the image template is obtained, including: Grayscale processing is performed on the required main interface and image template to obtain grayscale matrices respectively; The starting point coordinates of the required main interface, the starting point coordinates of the image template, the grayscale matrix of the required main interface and the grayscale matrix of the image template are input into the correlation coefficient matching algorithm to obtain the similarity between each position in the required main interface and the image template.
5. The method according to claim 1, characterized in that Automatically restore the blocked part of the sub-image or the part beyond the screen boundary. If the HSV matching confidence of the restored sub-image is greater than or equal to a preset threshold, the restored sub-image is the target sub-interface, including: Click the unobstructed part of the sub-image by using the keyboard and mouse controller to obtain the restored sub-image; Calculating the HSV matching confidence between the restored sub-image and the image template, if the HSV matching confidence is greater than or equal to a preset threshold, the restored sub-image is the target sub-interface; If the HSV matching confidence is less than a preset threshold, the required main interface is evenly divided into four quadrants with the center of the required main interface as the origin, and the sub-image is dragged by a keyboard and mouse controller to move a preset distance in the quadrant direction symmetrical to the center of the quadrant where it is located to obtain a restored sub-image; The HSV matching confidence between the restored sub-image and the image template is calculated. If the HSV matching confidence is greater than or equal to a preset threshold, the restored sub-image is the target sub-interface.
6. The method according to claim 1, characterized in that Intelligently segmenting the target sub-interface to obtain a single character image includes: Converting the target sub-interface into a grayscale image to obtain a grayscale matrix corresponding to the target sub-interface; According to the grayscale matrix corresponding to the target sub-interface, the grayscale image is converted into a binary image; The pixel values of each column and each row in the binary image are accumulated, and the intersection of the mutation point of the accumulated pixel value of each column and the mutation point of the pixel value of each row is used as a segmentation point to segment a single character image.
7. A data acquisition device for a device interface, characterized in that: include: The automatic interface recognition and presentation module is used to recognize the required main interface according to the pre-trained main interface discriminator, and automatically present the required main interface on the device to be collected; The target sub-interface determination module is used to determine the image template to be collected, locate the target sub-interface similar to the image template on the desired main interface, and automatically restore the blocked part of the target sub-interface or the part beyond the screen boundary; including: intercepting the image template to be collected on the desired main interface and obtaining the length and width of the image template; obtaining the similarity between each position in the desired main interface and the image template according to the desired main interface, the image template and the correlation coefficient matching algorithm; taking the position coordinates with the greatest similarity as the starting coordinates, and taking the starting coordinates as the starting point, intercepting a sub-image with the same length and width as the image template on the desired main interface; calculating the HSV matching confidence of the sub-image and the image template, if the HSV matching confidence is greater than or equal to a preset threshold, the sub-image is the target sub-interface; if the HSV matching confidence is less than the preset threshold, the sub-image is blocked or exceeds the screen boundary, and the blocked part of the sub-image or the part beyond the screen boundary is automatically restored, and if the HSV matching confidence of the restored sub-image is greater than or equal to the preset threshold, the restored sub-image is the target sub-interface; The interface data acquisition module is used to intelligently segment the target sub-interface to obtain a single character image, and input the single character image into a pre-trained character classification discriminator to obtain recognized data to be collected.
8. A data acquisition device for a device interface, characterized in that: It comprises a processor and a memory storing program instructions, wherein the processor is configured to execute the data acquisition method for the device interface as described in any one of claims 1 to 6 when executing the program instructions.
9. A computer-readable medium, characterized in that Computer-readable instructions are stored thereon, and the computer-readable instructions are executed by a processor to implement a data collection method for a device interface as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Window management method and system
CN104537221A
Target occlusion intensity evaluation method in image identification and tracking
CN106023250A