Data extraction RPA robot manufacturing method based on position positioning and medium

By obtaining and identifying preset fields in screenshots, identifying their content data locations, and creating RPA robots, the problem of heavy workload in the development of RPA robots is solved, and automated data extraction and cost reduction are achieved.

CN120279561AActive Publication Date: 2025-07-08四川互慧软件有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510779267.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-08
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

During the development of RPA robots, customized development is required for different manufacturers and data source software versions, resulting in a sharp increase in manual workload. How to reduce the workload of RPA robot production has become an urgent problem.

Method used

By obtaining screenshots with preset fields, identifying the preset fields based on the screenshot, identifying the location of the content data corresponding to the preset fields, and creating an RPA robot, including the screenshot ID, the corresponding preset fields and the location of the content data.

Benefits of technology

It reduces the workload of RPA robot production, greatly reduces labor costs, and improves the efficiency of automatic positioning and data extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279561A_ABST
    Figure CN120279561A_ABST
Patent Text Reader

Abstract

The invention relates to a data extraction RPA robot manufacturing method based on position positioning and a medium, and relates to the technical field of computer data processing. A screenshot with a preset field is obtained, the preset field is recognized based on the screenshot, a screenshot ID corresponding to the preset field is obtained, and the preset field is obtained through recognition based on the obtained screenshot. And based on the identified preset field, identifying the position of the content data corresponding to the preset field, and obtaining the position of the data content corresponding to the to-be-extracted preset field. Therefore, the RPA robot can be manufactured based on the screenshot ID, the corresponding preset field and the position of the content data corresponding to the corresponding preset field, so that the data extraction RPA robot based on position positioning is obtained, the manufacturing workload of the RPA robot is reduced, and the labor cost is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer data processing, and particularly relates to a method and medium for manufacturing a data extraction RPA robot based on location positioning. Background Art

[0002] Currently, after stroke inpatients are discharged from the hospital, the hospital must upload 357 treatment data indicators during the hospitalization of stroke patients to the specified stroke database in a timely manner for further quality control and data analysis of stroke cases, optimize treatment measures, improve the treatment rate, and reduce the disease burden.

[0003] With the rapid development of information technology and the wide application of artificial intelligence, it is necessary to build a software robot for automatic extraction, automatic review, and automatic reporting of stroke center data based on RPA (Robotic Process Automation) technology, which is targeted at thousands of stroke centers in hospitals across the country to assist hospitals in reporting stroke data. However, during the RPA development process, customized development needs to be carried out for different vendors and / or software versions of data sources. There are many versions of the same system in each company, and a new robot needs to be made for each version of the software, resulting in a sharp increase in the workload of manual RPA production. How to reduce the workload of RPA robot production has become an urgent problem to be solved. Summary of the Invention

[0004] The technical problem to be solved by this application is to provide a method and medium for manufacturing a data extraction RPA robot based on location positioning, which has the characteristics of reducing the workload of RPA robot production and reducing labor costs.

[0005] In a first aspect, in one embodiment, a method for manufacturing a data extraction RPA robot based on location positioning is provided, including: Obtain a screenshot with preset fields and identify the preset fields based on the screenshot; Based on the identified preset fields, identify the positions where the content data corresponding to the preset fields is located; Manufacture an RPA robot, including the screenshot ID, the corresponding preset fields, and the positions where the content data corresponding to the preset fields is located.

[0006] In one embodiment, the obtaining of the screenshot with preset fields includes: Obtain a screenshot and, based on the obtained screenshot, identify the preset fields existing in the screenshot through OCR to obtain a screenshot with preset fields.

[0007] In one embodiment, the obtaining of the screenshot with preset fields includes: Based on the correspondence between the preset fields and the screenshot, obtain the screenshot where the content data corresponding to the preset fields needs to be extracted.

[0008] In one embodiment, identifying the location where the content data corresponding to the identified preset field is located includes: For any identified preset field, identify whether there is a first text box in which the first text content only includes the preset field, and the first text content does not include punctuation marks; If there is a first text box, search to the right in the first same content area for an adjacent second text box on the right. If there is an adjacent second text box on the right, take the location where the second text box is located as the location where the content data of any one of the preset fields is located; if there is no adjacent second text box on the right, search downward for an adjacent second text box below that is aligned with the left end of the first text box, and take the location where the adjacent second text box below is located as the location where the content data of any one of the preset fields is located; the location where the second text box is located includes: the upper left corner coordinate point and the lower right corner coordinate point of the second text box, or the lower left corner coordinate point and the upper right corner coordinate point of the second text box; the first same content area refers to an area within a first preset distance threshold from the right edge of the first text box.

[0009] If there is no first text box, identify the first start coordinate point and the first end coordinate point of the preset field; based on the first end coordinate point, search to the right in the second same content area for an adjacent first character, and determine whether the adjacent first character is a preset separator character. If so, based on the second end coordinate point of the first character, search to the right in the second same content area for an adjacent second character on the right. If the second character is found, take the start coordinate position of the found second character as the start coordinate position where the content data of any one of the preset fields is located; if the first character is not a preset separator character, take the start coordinate position of the first character as the start coordinate position where the content data of any one of the preset fields is located; if the first character or the second character is not found, search downward for an adjacent third character to the right with the X-axis coordinate of the first start coordinate point as the start abscissa of the next line, and take the start coordinate position of the third character as the start coordinate position where the content data of any one of the preset fields is located; the second same content area refers to a preset field equal-height area within a second preset distance threshold from the based coordinate point, and the separator characters include ":" and / or "—".

[0010] In one embodiment, in the presence of a first text box, the method of creating an RPA robot includes a screenshot ID, corresponding preset fields, and the location of the content data corresponding to the preset fields, including: creating an RPA robot, including a screenshot ID, corresponding preset fields, a second text box content extraction algorithm, and the location of the second text box corresponding to the preset fields.

[0011] In one embodiment, in the absence of a first text box, the method of creating an RPA robot includes a screenshot ID, corresponding preset fields, and the location of the content data corresponding to the preset fields, including: creating an RPA robot, including a screenshot ID, corresponding preset fields, the starting coordinate position of the content data corresponding to the preset fields, a text dynamic extension recognition algorithm, and semantic recognition; The text dynamic extension recognition includes extending and recognizing adjacent characters to the right and downward based on the starting coordinate position of the content data, and determining whether the adjacent characters belong to the content data to be extracted based on semantic recognition, so as to obtain the content data to be extracted.

[0012] In one embodiment, the starting coordinate position of the first character includes the upper left corner coordinate position and / or the lower left corner coordinate position of the first character; the starting coordinate position of the second character includes the upper left corner coordinate position and / or the lower left corner coordinate position of the second character; the starting coordinate position of the third character includes the upper left corner coordinate position and / or the lower left corner coordinate position of the third character.

[0013] In one embodiment, the first starting coordinate point is the upper left corner of the starting position, and the first ending coordinate point is the lower right corner of the ending position; or, the first starting coordinate point is the lower left corner of the starting position, and the first ending coordinate point is the upper right corner of the ending position.

[0014] In one embodiment, the screenshot ID retains the serial number, name, or other unique identifier of the screenshot.

[0015] In a second aspect, in one embodiment, a computer-readable storage medium is provided, in which a program is stored, and the program can be loaded and executed by a processor to perform the method for creating a data extraction RPA robot according to any one of the above embodiments.

[0016] The beneficial effects of the present invention are: By obtaining screenshots with preset fields and identifying the preset fields based on the screenshots, it is possible to obtain the screenshot IDs corresponding to the preset fields and identify the preset fields from the recognized screenshots. Since the location of the content data corresponding to the preset fields is identified based on the recognized preset fields, it is possible to obtain the location of the data content corresponding to the preset fields to be extracted. Thus, an RPA robot can be made based on the screenshot ID, the corresponding preset fields, and the location of the content data corresponding to the preset fields, thereby obtaining an RPA robot for data extraction based on location positioning, reducing the workload of making the RPA robot and greatly reducing the labor cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a schematic flowchart of a method for making an RPA robot for data extraction based on location positioning according to an embodiment of the present application; Figure 2 is a schematic diagram of the arrangement of field names and content data according to the first embodiment of the present application; Figure 3 is a schematic diagram of the arrangement of field names and content data according to the second embodiment of the present application; Figure 4 is a schematic diagram of the arrangement of field names and content data according to the third embodiment of the present application; Figure 5 is a schematic diagram of the arrangement of field names and content data according to the fourth embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The present invention will be further described in detail below in conjunction with the drawings through specific embodiments. Similar elements in different embodiments are labeled with related similar element numbers. In the following embodiments, many details are described to enable a better understanding of the present application. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other elements, materials, or methods. In some cases, some operations related to the present application are not shown or described in the specification to avoid the core part of the present application being overwhelmed by excessive description. For those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations based on the description in the specification and general technical knowledge in the art.

[0019] In addition, the features, operations or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can also be reordered or adjusted in a manner obvious to those skilled in the art. Therefore, the various sequences in the specification and drawings are only for clearly describing a certain embodiment and do not mean a necessary sequence, unless it is stated that a certain sequence must be followed.

[0020] The serial numbers assigned to the components herein, such as "first", "second", etc., are only used to distinguish the described objects and do not have any sequential or technical meaning.

[0021] In the current data extraction of stroke centers, for each version of the data source software, the data extraction process needs to be manually operated. There are a total of 357 fields required by stroke centers, and the data of these fields comes from various different systems. It is necessary to use RPA to create robots for the collection of each field. It is very likely that the information systems used by different hospitals are provided by different information companies, and these companies may also have multiple versions of the same system. There are more than a hundred HIS software companies across the country, resulting in extremely high labor costs in finding which software interface a field is in. If RPA can also automatically locate the software system and interface where the field is located, then the labor cost will be greatly reduced.

[0022] The applicant found in the research that if RPA is to automatically locate the software system and interface where the field is located, first finding the screenshot of the software interface where the field is located and automatically locating from the screenshot is a feasible option. However, to further improve work efficiency and reduce the workload of creating RPA robots, it becomes an urgent problem to implement automatic location of data related to preset fields based on the obtained screenshot to obtain an RPA robot for extracting data through automatic location.

[0023] In view of this, the embodiments of the present application provide a method and medium for creating an RPA robot for data extraction based on location positioning. By obtaining a screenshot with a preset field and identifying the preset field based on the screenshot, the screenshot ID corresponding to the preset field can be obtained, and the preset field can be recognized from the obtained screenshot. Based on the recognized preset field, the location where the content data corresponding to the preset field is located can be recognized, and the location where the data content corresponding to the preset field to be extracted is located can be obtained. Thus, an RPA robot can be created based on the screenshot ID, the corresponding preset field, and the location where the content data corresponding to the preset field is located, thereby obtaining an RPA robot for data extraction based on location positioning, reducing the workload of creating RPA robots, and greatly reducing the labor cost.

[0024] The method for manufacturing a location-based data extraction RPA robot according to an embodiment of the present application can be applied to a computer system, which includes a processor and a computer storage medium. At least one processor is used to load and execute a program of the method for manufacturing a location-based data extraction RPA robot and perform processing.

[0025] For a method for manufacturing a location-based data extraction RPA robot according to an embodiment of the present application, please refer to Figure 1 , including: Step S10, obtain a screenshot with a preset field, and identify the preset field based on the screenshot.

[0026] Through the obtained screenshot with a preset field, the preset field existing in the screenshot can be identified based on this screenshot first.

[0027] In one embodiment, when the correspondence between the preset field and the screenshot is not clear, in order to obtain a screenshot with a preset field, the screenshot can be obtained first, and then based on the obtained screenshot, the preset field existing in the screenshot can be obtained through OCR recognition to obtain a screenshot with a preset field. For example, it may be known that there is a preset field on a certain screenshot, but it is not known which preset fields exist, or it may not be known whether there is a preset field on the screenshot, but it can be obtained through OCR recognition whether there is a preset field on the screenshot and which preset fields exist.

[0028] In one embodiment, when the correspondence between the preset field and the screenshot is clear, in order to obtain a screenshot with a preset field, the screenshot where the content data corresponding to the preset field to be extracted is located can be directly obtained. For example, if both preset field A and preset field B have a correspondence with screenshot M, if the content data corresponding to preset field A is to be extracted, then based on this correspondence, screenshot M can be directly obtained.

[0029] Step S20, based on the identified preset field, identify the location where the content data corresponding to the preset field is located.

[0030] Based on the identified preset field, by identifying the location where the content data corresponding to the preset field is located, the location where the data content corresponding to the preset field to be extracted is located can be obtained.

[0031] The applicant found in the research that for the preset field to be identified and the corresponding content data, there may be respective corresponding text boxes, or there may be no corresponding text boxes. Please refer to Figures 2 to 4 , both of which are expression forms of the patient's basic information. Among them, name and gender are preset fields, Zhang San is the content data of the name, and male is the content data of the gender. Figure 2 and Figure 3In the case where there are preset fields and their corresponding content data, and there are respective corresponding text boxes, Figure 4 In the case where there are preset fields and their corresponding content data, and there are no respective corresponding text boxes. Therefore, how to obtain a solution that takes into account the above two cases becomes a technical difficulty. In view of this problem, in an embodiment of the present application, step S20 includes: For any one of the identified preset fields, such as "Name", identify whether there is a first text box in which the first text content only includes the preset field, where the first text content does not include punctuation marks. If so, Figure 2 and Figure 3 As shown, if there is a first text box, search to the right in the first same content area for an adjacent second text box on the right. The first same content area refers to an area within a first preset distance threshold range from the right edge of the first text box. For example, an area within a first preset distance threshold range from the right end of the first text box of "Name". If as Figure 2 shown, the border of an adjacent second text box on the right is found by searching to the right in the first same content area, it is considered that there is an adjacent second text box on the right, and the position where the second text box is located is used as the position where the content data of any one of the above preset fields (such as "Name") is located.

[0032] If as Figure 3 shown, the border of an adjacent second text box on the right is not found by searching in the first same content area, it means that the second text box is not on the right side of the first text box but below the first text box. Then search downward for an adjacent second text box below that is aligned with the left end of the first text box, and use the position where the adjacent second text box below is located as the position where the content data of any one of the above preset fields is located.

[0033] In an embodiment, the position where the second text box is located can be represented as diagonal coordinate points, or the coordinate points of three or four corners. In the case of diagonal coordinate points, it can be the upper left corner coordinate points and the lower right corner coordinate points, or the lower left corner coordinate points and the upper right corner coordinate points of the second text box.

[0034] It can be understood that for those skilled in the art, the first preset distance threshold can be set according to actual needs. In an embodiment, the first preset distance threshold is 10px.

[0035] If there is no first text box, it may be Figure 4In the case shown, in the absence of the first text box, the content data corresponding to the preset field may be on the right side of the preset field or below the preset field. Therefore, in one embodiment, if the first text box does not exist, the first starting coordinate point and the first ending coordinate point of the preset field are identified. For example, the starting coordinate point of "surname" in the "Name" preset field and the ending coordinate point of "given name" are identified. In one embodiment, taking "Name" as an example, the first starting coordinate point is the upper left corner of the starting position, that is, the upper left corner of "surname", and the first ending coordinate point is the lower right corner of the ending position, that is, the lower right corner of "given name". In one embodiment, the first starting coordinate point is the lower left corner of the starting position, that is, the lower left corner of "surname", and the first ending coordinate point is the upper right corner of the ending position, that is, the upper right corner of "given name".

[0036] First, it is determined whether there is corresponding content data on the right side, that is, based on the first ending coordinate point, the first adjacent character on the right is searched in the second same content area, and it is determined whether the first adjacent character on the right is a preset separator character. If so, based on the second ending coordinate point of the first character, the second adjacent character on the right is searched in the second same content area. If the second character is found, the starting coordinate position of the found second character is used as the starting coordinate position where the content data of any of the above preset fields is located, and the starting coordinate position of the content data corresponding to the preset field is found. The separator characters include ":" (colon) and / or "—" (dash). The applicant found in the research that since there may be ":" or "—" following the preset field, in this case, the first character recognized is not the corresponding content data. Therefore, it is necessary to further search to the right in the same way to see if there is a second adjacent character on the right. If the second character is found, the starting coordinate position of the found second character is used as the starting coordinate position where the content data of any of the above preset fields is located. If the first character is not a preset separator character, the starting coordinate position of the first character is used as the starting coordinate position where the content data of any of the above preset fields is located. The second same content area refers to the preset field equal-height area within the second preset distance threshold from the based coordinate point. For example, when searching for the first character above, the first ending coordinate point is based, and at this time, the second same content area refers to the preset field equal-height area within the second preset distance threshold from the right of the first ending coordinate point. When searching for the second character, the second ending coordinate point is based, and at this time, the second same content area refers to the preset field equal-height area within the second preset distance threshold from the right of the second ending coordinate point.

[0037] It can be understood that for those skilled in the art, the second preset distance threshold can be set according to actual needs. In one embodiment, the second preset distance threshold is 10px.

[0038] In some embodiments, the starting coordinate position of the first character includes the upper left corner coordinate position and / or the lower left corner coordinate position of the first character; the starting coordinate position of the second character includes the upper left corner coordinate position and / or the lower left corner coordinate position of the second character.

[0039] If the first character or the second character is not found, it means that the corresponding content data is not on the right side of the preset field, but below. Then, use the X-axis coordinate of the first starting coordinate point as the starting abscissa of the next line to search for the adjacent third character below to the right, and use the starting coordinate position of the third character as the starting coordinate position where the content data of any of the above preset fields is located. Since the starting position of the content data and the starting position of the preset field may not be aligned vertically, therefore, use the X-axis coordinate of the first starting coordinate point as the starting abscissa of the next line to search for the adjacent third character below to the right.

[0040] In one embodiment, the starting coordinate position of the third character includes the upper left corner coordinate position and / or the lower left corner coordinate position of the third character.

[0041] Step S30, make an RPA robot, including the screenshot ID, the corresponding preset field, and the position where the content data corresponding to the corresponding preset field is located.

[0042] In one embodiment, making an RPA robot includes the screenshot ID, the corresponding preset field, the algorithm for extracting content based on the second text box content, and the position where the second text box corresponding to the corresponding preset field is located. Based on the description of the above step S20, for example, as Figure 2 and Figure 3 shown, based on the obtained screenshot ID, the corresponding preset field "Name", the algorithm for extracting content based on the second text box content, that is, the algorithm for extracting all content data in the second text box based on the second text box where "Zhang San" is located, and the position where the second text box corresponding to the corresponding preset field is located, that is, the position where the second text box where "Zhang San" is located, to make a data extraction RPA corresponding to the content data of the preset field "Name". In this way, the production of the data extraction RPA robot is realized.

[0043] In one embodiment, in the case where the first text box does not exist, an RPA robot can be made based on the screenshot ID, the corresponding preset field, the starting coordinate position where the content data corresponding to the corresponding preset field is located, the text dynamic extension recognition algorithm, and semantic recognition. Among them, text dynamic extension recognition includes extending and recognizing adjacent characters to the right and downward based on the starting coordinate position where the content data is located, and judging whether the adjacent characters belong to the content data to be extracted based on semantic recognition, so as to obtain the content data to be extracted. For example, please refer to Figure 5, based on the obtained screenshot ID and the corresponding preset field "department name", there is a second character on the right. Therefore, the starting coordinate position of the second character "new" is used as the starting coordinate position where the content data is located, and adjacent characters are identified by extending to the right and downwards. In one embodiment, the rightward extension is based on the end coordinate point of the current character, and it searches for adjacent characters on the right in the second same content area until the end. In Figure 5 the shown embodiment, the rightward extension identifies until "three" ends. The downward extension is based on the starting coordinate position of the content data, that is, the X-axis coordinate of the starting coordinate position of "new" is used as the starting abscissa of the next line to search for the third adjacent character below. Combining the above rightward and downward extensions to identify adjacent characters, and semantic recognition to obtain the content data to be extracted. In Figure 5 the embodiment, after the RPA robot made obtains the starting coordinate position of the content data, it can obtain the content data to be extracted, "Neonatal and Maternal and Child Health Department", based on the text dynamic extension recognition algorithm and semantic recognition plugin set inside the RPA robot. Those skilled in the art can understand that since the semantic recognition model can adopt existing technology models and can be used as a plugin to be called during the execution of the RPA robot, the specific semantic recognition algorithm will not be elaborated.

[0044] Based on the above method for making a location-based data extraction RPA robot in one embodiment, by obtaining a screenshot with a preset field and identifying the preset field based on the screenshot, the screenshot ID corresponding to the preset field can be obtained, and the preset field can be recognized from the screenshot obtained. Based on the recognized preset field, the location where the content data corresponding to the preset field is located can be recognized, and the location where the data content corresponding to the preset field to be extracted is located can be obtained. Thus, an RPA robot can be made based on the screenshot ID, the corresponding preset field, and the location where the content data corresponding to the preset field is located, thereby obtaining a location-based data extraction RPA robot, reducing the workload of making the RPA robot, and greatly reducing the labor cost.

[0045] In one embodiment of the present application, a computer-readable storage medium is provided, and a program is stored on the storage medium. The stored program includes the method that can be loaded and processed by a processor in any of the above embodiments.

[0046] Those skilled in the art can understand that all or part of the functions of the various methods in the above embodiments can be implemented in a hardware manner or in a computer program manner. When all or part of the functions in the above embodiments are implemented in a computer program manner, the program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory, random access memory, magnetic disk, optical disk, hard disk, etc. The above functions can be realized by a computer executing the program. For example, the program is stored in the memory of the device, and when the processor executes the program in the memory, the above all or part of the functions can be realized. In addition, when all or part of the functions in the above embodiments are implemented in a computer program manner, the program can also be stored in a storage medium such as a server, another computer, magnetic disk, optical disk, flash drive or mobile hard disk, and saved to the memory of the local device by downloading or copying, or the system of the local device is updated with a version. When the processor executes the program in the memory, all or part of the functions in the above embodiments can be realized.

[0047] The above uses specific examples to elaborate on the present invention, which is only used to help understand the present invention and is not intended to limit the present invention. For those skilled in the art of the present invention, according to the idea of the present invention, several simple deductions, deformations or substitutions can also be made.

Claims

1. A method for manufacturing a data extraction RPA robot based on position positioning, characterized in that, Including: Obtain a screenshot with a preset field and identify the preset field based on the screenshot; Based on the identified preset field, identify the location where the content data corresponding to the preset field is located; Create an RPA robot, including the screenshot ID, the corresponding preset field, and the location where the content data corresponding to the corresponding preset field is located.

2. The method for manufacturing a data extraction RPA robot according to claim 1, wherein, The obtaining of the screenshot with a preset field includes: Obtain a screenshot and, based on the obtained screenshot, identify the preset field existing in the screenshot through OCR to obtain a screenshot with a preset field.

3. The method for manufacturing a data extraction RPA robot according to claim 1, wherein, The obtaining of the screenshot with a preset field includes: Based on the correspondence between the preset field and the screenshot, obtain the screenshot where the content data corresponding to the preset field needs to be extracted.

4. The method for manufacturing a data extraction RPA robot according to claim 1, wherein, The identifying of the location where the content data corresponding to the identified preset field is located includes: For any identified preset field, identify whether there is a first text box whose first text content only includes the preset field, and the first text content does not include punctuation marks; If there is a first text box, search to the right in the first same content area for an adjacent second text box on the right. If there is an adjacent second text box on the right, use the location where the second text box is located as the location where the content data of any preset field is located; if there is no adjacent second text box on the right, search downward for an adjacent second text box below that is aligned with the left end of the first text box, and use the location where the adjacent second text box below is located as the location where the content data of any preset field is located; the location where the second text box is located includes: the upper left corner coordinate point and the lower right corner coordinate point of the second text box, or the lower left corner coordinate point and the upper right corner coordinate point of the second text box; the first same content area refers to the area within a first preset distance threshold from the right edge of the first text box. If there is no first text box, identify the first starting coordinate point and the first ending coordinate point of the preset field; based on the first ending coordinate point, search to the right in the second same content area for an adjacent first character, and determine whether the adjacent first character is a preset separator character. If so, based on the second ending coordinate point of the first character, search to the right in the second same content area for an adjacent second character on the right. If the second character is found, use the starting coordinate position of the found second character as the starting coordinate position where the content data of any preset field is located; if the first character is not a preset separator character, use the starting coordinate position of the first character as the starting coordinate position where the content data of any preset field is located; if the first character or the second character is not found, use the X-axis coordinate of the first starting coordinate point as the starting abscissa of the next line to search for an adjacent third character below, and use the starting coordinate position of the third character as the starting coordinate position where the content data of any preset field is located; the second same content area refers to the area with the same height as the preset field within a second preset distance threshold from the based coordinate point, and the separator character includes ":" and / or "—".

5. The method for manufacturing a data extraction RPA robot according to claim 4, wherein In the case of the existence of the first text box, the RPA robot production includes a screenshot ID, corresponding preset fields, and the location of the content data corresponding to the corresponding preset fields, including: the production of the RPA robot, including a screenshot ID, corresponding preset fields, a second text box content extraction algorithm, and the location of the second text box corresponding to the corresponding preset fields.

6. The method for manufacturing a data extraction RPA robot according to claim 4, wherein, In the case of the non-existence of the first text box, the RPA robot production includes a screenshot ID, corresponding preset fields, and the location of the content data corresponding to the corresponding preset fields, including: the production of the RPA robot, including a screenshot ID, corresponding preset fields, the starting coordinate position of the content data corresponding to the corresponding preset fields, a text dynamic extension recognition algorithm, and semantic recognition; The text dynamic extension recognition includes performing rightward and downward extension recognition of adjacent characters based on the starting coordinate position of the content data, and judging whether the adjacent characters belong to the content data to be extracted based on semantic recognition, so as to obtain the content data to be extracted.

7. The method for manufacturing a data extraction RPA robot according to claim 4, characterized in that, The starting coordinate position of the first character includes the upper left corner coordinate position and / or the lower left corner coordinate position of the first character; the starting coordinate position of the second character includes the upper left corner coordinate position and / or the lower left corner coordinate position of the second character; the starting coordinate position of the third character includes the upper left corner coordinate position and / or the lower left corner coordinate position of the third character.

8. The method for manufacturing a data extraction RPA robot according to claim 4, wherein The first starting coordinate point is the upper left corner of the starting position, and the first ending coordinate point is the lower right corner of the ending position; or, the first starting coordinate point is the lower left corner of the starting position, and the first ending coordinate point is the upper right corner of the ending position.

9. The method for manufacturing a data extraction RPA robot according to claim 5 or 6, characterized in that, The screenshot ID retains the serial number, name, or other unique identification code of the screenshot.

10. A computer-readable storage medium, characterized in that, A program is stored in the medium, and the program can be loaded and executed by a processor to perform the data extraction RPA robot production method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Text recognition method, electronic equipment and storage medium

    CN116580416A

  • Stroke data robot manufacturing method, device and equipment and storage medium

    CN119166114A