Business carrier automatic registration and login and fund account acquisition method and system
Through the visually driven UI recognition mechanism and object detection model, the problem of automation operation robustness under dynamic pages is solved, the accuracy of cross-frame elements and the intelligent processing of complex verification codes is realized, and the automation operation robustness of dynamic pages and the complete rate of extracting key account information is improved.
Patent Information
- Application Number
- CN202510747789.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-06
AI Technical Summary
When existing automated web page interaction technology deals with dynamic page environments and composite interaction scenarios, there are problems with page structure adaptability and cross-frame operation compatibility, making it difficult to accurately calculate the location of elements in the subframe, and traditional tools lack processing capabilities in verification code scenarios.
The visually driven UI recognition mechanism is adopted to analyze the page image through the object detection model, establish a dynamic conversion chain of the coordinate system, and generate an interactive task queue in combination with preset priority rules to realize cross-frame element operations and intelligent processing of complex verification codes, and use regular expressions to extract payment account information.
It improves the robustness of automation operations under dynamic pages, solves the process interruption problem caused by path failure in traditional scripts, ensures the accuracy of cross-frame element operations and the recognition rate of complex verification codes, and improves the complete rate of extracting key account information.
Smart Images

Figure CN120256010A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method and system for automatic registration and login of service carriers and acquisition of fund accounts. Background Art
[0002] In the field of automated web interaction technologies, there is still room for optimization in existing solutions when dealing with dynamic page environments and complex interaction scenarios, which are mainly manifested in the following aspects: Limitations in page structure adaptability. Current mainstream automated tools (such as script engines based on DOM parsing) usually rely on static element positioning paths (such as XPath / CSS selectors). When the target page exhibits: Common changes such as interface layout version iteration, dynamically generated element identifiers, or adaptive adjustment of the positions of interactive components are likely to cause positioning failures and process interruptions.
[0003] For those adopting framework nesting technologies (such as <iframe>Page elements of ( ), in the existing solutions, due to the lack of a coordinate system conversion mechanism, it is easy to cause difficulties in accurately calculating the positions of elements within sub - frames, lack of an automated context - switching function, or compatibility issues in cross - frame operations. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method for automatic registration, login of business carriers, and acquisition of fund accounts, which can improve the robustness of automated operations under dynamic pages.
[0005] To solve the above - mentioned technical problems, the technical solution of the present invention is as follows: A method for automatic registration, login of business carriers, and acquisition of fund accounts, the method comprising: Step S1: Access the target website domain name through a headless browser, and after loading is completed, take a screenshot of the current visible area to generate a page image; Step S2: Input the page image into a pre - trained target detection model to identify the UI element categories, position coordinates, and confidence levels in the page; sort the recognition results based on a preset priority rule to generate an interaction task queue arranged in descending order of priority; Step S3: For the elements located within the frame in the interaction task queue, convert their coordinates to coordinates relative to the frame, and control the browser to switch to the corresponding frame; Step S4: Traverse the interaction task queue and perform the following operations in sequence: perform a simulated click operation on button - type elements; inject virtual identity information into input - box - type elements; call the corresponding recognition model to complete the verification of verification - code - type elements; Step S5: Page scrolling and submission detection. After completing the operations of the interaction task queue, detect whether there is a visible submission button on the current page. If not, scroll the page and re - execute steps S1 to S2 until a visible submission button is recognized. If visible, trigger the submission operation; Step S6: After submission, if it jumps to the recharge page, extract the payment account information through a regular expression and associate it with the virtual identity information for storage in the database.
[0006] An automated registration, login, and fund account acquisition system for business carriers, comprising:An acquisition module, configured to access the target website domain name through a headless browser, take a screenshot of the current visible area after loading is completed, and generate a page image;A generation module, configured to input the page image into a pre-trained object detection model to identify the UI element categories, position coordinates, and confidence levels in the page; sort the recognition results based on a preset priority rule to generate an interaction task queue arranged in descending order of priority;A conversion module, configured to convert the coordinates of the elements within the frame in the interaction task queue into coordinates relative to the frame, and control the browser to switch to the corresponding frame;A verification module, configured to traverse the interaction task queue and sequentially perform the following operations: perform a simulated click operation on button-type elements; inject virtual identity information into input box-type elements; call the corresponding recognition model to complete the verification of verification code-type elements;A processing module, configured to perform page scrolling and submission detection. After completing the operation of the interaction task queue, detect whether there is a visible submission button on the current page until a visible submission button is recognized. If visible, trigger the submission operation; after submission, if redirected to a recharge page, extract the payment account information through a regular expression, and associate and store it with the virtual identity information in a database.
[0007] The above solution of the present invention has at least the following beneficial effects:Adopting a vision-driven UI recognition mechanism, directly parsing the page image through an object detection model, eliminating the dependence on the DOM structure, effectively coping with random changes in page layout, dynamic generation of element identifiers, and adaptive adjustment of component positions, and can solve the problem of process interruption caused by path failure of traditional scripts.
[0008] Establishing a dynamic coordinate system conversion chain can automatically detect the frame nesting structure, calculate relative coordinates in real time, and intelligently switch the browser context, which can break through the operation error rate of cross-frame elements.
[0009] Constructing a classification recognition - dynamic execution system can improve the average recognition rate of complex verification codes; by designing a scroll - re-detection closed-loop mechanism, scrolling can be triggered based on the visible area status, and after scrolling, automatic re-screenshot and element recognition can be performed, and a locking rule for operated elements can be established to ensure that the bottom operation components can be triggered 100%, avoiding page state conflicts.
[0010] Developing a hierarchical parsing - standardized storage process, regular expressions adapting to fragmented texts, automatic association of virtual identities and payment accounts, and background parsing of unstructured data (such as two-dimensional codes) can improve the extraction integrity rate of key account information.Description of the Drawings
[0011] Figure 1 is a schematic flowchart of a method for automatically registering and logging in a service carrier and obtaining a fund account provided by an embodiment of the present invention.
[0012] Figure 2 is a schematic diagram of a system for automatically registering and logging in a service carrier and obtaining a fund account provided by an embodiment of the present invention.Detailed Embodiments
[0013] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0014] As shown in Figure 1, an embodiment of the present invention provides a method for automated registration and login of business carriers and obtaining a fund account. The method includes the following steps: Step S1, access the target website domain name through a headless browser, take a screenshot of the current visible area after loading is completed, and generate a page image; Step S2, input the page image into a pre-trained object detection model to identify the UI element categories, position coordinates, and confidence levels in the page; sort the recognition results based on a preset priority rule to generate an interaction task queue arranged in descending order of priority; Step S3, for the elements located within the frame in the interaction task queue, convert their coordinates to coordinates relative to the frame, and control the browser to switch to the corresponding frame; Step S4, traverse the interaction task queue and perform the following operations in sequence: perform a simulated click operation on button-type elements; inject virtual identity information into input box-type elements; call the corresponding recognition model to complete the verification for verification code-type elements; Step S5, page scrolling and submission detection. After completing the operations in the interaction task queue, detect whether there is a visible submission button on the current page. If not, scroll the page and re-execute Steps S1 to S2 until a visible submission button is recognized. If visible, trigger the submission operation; Step S6, after submission, if it jumps to the recharge page, extract the payment account information through a regular expression and associate and store it with the virtual identity information in the database.
[0015] In the embodiment of the present invention, the full-process automation is executed by simulating manual operations through a headless browser, combined with an object detection model to identify page elements, realizing full-process automation from accessing the website, filling in information to submitting the form, reducing the cost of manual intervention, and is particularly suitable for scenarios such as batch registration and multi-account management. The intelligent sorting of task priorities sorts UI elements according to priorities based on preset rules (such as required fields first), ensuring that interaction tasks are executed in a logical order, avoiding operation errors caused by chaotic element loading orders, and improving the stability of the process. The coordinate conversion of elements within the frame is aimed at nested frame (such as iframe) scenarios, converting element coordinates to coordinates relative to the frame and switching the context, solving the technical difficulties of cross-frame interaction, and ensuring that operations are accurately positioned to the target elements. Automatically trigger clicks on button-type elements, inject virtual identity information into input boxes, and achieve standardized data filling; for verification codes, call dedicated recognition models (such as OCR, machine learning models), breaking through the limitations of traditional automation tools in verification code scenarios and improving the processing ability of complex pages.Dynamic page scrolling detection automatically detects the visibility of the submit button, refreshes the visible area by scrolling the page and re-identifies the elements, adapts to scenarios with long pages or dynamically loaded content, and avoids process interruptions caused by hidden elements. If you jump to the recharge page after submitting multiple page jumps and data associations, you can extract payment account information through regular expressions and bind it to the virtual identity for storage, so as to achieve automatic collection and structured storage of cross-page data. Model-driven scalable target detection models and verification code recognition models can be continuously optimized through training to adapt to UI changes of different websites, reduce the risk of tool failure caused by website revisions, and improve cross-platform adaptability. Data security and compliance simulate operations by injecting virtual identity information (rather than real data) to reduce the risk of real data leakage, while supporting business needs such as automated testing and data collection in compliance scenarios.
[0016] In a preferred embodiment of the present invention, step S1, accessing the target website domain name through a headless browser, taking a screenshot of the current visible area after loading is completed, and generating a page image, including: Step S11, after the page is loaded, detect whether there is a pop-up window element blocking the key area in the current visible area; if so, automatically calculate the coordinate position of the pop-up window close button and trigger a simulated click operation to expose the underlying page elements; Step S12, based on the page status of the cleared pop-up window, identify the position of the registration button in the page; calculate the coordinates of the center point of the registration button, and execute the coordinate positioning click through the browser driver to trigger the page to jump to the registration page; Step S13, after the registration page is loaded, control the browser to perform a vertical scrolling operation: the initial scrolling position is the top of the page, the scrolling step is set according to the screen height, and scrolls down gradually until the bottom of the page; after scrolling to the bottom, a screenshot operation is performed to capture a complete image including all form input boxes and the submit button at the bottom of the page; Step S14, after completing the registration form submission operation, monitor the new page URL that the browser jumps to. When the page URL contains characteristic keywords, it is determined to be a recharge page. After the page is loaded, directly capture the current visible area image.
[0017] In the embodiments of the present invention, when specifically applied, the above steps can be implemented through the following specific solutions. For example: In step S11 above, the conventional positions of core interactive elements such as the preset registration button (such as the middle and lower part of the page) are defined as the coordinate range of the key area (such as the interval of 30% - 70% of the page height and 20% - 80% of the width); analyze the coordinate range of the pop-up window element in the screenshot, and determine whether it overlaps with the coordinate of the key area. If it overlaps, it is determined that the pop-up window is blocked; based on the coordinates (left, top, width, height) of the pop-up window element, calculate its upper right corner area (usually the position of the close button), and take the upper left corner coordinate of the pop-up window + (width - 20 pixels, height - 20 pixels) as the center point coordinate of the close button, and send a click event to this coordinate through the browser driver to close the pop-up window. The present invention automatically clears the interfering pop-up window, ensures that key elements such as the underlying registration button are visible, avoids the time-consuming operation of manually closing the pop-up window, and improves the degree of process automation.
[0018] In step S12 above, the bounding box coordinates (left, top, width, height) of the registration button are output by the target detection model. According to the bounding box coordinates, calculate the center point coordinate of the button as (left + width / 2, top + height / 2); map the center point coordinate to the browser viewport coordinate system, and drive the mouse to move to this coordinate and trigger a click event through the driver to trigger page jump. The present invention accurately locates the registration entry, avoids manual misoperation caused by page layout changes, ensures that the process is executed according to the preset path, and improves the interaction accuracy.
[0019] In step S13 above, obtain the vertical height (H, unit: pixel) of the browser viewport, divide the total page height (which can be obtained through document.body.scrollHeight) equally by the viewport height, and set the step size to H, that is, scroll one screen height each time, starting from the top of the page (scroll position 0), scroll down H pixels each time until the scroll position reaches or exceeds the total page height to ensure that all content is loaded. After scrolling to the bottom of the page, call the screenshot interface to capture the current complete page image, including all form input boxes and submit buttons. The present invention adapts to long form pages or dynamically loaded content, ensures that all input fields and submit buttons required for registration are captured, and avoids information omission caused by incomplete display of page elements.
[0020] In the above step S14, characteristic keywords of the preset recharge page (such as "recharge", "pay", "fund", etc.) are preset, the path or parameter part of the new page URL is parsed, and it is determined whether any keyword is included. If the URL contains the characteristic keyword, it is determined as the recharge page, and the screenshot operation is triggered; otherwise, the monitoring continues, and the image within the current browser viewport is directly intercepted for subsequent extraction of account information. The present invention automatically identifies the jump stage of the registration process, accurately locates the recharge page, and timely captures the key information page, laying a foundation for subsequent extraction of payment account information through regular expressions and improving the efficiency of cross-page data collection.
[0021] In a preferred embodiment of the present invention, in step S2, the page image is input into a pre-trained object detection model to identify the UI element categories, position coordinates, and confidence levels in the page; the recognition results are sorted based on a preset priority rule to generate an interactive task queue arranged in descending order of priority, including:Step S21, input the page images generated in steps S13 and S14 into the pre-trained object detection model, and output a set of recognition results for each UI element; each recognition result includes a class label, position coordinates, and confidence level;Step S22, assign priorities to each recognition result output in step S21 according to the Euclidean distance between the center point of the element and the center of the page to obtain the weights corresponding to the class label, position coordinates, and confidence level respectively; according to the weights corresponding to the class label, position coordinates, and confidence level respectively, obtain the priority scores, and divide the elements into five levels from high to low according to the scores to obtain a set of elements with priority labels;Step S23, arrange the set of elements with priority labels output in step S22 in descending order of priority, and arrange the elements with the same priority in descending order of confidence level to generate a structured task queue;Step S24, traverse the structured task queue in step S23, calculate the normalized coordinates of each element, that is, calculate its absolute pixel coordinates in the browser window; through the DOM position inverse deduction algorithm, convert the absolute pixel coordinates into XPath paths.
[0022] In the embodiments of the present invention, when specifically applied, the above steps can be implemented by the following specific solutions. For example: In the above step S21, the complete page image or visual area image generated in steps S13 and S14 is scaled to the model input size (such as 640×480 pixels), and the pixel ratio is maintained to avoid distortion. The pre-trained object detection model (such as YOLO, FasterR-CNN) performs convolutional calculations on the image to identify the bounding box coordinates (left, top, width, height) of all visible UI elements (such as input boxes, buttons, labels), and matches the predefined class labels (such as "input", "button", "captcha"). The model generates a confidence score (between 0 and 1) for each recognition result, indicating the probability that the element is correctly recognized (for example, a confidence level ≥0.8 is regarded as a reliable recognition). The present invention automatically extracts the key information of all UI elements in the page, replaces manual annotation and manual positioning, and improves the element recognition efficiency; outputs structured data including coordinates, categories, and confidence levels, providing an accurate basis for subsequent task sorting and interaction operations.
[0023] In the above step S22, the page center coordinates are preset as (page width / 2, page height / 2); calculate the Euclidean distance between the center point of each element bounding box (left + width / 2, top + height / 2) and the page center. The closer the distance, the higher the priority base value (for example, elements with a distance ≤ 1 / 4 of the page height are regarded as "high-priority areas"); preset the priority weight of element categories (such as input box weight +30, button weight +20, hint text weight +10), and elements of key operation classes (such as registration button, submission button) have a higher weight; elements with a confidence level ≥0.9 are additionally increased in priority score (such as +10 points), and elements with a confidence level <0.5 are marked as "to be confirmed" and the score is reduced (such as -20 points); accumulate the distance, category, and confidence scores to obtain the total priority score, and divide it into five levels according to the score (level 1 is the highest, level 5 is the lowest), for example: Level 1 (score ≥ 80): required input box, submission button; Level 3 (score 40 - 60): optional items, ordinary hint labels; Level 5 (score < 20): advertisement elements, non-key decorative patterns.
[0024] The present invention combines the element position, functional importance, and recognition reliability to ensure that core interaction elements (such as registration form input boxes) are processed first, avoiding interference from secondary elements in the process; adapts to different page layouts (such as differences between mobile and PC), automatically adjusts the element priority, and improves the flexibility of the process.
[0025] In the above step S23, the first-level sorting (descending priority) is to sort from 1 to 5 according to the priority level in step S22 first, ensuring that high-level elements (such as required items) are in the front.Secondary sorting (confidence descending order), among the elements of the same level, sort them from high to low according to the confidence score (for example, for two input boxes of the same level 2, the element with a confidence score of 0.9 is ranked before the element with a confidence score of 0.8); the sorted elements are converted into JSON format task items containing "operation type (click / input / verification), element coordinates, priority label, confidence", and generate an executable queue structure (such as [task 1, task 2..., task N]).
[0026] The present invention ensures that the automated process is executed according to the principle of "importance first, reliability first" through dual sorting of "priority + confidence", reducing logical errors caused by disordered element order; the structured queue facilitates debugging and log recording, and supports advanced functions such as midway pause and breakpoint resumption.
[0027] In the above step S24, the normalized coordinates are converted to absolute pixels: If the element coordinates are normalized values (0-1 interval, such as the center point coordinate x=0.5 represents the horizontal center line of the page), then multiply by the actual size of the browser window (such as the window width is 1920 pixels, x=0.5×1920=960 pixels) to get the absolute pixel coordinates (XY); obtain the page DOM tree through the browser driver, starting from the root node (html), and traverse to find the element according to the following logic: Based on the absolute pixel coordinates (XY), the elementFromPoint(XY) method is used to locate the topmost visible element. The parent node is recursively traced upward, and the tag name and attributes (such as / / input[@id='username']) are concatenated to generate a unique XPath path. Redundant levels are removed (such as skipping the div container layer) to retain the shortest valid path (such as / / form[@class='register-form'] / input[1]).
[0028] The present invention solves the matching problem of "visually visible elements" and "underlying code elements" through bidirectional mapping of pixel coordinates and DOM structures, ensuring that the simulated operation accurately acts on the target control (such as avoiding clicking on a hidden button that is blocked); the absolute pixel coordinate conversion mechanism adapts to different screen sizes (such as mobile phones, tablets, and desktops), and the XPath path supports dynamic page structure changes, reducing the risk of positioning failure caused by page revision.
[0029] In a preferred embodiment of the present invention, the pre-trained object detection model includes: extracting the domain names of target websites from the abusive domain name library, accessing the domain names through an automated script and performing page loading status verification; for the websites that pass the verification, automatically triggering pop-up window closing, registration button clicking, and page scrolling operations, and capturing complete page screenshots containing key interactive elements; based on the complete page screenshots, using the hierarchical clustering algorithm to group images with similar visual features to obtain deduplicated images; for the deduplicated images, annotating the class labels and position information of 17 types of UI elements to generate an annotation file that conforms to the object detection format; integrating the annotation files and dividing them into a training set, a validation set, and a test set according to a preset ratio; based on the training set, the validation set, and the test set, performing the following automatic optimization process: initializing the parameter population, where each group of parameters includes a learning rate, anchor box size, and network structure configuration; training the model for each parameter combination, and calculating the fitness using the validation set accuracy and loss value; iteratively generating a new population through selection, crossover, and mutation operations until the fitness converges to obtain optimized parameters; using the optimized parameters, performing full-network fine-tuning and fusing incremental sample training to obtain the pre-trained object detection model.
[0030] In the embodiments of the present invention, when specifically applied, the above steps can be implemented through the following specific solutions. For example: screening out the domain names in the active state (such as having DNS resolution records in the past 30 days) from the abusive domain name library, and excluding the domain names that have been marked as inaccessible. Sending an HTTP request through an automated script (such as the requests library in Python), and detecting the response status code (such as 200 indicating success); parsing the response content to verify whether it contains a complete HTML structure (such as the existence of the <!DOCTYPEhtml> declaration and the <html> root tag). Marking the domain names with timeouts (such as not responding within more than 5 seconds) or returning error status codes (such as 404, 500) as invalid and removing them from the candidate set. The present invention can ensure that the page data input for model training are all valid samples that can be normally accessed, avoiding the waste of computing resources by invalid domain names; filtering abnormal pages automatically through structured verification rules to improve the reliability of data collection.
[0031] Preset common pop-up window close button features (such as the "×" icon in the upper right corner, with coordinates in the 10% - 20% area at the top of the page and 10% - 20% area on the right side), locate and trigger a click through image matching technology (such as template matching); locate the button based on page text features (such as "Register", "SignUp") or visual features (such as orange buttons, larger fonts), calculate the center point coordinates and then simulate a click to trigger the registration process; scroll the page step by step (each time scrolling 1 / 2 of the viewport height) until the content at the bottom of the page no longer updates (such as the scrollHeight remaining unchanged after two consecutive scrolls), to ensure capturing a complete image containing the footer submit button. The present invention clears interference elements (pop-ups) and triggers key interactions (registration) by simulating manual operations, ensuring that the screenshot contains the complete registration form structure; the scroll loading mechanism can capture form elements on long pages or asynchronously loaded, avoiding data loss.
[0032] Extract visual features (such as color histogram, Histogram of Oriented Gradients HOG, CSS style features) from each screenshot to generate a feature vector with a fixed dimension (such as 512 dimensions); calculate the cosine similarity between images (similarity > 0.8 is considered similar), and construct a tree-like clustering structure; select the central image of each cluster as a representative sample, remove other similar images within the cluster, and retain about 30% - 50% of the original images, ensuring that only 1 - 2 typical samples are retained for the same type of page. By clustering, duplicate or highly similar page images (such as registration pages with different domain names but the same template) are removed, reducing the training data volume while retaining feature diversity;Improve training efficiency: After deduplication, the data volume decreases, which can shorten the model training time and avoid repeated learning of redundant information.
[0033] Manually annotate 17 types of UI elements (such as input boxes, radio boxes, dropdown menus, verification code boxes, etc.) on the deduplicated images using annotation tools (such as LabelMe, RectLabel), draw bounding boxes and associate category labels; convert the annotation data into the format required by the object detection model (such as the.txt format of YOLO, with each line containing the category index and normalized coordinates).
[0034] Dataset division:Randomly divide the training set, validation set, and test set in a ratio of 7:2:1; ensure that the distribution of each type of element in each set is balanced (such as the proportion of input boxes in the training set is the same as the overall proportion). The present invention's manual annotation ensures the element positioning accuracy (pixel-level bounding box) and category accuracy, providing a reliable supervision signal for the model. The staged dataset supports performance verification (tuning parameters for the validation set) and final generalization ability testing (testing set) during the model training process.Model parameter optimization process based on genetic algorithm : Parameter population initialization: Generate 100 sets of initial parameter combinations, set the learning rate range to 0.001~0.1, and use the anchor box size preset according to the COCO dataset (such as (10,13), (16,30), etc.). Select YOLOv3-spp or FasterR-CNN variant as the network structure. Train the model for 50 rounds for each set of parameters, and record the mAP (mean average precision) and loss value of the validation set. The fitness formula is: fitness = mAP - 0.5 × loss value. The higher the value, the better the performance. Genetic manipulation: Select and keep the parameter combinations with the top 20% fitness as the parents; Crossover, the parent parameters randomly exchange some dimensions (such as the learning rate and the anchor box size crossover); Mutation, adding random perturbations to the parameters after crossover (such as ±10% fluctuations in learning rate). The iteration is terminated, and the iteration is stopped when the fitness improvement is less than 1% for 5 consecutive generations, and the optimal parameter combination is selected.
[0038] Model fine-tuning: Use the optimized parameters to train the model and gradually add newly collected incremental samples (such as 1,000 new screenshots per week) for online learning.
[0039] The present invention can replace manual trial-and-error parameter adjustment, quickly search for the optimal parameter combination through a bionic algorithm, and improve the model training efficiency (the parameter adjustment time is shortened by more than 70%); the incremental sample training mechanism enables the model to continuously learn new page design patterns (such as new verification code styles), delaying the model failure cycle caused by website revision. In a preferred embodiment of the present invention, fitness is calculated using verification set accuracy and loss value, including: Based on the model training output of the current parameter combination, load its prediction results on the validation set; Count the matches between the predicted box and the real box, that is, for each image in the validation set, calculate the overlap between the predicted box of the target detection model and the labeled box; According to the 17 categories of UI elements, the proportion of the number of correct detections to the total number of corresponding category annotations is counted, and the arithmetic average of the accuracy values of all categories is taken to obtain the average accuracy index of the verification set; Extract the CIoU loss values of all samples from the object detection model verification log, and take the arithmetic average of the CIoU loss values of all samples to obtain the mean CIoU loss; The accuracy weight coefficient and the loss weight coefficient are defined, the average accuracy index is weighted according to the positive correlation, and the CIoU loss mean is weighted according to the negative correlation to generate a fitness score.
[0041] In the embodiments of the present invention, when specifically applied, the above steps can be implemented through the following specific solutions. For example: Use the target detection model trained with the current parameter combination to perform forward propagation on each image in the validation set, and output the coordinates (x1, y1, x2, y2), class labels, and confidence scores of the predicted bounding boxes. Store the prediction results in a structured format (such as JSON) according to the image ID. Each predicted bounding box contains information [class, confidence, coordinates]. Read the true annotation bounding box information of the corresponding image from the validation set annotation file to ensure that the prediction results and the true labels correspond one by one according to the image ID. The present invention stores the prediction and true results in a unified format, providing a basis for subsequent matching calculations, avoiding evaluation errors caused by inconsistent data structures, and saving intermediate results to support subsequent detailed analysis (such as false detection / missing detection case backtracking).
[0042] For each image in the validation set, traverse all its predicted bounding boxes and true annotation bounding boxes.
[0043] IoU calculation: For each predicted bounding box, calculate its intersection over union (IoU) with all true bounding boxes:Intersection area: The number of pixels in the overlapping region between the predicted bounding box and the true bounding box;Union area: The area of the predicted bounding box + the area of the true bounding box - the intersection area;IoU = Intersection area / Union area.
[0044] Matching determination: If the IoU of a certain predicted bounding box and a certain true bounding box ≥ 0.5 (the threshold can be adjusted) and the classes are the same, it is determined as a valid match.
[0045] In the present invention, IoU is used as a standard metric, which intuitively reflects the spatial coincidence degree between the predicted bounding box and the true target, avoiding the limitations of simply relying on coordinate errors; by traversing all possible combinations of box pairs, it ensures that each true target has the opportunity to be correctly matched, improving the evaluation accuracy in multi-target scenarios.
[0046] The calculation process of calculating the classification accuracy is as follows:Class-level statistics: Statistically calculate separately according to 17 types of UI elements (such as input boxes, buttons, verification codes):Number of correct detections: The number of predicted bounding boxes and true bounding boxes with IoU ≥ 0.5 and the same class;Total number of annotations: The total number of true annotation bounding boxes of this class in the validation set.
[0047] Class accuracy calculation: Accuracy of each class = Number of correct detections / Total number of annotations × 100%.
[0048] Mean average precision calculation: Take the arithmetic mean of the accuracy values of 17 classes to obtain the mean average precision (mAP@0.5) of the validation set.The present invention differentiates the detection effects of different UI elements, identifies the advantages and disadvantages of the recognition model (such as continuous missed detection of a certain type of element), avoids evaluation bias caused by uneven class distribution (such as the dominant result of the class with a large proportion), and ensures that all classes contribute equally to the final index.
[0049] Extract the CIoU (Complete IoU) loss value of each sample from the model validation log. CIoU loss comprehensively considers the overlap rate, center point distance, and aspect ratio of the predicted box and the ground truth box. Extreme loss values (such as values exceeding 3 times the standard deviation) are removed to avoid the influence of outliers on the overall mean. After summing the CIoU loss values of all valid samples and taking the average, the mean CIoU loss is obtained. The CIoU loss of the present invention simultaneously optimizes the overlap rate, position accuracy, and shape consistency, and can more comprehensively reflect the quality of the predicted box than the traditional IoU; noise interference is reduced through outlier filtering, making the loss mean more representative of the true performance of the model and avoiding misjudgment caused by the fluctuation of individual samples.
[0050] The calculation process of the weighted generation of fitness scores is as follows: Coefficient setting: Precision weight coefficient (such as 0.7): Reflects the importance of average precision; Loss weight coefficient (such as 0.3): Reflects the importance of CIoU loss.
[0051] Normalization processing: The average precision index is scaled to the interval [0,1] (for example, if the original precision is 85%, it is 0.85 after normalization); The mean CIoU loss is mapped to the interval [0,1] through 1 / (1 + loss mean) (the smaller the loss, the larger the mapped value).
[0052] Weighted summation: Fitness score = precision weight coefficient × normalized precision + loss weight coefficient × normalized loss mapped value.
[0053] The present invention simultaneously considers the accuracy (precision) and positioning quality (loss) of the model, avoids one-sided optimization caused by a single index (such as only pursuing high recall rate at the expense of positioning accuracy), and can flexibly adapt to different business requirements by adjusting the weight coefficient (such as increasing the loss weight in scenarios with extremely high requirements for positioning accuracy).
[0054] In a preferred embodiment of the present invention, in step S3, for the elements located within the frame in the interaction task queue, convert their coordinates to coordinates relative to the frame, and control the browser to switch to the corresponding frame, including:Step S31, analyze the XPath path features generated in S24. If the path contains HTML tags or frame hierarchical structures, determine that the corresponding element is located within the frame, and extract the frame identifier in the parent page, that is, the ID or Name attribute;Step S32, relative coordinate conversion calculation. Read the pixel coordinates (X_frame, Y_frame) of the upper left corner of the frame in the parent page through the browser interface, and read the width W_frame and height H_frame of the frame itself;Calculate the relative position of the element within the frame:Relative abscissa = (absolute abscissa of the element - X_frame) ÷ W_frame;Relative ordinate = (absolute ordinate of the element - Y_frame) ÷ H_frame;Output the normalized relative coordinates;Step S33, based on the frame identifier extracted in S31, locate the DOM node of the corresponding frame in the browser, and switch the WebDriver operation context to the inside of the frame.
[0055] In the embodiments of the present invention, when specifically applied, the above steps can be implemented through the following specific solutions. For example:In the above step S31, XPath path parsing. Traverse the XPath path of each element in the task queue (such as / / html / body / iframe[@id='register'] / div / input), and check whether it contains iframe or frame tags. If there is an iframe tag in the path, extract the id or name attribute value of the tag (such as register) as the unique identifier of the frame. If the path contains nested frames (such as parent frame → child frame), recursively extract the identifiers of each level to generate a frame hierarchy chain (such as ["parent_frame","child_frame"]).
[0056] The present invention does not require manual annotation, automatically discovers the frame hierarchy in the page through XPath syntax features, and solves the cross-frame positioning problem that is difficult to handle by traditional automation tools; extracts the id / name as the positioning basis, which is more stable than coordinates, and reduces the risk of positioning failure caused by page layout changes.
[0057] In the above step S32, the upper left corner coordinates (X_frame, Y_frame) of the target frame in the parent page are obtained through the browser driver, the width (W_frame) and height (H_frame) of the frame are obtained, and the absolute pixel coordinates (X_abs, Y_abs) of the element are read from the task queue. The offset of the element relative to the upper left corner of the frame is calculated as follows: Horizontal offset = X_abs - X_frame; Vertical offset = Y_abs - Y_frame.
[0058] Normalization processing: The offset is divided by the frame size to obtain the normalized relative coordinates: Relative abscissa = Horizontal offset / W_frame; Relative ordinate = Vertical offset / H_frame.
[0059] The normalized coordinates of the present invention are independent of the specific screen size, and the same set of coordinates can be reused at different resolutions, improving the compatibility of the automation script; Dynamic frame adaptation: Even if the frame position changes due to the responsive design of the page, as long as the internal structure of the frame remains unchanged, the relative coordinates can still accurately locate the element, reducing the maintenance cost.
[0060] In the above step S33, the frame identifier extracted in S31 (such as id="register") is used to locate the frame DOM node through the switch_to.frame() method of the browser driver (such as Selenium); the context switching operation is performed to limit the scope of all subsequent operations (such as clicking, inputting) within the target frame; if there are multiple nested frames, switch them in sequence according to the hierarchical chain (such as first switching to the parent frame and then to the child frame); after switching, the element query is executed through the driver (such as find_element_by_xpath()) to verify whether the elements within the frame can be correctly located to ensure the success of the switching.
[0061] Through context switching, the present invention enables the automation tool to operate the elements within the frame just like operating a normal page, breaking through the limitation that traditional tools can only operate on the top-level page, ensuring that all operations on the elements within the frame are executed in the correct context, and avoiding operation failures or misoperations caused by context confusion.
[0062] In a preferred embodiment of the present invention, in step S4, traverse the interactive task queue and perform the following operations in sequence: perform a simulated click operation on button-type elements; inject virtual identity information into input box-type elements; call the corresponding recognition model to complete the verification for verification code-type elements, including:Step S41, start reading task items sequentially from the head of the structured task queue generated in S23, and determine the execution order according to the priority label in the task item;Step S42, for the current task item:If the element is within the frame, use the relative coordinates calculated in step S32, and if it is in the main document, use the absolute coordinates converted in step S24; dynamically calculate the operation coordinates in the browser viewport:Main document: directly use the absolute pixel coordinates;Within the frame: relative coordinates × actual frame size + frame offset;Step S43, perform the corresponding operation according to the element category, button click operation, wait for 500 ms for the page response time after clicking, form input operation, branch processing according to the verification code type, sliding verification code;Step S44, after each task item operation is completed, re-capture the visible area of the current page and call the target detection model to verify the operation effect; if the operation fails, record the current element positioning information and re-insert the task at the end of the queue.
[0063] In the embodiments of the present invention, when specifically applied, the above steps can be implemented through the following specific solutions. For example:In the above step S41, obtain the sorted structured task queue (sorted in descending order of priority, and in descending order of confidence for the same priority) from step S23, extract task items one by one in the queue order, each task item includes the element type (button / input box / verification code), coordinate information, and priority label. If there are high-priority tasks (such as required input boxes) in the queue, process them first; low-priority tasks (such as optional items) are executed after the high-priority tasks are completed. The present invention ensures that core operations (such as clicking the registration button) are executed first, avoiding process interruption caused by the processing of secondary elements. When a high-priority task fails, the process can be quickly terminated, reducing the resources consumed by ineffective operations.
[0064] In the above step S42, check whether the XPath path of the task item contains the iframe tag to determine whether the element is within the frame. If it is within the frame, use the relative coordinates of S32; if it is in the main document, use the absolute coordinates of S24.
[0065] Coordinate mapping:For the main document elements, directly use the absolute pixel coordinates (X_abs, Y_abs) as the operation points;For the elements within the frame:Obtain the real-time dimensions (W_frame, H_frame) and the upper-left offset (X_frame, Y_frame) of the current frame;Calculate the operation coordinates: X = relative horizontal coordinate × W_frame + X_frame, Y = relative vertical coordinate × H_frame + Y_frame; If the calculated coordinates exceed the current viewport range, perform page scrolling to make the element enter the visible area. The present invention uniformly processes the positioning of the main document and the elements within the frame, without the need to write special logic for different contexts; even if the page size changes due to interactions during operation, the elements can still be accurately operated through real-time calculation.
[0066] In the above step S43, use a browser driver (such as Selenium) to send a mouse click event to the calculated coordinates, wait for 500 ms after clicking to leave a response time for the page (such as loading new content, displaying a prompt box), locate the input box element, clear the original content, inject virtual identity information according to preset rules (such as randomly generated name, mobile phone number, email), call the OCR model to recognize the characters in the picture, input them into the corresponding input box, recognize the positions of the slider and the notch, calculate the sliding distance, simulate the human operation trajectory (accelerate first and then decelerate) to drag the slider, recognize the prompt text, and click on the corresponding picture area.
[0067] The present invention simplifies complex UI interactions into a unified operation interface through classification processing, reducing the development cost; customizes solutions for different types of verification codes, enabling the automation process to handle more than 80% of common verification mechanisms.
[0068] In the above step S44, after completing the current task item, capture the visible area of the page to generate a new image, input the new image into the target detection model to recognize the state changes of the operation elements (such as whether the button changes to the "clicked" style, whether the input box is filled with content), compare the attribute changes of the elements before and after the operation (such as the disabled state of the button, the value attribute of the input box), if the verification fails (such as the input box is still empty), record the positioning information such as the XPath and coordinates of the element, re-insert the task at the end of the queue, and mark it as needing to be retried (up to 3 times). The present invention uses model recognition to verify the operation effect in real time, avoids the hidden error of "operation successful but not effective", automatically retries failed tasks, and improves the success rate of the process (tests show that it can increase the overall passing rate by more than 25%).
[0069] In a preferred embodiment of the present invention, in step S5, page scrolling and submission detection, after completing the interactive task queue operation, it is detected whether there is a visible submission button on the current page. If it is not visible, the page is scrolled and steps S1 to S2 are re-executed until a visible submission button is recognized. If it is visible, a submission operation is triggered, including:Step S51, based on the latest page screenshot after the operation in S44, call the target detection model to identify whether there is an element of the "submission button" type in the current visible area;Step S52, when the detection result in step S51 is invisible, execute:Obtain the total height of the current page and the height of the browser window, and set the single scroll amount to 80% of the window height; perform the scroll operation:Record the current scroll bar position as the initial position, trigger the browser to scroll down by one scroll amount unit, and wait for 500 ms for the page to redraw;Status update and loop:Re-execute step S1 to capture the image of the new visible area; re-execute step S2 to generate a new interactive task queue; return to step S51 for submission button detection, and loop until any of the following conditions is met:A visible submission button is detected, or the cumulative scroll amount exceeds the total height of the page;Step S53, during each scroll and re-identification process, establish a library of marked elements that have been operated on, and filter the new interactive task queue generated in S2;Step S54, when a visible submission button is detected in step S51, locate the button coordinates.
[0070] In the embodiments of the present invention, the above steps can be implemented in the following ways. For example, input the latest page screenshot generated in S44 into the target detection model. The model classifies all elements in the screenshot, filters out the elements with the label "submission button" (such as text buttons like "register" and "submit immediately"), and checks whether the button coordinates are within the current window range; exclude the occluded elements (such as pop-up windows or floating layers above the button). The present invention automatically identifies the submission button through the model, avoiding hard-coding the button position manually, adapting to the UI design differences of different websites, and combining coordinate and occlusion relationship judgments to ensure that only the truly operable submission button is recognized, reducing the risk of mis-triggering.
[0071] The above step S51 can be implemented in the following ways. For example, obtain the total height of the page (such as document.body.scrollHeight) and the window height; set the single scroll amount to 80% of the window height (to avoid missing elements due to excessive scrolling).
[0072] Scroll execution:Record the current scroll bar position as the initial position, trigger the browser to scroll down by one scroll unit (e.g., the current position + 80% of the viewport height), wait for 500 ms to ensure that the page finishes redrawing and dynamic loading. After each scroll, re-execute step S1 (take a screenshot) and step S2 (generate a task queue), and call step S51 again to detect the submit button until the termination condition is met, i.e., the cumulative scroll amount exceeds the total height of the page (indicating that all content has been traversed).
[0073] Through the adaptive scrolling strategy of the present invention, it is ensured that even if the submit button is located at the bottom of a long page or in a dynamically loaded area, it can be detected. The reasonable scroll step and waiting time balance the detection efficiency and the page loading integrity, and the average detection time is shortened by 30%.
[0074] In step S53, when initially executing step S4, establish an empty library for marking processed elements (such as a set processed_elements). For each completed task item, add the XPath or unique identifier of the element to the library. When generating a new task queue after each scroll, traverse each element in the queue. If the element identifier exists in the library, remove the element from the queue. The present invention prevents repeated operations on filled input boxes or clicked buttons, reduces invalid interactions, and avoids page state chaos caused by repeated operations (such as multiple clicks on the submit button triggering repeated requests).
[0075] In the above step S54, extract the bounding box coordinates (left, top, width, height) of the submit button from the output of the target detection model, and calculate the center point coordinates of the button as (left + width / 2, top + height / 2); if the button is within a frame, use the relative coordinate conversion method of S32 to map the center point coordinates to the browser viewport. If it is in the main document, directly use the absolute coordinates. By calculating the center point through the bounding box output by the model, the present invention is more accurate than traditional positioning based on text or CSS selectors, and the success rate is increased to 98%. Even if the website modifies the class or ID of the button, as long as the visual style remains unchanged, it can still be accurately positioned, and the maintenance cost is reduced by 50%.
[0076] In a preferred embodiment of the present invention, in step S6, after submission, if it jumps to the recharge page, extract the payment account information through regular expressions and store it in association with the virtual identity information in the database, including:Step S61, based on the page jump result triggered by step S54, monitor whether the new page URL contains a preset keyword, intercept the complete HTML source code of the current page, and perform a dual verification mechanism to obtain a verified recharge page;Step S62, perform hierarchical parsing on the recharge page verified in step S61, standardize and store the data in association. When the recharge page verification in step S61 fails, it is packaged and transmitted to the data construction module.
[0077] In the embodiment of the present invention, the above steps can be implemented in the following ways. For example,In the above step S61, preset recharge page feature keywords (such as "recharge", "payaccount", "fund-transfer"), parse the path or parameter part of the new page URL, determine whether it contains any keyword, call the target detection model to identify whether there are UI elements related to recharge in the page (such as text labels of "payment account" and "bank of deposit" or account input boxes), check whether the HTML source code contains a form structure related to payment (such as <formaction=" / pay">) or a specific JS file reference (such as a payment SDK script). If the dual verification is passed, intercept the complete HTML source code and remove redundant content such as comments and blank lines. If it fails (such as jumping to an error page or an advertisement page), record the error URL and terminate the subsequent operations. The present invention combines triple-dimensional verification of URL, visual elements, and source code structure to avoid misjudging similar pages as recharge pages (such as misjudging the "account settings" page), and only performs subsequent processing on the pages that pass the verification, reducing interference from invalid data and improving the accuracy of account information extraction.
[0078] In the above step S62, the text content in the HTML tag is matched through regular expressions (such as 123456789), relevant fields of account information are extracted, and the target data is located by combining the context semantics (such as keywords like "bank account number", "payment account number", etc.), and interference information (such as false accounts in advertisements) is excluded; the format of the extracted account information is cleaned: for the bank account number, spaces and special symbols are removed and unified into pure numbers, the name of the bank where the account is held is mapped to the standard name (for example, "Industrial and Commercial Bank of China" is uniformly abbreviated as "ICBC"), the virtual identity information of the current operation is obtained from the task queue (such as the generated random name and mobile phone number), and a one-to-one association is established between the account information and the identity information through the task ID; the associated data is written into a structured database (such as MySQL), and the fields include: virtual name, mobile phone number, payment account, bank where the account is held, association timestamp, etc. If the page fails to pass the recharge page verification, the original URL, screenshot, and source code are packaged and transmitted to the data construction module for expanding the training set or optimizing the verification rules.
[0079] Through hierarchical parsing and standardized processing, the present invention converts unstructured page data into directly usable business data, adapts to the data format requirements of downstream systems, and the data that fails to pass the verification can feed back to model training to form a data closed-loop of "recognition - verification - optimization", continuously improving the system's recognition ability for new recharge pages; the strong association between virtual identity and account information provides complete context data for subsequent operations such as batch recharge and fund management, reducing the manual data matching cost.
[0080] A business carrier automated registration, login and fund account acquisition system, comprising:An acquisition module, configured to access the target website domain name through a headless browser, take a screenshot of the current visible area after loading is completed, and generate a page image;A generation module, configured to input the page image into a pre-trained target detection model to identify the UI element categories, position coordinates and confidence levels in the page; sort the recognition results based on a preset priority rule, and generate an interaction task queue arranged in descending order of priority;A conversion module, configured to convert the coordinates of the elements within the frame in the interaction task queue into coordinates relative to the frame, and control the browser to switch to the corresponding frame;A verification module, configured to traverse the interaction task queue and sequentially perform the following operations: perform a simulated click operation on button-type elements; inject virtual identity information into input box-type elements; call the corresponding recognition model to complete verification on verification code-type elements;A processing module, configured to perform page scrolling and submission detection. After completing the operation of the interaction task queue, detect whether there is a visible submission button on the current page until a visible submission button is recognized. If visible, trigger a submission operation; after submission, if it jumps to a recharge page, extract payment account information through a regular expression, and associate and store it with the virtual identity information in a database.
[0081] Dataset preparation and model training:In this embodiment, a structured image dataset covering 17 types of UI elements is constructed by means of screenshotting, screening, clustering and duplicate removal, and manual annotation of target domain website pages, providing high-quality training samples for the target detection model. In this example, the YOLOv series of models is selected as the UI recognition framework. During the training process, a batch incremental training strategy, CIoU bounding box loss optimization, and a multi-scale feature fusion structure are combined, and a genetic algorithm is introduced to automatically optimize the hyperparameters and model structure to improve the detection accuracy, training efficiency and model generalization ability. The specific steps are as follows:Step 1, Target website collection, screening and access verification Randomly extract approximately 20,000 suspicious business website domain names from the domain name database. The system uses an automatic script to access each domain name and performs a preliminary screening according to page characteristics, removing the following types of sites: 1) Pages that cannot be loaded or have a 404 error; 2) Content is empty and irrelevant to the business; 3) Those with a non-standard structure and cannot be screenshot.
[0082] Step 2, Automatic screenshot process Automatically execute the screenshot process on the screened websites in the following order to ensure capturing the key registration path pages: 1) Close the home page pop-up window (if any); 2) Click the registration button after the home page is loaded; 3) After the registration page is fully loaded, scroll to the bottom of the page and capture the page containing all form fields and the "Submit" button; 4) After registration, jump to the business operation page and capture the page containing operation method information.
[0083] Step 3: Image Duplicate Removal and Clustering for Redundancy Reduction To improve the training efficiency, the Ward-based hierarchical clustering method is used to remove duplicates from the screenshot images, eliminating duplicate samples and retaining the central images of the clusters to ensure data diversity and representativeness.
[0084] Step 4: UI Element Annotation The Yolo_mark tool is used to manually annotate 17 types of UI elements in the screenshot images, and the annotation format conforms to the txt format required by the YOLOv series. Each file contains the element category number and its normalized coordinates in the image.
[0085] Step 5: Data Organization and Partitioning After constructing the dataset, it is partitioned into a training set, a validation set, and a test set in the ratio of 8:1:1 to ensure an even distribution of samples for the three types of pages (closed pop-up windows, home pages, registration pages).
[0086] Step 6: Parameter Optimization Training Incorporating Genetic Algorithm Before the formal training of the model, the genetic algorithm is introduced into the system to globally search for and automatically optimize the key training parameters of the YOLOv series model to improve the performance and generalization ability of the model in the task of web UI element recognition. It specifically includes the following steps: (1) Genetic algorithm initialization: Define the fitness function, using the performance metrics of the model on the validation set as the evaluation criterion. The fitness is used to measure the quality of individuals and is the core basis for the genetic algorithm to perform "selection - evolution". The optimization parameter space includes: learning rate, batch size, anchor box size, multi-scale feature fusion structure configuration (FPN level), attention mechanism insertion position, etc. (2) Population generation and evolution: The system initializes and generates multiple individuals (each individual is a set of parameter combinations). In each generation, a small number of rounds of model training (such as 10 - 20 rounds) are performed on each individual, and their fitness values are calculated. Subsequently, standard genetic operations (selection, crossover, mutation) are executed. (3) Iterative optimization: The algorithm iterates generation by generation according to the change of fitness until the stop condition is met (such as the fitness improvement amplitude is lower than the threshold or the maximum number of generations is reached). Finally, the individual parameters with the highest fitness are retained for formal training.
[0087] Step 7: Three-Stage Model Training Based on Optimized Parameters Based on the optimization results of the genetic algorithm, the model training is performed in a three-stage manner: Freeze the backbone network and only train the detection head to stabilize the basic recognition ability; Thaw some backbone layers and introduce a multi-scale training mechanism to improve the detection accuracy for UI components of different sizes; Fine-tune the entire network and further optimize the model performance by combining strategies such as incremental samples and hard example mining.
[0088] Automatic registration process:In this embodiment, through the deep linkage between model recognition and browser execution, the system automatically parses the web page structure and executes simulated user behaviors based on the priority mechanism, breaking through the limitations of the fixed path of traditional scripts. The specific steps are as follows:Step 1, Initialize access and page screenshot. The system calls Selenium WebDriver through the automation control module, starts a headless browser, and accesses the domain name of the target business website. After the page is loaded, it automatically takes a screenshot of the current visible area and passes it to the target detection model that has been trained in Embodiment 1 for processing.
[0089] Step 2, UI element recognition and priority sorting. The model recognizes 17 types of UI elements on the page and outputs the category label, normalized position coordinates, and confidence value. The system divides the elements into five levels of priority (immediate, high, moderate, low, last) according to the preset priority mechanism:Immediate priority: "Close pop-up window button" and covering the key area;High priority: Jump components such as "Register button", "Login button", etc.;Moderate priority: Form input components (username, password input box, etc.);Low priority: Verification code components (graphical verification code, sliding verification code);Last priority: Non-immediate interaction components such as "Submit button", "Contact customer service". After the priority is divided, an interaction task queue is generated in descending order, including information such as element category, coordinates, confidence, and XPath path.
[0090] Step 3, Coordinate conversion of elements within the frame. If the element is located in or, the system converts the coordinates to relative positions, switches the WebDriver to the corresponding frame, and uses a custom script to convert the coordinates to the XPath path to ensure operation accuracy.
[0091] Step 4, Simulated operation execution. Traverse the task queue and perform corresponding operations according to the category:Click operation: Simulate clicking buttons such as "Register", "Close", "Submit", etc.;Input operation: Enter compliant identity information generated from the virtual information pool in the form field;Verification code processing:Graphical verification code: Use the OCR model to recognize and input;Sliding verification code: Calculate the sliding distance through image processing and simulate dragging. After the operation is completed, automatically take a screenshot and repeat the recognition, and enter the next process.
[0092] Step 5, after the page scrolling and submission detection complete the interaction operation, the system analyzes the page status:If the "Submit Button" within the visible area is recognized, directly execute the submission;If not recognized or the button is outside the viewport, automatically scroll the page and re-recognize until the button is detected;Establish an interaction locking mechanism to mark the operated elements to avoid repeated triggering.
[0093] Step 6, determine whether to jump to the business operation page after the registration status is submitted:Not successful: Record the status and error screenshots, mark as failed and transfer to the data module;Successful: Enter the information extraction process.
[0094] Business information extraction:After successfully entering the business operation page, the system parses the content by matching keyword fields through regular expressions, extracts information and stores it in a standardized manner. The specific steps are as follows:Step 1, page parsing and information extraction For the page screenshot or HTML source code, identify the operation method and account information by matching regular expressions with keywords, including:Payment account and account holder information; Third-party payment identifier; Electronic payment account; Other compliant business fields.
[0095] Step 2, information classification and standardization Clean the format of the extracted information (such as removing spaces, mapping standard names), and classify it according to tags such as type, source, time, etc.
[0096] Step 3, data association and storage Associate the virtual identity information used for registration with the extracted business operation information and store it in the backend database to provide support for subsequent business analysis.
[0097] Example 4: Exception and feedback mechanismThrough task status monitoring and feedback of failed samples, realize exception recognition and data reflux, and support model optimization and system iteration. The specific steps are as follows:Step 1, exception handling If the following exceptions are encountered during the registration process, interrupt the operation and record the reasons:Element recognition failure or misjudgment; Page jump exception; Captcha processing failure; Logic interruption caused by repeated operations.
[0098] Step 2, failed sample feedback The failed tasks are passed back to the data module, and screenshots and failure classifications are automatically stored for incremental model training and system optimization.
[0099] Step 3, locking mechanism and maximum number of attempts To prevent infinite loops, set the maximum number of operations for each type of UI element (the default is no more than 2 times, except for the close button, which allows up to 5 times).
[0100] The above embodiments take the business automation process as the scenario and demonstrate the technical solutions through compliance descriptions, with the key points being:Improving the generalization ability of the model through clustering and deduplication, hierarchical annotation, and genetic algorithm optimization, achieving full-process automation through priority-based task scheduling, cross-frame operations, and dynamic rolling detection, and continuously improving the system stability and adaptability through the return of failed samples. The solution can be widely applied to compliance business scenarios that require automated page interaction and data collection (such as enterprise services, digital operations, etc.).
[0101] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.< / iframe>
Claims
1. Method for automatic registration and login of business carriers and obtaining of fund accounts, characterized in that, The method includes: Step S1, accessing the target website domain name through a headless browser, taking a screenshot of the current visible area after loading is completed, and generating a page image; Step S2, inputting the page image into a pre-trained object detection model to identify the UI element categories, position coordinates, and confidence levels in the page; sorting the recognition results based on a preset priority rule to generate an interaction task queue arranged in descending order of priority; Step S3, for the elements within the frame in the interaction task queue, converting their coordinates to coordinates relative to the frame, and controlling the browser to switch to the corresponding frame; Step S4, traversing the interaction task queue and sequentially performing the following operations: performing a simulated click operation on button-type elements; injecting virtual identity information into input box-type elements; calling the corresponding recognition model to complete the verification for verification code-type elements; Step S5, page scrolling and submission detection. After completing the operation of the interaction task queue, detecting whether there is a visible submission button on the current page. If not, scrolling the page and re-executing steps S1 to S2 until a visible submission button is recognized. If visible, triggering the submission operation; Step S6, after submission, if redirected to a recharge page, extracting the payment account information through a regular expression and associating and storing it with the virtual identity information in the database.
2. The method for automatically registering and logging in a service carrier and obtaining a fund account according to claim 1, wherein Step S1, accessing the target website domain name through a headless browser, taking a screenshot of the current visible area after loading is completed, and generating a page image, including: Step S11, after the page loading is completed, detecting whether there is a pop-up window element in the current visible area that blocks the key area; if so, automatically calculating the coordinate position of the pop-up window close button and triggering a simulated click operation to expose the underlying page elements; Step S12, based on the page state with the pop-up window cleared, identifying the position of the registration button on the page; calculating the center point coordinates of the registration button, and performing coordinate positioning and clicking through the browser driver to trigger the page to jump to the registration page; Step S13, after the registration page loading is completed, controlling the browser to perform a vertical scrolling operation: the initial scrolling position is the top of the page, setting the scrolling step size by equally dividing the screen height, and scrolling down successively until the bottom of the page; performing a screenshot operation after scrolling to the bottom to capture a complete image including all form input boxes and the submission button at the bottom of the page; Step S14, after completing the registration form submission operation, monitoring the new page URL that the browser jumps to. When the page URL contains the feature keyword, it is determined as a recharge page. After the page loading is completed, directly capturing the image of the current visible area.
3. The method for automatic registration and login of a service carrier and obtaining a fund account according to claim 2, characterized in that, Step S2, inputting the page image into a pre-trained object detection model to identify the UI element categories, position coordinates, and confidence levels in the page; Sorting the recognition results based on a preset priority rule to generate an interaction task queue arranged in descending order of priority, including: Step S21, inputting the page images generated in steps S13 and S14 into a pre-trained object detection model, and outputting the recognition result set of each UI element; each recognition result includes a category label, position coordinates, and confidence level; Step S22: Assign priorities to each recognition result output in Step S21 according to the Euclidean distance between the element center point and the page center to obtain the weights corresponding to the category label, position coordinates, and confidence level respectively; according to the weights corresponding to the category label, position coordinates, and confidence level respectively, obtain the priority scores, and divide the elements into five levels from high to low according to the scores to obtain the element set with priority labels. Step S23: Sort the element set with priority labels output in Step S22 in descending order of priority, and sort the elements with the same priority in descending order of confidence level to generate a structured task queue. Step S24: Traverse the structured task queue in Step S23, calculate the normalized coordinates of each element, that is, calculate its absolute pixel coordinates in the browser viewport; through the DOM position reverse deduction algorithm, convert the absolute pixel coordinates into XPath paths.
4. The method for automatically registering and logging in a service carrier and obtaining a fund account according to claim 3, wherein Step S3: For the elements located within the frame in the interaction task queue, convert their coordinates to coordinates relative to the frame and control the browser to switch to the corresponding frame, including: Step S31: Analyze the XPath path features generated in S24. If the path contains HTML tags or frame hierarchical structures, determine that the corresponding element is located within the frame, and extract the frame identifier in the parent page, that is, the ID or Name attribute. Step S32: Relative coordinate conversion calculation. Read the pixel coordinates (X_frame, Y_frame) of the upper left corner of the frame in the parent page through the browser interface, and read the width W_frame and height H_frame of the frame itself. Calculate the relative position of the element within the frame: Relative abscissa = (element absolute abscissa - X_frame) ÷ W_frame; Relative ordinate = (element absolute ordinate - Y_frame) ÷ H_frame; Output the normalized relative coordinates. Step S33: Based on the frame identifier extracted in S31, locate the DOM node of the corresponding frame in the browser and switch the WebDriver operation context to the inside of this frame.
5. The method for automatic registration and login of a service carrier and obtaining a fund account according to claim 4, wherein, Step S4: Traverse the interaction task queue and perform the following operations in sequence: perform a simulated click operation on button elements; inject virtual identity information into input box elements; Call the corresponding recognition model to complete the verification for verification code elements, including: Step S41: Start reading task items sequentially from the head of the structured task queue generated in S23, and determine the execution order according to the priority labels in the task items. Step S42: For the current task item: If the element is located within the frame, use the relative coordinates calculated in Step S32. If it is located in the main document, use the absolute coordinates converted in Step S24; dynamically calculate the operation coordinates in the browser viewport: Main document: Directly use the absolute pixel coordinates. Within the frame: Relative coordinates × actual frame size + frame offset. Step S43: Perform corresponding operations according to the element category, button click operation, wait for 500ms page response time after clicking, form input operation, branch processing according to the verification code type, sliding verification code. Step S44, after each task item operation is completed, re-capture the visible area of the current page, and call the target detection model to verify the operation effect; if the operation fails, record the current element positioning information and re-insert the task at the end of the queue.
6. The method for automatically registering and logging in a service carrier and obtaining a fund account according to claim 5, characterized in that, Step S5, page scrolling and submission detection, after completing the interactive task queue operation, detect whether there is a visible submission button on the current page. If it is not visible, scroll the page and re-execute Step S1 to Step S2 until a visible submission button is recognized. If it is visible, trigger the submission operation, including: Step S51, based on the latest page screenshot after the operation in S44, call the target detection model to identify whether there is an element of the "submission button" type in the current visible area; Step S52, execute when the detection result in Step S51 is invisible: Obtain the total height of the current page and the height of the browser viewport, and set the single scroll amount to 80% of the viewport height; perform the scroll operation: Record the current scroll bar position as the initial position, trigger the browser to scroll down by one scroll amount unit, and wait for 500ms for the page to redraw; Status update and loop: Re-execute Step S1 to capture the new visible area image; re-execute Step S2 to generate a new interactive task queue; return to Step S51 for submission button detection, and loop until any of the following conditions is met: A visible submission button is detected, and the cumulative scroll amount exceeds the total height of the page; Step S53, during each scroll re-identification process, establish a library of marked elements that have been operated on, and filter the new interactive task queue generated in S2; Step S54, when the visible submission button is detected in Step S51, locate the button coordinates.
7. The method for automatically registering and logging in a service carrier and obtaining a fund account according to claim 6, wherein Step S6, if it jumps to the recharge page after submission, extract the payment account information through regular expressions and associate it with the virtual identity information and store it in the database, including: Step S61, based on the page jump result triggered in Step S54, monitor whether the new page URL contains a preset keyword, capture the complete HTML source code of the current page, and perform a dual verification mechanism to obtain a verified recharge page; Step S62, perform hierarchical parsing on the recharge page verified in Step S61, standardize and associate the data for storage. When the recharge page verification in Step S61 fails, package and transmit it to the data construction module.
8. The method for automatically registering and logging in a service carrier and obtaining a fund account according to claim 1, characterized in that The pre-trained target detection model includes: Extract the target website domain names from the abusive domain name library, access the domain names through an automated script and perform page loading status verification; for the websites that pass the verification, automatically trigger pop-up window closing, registration button clicking, and page scrolling operations, and capture the complete page screenshots containing key interactive elements; Based on the complete page screenshots, use the hierarchical clustering algorithm to group images with similar visual features to obtain de-duplicated images; For the de-duplicated images, label the category labels and position information of 17 types of UI elements to generate an annotation file in the target detection format; Integrate the annotation files and divide them into a training set, a validation set, and a test set according to a preset ratio; Based on the training set, validation set, and test set, perform the following automatic optimization process: Initialize the parameter population, where each set of parameters includes the learning rate, anchor box size, and network structure configuration; train the model for each parameter combination, and calculate the fitness using the validation set accuracy and loss value; iteratively generate a new population through selection, crossover, and mutation operations until the fitness converges to obtain optimized parameters; use the optimized parameters to perform fine-tuning of the entire network and fuse incremental sample training to obtain a pre-trained object detection model.
9. The method for automatically registering and logging in a service carrier and obtaining a fund account according to claim 8, wherein Calculate the fitness using the validation set accuracy and loss value, including: Based on the model training output of the current parameter combination, load its prediction results on the validation set; Count the matching situation between the predicted boxes and the ground truth boxes, that is, for each image in the validation set, calculate the overlap degree between the predicted boxes of the object detection model and the annotated boxes; Count the proportion of the number of correctly detected elements in each of the 17 UI element categories to the total number of category annotations corresponding to each category separately, and take the arithmetic mean of the accuracy values of all categories to obtain the average precision index of the validation set; Extract the CIoU loss values of all samples from the object detection model validation log, and take the arithmetic mean of the CIoU loss values of all samples to obtain the CIoU loss mean; Define the accuracy weight coefficient and the loss weight coefficient, weight the average precision index in a positive correlation, and weight the CIoU loss mean in a negative correlation to generate the fitness score.
10. An automated registration, login and fund account acquisition system for a service carrier, characterized in that, This system is used to execute the method according to any one of claims 1 to 9, including: An acquisition module for accessing the target website domain name through a headless browser, taking a screenshot of the current visible area after loading is completed, and generating a page image; A generation module for inputting the page image into the pre-trained object detection model to identify the UI element categories, position coordinates, and confidence levels in the page; sorting the recognition results based on a preset priority rule to generate an interaction task queue sorted in descending order of priority; A conversion module for converting the coordinates of the elements within the frame in the interaction task queue into coordinates relative to the frame, and controlling the browser to switch to the corresponding frame; A verification module for traversing the interaction task queue and sequentially performing the following operations: performing a simulated click operation on button elements; injecting virtual identity information into input box elements; calling the corresponding recognition model to complete the verification of verification code elements; A processing module for page scrolling and submission detection. After completing the operations of the interaction task queue, detect whether there is a visible submission button on the current page until a visible submission button is recognized. If visible, trigger the submission operation; after submission, if it jumps to the recharge page, extract the payment account information through a regular expression and associate it with the virtual identity information and store it in the database.
Citation Information
Patent Citations
Method with verification for intelligently crawling network information in distributed way
CN106897357A
Automatic login method and device, medium and equipment
CN111460356A
Bank security control automatic login method, system and device and storage medium
CN114743187A
Webpage login method and system based on fixed candidate box target detection algorithm
CN115905767A
Automatic task page popup window automatic closing method and device, equipment and medium
CN119597388A
Cited By
RPA verification code interaction method and system based on asynchronous man-machine cooperation
CN121167704A