A model training and page detection method, device, medium and equipment
By training the coordinate decoder to correct the predicted coordinates of the large language model output, the coordinate inaccuracy problem caused by the large language model being affected by language participle in page detection is solved, and the accuracy of page detection is improved.
Patent Information
- Application Number
- CN202411535888.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-30
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-10-30
AI Technical Summary
When using large language models for page detection, the prior art is affected by language participle, resulting in a large difference between the predicted coordinates and the actual coordinates, and the detection task cannot be completed accurately.
By obtaining the sample page image and its corresponding navigation text and label text, input it into the large language model to generate a predicted coordinate representation, and input it to the coordinate decoder to be trained for correction, determine the comprehensive loss value based on the difference between the predicted coordinates and the actual coordinates, and train the coordinate decoder.
The trained coordinate decoder corrects the predicted coordinates output from the large language model, which significantly improves the accuracy of page detection and reduces the difference between the predicted coordinates and the actual coordinates.
Smart Images

Figure CN119046174B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a model training and page detection method, device, medium and equipment. Background Art
[0002] At present, with the increasing popularity of smartphones, various applications have penetrated into people's daily lives. These applications provide people with various services, greatly enhance the convenience and efficiency of people's daily lives, and change the way people obtain information, social interaction, work and study, and entertain themselves. However, some applications also damage users' privacy and other rights through false advertising and other means.
[0003] Therefore, in order to protect the rights and interests of users, it is particularly necessary to detect the pages displayed by the application. The traditional detection method is to manually operate on the page displayed by the application to determine whether the result displayed by the application after the operation matches the content displayed on the application page before the operation, but the efficiency of manual detection is low.
[0004] In the current technology, a large language model can be trained first through artificial intelligence technology and massive data, and then the trained large language model can be used to automatically navigate and detect the detection task to improve the efficiency of patrol detection. Through the large language model, the description text and page image that need to be operated on the page displayed by the application are input into the large language model. The description text is used to represent the target control that needs to be touched in the page when performing automatic detection on the page corresponding to the page image, and the large language model outputs the predicted coordinates of the target control position touched in the page corresponding to the page image, and then the coordinates of the determined position can be sent to the detection program to realize automatic detection of the page corresponding to the page image. However, in the process of using the large language model to perform automatic navigation, when the large language model predicts the coordinates of the target control position that should be touched on the page image, it will be affected by the language segmentation, resulting in a large difference between the predicted coordinates and the actual coordinates, and the detection task cannot be completed accurately.
[0005] To this end, this specification provides a model training and page detection method, device, medium and equipment. Summary of the invention
[0006] This specification provides a model training and page detection method, device, medium and equipment to partially solve the above-mentioned problems existing in the prior art.
[0007] This manual adopts the following technical solutions:
[0008] This specification provides a model training method, including:
[0009] Acquire a sample page image, navigation text and label text corresponding to the sample page image, wherein the navigation text is used to indicate a target control that needs to be touched in the page when automatic detection is performed on the page corresponding to the sample page image, and the label text records the actual page coordinates of the target control in the page;
[0010] Inputting the sample page image and the navigation text into a preset large language model, so that the large language model determines an output text according to the navigation text, wherein the output text includes a predicted coordinate representation of the position of the target control in the page;
[0011] Inputting the predicted coordinate representation into a coordinate decoder to be trained to obtain the predicted coordinates of the position of the target control in the page;
[0012] A comprehensive loss value is determined based on the difference between the predicted coordinates and the actual page coordinates, so as to train the coordinate decoder based on the comprehensive loss value, and there is a positive correlation between the difference and the comprehensive loss value.
[0013] Optionally, determining a comprehensive loss value according to a difference between the predicted coordinates and the actual page coordinates specifically includes:
[0014] Determine the remaining text in the output text except the predicted coordinate representation as the first remaining text, and determine the remaining text in the label text except the actual page coordinates as the second remaining text;
[0015] Determining a first loss value according to a difference between the predicted coordinates and the actual page coordinates, and determining a second loss value according to a difference between the first remaining text and the second remaining text;
[0016] Determining the comprehensive loss value according to the first loss value and the second loss value;
[0017] According to the comprehensive loss value, the coordinate decoder is trained, specifically including:
[0018] The coordinate decoder and the large language model are jointly trained according to the comprehensive loss value.
[0019] Optionally, determining a comprehensive loss value according to a difference between the predicted coordinates and the actual page coordinates specifically includes:
[0020] Inputting the actual page coordinates into a preset coordinate encoder to obtain an actual coordinate representation corresponding to the actual page coordinates;
[0021] Determining a first loss value based on a difference between the predicted coordinates and the actual page coordinates, and determining a third loss value based on a difference between the actual coordinate representation and the predicted coordinate representation;
[0022] The comprehensive loss value is determined according to the first loss value and the third loss value.
[0023] Optionally, training the coordinate decoder according to the comprehensive loss value specifically includes:
[0024] The coordinate decoder and the coordinate encoder are jointly trained according to the comprehensive loss value.
[0025] Optionally, inputting the sample page image and the navigation text into a preset large language model specifically includes:
[0026] Inputting the sample page image into a preset image encoder to obtain image features of the sample page image, and inputting the navigation text into a preset text encoder to obtain text features of the navigation text;
[0027] splicing the text features with the image features to determine comprehensive features;
[0028] The comprehensive features are input into a preset large language model.
[0029] Optionally, training the coordinate decoder according to the comprehensive loss value specifically includes:
[0030] The coordinate decoder and the text encoder are jointly trained according to the comprehensive loss value.
[0031] Optionally, before combining the text feature with the image feature to determine the comprehensive feature, the method further includes:
[0032] Inputting the image features into a preset multi-layer perceptron to obtain converted features output by the multi-layer perceptron;
[0033] The text features and the image features are combined to determine comprehensive features, specifically including:
[0034] The text features are concatenated with the converted features to determine comprehensive features.
[0035] Optionally, training the coordinate decoder according to the comprehensive loss value specifically includes:
[0036] The coordinate decoder and the multi-layer perceptron are jointly trained according to the comprehensive loss value.
[0037] Optionally, inputting the sample page image and the navigation text into a preset large language model specifically includes:
[0038] Inputting the sample page image into a preset image encoder to obtain image features of the sample page image, and inputting the navigation text into a preset text encoder to obtain text features of the navigation text;
[0039] Inputting the image features into a preset multi-layer perceptron to obtain converted features output by the multi-layer perceptron;
[0040] splicing the text feature with the converted feature to determine a comprehensive feature;
[0041] Inputting the comprehensive features into a preset large language model;
[0042] Determining a comprehensive loss value according to the difference between the predicted coordinates and the actual page coordinates, specifically including:
[0043] Determine the remaining text in the output text except the predicted coordinate representation as the first remaining text, determine the remaining text in the label text except the actual page coordinates as the second remaining text, and input the actual page coordinates into a preset coordinate encoder to obtain the actual coordinate representation corresponding to the actual page coordinates;
[0044] Determining a first loss value based on a difference between the predicted coordinates and the actual page coordinates, determining a second loss value based on a difference between the first remaining text and the second remaining text, and determining a third loss value based on a difference between the actual coordinate representation and the predicted coordinate representation;
[0045] Determining the comprehensive loss value according to the first loss value, the second loss value and the third loss value;
[0046] The coordinate decoder, the text encoder, the multi-layer perceptron, the coordinate encoder, and the large language model are jointly trained according to the comprehensive loss value.
[0047] This specification provides a page detection method, including:
[0048] Acquire a page image and a navigation text corresponding to the page image, wherein the navigation text is used to indicate a target control that needs to be touched in the page when automatic detection is performed on the page corresponding to the page image;
[0049] Inputting the page image and the navigation text into a preset large language model, so that the large language model determines an output text according to the navigation text, wherein the output text includes a predicted coordinate representation of the position of the target control in the page;
[0050] Inputting the predicted coordinate representation into a pre-trained coordinate decoder to obtain the predicted coordinates of the position of the target control in the page, wherein the coordinate decoder is trained by the above-mentioned model training method;
[0051] The output text is adjusted according to the predicted coordinates, so as to automatically detect the page corresponding to the page image according to the adjusted output text.
[0052] This specification provides a model training device, including:
[0053] A sample acquisition module, used to acquire a sample page image, navigation text and label text corresponding to the sample page image, wherein the navigation text is used to indicate a target control that needs to be touched in the page when automatic detection is performed on the page corresponding to the sample page image, and the label text records the actual page coordinates of the target control in the page;
[0054] A sample prediction module, used for inputting the sample page image and the navigation text into a preset large language model, so that the large language model determines an output text according to the navigation text, wherein the output text contains a predicted coordinate representation of the position of the target control in the page;
[0055] A sample decoding module, used for inputting the predicted coordinate representation into a coordinate decoder to be trained to obtain the predicted coordinates of the position of the target control in the page;
[0056] A training module is used to determine a comprehensive loss value according to the difference between the predicted coordinates and the actual page coordinates, so as to train the coordinate decoder according to the comprehensive loss value, and there is a positive correlation between the difference and the comprehensive loss value.
[0057] This specification provides a page detection device, including:
[0058] A data acquisition module, used to acquire a page image and a navigation text corresponding to the page image, wherein the navigation text is used to indicate a target control that needs to be touched in the page when automatic detection is performed on the page corresponding to the page image;
[0059] A data prediction module, used for inputting the page image and the navigation text into a preset large language model, so that the large language model determines an output text according to the navigation text, wherein the output text contains a predicted coordinate representation of the position of the target control in the page;
[0060] A data decoding module, used for inputting the predicted coordinate representation into a pre-trained coordinate decoder to obtain the predicted coordinates of the position of the target control in the page, wherein the coordinate decoder is trained by the above-mentioned model training method;
[0061] A detection module is used to perform automatic detection on the page corresponding to the page image according to the predicted coordinates.
[0062] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned model training and page detection methods are implemented.
[0063] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements a model training and page detection method when executing the program.
[0064] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0065] The model training method provided in this specification obtains a sample page image, a navigation text corresponding to the sample page image, and a label text, wherein the navigation text is used to indicate the target control that needs to be touched in the page when automatic detection is performed on the page corresponding to the sample page image, and the label text records the actual page coordinates of the target control in the page. The sample page image and the navigation text are input into a preset large language model so that the large language model determines the output text according to the navigation text, and the output text contains a predicted coordinate representation of the position of the target control in the page. The predicted coordinate representation is input into the coordinate decoder to be trained to obtain the predicted coordinates of the position of the target control in the page. Based on the difference between the predicted coordinates and the actual page coordinates, a comprehensive loss value is determined, and the coordinate decoder is trained based on the comprehensive loss value, and the difference is positively correlated with the comprehensive loss value.
[0066] Through the above-mentioned model training method, the trained coordinate decoder can correct the predicted coordinate representation output by the large language model, avoid the influence of language segmentation on the large language model, and cause a large difference between the predicted coordinates and the actual coordinates, thereby improving the accuracy of page detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation on this specification. In the drawings:
[0068] Figure 1 A flowchart of a model training method provided in an embodiment of this specification;
[0069] Figure 2 A schematic diagram of a process for determining loss provided for this specification;
[0070] Figure 3 A schematic diagram of a flow chart of a page detection method provided in an embodiment of this specification;
[0071] Figure 4 A schematic diagram of a process for determining output text of a page image provided in this specification;
[0072] Figure 5 A schematic diagram of a model training device provided in this specification;
[0073] Figure 6 A schematic diagram of a page navigation detection device provided in this specification;
[0074] Figure 7 A method corresponding to the Figure 1 and Figure 3 Schematic diagram of the structure of an electronic device. DETAILED DESCRIPTION
[0075] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0076] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.
[0077] Figure 1 A flow chart of a model training method provided in an embodiment of this specification includes the following steps:
[0078] S100: Obtain a sample page image, navigation text and label text corresponding to the sample page image, wherein the navigation text is used to indicate a target control that needs to be touched in the page when automatic detection is performed on the page corresponding to the sample page image, and the label text records the actual page coordinates of the target control in the page.
[0079] In this specification, the process of positioning in model training and page navigation involves the processing of image data and text data. In the embodiments of this specification, the process of positioning in model training and page navigation can be performed by a server. Of course, this specification does not limit which device is used to perform the process of positioning in model training and page navigation. For example, devices such as personal computers and mobile terminals can also be used for model training and positioning in page navigation. For the convenience of description, the following description is based on the server as the execution subject.
[0080] In one or more embodiments of the present specification, the server obtains a sample page image, navigation text and label text corresponding to the sample page image, the navigation text is used to indicate a target control that needs to be touched in the page when automatic detection is performed on the page corresponding to the sample page image, and the label text records the actual page coordinates of the target control in the page.
[0081] For example, if the sample page image is a payment page for a product, the navigation text corresponding to the sample page image can be "Which control do you need to click to pay?" The label text can be "You need to click the position (x1, y1, x2, y2)", and "(x1, y1, x2, y2)" is the actual page coordinates.
[0082] S102: Inputting the sample page image and the navigation text into a preset large language model, so that the large language model determines an output text according to the navigation text, wherein the output text includes a predicted coordinate representation of the position of the target control in the page.
[0083] In one or more embodiments of the present specification, the server inputs the acquired sample page image and navigation text into a preset large language model, so that the large language model determines the answer output by the large language model based on the navigation text, that is, the output text, and the output text contains the predicted coordinate representation of the location of the target control in the page.
[0084] For example, if the sample page image is a payment page for a product, the navigation text may be "Which control do you need to click to pay?" The answer (i.e., output text) given by the large language model may be "You need to click <**> position," where <**> is the predicted coordinate representation of the target control's position on the page in vector form.
[0085] S104: Input the predicted coordinate representation to a coordinate decoder to be trained to obtain the predicted coordinates of the position of the target control in the page.
[0086] S106: Determine a comprehensive loss value based on the difference between the predicted coordinates and the actual page coordinates, and train the coordinate decoder based on the comprehensive loss value, wherein the difference is positively correlated with the comprehensive loss value.
[0087] In one or more embodiments of the present specification, due to the influence of the language segmentation on the large language model, the predicted coordinates differ greatly from the actual coordinates. For example, in the above example, the predicted coordinate representation of the position "<**>" given by the large language model is the coordinate representation given by the large language model itself, so it is necessary to correct the coordinate representation directly predicted by the large language model so that the predicted coordinates are as close to the actual coordinates as possible.
[0088] Therefore, in one or more embodiments of the present specification, the server inputs the predicted coordinate representation output by the large language model into the coordinate decoder to be trained, and obtains the predicted coordinates of the target control in the page decoded by the coordinate decoder. Then, according to the difference between the predicted coordinates and the actual page coordinates, the comprehensive loss value is determined, and the coordinate decoder is trained with the minimum comprehensive loss value as the optimization goal, and the difference between the predicted coordinates and the actual page coordinates is positively correlated with the comprehensive loss value.
[0089] It can be seen from the above method that the trained coordinate decoder can correct the predicted coordinate representation output by the large language model. Even if the large language model is affected by the language segmentation, the actual coordinates of the target control in the page can be accurately determined, thereby effectively improving the accuracy of page detection.
[0090] In addition to training the coordinate decoder, the large language model can also be trained in this specification, which can not only enable the trained coordinate decoder to have the ability to correct the predicted coordinate representation output by the large language model, but also reduce the representation deviation between the predicted coordinate representation output by the large language model and the actual coordinates. Based on this, in one or more embodiments of this specification, according to the difference between the predicted coordinates and the actual page coordinates, the means for determining the comprehensive loss value can be that the server determines the remaining text in the output text except the predicted coordinate representation as the first remaining text, and determines the remaining text in the label text except the actual page coordinates as the second remaining text.
[0091] The server can determine a first loss value based on the difference between the predicted coordinates and the actual page coordinates, and determine a second loss value based on the difference between the first remaining text and the second remaining text, and then determine a comprehensive loss value based on the first loss value and the second loss value.
[0092] The first loss value may be determined by calculating the Intersection over Union (IoU) loss between the predicted coordinates and the actual page coordinates, and the second loss value may be determined by calculating the Cross-entropy Loss between the first remaining text and the second remaining text. Of course, the method for determining the first loss value and the second loss value is not limited in this specification, and can be set according to actual conditions.
[0093] Afterwards, the server can jointly train the coordinate decoder to be trained and the large language model according to the determined comprehensive loss value.
[0094] In one or more embodiments of the present specification, the coordinate decoder can be trained through the above content, and a coordinate encoder can also be set in the present specification and jointly trained with the coordinate decoder. The coordinate encoder is used to convert the actual page coordinates into a representation in the form of a vector for comparison with the predicted coordinate representation output by the large language model. Specifically, the server inputs the actual page coordinates into a preset coordinate encoder to obtain the actual coordinate representation corresponding to the actual page coordinates. A first loss value is determined based on the difference between the predicted coordinates and the actual page coordinates, and a third loss value is determined based on the difference between the actual coordinate representation and the predicted coordinate representation. The server can determine a comprehensive loss value based on the first loss value and the third loss value.
[0095] The third loss value may be determined by calculating the mean square error (MSE) loss between the actual coordinate representation and the predicted coordinate representation. Of course, there is no limitation on determining the third loss value in this specification, and it can be set according to actual conditions.
[0096] After that, the server can jointly train the coordinate decoder and the coordinate encoder according to the comprehensive loss value. Of course, the server can also jointly train the coordinate decoder, the coordinate encoder and the large language model according to this comprehensive loss value. By determining the first loss value and the third loss value, the coordinate encoder and the coordinate decoder can supervise each other during the training phase to complete the training, so that the trained coordinate encoder can have a certain correction function for the predicted coordinate representation output by the large language model, thereby improving the accuracy of the coordinate decoder after decoding the predicted coordinate representation.
[0097] In one or more embodiments of the present specification, before the server inputs the sample page image and the navigation text into the large language model, the server may first input the sample page image into a preset image encoder to obtain the image features of the sample page image, and input the navigation text into a preset text encoder to obtain the text features of the navigation text. The text features and the image features are spliced to determine the comprehensive features, and the comprehensive features are input into the preset large language model. Of course, the navigation text may also be first segmented and then input into the preset text encoder to obtain the text features of the navigation text.
[0098] The pre-trained image encoder may be a VIT-large of Contrastive Language-Image Pre-Training (CLIP) or a Residual Neural Network (ResNet) model, which is not limited in this specification.
[0099] Afterwards, the server can jointly train the coordinate decoder and the text encoder based on the comprehensive loss value, so that the trained coordinate decoder can correct the output of the large language model, and can also make the trained text encoder encode the navigation text more accurately.
[0100] In one or more embodiments of the present specification, after obtaining the image features of the sample page image by inputting the sample page image into a preset image encoder, the image features are input into a preset multi-layer perceptron to obtain converted features output by the multi-layer perceptron. The text features are spliced with the converted features to determine the comprehensive features. The main function of the multi-layer perceptron here is that the image features and the text features belong to features of two different modalities, and in order to enable them to be spliced to obtain the comprehensive features, the image features need to be converted into features that can be spliced with the text features through the multi-layer perceptron.
[0101] Afterwards, the server can jointly train the coordinate decoder and the multi-layer perceptron according to the comprehensive loss value. Of course, the server can also jointly train the coordinate decoder, the multi-layer perceptron, and the text encoder according to this comprehensive loss value. This can make the trained multi-layer perceptron more accurate in converting the image features of the sample page image, and enable the large language model to better understand the comprehensive features.
[0102] In the process of training the above models, the training is mainly centered around the coordinate decoder, and the remaining different models can be jointly trained according to different needs. For example, in order to reduce the deviation between the predicted coordinate representation output by the large language model and the actual coordinate representation, it can be seen from the above content that the large language model and the coordinate decoder can be jointly trained. Or in order to make the encoding result of the navigation text by the text encoder more accurate, the text encoder and the coordinate decoder can be jointly trained. Of course, in practical applications, the above-mentioned text encoder, multi-layer perceptron, large language model, coordinate decoder, etc. can also participate in the training together, and the models of joint training will not be listed one by one here.
[0103] In one or more embodiments of the present specification, the server may also jointly train the above-mentioned models, specifically inputting the sample page image into a preset image encoder to obtain the image features of the sample page image, and inputting the navigation text into a preset text encoder to obtain the text features of the navigation text. The image features are input into a preset multi-layer perceptron to obtain the converted features output by the multi-layer perceptron. The text features are concatenated with the converted features to determine the comprehensive features. The comprehensive features are input into a preset large language model so that the large language model determines the output text based on the navigation text, and the output text contains the predicted coordinate representation of the position of the target control in the page. The predicted coordinate representation is input into the coordinate decoder to be trained to obtain the predicted coordinates of the position of the target control in the page.
[0104] Afterwards, the server determines the remaining text in the output text except the predicted coordinate representation as the first remaining text, and determines the remaining text in the label text except the actual page coordinates as the second remaining text, and inputs the actual page coordinates into a preset coordinate encoder to obtain the actual coordinate representation corresponding to the actual page coordinates.
[0105] Then, the server determines a first loss value based on the difference between the predicted coordinates and the actual page coordinates, determines a second loss value based on the difference between the first remaining text and the second remaining text, and determines a third loss value based on the difference between the actual coordinate representation and the predicted coordinate representation. Then, a comprehensive loss value is determined based on the first loss value, the second loss value, and the third loss value. Finally, based on the comprehensive loss value, the coordinate decoder, the text encoder, the multi-layer perceptron, the coordinate encoder, and the large language model can be jointly trained.
[0106] It is worth noting that in one or more embodiments of the present specification, the specific method of determining the second loss value according to the difference between the first remaining text and the second remaining text may be to segment the first remaining text and the second remaining text, determine the order of each segmentation in the first remaining text, and determine the order of each segmentation in the second remaining text. For each segmentation in the second remaining text, according to the position of the segmentation in the second remaining text, determine the segmentation in the first remaining text with the same order as the segmentation as the contrasting segmentation corresponding to the segmentation, and determine the second loss value of the segmentation and the contrasting segmentation corresponding to the segmentation according to the difference between the segmentation and the contrasting segmentation corresponding to the segmentation. Afterwards, the server may use the sum of the second loss values of each segmentation in the second remaining text and its corresponding contrasting segmentation as the difference between the first remaining text and the second remaining text to determine the second loss value.
[0107] In one or more embodiments of the present specification, the label text may be segmented to obtain each segmentation. Among these segmentations, some segmentations express coordinate information, that is, actual page coordinates, and the identification of such segmentations after text segmentation is the coordinate identification. Other segmentations express non-coordinate information, and the identification of such segmentations is not a coordinate identification. At the end of the label text, there is also a segmentation indicating the end of the label text, and the identification of such segmentation is the end identification.
[0108] Therefore, when determining the comprehensive loss value based on the difference between the predicted coordinates and the actual page coordinates, the server can also determine whether the identifier of each segmentation in the tag text is an end identifier. If so, the comprehensive loss value is no longer calculated. If not, it is determined whether the identifier of the segmentation is a coordinate identifier. If the identifier of the segmentation is not a coordinate identifier, the order of the segmentation in the second remaining text is determined, and the position of the segmentation in the second remaining text is determined according to the order of the segmentation, and the segmentation in the same position as the segmentation in the first remaining text is determined according to the position of the segmentation as the contrasting segmentation corresponding to the segmentation. Then, the server can determine the second loss value of the segmentation and the contrasting segmentation corresponding to the segmentation according to the difference between the segmentation and the contrasting segmentation corresponding to the segmentation.
[0109] If the identifier of the word segment is a coordinate identifier, the server can determine the first loss value based on the difference between the word segment and the actual page coordinates; determine the third loss value based on the difference between the actual coordinate representation and the predicted coordinate representation obtained by the word segment input coordinate decoder. Finally, the server can determine the comprehensive loss value based on the sum of the first loss value, the second loss value, and the third loss value.
[0110] Figure 2 A schematic diagram of a process for determining loss is provided in this manual. Figure 2As shown, for the i-th identifier of the label text, determine whether the identifier is an end identifier. If so, end the current training model process, and then select a new sample page image to continue training or completely end the training. If the identifier is not an end identifier, determine whether the identifier is a coordinate identifier. If not, calculate the cross entropy loss L1 (i.e., the second loss value mentioned above) between the segmentation corresponding to the identifier and the contrastive segmentation, and then add L1 to the sum of the losses of the identifiers before the i-th identifier to obtain a new loss sum. After adding i to 1, continue the same process for the next identifier. If the identifier is a coordinate identifier, input the segmentation corresponding to the identifier into the coordinate encoder to be trained to obtain the actual coordinate representation. Determine the mean square error loss L2 (i.e., the third loss value mentioned above) between the actual coordinate representation and the predicted coordinate representation, and then add L2 to the sum of the losses of the identifiers before the i-th identifier to obtain a new loss sum. The predicted coordinate representation is then input into the coordinate decoder to be trained to obtain the predicted coordinates, determine the intersection-over-union loss L3 (i.e., the first loss value mentioned above) between the predicted coordinates and the word corresponding to the identifier, add L3 to the sum of the new losses, update the sum of the losses, and finally add 1 to i and continue the same process for the next identifier.
[0111] Based on the above model training method, the coordinate decoder is trained, and then the trained coordinate decoder can be made capable of correcting the predicted coordinates output by the large language model. Therefore, the coordinate decoder can be deployed in the actual page detection process to perform page detection.
[0112] Figure 3 A flow chart of a page detection method provided in an embodiment of this specification includes the following steps:
[0113] S300: Acquire a page image and a navigation text corresponding to the page image, wherein the navigation text is used to indicate a target control that needs to be touched in the page when automatic detection is performed on the page corresponding to the page image.
[0114] In one or more embodiments of the present specification, the server obtains a page image and a navigation text corresponding to the page image, and the navigation text is used to indicate a target control that needs to be touched in the page when automatic detection is performed on the page corresponding to the page image. For example, if the page image is a payment page for a product, the navigation text corresponding to the page image may be "Which control needs to be clicked for payment".
[0115] S302: Input the page image and the navigation text into a preset large language model, so that the large language model determines an output text according to the navigation text, wherein the output text includes a predicted coordinate representation of the position of the target control in the page.
[0116] In one or more embodiments of the present specification, the server inputs the page image and the navigation text into a preset large language model, so that the large language model determines the output text according to the navigation text, and the output text contains the predicted coordinate representation of the location of the target control in the page. For example, for a payment page with a page image of a commodity, the navigation text corresponding to the page image may be "Which control do you need to click to pay", and the answer (i.e., the output text) given by the large language model may be "You need to click <**> position", where <**> is the predicted coordinate representation of the position.
[0117] S304: Input the predicted coordinate representation into a pre-trained coordinate decoder to obtain the predicted coordinates of the position of the target control in the page, wherein the coordinate decoder is trained by the above-mentioned model training method.
[0118] In one or more embodiments of the present specification, the server inputs the predicted coordinate representation into the trained coordinate decoder to obtain the predicted coordinates of the position of the target control in the page. For example, the predicted coordinate representation <**> is input into the coordinate decoder to obtain the predicted coordinates "(x1, y1, x2, y2)" of the position of the target control in the page, and then the output text is adjusted according to the decoded predicted coordinates or the page detection is directly performed.
[0119] S306: Performing automatic detection on a page corresponding to the page image according to the predicted coordinates.
[0120] In one or more embodiments of the present specification, the server performs automatic detection on the page corresponding to the page image based on the predicted coordinates. Specifically, after obtaining the predicted coordinates, an automatic detection instruction can be generated to perform an automatic detection task on the page corresponding to the page image, thereby obtaining a detection result, and the page can be subsequently repaired or adjusted based on the obtained detection result.
[0121] Figure 4 The following is a flow chart of determining the output text of a page image provided in this specification. Figure 4As shown, the page image is passed through the image encoder and the multi-layer perceptron to determine the image features of the page image. The navigation text is input into the text encoder to determine the text features of the navigation text. The text feature sequence and the image feature sequence are concatenated and input into the large language model to obtain the output text of the large language model, that is, "You need to click <**> position". The "<**>" coordinate representation in the output text is input into the coordinate decoder to obtain the decoding result of the coordinate decoder "(x1, y1, x2, y2)". After that, the "<**>" in the output text is replaced with the decoding result "(x1, y1, x2, y2)" to obtain the output text of the page "You need to click (x1, y1, x2, y2) position".
[0122] In one or more embodiments of the present specification, before training the multi-layer perceptron to be trained, the text encoder to be trained, the large language model to be trained, the coordinate encoder to be trained, and the coordinate decoder to be trained, the large language model, the multi-layer perceptron, and the image encoder involved in the training process can be initialized. The initialization parameter standard is not limited in the present specification. For example, the parameters of LLaVA1.6 can be used for initialization. Among them, the large language model can adopt the structure of LLaMA2, which is suitable for generation tasks. There is no limitation on the large language model in the present specification. It can also be a Generative Pre-Trained Transformer (GPT) model, such as GPT3, GPT3.5 and GPT4 models can be used in this solution. There is no limitation on the GPT model version, which can be set according to actual conditions.
[0123] Based on the same idea of a model training method provided in one or more embodiments of this specification, this specification also provides a corresponding model training device, such as Figure 5 shown.
[0124] Figure 5 A schematic diagram of a model training device provided in this specification, specifically including:
[0125] The sample acquisition module 500 is used to acquire a sample page image, a navigation text and a label text corresponding to the sample page image, wherein the navigation text is used to indicate a target control that needs to be touched in the page when automatic detection is performed on the page corresponding to the sample page image, and the label text records the actual page coordinates of the target control in the page;
[0126] The sample prediction module 502 is used to input the sample page image and the navigation text into a preset large language model, so that the large language model determines an output text according to the navigation text, and the output text contains a predicted coordinate representation of the position of the target control in the page;
[0127] A sample decoding module 504 is used to input the predicted coordinate representation into a coordinate decoder to be trained to obtain the predicted coordinates of the position of the target control in the page;
[0128] The training module 506 is used to determine a comprehensive loss value according to the difference between the predicted coordinates and the actual page coordinates, so as to train the coordinate decoder according to the comprehensive loss value, and there is a positive correlation between the difference and the comprehensive loss value.
[0129] Optionally, the training module 506 is used to determine the remaining text in the output text except the predicted coordinate representation as the first remaining text, and determine the remaining text in the label text except the actual page coordinates as the second remaining text, determine a first loss value based on the difference between the predicted coordinates and the actual page coordinates, and determine a second loss value based on the difference between the first remaining text and the second remaining text, determine the comprehensive loss value based on the first loss value and the second loss value, and jointly train the coordinate decoder and the large language model based on the comprehensive loss value.
[0130] Optionally, the training module 506 is used to input the actual page coordinates into a preset coordinate encoder to obtain an actual coordinate representation corresponding to the actual page coordinates, determine a first loss value based on the difference between the predicted coordinates and the actual page coordinates, and determine a third loss value based on the difference between the actual coordinate representation and the predicted coordinate representation, and determine the comprehensive loss value based on the first loss value and the third loss value.
[0131] Optionally, the training module 506 is used to jointly train the coordinate decoder and the coordinate encoder according to the comprehensive loss value.
[0132] Optionally, the sample prediction module 502 is used to input the sample page image into a preset image encoder to obtain image features of the sample page image, and to input the navigation text into a preset text encoder to obtain text features of the navigation text, concatenate the text features with the image features, determine comprehensive features, and input the comprehensive features into a preset large language model.
[0133] Optionally, the training module 506 is used to jointly train the coordinate decoder and the text encoder according to the comprehensive loss value.
[0134] Optionally, the training module 506 is used to input the image features into a preset multi-layer perceptron to obtain converted features output by the multi-layer perceptron, and to concatenate the text features with the converted features to determine comprehensive features.
[0135] Optionally, the training module 506 is used to jointly train the coordinate decoder and the multi-layer perceptron according to the comprehensive loss value.
[0136] Optionally, the sample prediction module 502 is used to input the sample page image into a preset image encoder to obtain image features of the sample page image, and input the navigation text into a preset text encoder to obtain text features of the navigation text, input the image features into a preset multi-layer perceptron to obtain converted features output by the multi-layer perceptron, concatenate the text features with the converted features to determine comprehensive features, and input the comprehensive features into a preset large language model;
[0137] The training module 506 is used to determine the remaining text in the output text except the predicted coordinate representation as the first remaining text, and determine the remaining text in the label text except the actual page coordinates as the second remaining text, and input the actual page coordinates into a preset coordinate encoder to obtain the actual coordinate representation corresponding to the actual page coordinates, determine a first loss value based on the difference between the predicted coordinates and the actual page coordinates, determine a second loss value based on the difference between the first remaining text and the second remaining text, and determine a third loss value based on the difference between the actual coordinate representation and the predicted coordinate representation, determine the comprehensive loss value based on the first loss value, the second loss value and the third loss value, and jointly train the coordinate decoder, the text encoder, the multi-layer perceptron, the coordinate encoder and the large language model based on the comprehensive loss value.
[0138] Based on the same idea as a page detection method provided in one or more embodiments of this specification, this specification also provides a corresponding page detection device, such as Figure 6 shown.
[0139] Figure 6 A schematic diagram of a page detection device provided in this specification specifically includes:
[0140] A data acquisition module 600 is used to acquire a page image and a navigation text corresponding to the page image, wherein the navigation text is used to indicate a target control that needs to be touched in the page when automatic detection is performed on the page corresponding to the page image;
[0141] A data prediction module 602 is used to input the page image and the navigation text into a preset large language model, so that the large language model determines an output text according to the navigation text, and the output text contains a predicted coordinate representation of the position of the target control in the page;
[0142] A data decoding module 604 is used to input the predicted coordinate representation into a pre-trained coordinate decoder to obtain the predicted coordinates of the position of the target control in the page, wherein the coordinate decoder is trained by the above-mentioned model training method;
[0143] The detection module 606 is used to perform automatic detection on the page corresponding to the page image according to the predicted coordinates.
[0144] This specification also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 and Figure 3 A model training and page detection method is provided.
[0145] This manual also provides Figure 7 The schematic structure diagram of the electronic device shown in FIG. Figure 7 As shown, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 and Figure 3 The model training and page detection methods described.
[0146] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the executor of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0147] In the 1990s, it was very clear whether the improvement of a technology was hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with hardware entity modules. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0148] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.
[0149] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0150] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0151] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0152] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0153] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0155] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0156] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0157] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0158] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0159] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0160] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0161] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0162] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.
Claims
1. A model training method, comprising: Acquire a sample page image, navigation text and label text corresponding to the sample page image, wherein the navigation text is used to indicate a target control that needs to be touched in the page when automatic detection is performed on the page corresponding to the sample page image, and the label text records the actual page coordinates of the target control in the page; Inputting the sample page image and the navigation text into a preset large language model, so that the large language model determines an output text according to the navigation text, wherein the output text includes a predicted coordinate representation of the position of the target control in the page; Inputting the predicted coordinate representation into a coordinate decoder to be trained to obtain the predicted coordinates of the position of the target control in the page; A comprehensive loss value is determined based on the difference between the predicted coordinates and the actual page coordinates, so as to train the coordinate decoder based on the comprehensive loss value, and there is a positive correlation between the difference and the comprehensive loss value.
2. The method according to claim 1, determining the comprehensive loss value according to the difference between the predicted coordinates and the actual page coordinates, specifically comprising: Determine the remaining text in the output text except the predicted coordinate representation as the first remaining text, and determine the remaining text in the label text except the actual page coordinates as the second remaining text; Determining a first loss value according to a difference between the predicted coordinates and the actual page coordinates, and determining a second loss value according to a difference between the first remaining text and the second remaining text; Determining the comprehensive loss value according to the first loss value and the second loss value; According to the comprehensive loss value, the coordinate decoder is trained, specifically including: The coordinate decoder and the large language model are jointly trained according to the comprehensive loss value.
3. The method according to claim 1 or 2, wherein determining a comprehensive loss value according to the difference between the predicted coordinates and the actual page coordinates comprises: Inputting the actual page coordinates into a preset coordinate encoder to obtain an actual coordinate representation corresponding to the actual page coordinates; Determining a first loss value based on a difference between the predicted coordinates and the actual page coordinates, and determining a third loss value based on a difference between the actual coordinate representation and the predicted coordinate representation; The comprehensive loss value is determined according to the first loss value and the third loss value.
4. The method according to claim 3, training the coordinate decoder according to the comprehensive loss value, specifically comprising: The coordinate decoder and the coordinate encoder are jointly trained according to the comprehensive loss value.
5. The method according to claim 1 or 2, inputting the sample page image and the navigation text into a preset large language model, specifically comprising: Inputting the sample page image into a preset image encoder to obtain image features of the sample page image, and inputting the navigation text into a preset text encoder to obtain text features of the navigation text; splicing the text features with the image features to determine comprehensive features; The comprehensive features are input into a preset large language model.
6. The method according to claim 5, training the coordinate decoder according to the comprehensive loss value, specifically comprising: The coordinate decoder and the text encoder are jointly trained according to the comprehensive loss value.
7. The method according to claim 1 or 2, before combining the text feature with the image feature to determine the comprehensive feature, the method further comprises: Inputting the image features into a preset multi-layer perceptron to obtain converted features output by the multi-layer perceptron; The text features and the image features are combined to determine comprehensive features, specifically including: The text features are concatenated with the converted features to determine comprehensive features.
8. The method according to claim 7, training the coordinate decoder according to the comprehensive loss value, specifically comprising: The coordinate decoder and the multi-layer perceptron are jointly trained according to the comprehensive loss value.
9. The method according to claim 1, inputting the sample page image and the navigation text into a preset large language model, specifically comprising: Inputting the sample page image into a preset image encoder to obtain image features of the sample page image, and inputting the navigation text into a preset text encoder to obtain text features of the navigation text; Inputting the image features into a preset multi-layer perceptron to obtain converted features output by the multi-layer perceptron; splicing the text feature with the converted feature to determine a comprehensive feature; Inputting the comprehensive features into a preset large language model; Determining a comprehensive loss value according to the difference between the predicted coordinates and the actual page coordinates, specifically including: Determine the remaining text in the output text except the predicted coordinate representation as the first remaining text, determine the remaining text in the label text except the actual page coordinates as the second remaining text, and input the actual page coordinates into a preset coordinate encoder to obtain the actual coordinate representation corresponding to the actual page coordinates; Determining a first loss value based on a difference between the predicted coordinates and the actual page coordinates, determining a second loss value based on a difference between the first remaining text and the second remaining text, and determining a third loss value based on a difference between the actual coordinate representation and the predicted coordinate representation; Determining the comprehensive loss value according to the first loss value, the second loss value and the third loss value; The coordinate decoder, the text encoder, the multi-layer perceptron, the coordinate encoder, and the large language model are jointly trained according to the comprehensive loss value.
10. A page detection method, comprising: Acquire a page image and a navigation text corresponding to the page image, wherein the navigation text is used to indicate a target control that needs to be touched in the page when automatic detection is performed on the page corresponding to the page image; Inputting the page image and the navigation text into a preset large language model, so that the large language model determines an output text according to the navigation text, wherein the output text includes a predicted coordinate representation of the position of the target control in the page; Inputting the predicted coordinate representation into a pre-trained coordinate decoder to obtain the predicted coordinates of the position of the target control in the page, wherein the coordinate decoder is trained by the method according to any one of claims 1 to 9; Automatically detecting a page corresponding to the page image according to the predicted coordinates.
11. A model training device, comprising: A sample acquisition module, used to acquire a sample page image, navigation text and label text corresponding to the sample page image, wherein the navigation text is used to indicate a target control that needs to be touched in the page when automatic detection is performed on the page corresponding to the sample page image, and the label text records the actual page coordinates of the target control in the page; A sample prediction module, used for inputting the sample page image and the navigation text into a preset large language model, so that the large language model determines an output text according to the navigation text, wherein the output text contains a predicted coordinate representation of the position of the target control in the page; A sample decoding module, used for inputting the predicted coordinate representation into a coordinate decoder to be trained to obtain the predicted coordinates of the position of the target control in the page; A training module is used to determine a comprehensive loss value according to the difference between the predicted coordinates and the actual page coordinates, so as to train the coordinate decoder according to the comprehensive loss value, and there is a positive correlation between the difference and the comprehensive loss value.
12. A page detection device, comprising: A data acquisition module, used to acquire a page image and a navigation text corresponding to the page image, wherein the navigation text is used to indicate a target control that needs to be touched in the page when automatic detection is performed on the page corresponding to the page image; A data prediction module, used for inputting the page image and the navigation text into a preset large language model, so that the large language model determines an output text according to the navigation text, wherein the output text contains a predicted coordinate representation of the position of the target control in the page; A data decoding module, used for inputting the predicted coordinate representation into a pre-trained coordinate decoder to obtain the predicted coordinates of the position of the target control in the page, wherein the coordinate decoder is trained by the method according to any one of claims 1 to 9; A detection module is used to perform automatic detection on the page corresponding to the page image according to the predicted coordinates.
13. A computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 10 when executing the program.
Citation Information
Patent Citations
Language model pre-training method and device, medium and electronic equipment
CN116502176A
Data annotation method and device, medium and equipment
CN118410191A