Visual semantic fusion page element positioning method based on large model
Through the visual semantic fusion method based on the big model, feature groups are constructed and confidence is evaluated, and the positioning processing inaccuracy caused by dynamic page element changes is solved, achieving higher recognition accuracy and reliability.
Patent Information
- Application Number
- CN202510713864.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-30
AI Technical Summary
In the prior art, due to changes in dynamic page elements in the front-end network interface, the reliability and accuracy of image positioning processing are difficult to meet the requirements.
The visual semantic fusion method based on the big model is adopted. By determining the page elements associated with user instructions, the features of different dimensions are extracted, multiple feature groups are constructed, the confidence and deviation of the trusted feature groups are evaluated, and the positioning processing method of page elements is determined based on the deviation of the associated page elements.
It improves the accuracy and reliability of page elements identification processing, takes into account interference risks and deviations of dynamic page elements, and improves the confidence and accuracy of positioning processing.
Smart Images

Figure CN120257213A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method for positioning page elements based on visual-semantic fusion of large models. Background Art
[0002] In order to implement the positioning process for the front-end web interface, in the invention patent application CN202410077624.7, "Method, Device, Equipment and Medium for Page Element Positioning", the first similarity between the target image in the target area and the first image is calculated to determine whether the element to be positioned is successfully located in the page image based on the first similarity. However, there are the following technical problems: The front-end page is not static. There are not only a large number of dynamic page elements, but also, with the revision of the front-end page, due to the change of image elements, it is often impossible to accurately implement the positioning process of page elements by using the image method, so that the reliability and accuracy of the positioning process are difficult to meet the requirements.
[0003] To solve the above technical problems, the present application provides a method for positioning page elements based on visual-semantic fusion of large models. Summary of the Invention
[0004] To achieve the object of the present invention, the present invention adopts the following technical solutions: Specifically, the present application provides a method for positioning page elements based on visual-semantic fusion of large models, which specifically includes: S1 Use a large model to determine the page elements associated with the user instruction, and based on the recognition result of the dynamic page elements of the front-end page corresponding to the page elements, and in combination with the deviation situation of the page elements during the change process of the dynamic page elements, when it is determined that the recognition deviation situation of the page elements does not meet the requirements, proceed to the next step; S2 Based on the associated data of the front-end page, extract features of different dimensions of the page elements, construct multiple feature groups based on the features, and determine the credible feature groups in the feature groups according to the similarity of the features of different feature groups to other page elements; S3 Use the other page elements with similar features to the page elements in the credible feature groups as associated page elements, and determine the confidence of the features of different dimensions in the credible feature groups according to the similarity of the features of different dimensions of the associated page elements to the page elements in the credible feature groups; S4 Determine the deviation situation of the associated page elements between different credible feature groups, and in combination with the confidence of the features of different dimensions in the credible feature groups, determine the positioning process method of the page elements.
[0005] The beneficial effects of the present invention are as follows: Based on the similarity of the features of associated page elements in different dimensions within the trusted feature group, the confidence levels of the features in different dimensions in the trusted feature group are determined. Considering the deviation of the features of the associated page elements at risk of interference in a certain dimension, and the resulting differences in the recognition accuracy of the features for the associated page elements at risk of interference, it also lays a foundation for evaluating the confidence levels of the features based on the differences in the recognition accuracy of the associated page elements at risk of interference, improving the accuracy of the confidence level evaluation process.
[0006] Based on the deviation of the associated page elements between different trusted feature groups and the confidence levels of the features in different dimensions in the trusted feature group, the positioning processing method of the page elements is determined. It not only considers the differences in the confidence levels of the recognition processing of the trusted feature group alone, but also takes into account the recognition interference of the associated page elements when combined with other trusted feature groups, improving the reliability of the recognition processing of the page elements.
[0007] A further technical solution is that the associated page elements are determined according to the semantic recognition results of the large model based on the operation instructions of the user.
[0008] A further technical solution is that the deviation of the page elements during the change process of the dynamic page elements is determined according to the deviation of the page images of the page elements in different front-end pages during the change process.
[0009] A further technical solution is that determining that the recognition deviation of the page elements does not meet the requirements specifically includes: Based on the recognition results of the dynamic page elements of the front-end page corresponding to the page, determine the number of dynamic page elements in the front-end page; According to the change situation of the page elements of the dynamic page elements in the front-end page of different dynamic page elements during the change process, determine and identify the dynamic page elements whose image positions of the page elements in the front-end page have changed, and use them as the changed page elements; Based on the changed update data of the changed page elements, determine whether the recognition deviation of the page elements meets the requirements.
[0010] A further technical solution is that when there are changed page elements with a change update cycle less than the preset update cycle threshold, it is determined that the recognition deviation of the page elements does not meet the requirements.
[0011] A further technical solution is that the method for determining the positioning processing method of the page elements is: Based on the sum of the confidence levels of the features in different dimensions of different trusted feature groups, determine the group confidence levels of different trusted feature groups; Determine the same associated page elements among different trusted feature groups according to the deviation of the associated page elements among different trusted feature groups, and use them as the same page elements; Determine the positioning processing method of the page element according to the group confidence of different trusted feature groups and the same page elements among different trusted feature groups.
[0012] A further technical solution is that, according to the group confidence of different trusted feature groups and the same page elements among different trusted feature groups, determine the positioning processing method of the page element, specifically including: When there is a trusted feature group with a group confidence greater than the preset group confidence threshold, use the features corresponding to the trusted feature group with the maximum group confidence to perform the positioning processing of the page element; When there is no trusted feature group with a group confidence greater than the preset group confidence threshold, based on the preset group confidence threshold, freely combine the trusted feature groups to obtain alternative positioning schemes, use the trusted feature groups corresponding to the alternative positioning schemes as alternative feature groups, and determine the positioning processing method of the page element according to the same page elements among the alternative feature groups.
[0013] A further technical solution is that the alternative positioning scheme is constructed based on the sum of the group confidences of the composed trusted feature groups being greater than the preset group confidence threshold.
[0014] A further technical solution is that, according to the same page elements among the alternative feature groups, determine the positioning processing method of the page element, specifically including: Determine the same page elements among the alternative feature groups, and use the average value of the feature similarity coefficients of the same page elements between the features corresponding to different alternative feature groups to determine the similarity coefficient mean of the same page elements in different alternative feature groups; Aim to minimize the number of the same page elements with the minimum similarity coefficient mean in different alternative feature groups being greater than the preset value of the similarity feature coefficient, determine the recognition scheme in the alternative positioning scheme, and use the recognition scheme to perform the positioning processing of the page element.
[0015] A further technical solution is that, using the recognition scheme to perform the positioning processing of the page element, specifically including: Based on the positioning results of the alternative feature groups in the recognition scheme, use the positioning result with the largest number of corresponding alternative feature groups as the positioning processing result of the page element.
[0016] Other features and advantages will be set forth in the following description, and the objects and other advantages of the invention are realized and attained by the structure particularly pointed out in the specification and the drawings.
[0017] To make the above objects, features, and advantages of the present invention more comprehensible, the following preferred embodiments are specifically exemplified and described in detail in conjunction with the accompanying drawings as follows. Brief Description of the Drawings
[0018] By referring to the drawings and describing in detail its exemplary embodiments, the above and other features and advantages of the present invention will become more apparent.
[0019] Figure 1 is a flowchart of a method for page element localization based on large model visual-semantic fusion; Figure 2 is a flowchart for determining that the recognition deviation situation of page elements does not meet the requirements; Figure 3 is a flowchart of a method for determining a reliable feature group in a feature group; Figure 4 is a flowchart of a method for determining a page element localization processing method. Detailed Description of the Embodiments
[0020] To enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0021] In this application, page element localization processing is performed based on visual features, text features, and structural features.
[0022] Specifically, for example, the user instruction: "Reduce the price of all products with inventory greater than 100 by 10%"; Execution process: User ->> System: Input a natural language instruction; System ->> Visual module: Capture the product list page; Visual module -->> System: Return the coordinates of the price input box (conf = 0.91); System ->> Semantic module: Parse "inventory > 100" and "price reduction of 10%"; Semantic module -->> System: Map to DOM attributes data-stock and data-price; System ->> Fusion Engine: α = 0.3, β = 0.6, γ = 0.1. Here, α is the weight of visual features, β is the weight of text features, and γ is the weight of structural features. Visual features and structural features are extracted using a visual recognition model, and text features are extracted according to a semantic recognition model; Fusion Engine -->> System: Comprehensive confidence level 0.87; System ->> RPA Executor: Perform batch price modification operations.
[0023] Dynamic page elements whose image positions of page elements on the front - end page change are regarded as changed page elements. When there are changed page elements with a change update period less than 3 seconds, it is determined that the recognition deviation situation of the page elements does not meet the requirements.
[0024] It should be noted that the large - model in this application is built using deepSeek R1.
[0025] The features of page elements in a feature group are used as group - matching features. Determine the feature similarity coefficient between page elements and other page elements among different group - matching features. Group - matching features with a feature similarity coefficient greater than a preset similarity threshold are used as similar features. Other page elements whose group - matching features all belong to similar features are used as interfering page elements. When the number of interfering page elements is less than 3, it is determined that the feature group is a credible feature group.
[0026] Associate page elements whose features belong to similar features are used as filtered associated elements. Based on the proportion of the number of filtered associated elements in the associated page elements of the credible feature group, determine the similarity interference coefficient of the feature in the credible feature group. Based on the average value of the feature similarity coefficients between the associated page elements of the credible feature group and the features of the page elements, determine the average similarity coefficient. Based on the product of the similarity interference coefficient and the average similarity coefficient, determine the interference value. The confidence level of the feature is based on the difference between 1 and the interference value for different features.
[0027] Based on the sum of the confidence levels of features in different dimensions of different credible feature groups, determine the group confidence level of different credible feature groups. According to the deviation situation of associated page elements between different credible feature groups, determine the same associated page elements between different credible feature groups and use them as the same page elements. Based on the group confidence level of different credible feature groups and the same page elements between different credible feature groups, determine the positioning processing method of the page elements.
[0028] As Figure 1 shown, this application provides a page - element positioning method based on visual - semantic fusion of a large - model, specifically including: S1 uses a large model to determine the page elements associated with the user's instruction, based on the recognition results of the dynamic page elements of the front-end page corresponding to the page elements, and combines the deviation of the page elements during the change process of the dynamic page elements. When it is determined that the recognition deviation of the page elements does not meet the requirements, proceed to the next step; Further, the associated page elements are determined according to the operation instruction of the user and the semantic recognition result of the large model.
[0029] Specifically, the deviation of the page elements during the change process of the dynamic page elements is determined according to the deviation of the page images of the page elements in different front-end pages during the change process.
[0030] Specifically, as Figure 2 shown, determining that the recognition deviation of the page elements does not meet the requirements specifically includes: Based on the recognition results of the dynamic page elements of the front-end page corresponding to the page, determine the number of dynamic page elements in the front-end page; According to the change situation of the page elements of the dynamic page elements in the front-end pages of different dynamic page elements during the change process, determine the dynamic page elements whose image positions of the page elements in the front-end page change during the change process, and use them as the changed page elements; Based on the changed data updated by the changed page elements, determine whether the recognition deviation of the page elements meets the requirements.
[0031] Further, when there are changed page elements with a change update period less than the preset update period threshold, it is determined that the recognition deviation of the page elements does not meet the requirements.
[0032] It can be understood that when the recognition deviation of the page elements meets the requirements, the positioning process of the page elements is performed by using the image of the front-end page.
[0033] In another possible embodiment, determining that the recognition deviation of the page elements does not meet the requirements specifically includes: Based on the recognition results of the dynamic page elements of the front-end page corresponding to the page, determine the number of dynamic page elements in the front-end page; According to the change situation of the page elements of the dynamic page elements in the front-end pages of different dynamic page elements during the change process, determine the number of times the image positions of the page elements in the front-end page change during the change process of different dynamic page elements, and use it as the number of changes; Based on the number of changes, determine whether the recognition deviation of the page elements meets the requirements.
[0034] Further, when the number of changes is greater than a preset change number threshold, it is determined that the recognition deviation situation of the page element does not meet the requirements.
[0035] S2 Extracts features of different dimensions of the page element based on the associated data of the front-end page, constructs multiple feature groups based on the features, and determines the credible feature groups in the feature groups according to the similarity of the features of different feature groups to the features of other page elements; Further, the associated data includes the page image and front-end code of the front-end page.
[0036] It should be noted that the features include image visual features, semantic features, and image structure features.
[0037] Table 1 is a comparison table of the features of the page elements of a certain page in multiple dimensions with the features of other page elements
[0038] It can be understood that the feature groups are determined by free combination according to different features.
[0039] Specifically, as Figure 3 shown, the method for determining the credible feature groups in the feature groups is: Taking the features of the page elements of the feature group as group matching features, determines the feature similarity coefficient between the page element and other page elements in different group matching features; Taking the group matching features with a feature similarity coefficient greater than the preset similarity threshold as similar features, and taking the other page elements whose group matching features all belong to the similar features as interference page elements; Determines whether the feature group is a credible feature group according to the number of the interference page elements.
[0040] Further, when the number of the interference page elements is greater than the preset interference page element number threshold, it is determined that the feature group does not belong to the credible feature group.
[0041] It can be understood that the feature similarity coefficient is determined according to the Euclidean distance function.
[0042] In another possible embodiment, the method for determining the credible feature groups in the feature groups is: Taking the features of the page elements of the feature group as group matching features, determines the feature similarity coefficient between the page element and other page elements in different group matching features; Determines the comprehensive similarity coefficient of the other page elements with the average value of the feature similarity coefficients between different group matching features; Determine whether the feature group is a credible feature group according to the comprehensive similarity coefficient of the other page elements.
[0043] Furthermore, when the number of its page elements with the comprehensive similarity coefficient within the preset similarity coefficient range does not meet the requirements, it is determined that the feature group does not belong to the credible feature group.
[0044] In another possible embodiment, the method for determining the credible feature group in the feature group is as follows: S21 Use the features of the page elements of the feature group as group matching features, and determine the feature similarity coefficient between the page elements and other page elements in different group matching features; Optionally, in the above steps, it is further necessary to determine that if there are no group matching features with a feature similarity coefficient greater than the preset similarity threshold for other page elements, the feature group can be directly determined as a credible feature group.
[0045] S22 Use the group matching features with a feature similarity coefficient greater than the preset similarity threshold as similar features, and use the other page elements whose group matching features all belong to the similar features as interfering page elements; It can be understood that in the above steps, it is also necessary to further determine that when the number of the interfering page elements does not meet the requirements, it is determined that the feature group does not belong to the credible feature group. Specifically, when it is greater than a certain quantity threshold, it is determined that it does not meet the requirements.
[0046] S23 Use the other page elements with similar features as screened interfering elements, and determine the comprehensive similarity coefficient of different screened interfering elements based on the average value of the feature similarity coefficients between different screened interfering elements in different group matching features; Specifically, the above steps include the following three situations: Situation 1: If the number of the screened interfering elements does not meet the requirements, that is, when it is greater than the threshold, it is determined that the feature group does not belong to the credible feature group; Situation 2: If the number of the screened interfering elements meets the requirements, and at this time, if there are screened interfering elements with a comprehensive similarity coefficient that does not meet the requirements, that is, when the screened interfering factor with a comprehensive similarity coefficient greater than a certain threshold exists, it is determined that the feature group does not belong to the credible feature group; Situation 3: When the number of the screened interfering elements with the comprehensive similarity coefficient within the preset similarity coefficient range is too large, that is, when it is greater than a fixed threshold, it is determined that the feature group does not belong to the credible feature group.
[0047] When the number of the screened interfering elements with the comprehensive similarity coefficient within the preset similarity coefficient range meets the requirements: S24 determines the recognition confidence coefficient of the feature group according to the comprehensive similarity coefficient of different screened interference elements and the number of similar features, and determines whether the feature group is a credible feature group based on the recognition confidence coefficient.
[0048] It can be understood that in one possible embodiment, the recognition confidence coefficient of the feature group is determined according to the difference between a preset value and the product of the average value of the comprehensive similarity coefficients of different screened interference elements and the number of similar features.
[0049] In one embodiment, when the recognition confidence coefficient of the feature group is greater than the preset recognition confidence coefficient threshold, it is determined that the feature group belongs to the credible feature group.
[0050] Specifically, the similar features are features whose similarity coefficient with the features of the page element is greater than the preset value of the feature similarity coefficient.
[0051] S3 uses other page elements with similar features in the credible feature group of the page element as associated page elements, and determines the confidence of features in different dimensions in the credible feature group based on the similarity between the features in different dimensions of the associated page elements and the page element. Furthermore, the method for determining the confidence of the feature is as follows: The associated page elements to which the feature belongs and are similar features are used as screened associated elements, and the similarity interference coefficient of the feature in the credible feature group is determined according to the proportion of the number of the screened associated elements in the associated page elements of the credible feature group. Based on the average value of the feature similarity coefficients between the associated page elements of the feature in the credible feature group and the page element, the average similarity coefficient is determined. Based on the product of the similarity interference coefficient and the average similarity coefficient, the interference value is determined, and the confidence of the feature in the credible feature group is determined based on the interference value.
[0052] It should be noted that the confidence of the feature is based on the difference between 1 and the interference value of different features.
[0053] S4 determines the deviation situation of the associated page elements between different credible feature groups, and determines the positioning processing method of the page element in combination with the confidence of features in different dimensions in the credible feature group.
[0054] Among them, the visual recognition model and the semantic recognition model are constructed based on the CNN model and the NLP semantic model. The input of the visual recognition model is the page image of the front-end page, and the input of the semantic recognition model is the front-end code.
[0055] Specifically, such asFigure 4 As shown, the method for determining the positioning processing method of the page element is as follows: Determine the group confidence of different trusted feature groups by summing the confidence levels of features in different dimensions of different trusted feature groups; Determine the same associated page elements between different trusted feature groups according to the deviation of the associated page elements between different trusted feature groups, and use them as the same page elements; Determine the positioning processing method of the page element according to the group confidence of different trusted feature groups and the same page elements between different trusted feature groups.
[0056] Furthermore, determine the positioning processing method of the page element according to the group confidence of different trusted feature groups and the same page elements between different trusted feature groups, specifically including: When there is a trusted feature group with a group confidence greater than the preset group confidence threshold, use the features corresponding to the trusted feature group with the maximum group confidence for the positioning processing of the page element; When there is no trusted feature group with a group confidence greater than the preset group confidence threshold, based on the preset group confidence threshold, freely combine the trusted feature groups to obtain alternative positioning schemes, use the trusted feature groups corresponding to the alternative positioning schemes as alternative feature groups, and determine the positioning processing method of the page element according to the same page elements between the alternative feature groups.
[0057] It can be understood that the alternative positioning scheme is constructed based on the sum of the group confidences of the composed trusted feature groups being greater than the preset group confidence threshold.
[0058] Furthermore, determine the positioning processing method of the page element according to the same page elements between the alternative feature groups, specifically including: Determine the same page elements between the alternative feature groups, and determine the similarity coefficient mean of the same page elements in different alternative feature groups according to the average value of the feature similarity coefficients of the same page elements between the features corresponding to different alternative feature groups; Aim to minimize the number of the same page elements with the minimum similarity coefficient mean in different alternative feature groups being greater than the preset value of the similarity feature coefficient, determine the recognition scheme in the alternative positioning scheme, and use the recognition scheme for the positioning processing of the page element.
[0059] It should be noted that using the recognition scheme for the positioning processing of the page element specifically includes: Based on the positioning results of the alternative feature groups in the recognition scheme, the positioning result with the largest number of corresponding alternative feature groups is used as the positioning processing result of the page element.
[0060] In another possible embodiment, the method for determining the positioning processing method of the page element is as follows: S41 Determine the group confidence of different credible feature groups based on the sum of the confidence levels of the features of different credible feature groups in different dimensions. It should be noted that in the above steps, if it is determined that there is a credible feature group with a group confidence greater than the preset group confidence threshold, then the features corresponding to the credible feature group with the largest group confidence are used for the positioning processing of the page element.
[0061] S42 Determine the same associated page elements between different credible feature groups according to the deviation of the associated page elements between different credible feature groups, and use them as the same page elements. Based on the preset group confidence threshold, freely combine the credible feature groups to obtain alternative positioning schemes, and use the credible feature groups corresponding to the alternative positioning schemes as alternative feature groups. Determine the average similarity coefficient of the same page elements in different alternative feature groups based on the average value of the feature similarity coefficients of the same page elements between different alternative feature groups. In one possible embodiment, the above steps include the following two situations: Situation 1: When the number of the same page elements between the alternative feature groups of the alternative positioning scheme is too large, that is, greater than the threshold, it is determined that the alternative positioning scheme does not belong to the recognition scheme. Situation 2: When the number of the same page elements between the alternative feature groups of the alternative positioning scheme is not too large, if the number of the same page elements with the minimum similarity coefficient mean greater than the preset similarity feature coefficient value in different alternative feature groups is too large, the reliability of its positioning processing is difficult to meet the requirements, and specifically, it can be determined by means of a threshold. At this time, it is directly determined that the alternative positioning scheme does not belong to the recognition scheme.
[0062] S43 Determine the recognition deviation probability of the alternative positioning scheme based on the number of the same page elements and the minimum similarity coefficient mean in different alternative feature groups, and determine the recognition scheme in the alternative positioning scheme based on the recognition deviation probability.
[0063] In one of the embodiments, based on the number of the same page elements and the mean value of the minimum similarity coefficients in different filing feature groups, the sum of the mean values of the minimum similarity coefficients of different same page elements in different filing feature groups is determined, and the recognition deviation probability is determined by multiplying the sum by a preset scaling factor.
[0064] Further, the recognition scheme is an alternative positioning scheme with the minimum recognition deviation probability.
[0065] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0066] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0067] The above description is only for one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to one or more embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of this specification.
Claims
1. A method for page element positioning based on visual-semantic fusion of large models, characterized in that Specifically include: Use a large model to determine the page elements associated with the user instruction, based on the recognition result of the dynamic page elements of the front-end page corresponding to the page elements, and combine the deviation of the page elements during the change process of the dynamic page elements. When the recognition deviation of the page elements does not meet the requirements, proceed to the next step; Based on the associated data of the front-end page, extract the features of the page elements in different dimensions, construct multiple feature groups based on the features, and determine the credible feature groups in the feature groups according to the similarity of the features of different feature groups to other page elements; Use the other page elements with similar features in the credible feature groups of the page elements as associated page elements, and determine the confidence of the features in different dimensions in the credible feature groups according to the similarity of the features in different dimensions of the associated page elements in the credible feature groups to the page elements; Determine the deviation of the associated page elements between different credible feature groups, and combine the confidence of the features in different dimensions in the credible feature groups to determine the positioning processing method of the page elements.
2. The page element positioning method based on large model visual semantic fusion according to claim 1, wherein, The associated page elements are determined according to the semantic recognition result of the large model based on the user's operation instruction.
3. The page element positioning method based on large model visual semantic fusion according to claim 1, characterized in that, The deviation of the page elements during the change process of the dynamic page elements is determined according to the deviation of the page images of the page elements in different front-end pages during the change process.
4. The page element positioning method based on large model visual semantic fusion according to claim 1, wherein, Determining that the recognition deviation of the page elements does not meet the requirements specifically includes: Based on the recognition result of the dynamic page elements of the front-end page corresponding to the page, determine the number of dynamic page elements in the front-end page; According to the change of the page elements during the change process of the dynamic page elements of the front-end page of different dynamic page elements, determine and use the dynamic page elements whose image positions of the page elements of the front-end page change as the changed page elements; Based on the changed update data of the changed page elements, determine whether the recognition deviation of the page elements meets the requirements.
5. The method for positioning page elements based on visual-semantic fusion of large models according to claim 4, wherein When the recognition deviation of the page elements meets the requirements, the positioning processing of the page elements is performed by using the image of the front-end page.
6. The method for positioning page elements based on visual-semantic fusion of large models according to claim 1, characterized in that, The associated data includes the page image and front-end code of the front-end page.
7. The page element positioning method based on large model-based visual semantic fusion according to claim 1, wherein, The features include image visual features, semantic features, and image structure features.
8. The method for positioning page elements based on visual-semantic fusion of large models according to claim 1, wherein, The similar features are the features whose similarity coefficient with the features of the page elements is greater than the preset value of the feature similarity coefficient.
9. The page element positioning method based on large model visual semantic fusion according to claim 1, characterized in that The method for determining the positioning processing method of the page elements is: Determine the group confidence of different credible feature groups based on the sum of the confidence of the features in different dimensions of different credible feature groups; According to the deviation of the associated page elements between different credible feature groups, determine the same associated page elements between different credible feature groups and use them as the same page elements; Determine the positioning processing method of the page elements according to the group confidence of different credible feature groups and the same page elements between different credible feature groups.
10. The page element positioning method based on large model-based visual semantic fusion according to claim 9, wherein, Determine the positioning processing method of the page element according to the group confidence of different trusted feature groups and the same page elements between different trusted feature groups, specifically including: When there is a trusted feature group with a group confidence greater than the preset group confidence threshold, use the features corresponding to the trusted feature group with the maximum group confidence to perform the positioning processing of the page element; When there is no trusted feature group with a group confidence greater than the preset group confidence threshold, based on the preset group confidence threshold, freely combine the trusted feature groups to obtain alternative positioning schemes, use the trusted feature groups corresponding to the alternative positioning schemes as alternative feature groups, and determine the positioning processing method of the page element according to the same page elements between the alternative feature groups.
Citation Information
Patent Citations
Page identification method and device, computer equipment and storage medium
CN113946365A
Method and device for acquiring target content information in webpage and server
CN116561402A
Page generation method and device
CN116737150A
Multimedia page generation method and device, equipment, medium and program product
CN117011875A
Page element positioning method, electronic equipment and storage medium
CN117472744A