Interface infringement detection method, device, equipment, medium and program product

CN122821173APending Publication Date: 2026-09-25INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610826688.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

人工比对依靠专业人员目视判断,虽方法直观但效率低下、主观性强,难以适应海量界面的审核需求;机器检测虽利用模板匹配、图像特征提取及深度学习等技术提升了自动化水平,但传统算法对界面尺寸、色彩及动态变化的鲁棒性较差,且深度学习方案通常依赖海量标注样本,计算成本高、部署难度大

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821173A_ABST
    Figure CN122821173A_ABST
Patent Text Reader

Abstract

The application provides an interface infringement detection method and device, equipment, medium and program product, which can be applied to the technical field of big data and artificial intelligence. The method comprises the following steps: identifying a dynamic area in a target interface image, the dynamic area being an area in which the display content changes over time; replacing corresponding pixels in the dynamic area with a preset background filling value to obtain a target static interface image; obtaining multi-dimensional heterogeneous features of the target static interface image, the heterogeneous features at least comprising global features, text features, contour structure features and color features; inputting the multi-dimensional heterogeneous features into a pre-trained feature mapping model to generate a comprehensive feature vector representing the target interface image, the feature mapping model constructing a loss function by using a metric learning strategy in a training process; calculating the similarity between the comprehensive feature vector of the target interface image and the comprehensive feature vector of a benchmark interface image in a preset comparison library, and performing infringement risk judgment based on the similarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of big data and artificial intelligence technology, and more specifically to a method, apparatus, device, medium, and program product for detecting interface infringement. Background Technology

[0002] Interface infringement detection is a core technical means to identify design plagiarism and protect the intellectual property rights of interface designs.

[0003] The relevant technologies are mainly divided into two categories: manual comparison and machine detection. Manual comparison relies on visual judgment by professionals. Although the method is intuitive, it is inefficient and highly subjective, making it difficult to meet the review needs of massive interfaces. Machine detection, while improving automation through template matching, image feature extraction, and deep learning, suffers from poor robustness to changes in interface size, color, and dynamic changes. Furthermore, deep learning solutions typically rely on massive amounts of labeled samples, resulting in high computational costs and deployment difficulties. Especially in complex scenarios with highly similar interface designs, diverse presentation forms, and frequent interference from dynamic content, existing technologies are prone to missed detections or false positives, making it difficult to balance accuracy and comprehensiveness in detection.

[0004] Therefore, there is an urgent need for an interface infringement detection solution to reduce reliance on human experience and massive sample sizes, overcome the challenges posed by the variability of interface morphology and dynamic interference, and achieve a balance between detection efficiency and accuracy. Summary of the Invention

[0005] In view of the above problems, embodiments of this application provide an interface infringement detection method, apparatus, device, medium, and program product.

[0006] According to a first aspect of this application, an interface infringement detection method is provided, comprising: identifying dynamic regions in a target interface image, wherein the dynamic regions are regions where the displayed content changes over time; replacing corresponding pixels in the dynamic regions with preset background fill values ​​to obtain a target static interface image; acquiring multidimensional heterogeneous features of the target static interface image, wherein the heterogeneous features include at least global features, text features, contour structure features, and color features; inputting the multidimensional heterogeneous features into a pre-trained feature mapping model to generate a comprehensive feature vector representing the target interface image, wherein the feature mapping model employs a metric learning strategy to construct a loss function during training to constrain the spacing between comprehensive feature vectors of similar interfaces to be less than the spacing between comprehensive feature vectors of dissimilar interfaces; calculating the similarity between the comprehensive feature vector of the target interface image and the comprehensive feature vector of a benchmark interface image in a preset comparison library, and determining the infringement risk based on the similarity.

[0007] According to embodiments of this application, identifying dynamic regions in a target interface image includes: acquiring image sequences of the same target interface at multiple consecutive time points, using the acquisition time of the target interface image as a time reference point; performing image registration on multiple frames in the image sequence to obtain a geometrically aligned registered image sequence; comparing the target interface image with each frame in the registered image sequence pixel by pixel and marking the differences; performing connected component analysis on the differences to obtain at least one candidate dynamic region; performing optical character recognition on the candidate dynamic region to obtain character recognition results; and based on the character recognition results, dividing the candidate dynamic region into dynamic text regions and dynamic visual regions, and using the dynamic text region as the dynamic region in the target interface image.

[0008] According to an embodiment of this application, acquiring an image sequence of the same target interface at multiple consecutive time points includes: acquiring images based on a preset initial acquisition frequency to obtain a frame of adjacent images that are temporally adjacent to the target interface; determining the pixel change rate based on the target interface image and the adjacent image; determining the corresponding frequency correction coefficient based on the difference between the pixel change rate and the preset change rate; multiplying the frequency correction coefficient by the initial acquisition frequency to obtain the corrected acquisition frequency; and acquiring an image sequence of the same target interface at multiple consecutive time points based on the corrected acquisition frequency.

[0009] According to embodiments of this application, calculating the similarity between the comprehensive feature vector of a target interface image and the comprehensive feature vector of a benchmark interface image in a preset comparison library, and determining infringement risk based on the similarity, includes: obtaining the front-end code file of the target interface corresponding to the target interface image, wherein the front-end code file includes at least a Hypertext Markup Language document, Cascading Style Sheets, and script files; generating a front-end feature vector based on the front-end code file, wherein the front-end feature vector includes a front-end text feature vector and a front-end structural feature vector; calculating the similarity between the front-end feature vector and the front-end feature vector of the benchmark interface to obtain a code similarity; weighting and fusing the code similarity and the similarity to obtain a comprehensive similarity; and determining that there is an infringement risk if the comprehensive similarity is greater than a preset infringement risk threshold.

[0010] According to an embodiment of this application, generating a front-end feature vector based on a front-end code file includes: inputting the front-end code file into a pre-trained language model to obtain a front-end text feature vector; performing lexical and syntactic parsing on the hypertext markup language document in the front-end code file to generate a document object model tree; obtaining the structural topology features of each node in the document object model tree, and generating a front-end structural feature vector based on the structural topology features, wherein the structural topology features include at least the node's level depth, the number of parent and child nodes, and the number of sibling nodes; and fusing the front-end text feature vector with the front-end structural feature vector to obtain a front-end feature vector.

[0011] According to an embodiment of this application, the step of obtaining global features includes: performing multi-view transformation on the target static interface image to generate multiple view samples, wherein the view samples include at least one global view and at least one local view; inputting the global view and the local view into a first encoder to obtain a first feature; inputting the global view into a second encoder to obtain a second feature, wherein the second encoder has the same network architecture as the first encoder; calculating the distribution consistency loss between the first feature and the second feature, wherein the distribution consistency loss includes a global representation loss and a feature distribution regularization weighted loss; updating the parameters of the first encoder and the parameters of the second encoder through backpropagation based on the distribution consistency loss, wherein the parameters of the second encoder are updated based on the parameters of the first encoder through an exponential moving average and the update rate is lower than that of the first encoder; and using the first encoder as a global feature extractor to extract global features from the target static interface image.

[0012] According to an embodiment of this application, the training steps of the feature mapping model include: acquiring a training dataset, which includes anchor samples, positive samples, and negative samples, wherein the positive samples are interface images that have an infringement association with the anchor samples; projecting the multidimensional heterogeneous features of the anchor samples, positive samples, and negative samples through a first fully connected layer to obtain dimension-aligned multidimensional heterogeneous feature vectors, wherein the first fully connected layer is used for feature dimension alignment; concatenating the multidimensional heterogeneous feature vectors of the anchor samples, positive samples, and negative samples to generate corresponding concatenated feature vectors; inputting the concatenated feature vectors into a second fully connected layer to obtain a comprehensive feature vector of the anchor samples, positive samples, and negative samples, wherein the second fully connected layer is used to fuse multi-source feature information; calculating a loss value based on a triplet loss function, wherein the triplet loss function is used to constrain the distance between the comprehensive feature vectors of the anchor samples and positive samples to be less than the distance between the comprehensive feature vectors of the anchor samples and negative samples; and ending the training in response to the loss value satisfying a preset convergence condition.

[0013] According to a second aspect of this application, an interface infringement detection device is provided, comprising: a dynamic region recognition module for recognizing dynamic regions in a target interface image, wherein the dynamic region is a region where the displayed content changes over time; a static interface image acquisition module for replacing corresponding pixels in the dynamic region with preset background fill values ​​to obtain a target static interface image; a heterogeneous feature acquisition module for acquiring multidimensional heterogeneous features of the target static interface image, wherein the heterogeneous features include at least global features, text features, contour structure features, and color features; a comprehensive feature vector generation module for inputting the multidimensional heterogeneous features into a pre-trained feature mapping model to generate a comprehensive feature vector representing the target interface image, wherein the feature mapping model employs a metric learning strategy to construct a loss function during training to constrain the spacing between comprehensive feature vectors of similar interfaces to be less than the spacing between comprehensive feature vectors of dissimilar interfaces; and an infringement judgment module for calculating the similarity between the comprehensive feature vectors of the target interface image and a reference interface image, and determining the infringement risk based on the similarity.

[0014] According to a third aspect of this application, an electronic device is provided, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0015] According to a fourth aspect of this application, a computer-readable storage medium is also provided, on which a computer program or instructions are stored, wherein the computer program or instructions, when executed by a processor, implement the steps of the above-described method.

[0016] According to a fifth aspect of this application, a computer program product is also provided, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method. Attached Figure Description

[0017] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0018] Figure 1 The illustrations depict application scenarios of the interface infringement detection method, apparatus, device, medium, and program product according to embodiments of this application.

[0019] Figure 2 A flowchart illustrating an interface infringement detection method according to an embodiment of this application is shown schematically.

[0020] Figure 3 A flowchart illustrating text feature extraction according to an embodiment of this application is shown schematically;

[0021] Figure 4A flowchart illustrating contour feature extraction according to an embodiment of this application is shown schematically;

[0022] Figure 5 A flowchart illustrating color feature extraction according to an embodiment of this application is shown schematically;

[0023] Figure 6 A flowchart illustrating dynamic region partitioning according to an embodiment of this application is shown schematically;

[0024] Figure 7 A flowchart illustrating the training of a feature mapping model according to an embodiment of this application is shown schematically.

[0025] Figure 8 This illustration schematically shows another flowchart of an interface infringement detection method according to an embodiment of this application;

[0026] Figure 9 This schematic diagram illustrates the structural block diagram of an interface infringement detection device according to an embodiment of this application;

[0027] Figure 10 A block diagram schematically illustrates an electronic device suitable for implementing an interface infringement detection method according to an embodiment of this application. Detailed Implementation

[0028] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0029] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0030] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0031] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0032] It's important to note that the term "neural network" can refer to a machine learning network based on deep learning. A neural network processes input and provides corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between them. Neural networks used in deep learning applications often include many hidden layers, increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer serves as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output becomes the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each processing the input from the layer above.

[0033] It should be understood that machine learning generally includes three phases: training, testing, and application (also known as inference). In the training phase, a given model is trained using a large amount of training data, iteratively updating parameter values ​​until the model can consistently generate inferences that meet the expected goals from the training data. Through training, the model can be considered to have learned the relationship between inputs and outputs (also known as the input-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether it can provide the correct output, thus determining the model's performance. In the application phase, the model can be used to process actual inputs based on the trained parameter values ​​to determine the corresponding output.

[0034] Figure 1 The illustrations depict application scenarios of the interface infringement detection method, apparatus, device, medium, and program product according to embodiments of this application. Figure 1As shown, application scenario 100 according to an embodiment of this application may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables. For example, a user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send information, etc.

[0035] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be electronic devices such as smartphones, wearable devices, personal computers, intelligent voice interaction devices, smart home appliances, intelligent vehicles, in-vehicle terminals, aircraft, unmanned vending terminals, and extended reality devices. Extended reality devices can include virtual reality devices, augmented reality devices, and mixed reality devices. A client application for the target application can be installed and run on the terminal devices. This target application can include, but is not limited to, financial transaction applications, payment applications, shopping applications, web browser applications, search applications, instant messaging tools, email clients, and social media platform software (these are just examples). Furthermore, this application embodiment does not limit the form of the target application, and it can include, but is not limited to, applications, mini-programs, etc., installed on the terminal devices, and can also be in the form of web pages.

[0036] Server 105 can be a server providing various services, such as a backend management server supporting websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process received user requests and other data, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services such as cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and basic cloud computing services such as big data. The server can be the backend server of the aforementioned target application, used to provide backend services to the clients of the target application.

[0037] It should be noted that the interface infringement detection method provided in this application embodiment can generally be executed by server 105 and / or terminal devices 101-103. Accordingly, the interface infringement detection device provided in this application embodiment can generally be set in server 105 and / or terminal devices 101-103.

[0038] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0039] Figure 2 A flowchart illustrating an interface infringement detection method according to an embodiment of this application is shown schematically. Figure 2 As shown, the interface infringement detection method 200 according to the embodiments of this application may include steps S210 to S250.

[0040] In step S210, dynamic regions in the target interface image are identified.

[0041] In the embodiments of this application, the target interface image can refer to a static image of a visual interface such as a software interface, webpage interface, or mini-program interface obtained through methods such as screenshots or screen recording frame extraction; it is the original image to be processed for infringement detection. The dynamic area is the area where the displayed content changes over time.

[0042] Dynamic areas in the target interface are mostly temporary and personalized content, not fixed content with protective value such as the core layout, structure, and color scheme of the interface. If infringement detection is directly based on images containing dynamic content, the randomness of the dynamic content will lead to feature extraction bias, significantly reducing the accuracy of infringement detection. Therefore, it is necessary to identify and process dynamic areas separately first.

[0043] For example, multiple frames of images of the same target interface at different time points can be continuously acquired, and the RGB (Red, Green, Blue) pixel values ​​and grayscale values ​​of adjacent frames can be compared pixel by pixel. Areas where the pixel difference exceeds a preset threshold are marked and the area is determined to be a dynamic area.

[0044] In another embodiment, pixel motion vectors of the interface image can be extracted based on a sparse optical flow algorithm, and the changes in pixel motion amplitude and direction of motion in each local region can be statistically analyzed. Regions where the fluctuation amplitude of motion vectors exceeds a fixed range can be identified as dynamic regions.

[0045] In step S220, the corresponding pixels in the dynamic area are replaced with preset background fill values ​​to obtain the target static interface image.

[0046] In the embodiments of this application, the preset background fill value may refer to a pre-configured fixed pixel value, which is used to uniformly replace the changing pixels in the dynamic area.

[0047] For example, the global main background color of the target interface image can be extracted, the RGB value corresponding to the main background color can be used as the preset background fill value, and the pixels of all dynamic areas can be completely replaced to generate a static interface image.

[0048] In another embodiment, the average value of neighboring pixels can be used for filling. Specifically: select static neighboring pixels at the edge of the dynamic area; calculate the average RGB value or grayscale value of the static neighboring pixels, and use the average value as the fill value to replace the content of the dynamic area pixel by pixel. In this way, it can be ensured that the filled area has a natural visual connection with the surrounding interface.

[0049] Pixel filling can eliminate dynamic content, filter out invalid and interfering information, preserve static features of the interface with copyright attributes, and improve the accuracy and effectiveness of subsequent feature extraction.

[0050] In step S230, the multidimensional heterogeneous features of the target static interface image are obtained.

[0051] In the embodiments of this application, heterogeneous features include at least global features, text features, contour structure features, and color features. During the extraction of multi-dimensional heterogeneous features, corresponding feature vectors can be output. Global features can refer to feature vectors that describe the overall content, structure, or semantic information of the entire image; they reflect the overall feature distribution of the image without focusing on local details or individual objects. Text features can refer to features corresponding to fixed text content, text layout, font style, and text distribution position within the interface. Contour structure features can refer to features corresponding to the contour shape, size ratio, layout structure, and hierarchical relationship of various functional modules, buttons, pop-ups, and partition borders of the interface. Color features can refer to visual features corresponding to the overall color scheme, main color, distribution of auxiliary colors, color matching ratio, and color saturation of the interface.

[0052] Figure 3 A flowchart illustrating text feature extraction according to an embodiment of this application is shown schematically.

[0053] like Figure 3As shown, the target image to be processed is input into a convolutional neural network, which extracts three different levels of features: basic features, intermediate semantic features, and deep features. Basic features are shallow features with high resolution, mainly representing low-level details of the image; intermediate semantic features have lower resolution and correspond to intermediate semantic information; deep features have extremely low resolution and correspond to high-level semantic information. Then, basic features are downsampled, intermediate semantic features are upsampled and downsampled, and deep features are upsampled. Next, channel alignment convolution is used to match channel dimensions, followed by element-wise addition or concatenation of features to generate a fused feature map integrating multi-layer information. This fused feature map is then input into a structural segmentation network to generate a text probability map. After thresholding, connected component extraction, and region dilation / scaling, polygonal text candidate regions are obtained. To address the redundancy of rectangular bounding boxes in irregularly shaped text, a mask is generated and multiplied with the alignment features to filter out irrelevant content. The features of each region are pooled and concatenated to obtain the text features.

[0054] Figure 4 A flowchart illustrating contour feature extraction according to an embodiment of this application is shown schematically.

[0055] like Figure 4 As shown, the process begins by inputting the original interface image, followed by preprocessing operations such as perspective correction, image denoising, and border trimming to eliminate image distortion and noise interference. Then, an instance segmentation algorithm is used to simultaneously analyze and obtain three key types of information for each interface component: boundary contours, classification attributes, and coordinate positions. Based on the component boundary data, a simplified stroke point sequence is created, and the optimized stroke data is then encoded in five dimensions. The encoding result is input into a structural encoder for deep feature processing, ultimately generating a structural feature vector that quantifies the layout and arrangement patterns of the interface, supporting subsequent interface similarity comparisons and infringement determinations.

[0056] Figure 5 A flowchart illustrating color feature extraction according to an embodiment of this application is shown schematically.

[0057] like Figure 5 As shown, the target image is first input and divided into equal blocks, splitting the entire image into multiple independent sub-regions. Next, the color space format is converted from RGB to HSV color space (H represents hue, S represents saturation, and V represents lightness) to focus on the distribution of hue, chroma, and color difference at the interface. The independent color features of each sub-region are extracted one by one. Then, the color features of each region are numerically normalized to eliminate the influence of total pixel count and local scale differences, resulting in a color feature vector for each region. Finally, these vectors are concatenated in spatial order to form the interface color feature vector.

[0058] Single-dimensional image features cannot fully represent the attributes of an interface. Multi-dimensional heterogeneous features can comprehensively depict interface features from multiple dimensions such as the overall picture, text, structure, and color, providing complete feature data support for subsequent accurate determination of infringement risks.

[0059] In step S240, the multidimensional heterogeneous features are input into a pre-trained feature mapping model to generate a comprehensive feature vector representing the target interface image.

[0060] In the embodiments of this application, the feature mapping model employs a metric learning strategy to construct a loss function during training, thereby constraining the spacing between the comprehensive feature vectors of similar interfaces to be less than the spacing between the comprehensive feature vectors of dissimilar interfaces.

[0061] For example, the feature mapping model can adopt a neural network architecture. First, it performs linear or non-linear dimension alignment on the input multi-dimensional heterogeneous features to eliminate inconsistencies in the feature space. Then, it concatenates the aligned feature vectors to form a unified joint feature representation. Next, it inputs this joint feature into a multi-layer fully connected network, where adaptive weighted combination and non-linear activation are performed using learnable weight matrices and bias terms to achieve deep fusion of multi-source features. Finally, the model outputs a highly discriminative comprehensive feature vector that effectively characterizes the visual properties of the target interface.

[0062] In step S250, the similarity between the comprehensive feature vector of the target interface image and the comprehensive feature vector of the benchmark interface image in the preset comparison library is calculated, and the infringement risk is determined based on the similarity.

[0063] In the embodiments of this application, the preset comparison library refers to a pre-constructed database of infringement benchmark interfaces, used to store massive amounts of benchmark interface images and their corresponding multi-dimensional comprehensive feature vectors. Benchmark interface images refer to standard interface images included in the comparison library, labeled with copyright ownership information, and used as reference benchmarks for infringement determination.

[0064] For example, the similarity between the comprehensive feature vector of the target interface image and the comprehensive feature vector of each benchmark interface image in the preset comparison library can be calculated, for example, using cosine similarity, Euclidean distance, etc. Based on this, a corresponding similarity judgment threshold is set: when the similarity is greater than the threshold (for indicators such as cosine similarity, the larger the value, the more similar the image) or less than the threshold (for indicators such as Euclidean distance, the smaller the value, the more similar the image), it is determined that the two are substantially similar, thereby triggering an infringement risk warning.

[0065] According to embodiments of this application, dynamic region identification and filling can remove dynamic interference information in interface images, retaining static regions as the basis for analysis. On this basis, multidimensional heterogeneous feature representations are constructed from multiple dimensions such as global features, text features, contour structure features, and color features. This allows for the accurate identification of design schemes that are highly similar to existing interfaces in visual expression, enabling earlier detection of potential infringing interfaces, reducing missed detections, and effectively reducing subjective bias caused by manual judgment. This significantly improves the accuracy, comprehensiveness, and objectivity of interface infringement detection and can efficiently adapt to various infringement detection scenarios involving dynamic and static mixed interfaces.

[0066] In embodiments of this application, the step of obtaining global features may include: performing multi-view transformation on the target static interface image to generate multiple view samples, wherein the view samples include at least one global view and at least one local view; inputting the global view and the local view into a first encoder to obtain a first feature; inputting the global view into a second encoder to obtain a second feature, wherein the second encoder has the same network architecture as the first encoder; calculating the distribution consistency loss between the first feature and the second feature, wherein the distribution consistency loss includes a global representation loss and a feature distribution regularization weighted loss; updating the parameters of the first encoder and the parameters of the second encoder through backpropagation based on the distribution consistency loss, wherein the parameters of the second encoder are updated based on the parameters of the first encoder through an exponential moving average and the update rate is lower than that of the first encoder; and using the first encoder as a global feature extractor to extract global features from the target static interface image.

[0067] Multi-view transformation refers to the process of generating image samples with different observation dimensions from a static interface image through image processing methods such as cropping, scaling, local cropping, and viewpoint fine-tuning. The first encoder, as the main feature extractor, needs to learn the ability to represent global and local multi-dimensional features. The second encoder can be an auxiliary encoding network with a completely identical network architecture to the first encoder but a different parameter update method, used to provide a standard reference representation of global features.

[0068] For example, the original complete interface image can be retained as a global view; the core area of ​​the interface is cropped at the center and corners according to a fixed ratio to generate at least one local view at a different location, forming a set of view samples. Then, the global view and each local view are input into the first encoder to obtain the corresponding first feature; simultaneously, the global view is input into the second encoder to obtain the second feature. To make global alignment the ultimate goal of learning and to enhance the discriminativeness and extraction capability of global features, the penalty weight of the global loss can be increased when calculating the total distribution consistency loss. The total distribution consistency loss is obtained by weighted fusion of the global representation loss and the feature distribution regularization weighted loss. Based on this, the parameters of the first encoder are updated through backpropagation based on the total loss, and the parameters of the second encoder are updated synchronously using an exponential moving average strategy. After training, the converged first encoder is used as a global feature extractor to extract discriminative and robust global feature representations from the target static interface image.

[0069] According to embodiments of this application, global features are extracted through contrastive learning based on dual encoders, and the penalty weight of the global loss is increased to strengthen the global alignment target. This significantly improves the representation accuracy and discrimination capability of global interface features, providing high-quality global feature support for the accurate determination of subsequent interface infringement. Furthermore, this feature extraction method eliminates the need for manual sample labeling; it relies on multi-view self-supervised learning to learn a unified feature representation of interface images from different observation perspectives, reducing data labeling costs.

[0070] Figure 6 A flowchart illustrating dynamic region division according to an embodiment of this application is shown schematically.

[0071] In embodiments of this application, identifying dynamic regions in a target interface image may include: acquiring image sequences of the same target interface at multiple consecutive time points, using the acquisition time of the target interface image as a time reference point; performing image registration on multiple frames in the image sequence to obtain a geometrically aligned registered image sequence; comparing the target interface image with each frame in the registered image sequence pixel by pixel and marking the differences; performing connected component analysis on the differences to obtain at least one candidate dynamic region; performing optical character recognition on the candidate dynamic region to obtain character recognition results; and based on the character recognition results, dividing the candidate dynamic region into dynamic text regions and dynamic visual regions, and using the dynamic text region as the dynamic region in the target interface image.

[0072] like Figure 6 As shown, the method 600 of this embodiment may include steps S610 to S670.

[0073] In step S610, an image sequence is acquired based on the target interface image time.

[0074] In the embodiments of this application, for the target interface, the target interface image can be acquired first, and then, based on the acquisition time of the target interface image, a preset number of consecutive frame images, such as frame 1, frame 2... frame N, can be acquired at fixed time intervals to form an ordered image sequence with a fixed time step.

[0075] A single frame of an interface image cannot determine whether a region is dynamically changing. Only by using a continuous sequence of images can we compare the temporal changes in pixels and the overall picture, providing temporal data support for subsequent dynamic region recognition.

[0076] In step S620, image registration is performed.

[0077] During image acquisition, slight image shifts, scaling, and viewing angle deviations can easily occur, leading to false differences in pixel contrast between frames and mistakenly identifying static areas as dynamic areas. Image registration can eliminate geometric deviations, ensuring that inter-frame differences originate solely from changes in interface content.

[0078] For example, local feature points of the target interface image and each frame image can be extracted separately to complete the matching and pairing of feature points between different images. The overall geometric transformation parameters can be solved using the paired features. Based on the transformation parameters, the overall coordinate mapping and pixel interpolation resampling can be performed on the sequence frame images to correct the overall offset caused by image translation, rotation and scaling, and achieve image geometric alignment.

[0079] In step S630, pixel comparison.

[0080] In the embodiments of this application, the difference pixel can refer to the pixel point in each frame of the target interface image and the registered image sequence where the pixel value at the same coordinate position has a significant deviation.

[0081] For example, the target interface image and each frame in the registered image sequence can be converted into grayscale images first, and then the grayscale values ​​at corresponding coordinates can be directly compared. If the grayscale difference exceeds a set threshold, it is determined to be a difference pixel.

[0082] In another embodiment, the R (Red), G (Green), and B (Blue) channel values ​​at the same position are compared between the target interface image and each frame in the registered image sequence. If the difference between any channel exceeds a preset threshold, it is marked as a difference pixel.

[0083] In step S640, the difference pixels are merged and marked.

[0084] In the embodiments of this application, after performing pixel-by-pixel comparison on the image, the difference pixels between the target interface image and frames 1, 2, ..., N are obtained respectively. Subsequently, a union operation is performed on the difference pixels corresponding to all frames to aggregate the variation regions scattered in each frame into a global difference mask for the target interface image.

[0085] In step S650, connected component analysis is performed.

[0086] In the embodiments of this application, based on a global difference mask, a connected component analysis algorithm (such as four-connected component analysis or eight-connected component analysis) is used to cluster discrete difference pixels according to their neighborhood relationships to form candidate dynamic regions with spatial continuity.

[0087] In step S660, optical character recognition is performed.

[0088] In the embodiments of this application, based on the candidate dynamic region, the corresponding dynamic region candidate sub-image can be segmented and extracted from the original target interface image; then, optical character recognition processing is performed on the selected candidate sub-images one by one to generate the corresponding optical character recognition results.

[0089] In step S670, the region is dynamically divided.

[0090] In the embodiments of this application, a dynamic text area can refer to an interface area where the content that changes is dynamic text-based content such as scrolling text, bullet screen text, or refresh text. A dynamic visual area can refer to an interface area where the content that changes is non-text-based dynamic visual content such as images, animations, or color block switching.

[0091] For example, based on the optical character recognition results, the area ratio of valid text pixels within the candidate region can be statistically analyzed. This statistical value is compared with a preset threshold. If the ratio exceeds the threshold, the region is determined to be a dynamic text region containing text content; otherwise, it is determined to be a dynamic visual region without text content. Finally, the dynamic text region is selected as the dynamic interference region to be removed from the target interface image.

[0092] According to the embodiments of this application, by dividing the interface into dynamic text areas and dynamic visual areas based on optical character recognition results, and identifying the dynamic text areas as interference areas to be removed, the inaccurate infringement detection caused by accidental deletion of dynamic design elements of the interface can be effectively avoided.

[0093] In embodiments of this application, acquiring image sequences of the same target interface at multiple consecutive time points may include: acquiring images based on a preset initial acquisition frequency to obtain a neighboring image that is temporally adjacent to the target interface; determining the pixel change rate based on the target interface image and the neighboring image; determining the corresponding frequency correction coefficient based on the difference between the pixel change rate and the preset change rate; multiplying the frequency correction coefficient by the initial acquisition frequency to obtain the corrected acquisition frequency; and acquiring image sequences of the same target interface at multiple consecutive time points based on the corrected acquisition frequency.

[0094] In the embodiments of this application, the initial acquisition frequency may refer to a pre-set, default interface image acquisition frequency (e.g., 10 frames / second). Based on the target interface image, a single frame of interface image that is temporally adjacent to it can be acquired according to the initial acquisition frequency. Optionally, multiple frames of interface images that are temporally adjacent to the target interface image can be acquired according to the initial acquisition frequency.

[0095] For example, the process can begin by identifying the difference pixels between the target interface image and its neighboring images. Then, the proportion of these difference pixels to the total number of pixels on the interface is calculated and used as the pixel change rate to characterize the intensity of dynamic changes in the interface. Next, the absolute difference between the pixel change rate and a preset change rate is calculated. Based on this absolute difference, a preset mapping table is consulted to determine the corresponding frequency correction coefficient. This mapping table predefines the correspondence between different difference ranges and frequency correction coefficients. For instance, a larger frequency correction coefficient corresponds to a higher absolute difference range, increasing the acquisition density of the interface image; conversely, a smaller or standard frequency correction coefficient corresponds to a lower absolute difference range or zero. Multiplying the initial acquisition frequency by the frequency correction coefficient yields the optimal acquisition frequency adapted to the dynamic changes in the interface. After dynamically correcting the initial acquisition frequency, the corrected acquisition frequency is used as the current sampling strategy to continuously monitor the same target interface, acquiring image sequences covering multiple time points.

[0096] According to the embodiments of this application, the dynamic intensity of the interface is quantified by the pixel change rate, and the acquisition frequency is dynamically adjusted accordingly, which greatly improves the accuracy and efficiency of image sequence acquisition and provides high-quality image data support for accurate identification of dynamic regions.

[0097] In embodiments of this application, calculating the similarity between the comprehensive feature vector of the target interface image and the comprehensive feature vector of a benchmark interface image in a preset comparison library, and determining the infringement risk based on the similarity, may include: obtaining the front-end code file of the target interface corresponding to the target interface image, wherein the front-end code file includes at least a Hypertext Markup Language document, Cascading Style Sheets, and script files; generating a front-end feature vector based on the front-end code file, wherein the front-end feature vector includes a front-end text feature vector and a front-end structural feature vector; calculating the similarity between the front-end feature vector and the front-end feature vector of the benchmark interface to obtain a code similarity; weighting and fusing the code similarity and the similarity to obtain a comprehensive similarity; and determining that there is an infringement risk if the comprehensive similarity is greater than a preset infringement risk threshold.

[0098] Front-end code files refer to the source code files that build the visual display effects of the interface, including HyperText Markup Language (HTML), Cascading Style Sheets (CSS), and JavaScript (JS). Front-end text feature vectors refer to feature vectors that characterize the text content, syntax, and keyword distribution of the front-end code. Front-end structural feature vectors refer to feature vectors that characterize the page layout, node hierarchy, and module structure of the front-end code.

[0099] For example, front-end text feature vectors and front-end structural feature vectors can be extracted from the front-end code file respectively; then, the front-end text feature vectors and front-end structural feature vectors are fused to obtain the front-end feature vector. Further, the similarity between the front-end feature vector and the front-end feature vector of the baseline interface can be calculated to obtain the code similarity; the code similarity is then weighted and fused with the similarity between the combined feature vectors of the target interface image and the baseline interface image to obtain the comprehensive similarity; if the comprehensive similarity is greater than a preset infringement risk threshold, then an infringement risk is determined to exist.

[0100] According to the embodiments of this application, infringement detection is performed by superimposing front-end code similarity. Even if the target interface has only undergone partial adjustments in visual effect, as long as its front-end code structure remains highly similar to the baseline interface, the infringement can still be accurately identified based on code similarity.

[0101] In embodiments of this application, generating a front-end feature vector based on a front-end code file may include: inputting the front-end code file into a pre-trained language model to obtain a front-end text feature vector; performing lexical and syntactic parsing on the hypertext markup language document in the front-end code file to generate a document object model tree; obtaining the structural topology features of each node in the document object model tree, and generating a front-end structural feature vector based on the structural topology features, wherein the structural topology features include at least the node's level depth, the number of parent and child nodes, and the number of sibling nodes; and fusing the front-end text feature vector with the front-end structural feature vector to obtain a front-end feature vector.

[0102] For example, for front-end text feature vectors, HTML, CSS, and JS code can be concatenated into a string and input into a pre-trained language model. The model outputs a high-dimensional semantic vector as the text feature vector through a multi-layer transformer structure. Optionally, HTML, CSS, and JS code can be input separately into a pre-trained language model for encoding to obtain their respective feature representations. Then, these feature vectors are fused to generate a complete front-end text feature vector. For front-end structural feature vectors, lexical and syntactic parsing of HTML code can be performed to generate a document object model tree. The level depth, number of parent-child nodes, and number of sibling nodes of each node in the tree can then be extracted as structural topological features. These features are then encoded into the final front-end structural feature vector using a graph neural network. Finally, the front-end text feature vector and the front-end structural feature vector are fused to generate a complete front-end feature vector.

[0103] According to the embodiments of this application, by extracting the text features and structural features of the front-end code in different dimensions and performing feature fusion, the feature representation of the front-end code can be obtained, providing accurate data support for the similarity comparison of the underlying code of the target interface, and effectively improving the comprehensiveness and accuracy of infringement detection by fusing front-end code.

[0104] Figure 5 A flowchart illustrating the training of a feature mapping model according to an embodiment of this application is shown.

[0105] In embodiments of this application, the training steps of the feature mapping model may include: acquiring a training dataset, which includes anchor samples, positive samples, and negative samples, wherein the positive samples are interface images that have an infringement association with the anchor samples; projecting the multidimensional heterogeneous features of the anchor samples, positive samples, and negative samples through a first fully connected layer to obtain dimension-aligned multidimensional heterogeneous feature vectors, wherein the first fully connected layer is used for feature dimension alignment; concatenating the multidimensional heterogeneous feature vectors of the anchor samples, positive samples, and negative samples to generate corresponding concatenated feature vectors; inputting the concatenated feature vectors into a second fully connected layer to obtain a comprehensive feature vector of the anchor samples, positive samples, and negative samples, wherein the second fully connected layer is used to fuse multi-source feature information; calculating a loss value based on a triplet loss function, wherein the triplet loss function is used to constrain the distance between the comprehensive feature vectors of the anchor samples and positive samples to be less than the distance between the comprehensive feature vectors of the anchor samples and negative samples; and ending the training in response to the loss value satisfying a preset convergence condition.

[0106] like Figure 7 As shown, the method 700 of this embodiment may include steps S710 to S760.

[0107] In step S710, the anchor samples, positive samples, and negative samples of the training dataset are obtained.

[0108] In the embodiments of this application, the anchor point sample can refer to the target interface image sample. A positive sample can refer to an interface image sample that has an infringement association with the anchor point sample and is highly similar to it. A negative sample can refer to an interface image sample that has no infringement association with the anchor point sample and has a significantly different interface style and structure.

[0109] In step S720, dimensional projection is performed through the first fully connected layer.

[0110] In the embodiments of this application, global feature vectors, text feature vectors, contour structure feature vectors, and color feature vectors of anchor point samples, positive samples, and negative samples can be obtained respectively; then, linear or nonlinear mapping is performed on each modal feature vector to project them onto a feature space of uniform dimension. The projection formula can be expressed as:

[0111] ;

[0112] ;

[0113] ;

[0114] ;

[0115] in, Represents the global feature vector; This represents the global feature vector after projection processing; Represents the text feature vector; This represents the text feature vector after projection processing; Represents the feature vector of the contour structure; This represents the feature vector of the contour structure after projection processing; Represents the color feature vector; This represents the color feature vector after projection processing; , , , These represent the corresponding weight parameters; , , and These represent the corresponding bias parameters; , , and These represent the corresponding activation functions. After dimensionality projection, these four types of feature vectors have been mapped to the same dimension m (e.g., m=256), i.e. .

[0116] In step S730, the feature vectors are concatenated.

[0117] In embodiments of this application, the feature vectors after projection processing can be further concatenated in different dimensions to form a comprehensive feature representation. The concatenation formula can be expressed as:

[0118] ;

[0119] in, This represents the combined features generated after splicing.

[0120] In step S740, multi-source feature information is fused through the second fully connected layer.

[0121] For example, the second fully connected layer can be composed of two cascaded fully connected layers to achieve deep feature fusion, and its formula can be expressed as:

[0122] ;

[0123] ;

[0124] in, This represents the weight matrix of the first fully connected layer; This represents the bias vector of the first fully connected layer; Indicates the activation function; This represents the features extracted by the first fully connected layer; This represents the weight matrix of the second fully connected layer; This represents the bias vector of the second fully connected layer; This represents the final composite feature vector.

[0125] In step S750, the loss value is calculated using the triplet loss function.

[0126] In step S760, it is determined that the loss value meets the preset convergence condition.

[0127] The triplet loss learns a feature embedding space that makes similar samples closer together and dissimilar samples farther apart.

[0128] In the embodiments of this application, the current loss value can be calculated based on the triplet loss function; if the loss value meets the preset convergence condition, the training process is terminated; otherwise, the backpropagation algorithm is executed based on the loss value, and the model parameters are updated. For example, the comprehensive feature vector obtained from anchor samples, positive samples, and negative samples can be represented as... , and The formula for triplet loss can be expressed as:

[0129] ;

[0130] in, Represents the feature space distance function; Indicates the preset interval.

[0131] According to the embodiments of this application, by using the triplet loss function to train the model, the problems of chaotic dimensions, unbalanced weights, and low discriminativeness of multidimensional heterogeneous features are effectively solved. The trained model can transform the multi-source heterogeneous features of the target interface into a comprehensive feature vector with high recognition and high discriminativeness, which greatly improves the accuracy and stability of subsequent interface similarity comparison and infringement determination.

[0132] Figure 8 This illustration shows another flowchart of an interface infringement detection method according to an embodiment of the present application.

[0133] like Figure 8As shown, the process begins by inputting a target interface image, which is then dynamically divided into dynamic and static regions. Pixels in the dynamic regions are uniformly replaced with preset background fill values, and the static regions are then concatenated to generate a static target interface image. Subsequently, four types of multidimensional heterogeneous features—global, text, contour, and color—are extracted from this static image. These features are then mapped, normalized, and fused to obtain a comprehensive feature vector. The feature mapping model for generating the comprehensive feature vector is trained using a triplet loss function. Next, the similarity between this vector and the feature vector of a benchmark image in a preset comparison database is calculated. Finally, an infringement determination is made based on a threshold: if the similarity is higher than the threshold, an infringement risk is identified; otherwise, there is no infringement risk.

[0134] Based on the above-described interface infringement detection method, embodiments of this application also provide an interface infringement detection device. The following will be combined with... Figure 9 The device is described in detail.

[0135] Figure 9 A schematic diagram of the structure of an interface infringement detection device according to an embodiment of this application is shown.

[0136] like Figure 9 As shown, the interface infringement detection device 900 of this embodiment includes a dynamic region recognition module 910, a static interface image acquisition module 920, a heterogeneous feature acquisition module 930, a comprehensive feature vector generation module 940, and an infringement judgment module 950.

[0137] The dynamic region recognition module 910 is used to identify dynamic regions in the target interface image, wherein the dynamic region is the region where the displayed content changes over time. In one embodiment, the dynamic region recognition module 910 can be used to perform step S210 described above, which will not be repeated here.

[0138] The static interface image acquisition module 920 is used to replace the corresponding pixels in the dynamic area with preset background fill values ​​to obtain the target static interface image. In one embodiment, the static interface image acquisition module 920 can be used to perform step S220 described above, which will not be repeated here.

[0139] The heterogeneous feature acquisition module 930 is used to acquire multidimensional heterogeneous features of the target static interface image. The heterogeneous features include at least global features, text features, contour structure features, and color features. In one embodiment, the heterogeneous feature acquisition module 930 can be used to perform step S230 described above, which will not be repeated here.

[0140] The comprehensive feature vector generation module 940 is used to input multi-dimensional heterogeneous features into a pre-trained feature mapping model to generate a comprehensive feature vector representing the target interface image. During training, the feature mapping model employs a metric learning strategy to construct a loss function, constraining the distance between comprehensive feature vectors of similar interfaces to be less than the distance between comprehensive feature vectors of dissimilar interfaces. In one embodiment, the comprehensive feature vector generation module 940 can be used to execute step S240 described above, which will not be repeated here.

[0141] The infringement determination module 950 is used to calculate the similarity between the comprehensive feature vectors of the target interface image and the reference interface image, and to determine the infringement risk based on the similarity. In one embodiment, the infringement determination module 950 can be used to perform step S250 described above, which will not be repeated here.

[0142] According to an embodiment of this application, the dynamic region recognition module 910 is further configured to acquire image sequences of the same target interface at multiple consecutive time points, using the acquisition time of the target interface image as the time reference point; perform image registration on multiple frames in the image sequence to obtain a geometrically aligned registered image sequence; compare the target interface image with each frame in the registered image sequence pixel by pixel and mark the difference pixels; perform connected component analysis on the difference pixels to obtain at least one candidate dynamic region; perform optical character recognition on the candidate dynamic region to obtain character recognition results; and based on the character recognition results, divide the candidate dynamic region into a dynamic text region and a dynamic visual region, and use the dynamic text region as the dynamic region in the target interface image.

[0143] According to an embodiment of this application, the dynamic region identification module 910 may further include a sampling frequency setting module.

[0144] The acquisition frequency setting module is used to acquire images based on a preset initial acquisition frequency to obtain a frame of adjacent images that are temporally adjacent to the target interface; determine the pixel change rate based on the target interface image and the adjacent image; determine the corresponding frequency correction coefficient based on the difference between the pixel change rate and the preset change rate; multiply the frequency correction coefficient by the initial acquisition frequency to obtain the corrected acquisition frequency; and acquire image sequences of the same target interface at multiple consecutive time points based on the corrected acquisition frequency.

[0145] According to an embodiment of this application, the infringement determination module 950 is further configured to obtain the front-end code file of the target interface corresponding to the target interface image, wherein the front-end code file includes at least a Hypertext Markup Language document, Cascading Style Sheets, and script files; generate a front-end feature vector based on the front-end code file, wherein the front-end feature vector includes a front-end text feature vector and a front-end structural feature vector; calculate the similarity between the front-end feature vector and the front-end feature vector of the baseline interface to obtain a code similarity; perform a weighted fusion of the code similarity and the similarity to obtain a comprehensive similarity; and determine that there is an infringement risk if the comprehensive similarity is greater than a preset infringement risk threshold.

[0146] According to an embodiment of this application, the infringement determination module 950 may further include a front-end feature vector generation module.

[0147] The front-end feature vector generation module is used to input the front-end code file into the pre-trained language model to obtain the front-end text feature vector; perform lexical and syntactic parsing on the hypertext markup language document in the front-end code file to generate a document object model tree; obtain the structural topology features of each node in the document object model tree, and generate the front-end structural feature vector based on the structural topology features, wherein the structural topology features include at least the node's level depth, the number of parent and child nodes, and the number of sibling nodes; and fuse the front-end text feature vector and the front-end structural feature vector to obtain the front-end feature vector.

[0148] According to an embodiment of this application, the interface infringement detection device 900 further includes a global feature acquisition module and a feature mapping model training module.

[0149] The global feature acquisition module is used to perform multi-view transformation on the target static interface image to generate multiple view samples, wherein the view samples include at least one global view and at least one local view; the global view and the local view are input into the first encoder to obtain the first feature; the global view is input into the second encoder to obtain the second feature, wherein the second encoder has the same network architecture as the first encoder; the distribution consistency loss between the first feature and the second feature is calculated, wherein the distribution consistency loss includes global representation loss and feature distribution regularization weighted loss; based on the distribution consistency loss, the parameters of the first encoder are updated through backpropagation, and the parameters of the second encoder are also updated, wherein the parameters of the second encoder are updated based on the parameters of the first encoder through exponential moving average and the update rate is lower than that of the first encoder; the first encoder is used as a global feature extractor to extract global features from the target static interface image.

[0150] The feature mapping model training module is used to acquire the training dataset, which includes anchor samples, positive samples, and negative samples. Positive samples are interface images that have an infringement association with the anchor samples. The multidimensional heterogeneous features of the anchor samples, positive samples, and negative samples are projected through a first fully connected layer to obtain dimension-aligned multidimensional heterogeneous feature vectors. The first fully connected layer is used for feature dimension alignment. The multidimensional heterogeneous feature vectors of the anchor samples, positive samples, and negative samples are concatenated to generate corresponding concatenated feature vectors. These concatenated feature vectors are input into a second fully connected layer to obtain a comprehensive feature vector of the anchor samples, positive samples, and negative samples. The second fully connected layer is used to fuse multi-source feature information. A loss value is calculated based on a triplet loss function, which constrains the distance between the comprehensive feature vectors of the anchor samples and positive samples to be less than the distance between the comprehensive feature vectors of the anchor samples and negative samples. Training ends when the loss value meets a preset convergence condition.

[0151] According to embodiments of this application, any multiple modules among the dynamic region recognition module 910, static interface image acquisition module 920, heterogeneous feature acquisition module 930, comprehensive feature vector generation module 940, and infringement judgment module 950 can be merged into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this application, at least one of the dynamic region recognition module 910, static interface image acquisition module 920, heterogeneous feature acquisition module 930, comprehensive feature vector generation module 940, and infringement judgment module 950 can be at least partially implemented as hardware circuits, such as field-programmable gate arrays, programmable logic arrays, systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits, or implemented by any other reasonable means of integrating or packaging circuits, or implemented by any one of software, hardware, and firmware, or by any appropriate combination of any of these. Alternatively, at least one of the dynamic region recognition module 910, static interface image acquisition module 920, heterogeneous feature acquisition module 930, comprehensive feature vector generation module 940, and infringement judgment module 950 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0152] Figure 10 A block diagram schematically illustrates an electronic device suitable for implementing an interface infringement detection method according to an embodiment of this application.

[0153] like Figure 10As shown, an electronic device 1000 according to an embodiment of this application includes a processor 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage portion 1008 into a random access memory 1003. The processor 1001 may include, for example, a general-purpose microprocessor, an instruction set processor and / or an associated chipset and / or a dedicated microprocessor. The processor 1001 may also include onboard memory for caching purposes. The processor 1001 may include a single processing unit or multiple processing units for executing different steps of the method flow according to an embodiment of this application.

[0154] Random access memory 1003 stores various programs and data required for the operation of electronic device 1000. Processor 1001, read-only memory 1002, and random access memory 1003 are interconnected via bus 1004. Processor 1001 executes various steps of the method flow according to embodiments of this application by executing programs in read-only memory 1002 and / or random access memory 1003. It should be noted that the programs may also be stored in one or more memories other than read-only memory 1002 and random access memory 1003. Processor 801 may also execute various steps of the method flow according to embodiments of this application by executing programs stored in said one or more memories.

[0155] According to embodiments of this application, the electronic device 1000 may further include an input / output interface 1005, which is also connected to a bus 1004. The electronic device 1000 may also include one or more of the following components connected to the input / output interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube, liquid crystal display, etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card, such as a local area network card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the input / output interface 1005 as needed. A removable medium 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1010 as needed so that computer programs read from it can be installed into the storage section 1008 as needed.

[0156] Embodiments of this application also provide a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.

[0157] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In embodiments of this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include the read-only memory 1002, and / or random access memory 1003, and / or one or more memories other than read-only memory 1002 and random access memory 1003 described above.

[0158] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the methods provided in the embodiments of this application.

[0159] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 1009, and / or installed from a removable medium 1011. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0160] In embodiments of this application, the computer program can be downloaded and installed from a network via communication section 1009, and / or installed from removable medium 1011. When the computer program is executed by processor 1001, it performs the functions defined in the system of this application embodiment. According to embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0161] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0162] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0163] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

Claims

1. A method for detecting interface infringement, characterized in that, The method includes: Identify dynamic regions in a target interface image, wherein the dynamic region is a region where the displayed content changes over time; The corresponding pixels within the dynamic area are replaced with preset background fill values ​​to obtain the target static interface image; The multidimensional heterogeneous features of the target static interface image are obtained, and the heterogeneous features include at least global features, text features, contour structure features, and color features. The multidimensional heterogeneous features are input into a pre-trained feature mapping model to generate a comprehensive feature vector representing the target interface image. The feature mapping model uses a metric learning strategy to construct a loss function during training to constrain the spacing between comprehensive feature vectors of similar interfaces to be smaller than the spacing between comprehensive feature vectors of dissimilar interfaces. The similarity between the comprehensive feature vector of the target interface image and the comprehensive feature vector of the benchmark interface image in the preset comparison library is calculated, and the infringement risk is determined based on the similarity.

2. The method according to claim 1, characterized in that, The dynamic regions identified in the target interface image include: Using the acquisition time of the target interface image as the time reference point, image sequences of the same target interface at multiple consecutive time points are collected; Image registration is performed on multiple frames of images in the image sequence to obtain a geometrically aligned registered image sequence. The target interface image is compared pixel by pixel with each frame in the registered image sequence, and the differences are marked. Perform connected component analysis on the difference pixels to obtain at least one candidate dynamic region; Optical character recognition is performed on the candidate dynamic region to obtain character recognition results; Based on the character recognition results, the candidate dynamic region is divided into a dynamic text region and a dynamic visual region, and the dynamic text region is taken as the dynamic region in the target interface image.

3. The method according to claim 2, characterized in that, The image sequence of the same target interface at multiple consecutive time points includes: Image acquisition is performed based on a preset initial acquisition frequency to obtain a frame of adjacent image that is temporally adjacent to the target interface; Based on the target interface image and the adjacent images, determine the pixel change rate; Based on the difference between the pixel change rate and the preset change rate, the corresponding frequency correction coefficient is determined; Multiply the frequency correction coefficient by the initial acquisition frequency to obtain the corrected acquisition frequency; Based on the corrected acquisition frequency, image sequences of the same target interface at multiple consecutive time points are acquired.

4. The method according to claim 1, characterized in that, The calculation of the similarity between the comprehensive feature vector of the target interface image and the comprehensive feature vector of the benchmark interface image in the preset comparison library, and the determination of infringement risk based on the similarity, includes: Obtain the front-end code file of the target interface corresponding to the target interface image, wherein the front-end code file includes at least a Hypertext Markup Language document, Cascading Style Sheets, and a script file; Based on the aforementioned front-end code file, a front-end feature vector is generated, wherein the front-end feature vector includes a front-end text feature vector and a front-end structural feature vector; Calculate the similarity between the front-end feature vector and the front-end feature vector of the baseline interface to obtain the code similarity; The code similarity and the overall similarity are weighted and fused together to obtain a comprehensive similarity score. If the overall similarity is greater than a preset infringement risk threshold, an infringement risk is determined to exist.

5. The method according to claim 4, characterized in that, The process of generating the front-end feature vector based on the front-end code file includes: The front-end code file is input into a pre-trained language model to obtain the front-end text feature vector; Lexical and syntactic parsing are performed on the hypertext markup language document in the front-end code file to generate a document object model tree; Obtain the structural topology features of each node in the document object model tree, and generate a front-end structural feature vector based on the structural topology features, wherein the structural topology features include at least the node's level depth, the number of parent and child nodes, and the number of sibling nodes; The front-end text feature vector and the front-end structural feature vector are fused to obtain the front-end feature vector.

6. The method according to claim 1, characterized in that, The steps for obtaining the global features include: The target static interface image is subjected to multi-view transformation to generate multiple view samples, wherein the view samples include at least one global view and at least one local view; The global view and the local view are input into the first encoder to obtain the first feature; The global view is input into the second encoder to obtain the second feature, wherein the second encoder has the same network architecture as the first encoder; Calculate the distribution consistency loss between the first feature and the second feature, wherein the distribution consistency loss includes the global representation loss and the feature distribution regularization weighted loss; Based on the distribution consistency loss, the parameters of the first encoder are updated through backpropagation, and the parameters of the second encoder are also updated. The parameters of the second encoder are updated based on the parameters of the first encoder through an exponential moving average and the update rate is lower than that of the first encoder. The first encoder is used as a global feature extractor to extract global features from the target static interface image.

7. The method according to claim 1, characterized in that, The training steps of the feature mapping model include: Obtain a training dataset, which includes anchor samples, positive samples, and negative samples, wherein the positive samples are interface images that have an infringement association with the anchor samples; The multidimensional heterogeneous features of the anchor point sample, the positive sample, and the negative sample are respectively projected through the first fully connected layer to obtain a dimension-aligned multidimensional heterogeneous feature vector, wherein the first fully connected layer is used for feature dimension alignment. The multidimensional heterogeneous feature vectors of the anchor point sample, the positive sample, and the negative sample are concatenated to generate corresponding concatenated feature vectors. The concatenated feature vector is input into the second fully connected layer to obtain the comprehensive feature vector of the anchor sample, the positive sample, and the negative sample. The second fully connected layer is used to fuse multi-source feature information. The loss value is calculated based on the triplet loss function, which is used to constrain the distance between the composite feature vector of the anchor sample and the positive sample to be less than the distance between the composite feature vector of the anchor sample and the negative sample. Training ends when the loss value meets the preset convergence condition.

8. An interface infringement detection device, characterized in that, The device includes: A dynamic region recognition module is used to identify dynamic regions in a target interface image, wherein the dynamic region is a region where the displayed content changes over time; The static interface image acquisition module is used to replace the corresponding pixels in the dynamic area with preset background fill values ​​to obtain the target static interface image; The heterogeneous feature acquisition module is used to acquire multidimensional heterogeneous features of the target static interface image, wherein the heterogeneous features include at least global features, text features, contour structure features and color features; A comprehensive feature vector generation module is used to input the multidimensional heterogeneous features into a pre-trained feature mapping model to generate a comprehensive feature vector representing the target interface image. The feature mapping model employs a metric learning strategy to construct a loss function during training to constrain the spacing between comprehensive feature vectors of similar interfaces to be less than the spacing between comprehensive feature vectors of dissimilar interfaces. The infringement determination module is used to calculate the similarity between the comprehensive feature vectors of the target interface image and the reference interface image, and to determine the infringement risk based on the similarity.

9. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 7.