Similar application detection method and similar application detection model training method

By acquiring the displayed images and text of the application, and using the feature extraction model to perform multimodal feature fusion and compression, the problems of large errors and low efficiency in the prior art are solved, and efficient and accurate automatic recognition of similar applications are achieved.

CN120257032APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410014720.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The application classification methods that rely on manual audits in the prior art have large errors and low efficiency, making it difficult to cover unpopular applications, and cannot efficiently identify similar applications from multiple service providers.

Method used

By obtaining the displayed images and text of the application, using the feature extraction model to extract images and text feature vectors, perform multimodal feature fusion and compression, and then perform clustering processing to automatically identify similar applications.

Benefits of technology

Significantly reduces manual participation, improves the accuracy and efficiency of detection of similar application, can better identify unpopular applications, reduce redundant features, and improve model generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257032A_ABST
    Figure CN120257032A_ABST
Patent Text Reader

Abstract

The invention provides a similar application detection method. The similar application detection method comprises the steps of obtaining a display image and a display text corresponding to each application in a plurality of applications including a target application; respectively inputting the display image and the display text corresponding to each application into a first sub-model and a second sub-model in a feature extraction model, and fusing the obtained image feature vector and the text feature vector corresponding to each application to obtain a multi-modal feature vector corresponding to the application; and inputting the multi-modal feature vector corresponding to each application into a third sub-model in the feature extraction model, and performing clustering processing on the fusion feature vector corresponding to each application to determine an application similar to the target application in the plurality of applications. Wherein the fusion feature vector corresponding to each application is obtained at least based on the compression feature vector corresponding to the application obtained by the third sub-model. According to the application, the applications similar to the target application can be accurately identified by using the multi-modal features and clustering processing of the applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and in particular, to a method, apparatus, computing device, computer-readable storage medium, and computer program product for detecting similar applications. In addition, the present application also relates to a method for training a similar application detection model. Background Art

[0002] With the continuous development of computer and Internet technologies, more and more service providers have emerged, which provide a variety of services with rich types for various users to meet the diverse needs of users. These services often exist in the form of application programs (hereinafter simply referred to as "applications"). In order to ensure that users can obtain a better service experience, appropriate methods need to be used to classify these applications to meet the needs of further analysis, review, or other processing.

[0003] In the related art, applications belonging to the same or different categories can be found through manual review, and then a corresponding data set can be constructed. For example, for a target application deployed in the Android system, information such as its corresponding package name, software name, certificate, size, etc. can be obtained, and then other applications with the same or similar package names, certificates, software names, etc. can be searched for, and then it is manually determined whether these applications belong to the same type as the target application (for example, both belong to social applications). However, this method of classifying applications relying on manual labor has uncontrollable errors, low efficiency, and it is difficult to cover those relatively unpopular but numerous applications. Summary of the Invention

[0004] In view of this, the present application provides a method, apparatus, computing device, computer-readable storage medium, and computer program product for detecting similar applications to alleviate, reduce, or even eliminate the above problems.

[0005] According to one aspect of the present application, a method for detecting similar applications is provided. The method includes: obtaining a display image and display text corresponding to each of a plurality of applications including a target application, where the display image corresponding to each application includes at least a part of the image displayed during the operation of the application, and the display text corresponding to each application includes at least a part of the text in the display image corresponding to the application; respectively inputting the display image and display text corresponding to each application into a first sub-model and a second sub-model in a feature extraction model to obtain an image feature vector and a text feature vector corresponding to each application; fusing the image feature vector and the text feature vector corresponding to each application to obtain a multi-modal feature vector corresponding to the application; inputting the multi-modal feature vector corresponding to each application into a third sub-model in the feature extraction model to obtain a compressed feature vector corresponding to the application, where the dimension of the compressed feature vector corresponding to each application is smaller than the multi-modal feature vector corresponding to the application; and performing clustering processing on the fused feature vector corresponding to each application, and determining an application similar to the target application among the plurality of applications according to a plurality of clustering sub-spaces obtained by the clustering processing, where the fused feature vector corresponding to each application is obtained based at least on the compressed feature vector corresponding to the application.

[0006] According to another aspect of the present application, a method for training a similar application detection model is provided. The similar application detection model includes a first sub-model, a second sub-model, and a third sub-model. The method includes: obtaining a display image and display text corresponding to each sample application among a plurality of sample applications, where the display image corresponding to each sample application includes at least a part of the image displayed during the operation of the sample application, and the display text corresponding to each sample application includes at least a part of the text in the display image corresponding to the sample application; respectively inputting the display image and display text corresponding to each sample application into the first sub-model and the second sub-model to obtain an image feature vector and a text feature vector corresponding to each sample application; fusing the image feature vector and the text feature vector corresponding to each sample application to obtain a multi-modal feature vector corresponding to the sample application; inputting the multi-modal feature vector corresponding to each sample application into the third sub-model to obtain a compressed feature vector corresponding to the sample application output from the middle layer of the third sub-model, and obtaining an uncompressed feature vector corresponding to the sample application output from the output layer of the third sub-model, where the dimension of the compressed feature vector corresponding to each sample application is smaller than the multi-modal feature vector corresponding to the sample application, and the dimension of the uncompressed feature vector corresponding to each sample application is the same as the multi-modal feature vector corresponding to the sample application; determining a graphic-text loss based on the multi-modal feature vector and the uncompressed feature vector corresponding to each sample application; performing clustering processing on the fused feature vector corresponding to each sample application to determine a clustering loss, where the fused feature vector corresponding to each sample application is at least obtained based on the compressed feature vector corresponding to the sample application; determining an objective loss based on the graphic-text loss and the clustering loss; and iteratively updating the parameters of the similar application detection model such that the objective loss satisfies a preset condition.

[0007] According to another aspect of the present application, a similar application detection device is provided. The device includes: a data acquisition module configured to acquire a display image and display text corresponding to each of a plurality of applications including a target application, where the display image corresponding to each application includes at least a part of the image displayed during the operation of the application, and the display text corresponding to each application includes at least a part of the text in the display image corresponding to the application; a feature extraction module configured to input the display image and display text corresponding to each application into a first sub-model and a second sub-model in a feature extraction model respectively to obtain an image feature vector and a text feature vector corresponding to each application; a feature fusion module configured to fuse the image feature vector and the text feature vector corresponding to each application to obtain a multi-modal feature vector corresponding to the application; a feature compression module configured to input the multi-modal feature vector corresponding to each application into a third sub-model in the feature extraction model to obtain a compressed feature vector corresponding to the application, where the dimension of the compressed feature vector corresponding to each application is smaller than that of the multi-modal feature vector corresponding to the application; and an application detection module configured to perform clustering processing on the fusion feature vector corresponding to each application, and determine an application similar to the target application among the plurality of applications according to a plurality of clustering sub-spaces obtained by the clustering processing, where the fusion feature vector corresponding to each application is obtained at least based on the compressed feature vector corresponding to the application.

[0008] According to another aspect of the present application, there is provided an apparatus for training a similar application detection model, where the similar application detection model includes a first sub-model, a second sub-model, and a third sub-model. The apparatus includes: a sample data acquisition module configured to acquire a display image and display text corresponding to each sample application among a plurality of sample applications, where the display image corresponding to each sample application includes at least a part of the image displayed during the operation of the sample application, and the display text corresponding to each sample application includes at least a part of the text in the display image corresponding to the sample application; a sample feature extraction module configured to input the display image and display text corresponding to each sample application into the first sub-model and the second sub-model respectively to obtain an image feature vector and a text feature vector corresponding to each sample application; a sample feature fusion module configured to fuse the image feature vector and the text feature vector corresponding to each sample application to obtain a multi-modal feature vector corresponding to the sample application; a sample feature compression module configured to input the multi-modal feature vector corresponding to each sample application into the third sub-model to obtain a compressed feature vector corresponding to the sample application output from the intermediate layer of the third sub-model, and obtain an uncompressed feature vector corresponding to the sample application output from the output layer of the third sub-model, where the dimension of the compressed feature vector corresponding to each sample application is less than the dimension of the multi-modal feature vector corresponding to the sample application, and the dimension of the uncompressed feature vector corresponding to each sample application is the same as the dimension of the multi-modal feature vector corresponding to the sample application; a text-image loss determination module configured to determine a text-image loss based on the multi-modal feature vector and the uncompressed feature vector corresponding to each sample application; a clustering loss determination module configured to perform clustering processing on the fused feature vector corresponding to each sample application to determine a clustering loss, where the fused feature vector corresponding to each sample application is obtained at least based on the compressed feature vector corresponding to the sample application; a target loss determination module configured to determine a target loss based on the text-image loss and the clustering loss; and a model training module configured to iteratively update the parameters of the similar application detection model such that the target loss satisfies a preset condition.

[0009] According to another aspect of the present application, there is provided a computing device, including: a memory configured to store computer-executable instructions; and a processor configured to execute any method provided in the foregoing aspects of the present application when the computer-executable instructions are executed by the processor.

[0010] According to another aspect of the present application, there is provided a computer-readable storage medium storing computer-executable instructions, which when executed, execute any method provided in the foregoing aspects of the present application.

[0011] According to another aspect of the present application, there is provided a computer program product including computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, any method provided according to the foregoing aspects of the present application is executed.

[0012] According to the similar application detection method provided by the present application, the display image and display text corresponding to each application in the multiple applications including the target application obtained respectively can be input into the first sub-model and the second sub-model in the feature extraction model to obtain the image feature vector and text feature vector corresponding to each application; the image feature vector and text feature vector corresponding to each application are fused to obtain the multi-modal feature vector corresponding to the application, so that the multi-modal information of the application can be fully utilized to extract application features; in addition, the multi-modal feature vector corresponding to each application is input into the third sub-model in the feature extraction model to obtain the compressed feature vector corresponding to the application, and the dimension of the compressed feature vector corresponding to each application is smaller than that of the multi-modal feature vector corresponding to the application, so as to remove redundant features and reduce the amount of computation; finally, clustering processing is performed on the fused feature vector corresponding to each application, and the applications similar to the target application among the multiple applications are determined according to the multiple clustering sub-spaces obtained by the clustering processing. Compared with the method of manually detecting similar applications in the related art, the similar application detection method provided by the present application can significantly reduce or even eliminate the need for manual participation in the similar application detection process, and helps to obtain more accurate similar application detection results.

[0013] According to the embodiments described hereinafter, these and other aspects of the present application will be apparent and will be elucidated with reference to the embodiments described hereinafter. Description of the Drawings

[0014] In the following description of exemplary embodiments with reference to the drawings, more details, features and advantages of the technical solution of the present application are disclosed. In the drawings:

[0015] Figure 1 Schematically shows an example scenario to which the technical solution provided by some embodiments of the present application can be applied;

[0016] Figure 2 Schematically shows an example flowchart of a similar application detection method according to some embodiments of the present application;

[0017] Figure 3A Schematically shows an example schematic diagram of clustering processing in a similar application detection method according to some other embodiments of the present application;

[0018] Figure 3B Schematically shows Figure 3A An example flowchart of the similar class merging process in the similar application detection method in

[0019] Figure 4 Schematically shows an example schematic diagram of a method for obtaining a fusion feature vector corresponding to an application or a sample application according to some embodiments of the present application;

[0020] Figure 5 Schematically shows an example flowchart of a method for training a similar application detection model according to some embodiments of the present application;

[0021] Figure 6A Schematically shows Figure 5 an example training architecture of the method for training the similar application detection model in

[0022] Figure 6B Schematically shows Figure 5 an example architecture of a first sub - model in the method for training the similar application detection model in

[0023] Figure 7 Schematically shows an example flowchart of steps for obtaining a display image and display text corresponding to an application or a sample application according to some embodiments of the present application;

[0024] Figure 8 Schematically shows an example flowchart of a method for training a similar application detection model according to some other embodiments of the present application;

[0025] Figure 9 Schematically shows an example schematic diagram of an image augmentation operation in the method for training a similar application detection model according to some other embodiments of the present application;

[0026] Figure 10 Schematically shows an example training architecture of the method for training a similar application detection model according to some other embodiments of the present application;

[0027] Figure 11 Schematically shows an example block diagram of a similar application detection device according to some embodiments of the present application;

[0028] Figure 12 Schematically shows an example block diagram of a device for training a similar application detection model according to some embodiments of the present application; and

[0029] Figure 13 Illustrates an example system that includes an example computing device representing one or more systems and / or devices that can implement the various techniques described herein. Detailed implementation manners

[0030] Several embodiments of the present application will be described in more detail below with reference to the accompanying drawings so that those skilled in the art can implement the technical solutions of the present application. The technical solutions of the present application can be embodied in many different forms and purposes, and should not be limited to the embodiments described herein. These embodiments are provided to make the technical solutions of the present application clear and complete, but the described embodiments do not limit the protection scope of the present application.

[0031] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. It will be further understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the relevant art and / or the context of this specification, and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0032] Before introducing the embodiments of the present application in detail, some related concepts will be explained first.

[0033] 1. Artificial Intelligence (AI): It is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0034] Artificial intelligence technology is a comprehensive discipline involving a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include, for example, sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0035] 2. Natural Language Processing (NLP): It is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers in natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistics research; at the same time, it involves computer science and mathematics.

[0036] Figure 1 Exemplarily shows an example scenario 100 to which the technical solutions provided according to some embodiments of the present application can be applied. As Figure 1 shown, the scenario 100 may include a user 110, a terminal device 120 (for example, a computer), a terminal device 130 (for example, a tablet), a network 140, and a remote facility 150. As an example, the remote facility 150 includes a server 151 and optionally also includes a database device 152 for storing relevant data, and these servers or devices can communicate via the network 140.

[0037] Exemplarily, on the side of the remote facility 150, display images and display texts corresponding to each of multiple applications including a target application can be obtained. The display image corresponding to each application includes at least a part of the image displayed during the running of the application, and the display text corresponding to each application includes at least a part of the text in the display image corresponding to the application. At least a part of the multiple applications can be deployed on the devices on the user side, for example, deployed on the terminal device 120 or the terminal device 130. Alternatively, none of the multiple applications are deployed on the terminal device 120 or the terminal device 130, and the server 151 in the remote facility 150 can obtain the display images and display texts corresponding to these applications from other data sources (for example, open-source data sets in related fields, or service providers of corresponding applications). The display images and display texts corresponding to these applications can be stored in the database device 152, or alternatively, they can be stored in other storage devices or even in the cloud.

[0038] It should be noted that the expression "running" can mean the actual running of the corresponding application on the device where it is deployed, or it can mean the simulated running of the corresponding application. For example, for an application developed based on the Android system, at least a part of the images displayed (for example, output to the user 110) during its simulated running in an Android emulator can be obtained as the display image corresponding to the application. Alternatively, at least a part of the images displayed during the actual running of the application deployed on the terminal device 120, for example, can be obtained as the display image corresponding to the application. As another example, the images obtained in the above two cases can be jointly used as the display image corresponding to the application.

[0039] Next, the display image and display text corresponding to each application can be respectively input into the first sub-model and the second sub-model in the feature extraction model to obtain the image feature vector and text feature vector corresponding to each application. Specifically, the display image corresponding to each application can be input into the first sub-model in the feature extraction model to obtain the image feature vector corresponding to the application, and the display text corresponding to each application can be input into the second sub-model in the feature extraction model to obtain the text feature vector corresponding to the application. The first sub-model can be various models capable of extracting image features, including but not limited to pre-trained convolutional neural networks and their various variants (e.g., residual networks), Transformer models, or BERT models. The second sub-model can be various models capable of extracting text features, including but not limited to Word2Vec models, Transformer models, BERT models, or TextCNN models. Among them, the main principle of the TextCNN model is to extract features from text through a convolutional neural network, perform convolutional operations on text using a one-dimensional convolutional layer, and use a max-pooling operation to select the most significant features. TextCNN can process text in parallel, thereby improving the data processing speed and thus being applicable to large-scale text processing tasks.

[0040] After obtaining the image feature vector and text feature vector corresponding to each application, the image feature vector and text feature vector corresponding to each application can be fused to obtain the multi-modal feature vector corresponding to the application. Exemplarily, the image feature vector and text feature vector corresponding to each application can be concatenated to obtain the multi-modal feature vector corresponding to the application. Alternatively, the average vector of the image feature vector and text feature vector corresponding to each application can be calculated as the multi-modal feature vector corresponding to the application.

[0041] Next, the multi-modal feature vector corresponding to each application is input into the third sub-model in the feature extraction model to obtain the compressed feature vector corresponding to the application, and the dimension of the compressed feature vector corresponding to each application is smaller than that of the multi-modal feature vector corresponding to the application. The third sub-model can perform feature compression on the multi-modal feature vector corresponding to each application. Exemplarily, the third sub-model can be a fully connected neural network, which can have an appropriate number of layers (e.g., three layers, four layers, five layers, or more layers). Among them, exemplarily, the compressed feature vector can be the feature vector output from the middle layer of the third sub-model, and its dimension is smaller than the corresponding multi-modal feature vector, that is, the compressed feature vector can be regarded as the result of the multi-modal feature vector compressed by the third sub-model. This kind of feature compression processing helps to reduce the data complexity and improve the generalization ability of the model.

[0042] Then, clustering processing is performed on the fused feature vectors corresponding to each application, and applications similar to the target application among the multiple applications are determined according to the multiple clustering subspaces obtained by the clustering processing, where the fused feature vector corresponding to each application is obtained based at least on the compressed feature vector corresponding to the application. Exemplarily, the fused feature vectors corresponding to each application can be clustered by means of partition clustering. Exemplarily, the fused feature vector corresponding to each application can be the same as the compressed feature vector corresponding to the application.

[0043] It should be noted that the present application does not limit the specific sizes of the display image, display text, image feature vector, text feature vector, multimodal feature vector, compressed feature vector, and fused feature vector, which depend on the actual application. They can have appropriate sizes and satisfy the relationships between them disclosed above (for example, the dimension of the compressed feature vector is less than that of the multimodal feature vector). To facilitate a more intuitive understanding of the present application by those skilled in the art, an example of the specific sizes of these objects is provided here: the size of the display image is 512*512 (for example, in pixels, and may be adjusted to be unified after necessary size adjustment), the size of the display text is 512 (for example, in characters), the dimensions of both the image feature vector and the text feature vector are 128, the dimension of the multimodal feature vector is 256 (in this example, it is obtained by concatenating the image feature vector and the text feature vector), the third sub-model adopts a four-layer fully connected neural network, and the dimensions of its input layer, first intermediate layer, second intermediate layer, and output layer are 256, 128, 128, and 256 respectively. In this example, the compressed feature vector can be output by the second intermediate layer and thus has a dimension of 128, and in this example, the fused feature vector can be the same as the compressed feature vector and thus has a dimension of 128.

[0044] In addition, it should be noted that although the above steps can all be executed on the side of the remote facility 150 (for example, in the form of a program or system deployed on the server 151). Those skilled in the art should understand that these steps can be executed on the side of the terminal device (for example, executed on the terminal device 120). Alternatively, these steps can be executed on the side of the terminal device 120 and the side of the remote facility 150 respectively. For example, the steps of obtaining the display image and display text corresponding to each application among the multiple applications can be executed on the side of the terminal device 120, and the terminal device 120 can send the obtained display image and display text to the remote facility 150 via the network 140, and then the remote facility 150 executes the subsequent other steps. In this case, the above feature extraction model (including the first sub-model, the second sub-model, and the third sub-model) is deployed on the server 151.

[0045] In this application, the server 151 in the remote facility 150 can be a single server or a server cluster, and the database device 152 in the remote facility 150 can store various data required in the similar application detection process (for example, the display images and display texts corresponding to each of the above-mentioned multiple applications). Exemplarily, the user 110 can access the remote facility 150 in the form of a web page through the terminal device 120 or the terminal device 130. Alternatively, the user can communicate with the remote facility 150 through the client installed on the terminal device 120 or the terminal device 130 to participate in the similar application detection process. Optionally, the server 151 can also run other application programs and store other data. For example, the server 151 can include multiple virtual hosts for running different application programs and providing different services.

[0046] In this application, the terminal devices 120 and 130 can be various types of devices, such as mobile phones, tablet computers, laptop computers, in-vehicle devices, etc. A client can be deployed on the terminal devices 120 and 130, and this client can be used to perform operations related to similar application detection (for example, selecting an emulator, selecting an application) and optionally provide other services, and can take any of the following forms: a locally installed application program, a small program accessed via other application programs, a Web program accessed via a browser, etc. (It should be understood that the "application" in this application can include any of these various forms of programs, and can even include other types of programs not listed here). The user 110 can view the information presented by the client and perform corresponding interaction operations through the input / output interfaces of the terminal devices 120 and 130. Optionally, the terminal devices 120 and 130 can be integrated with the server 151.

[0047] In this application, the database device 152 can be regarded as an electronic filing cabinet, that is, a place for storing electronic files, and users can perform operations such as adding, querying, updating, and deleting data in the files. The so-called "database" is a data set stored together in a certain way, shared by multiple objects, having the smallest possible redundancy, and independent of application programs.

[0048] In addition, in this application, the network 140 can be a wired network connected via, such as cables, optical fibers, etc., or a wireless network such as 2G, 3G, 4G, 5G, Wi-Fi, Bluetooth, ZigBee, Li-Fi, etc.

[0049] It should be noted that the term "user" used herein refers to any party that can perform data interaction with a system (for example, the client system deployed on the terminal device 120), including but not limited to people, program software, network platforms, and even machines.

[0050] Figure 2 Schematically shown is an example flowchart of a similar application detection method 200 (hereinafter simply referred to as the detection method 200 for brevity) according to some embodiments of the present application. Exemplarily, the detection method 200 may be implemented by Figure 1 the remote facility 150 shown, although this is not restrictive. The principle of the detection method 200 will be described in detail below in conjunction with Figure 2 and Figure 7 .

[0051] Specifically, in step 210, the display image and display text corresponding to each application among a plurality of applications including the target application can be obtained. The display image corresponding to each application includes at least a part of the image displayed during the operation of the application, and the display text corresponding to each application includes at least a part of the text in the display image corresponding to the application. Depending on the device and system on which the application is deployed, the process of obtaining the display image and display text may be different. Taking the Android system and Android devices supported thereby as an example, step 210 can be executed through the example scheme 700 shown in Figure 7 . As shown in Figure 7 , in step 710, the environment and the application are prepared. For example, when this step is executed in the server 151 in the remote facility 150, an Android emulator can be deployed in the server 151 to run a plurality of applications, and these applications can be obtained from the database device 152 in the remote facility 150 (in this example, these applications can be stored in the database device 152 in the form of Android executable files); in step 720, these applications are run using the Android emulator. During the operation of each application, the output image of the application can be intercepted as the display image of the application (step 730), and the output text of each application can be obtained as the display text of the application (step 740).

[0052] Those skilled in the art should understand that depending on the system types on which the various applications among the plurality of applications can run, simulators corresponding to these system types can be flexibly selected to run the corresponding applications. In one example, if a part of the plurality of applications runs under the Android system, the Android emulator can be used to run this part of the applications; if another part of the plurality of applications runs under the iOS system, the iOS emulator can be used to run this part of the applications; if yet another part of the plurality of applications runs under the Windows system, the Windows emulator can be used to run this part of the applications. In another example, if all of the plurality of applications run under the Android system, correspondingly, the Android emulator is used to run the plurality of applications. As reflected by these examples, according to the system types on which the various applications among the plurality of applications can run, using the corresponding simulator to simulate and run the corresponding applications among the plurality of applications, the display images and display texts corresponding to the applications under various system types can be obtained.

[0053] It should be noted that the displayed text may include text obtained from the intercepted output image (e.g., through OCR recognition technology). Additionally, the displayed text may further include text directly output during the operation of the application (i.e., in text form rather than image form). In some embodiments, all child nodes of the root node of the window component corresponding to each application may be traversed, and in response to triggering a specified event, a screenshot operation may be performed to obtain the display image corresponding to the application. Exemplarily, in the case of using an Android emulator to run the corresponding application, all child nodes of the root node of the Activity component may be traversed, and after triggering the specified event, a screenshot operation may be performed to obtain the display image corresponding to the application. Similarly, to obtain the displayed text, all child nodes of the root node of the Activity component corresponding to each application may be traversed, and in response to triggering the specified event, the View component information of the target window may be extracted to obtain the displayed text corresponding to the application; or the text may be extracted from the display image corresponding to each application as the displayed text corresponding to the application. Among them, Activity is one of the four major components of Android, which is a container for storing Views and also a carrier for displaying interfaces, and can be used to display an interface. A View is the basic building block in the user interface, representing a rectangular area on the screen, and the View is responsible for drawing and event handling in this area. Additionally, the above-mentioned specified event may be set according to actual business requirements. For example, when it is not desired that the application has the behavior of frequently triggering advertisements, a screenshot operation may be performed when triggering the viewing of an advertisement; when it is not desired that the application frequently pops up pop-up windows, a screenshot operation may be performed when triggering the pop-up of a pop-up window. Alternatively, random screenshots or timed screenshots may be taken when the application is in simulated operation.

[0054] To reduce unnecessary data processing volume for subsequent operations, preprocessing operations (step 750) may be performed on the image or text data obtained in steps 730 and 740 to generate the display image and display text corresponding to each application among multiple applications (step 760). Taking the Android system as an example, exemplarily, cleaning operations may be performed on the obtained View class text to remove common text in the application, such as text like "Confirm", "Cancel", "Exit", "Login", etc. The same cleaning may also be performed on the text recognized by OCR. For example, only text in the form of text boxes, pop-up boxes, link text, etc. may be retained. Exemplarily, the data generated after the preprocessing operations may be stored in the database device 152 of the remote facility 150 described above regarding Figure 1 described.

[0055] In step 220, the display image and display text corresponding to each application can be respectively input into the first sub-model and the second sub-model in the feature extraction model to obtain the image feature vector and text feature vector corresponding to each application. As described above, the first sub-model can be various models capable of extracting image features. The second sub-model can be various models capable of extracting text features. Exemplarily, the first sub-model can be a residual network. The second sub-model can be a TextCNN model.

[0056] In step 230, the image feature vector and text feature vector corresponding to each application can be fused to obtain the multi-modal feature vector corresponding to the application. As described above, exemplarily, the image feature vector and text feature vector corresponding to each application can be concatenated, and the obtained new feature vector can be used as the multi-modal feature vector corresponding to the application.

[0057] In step 240, the multi-modal feature vector corresponding to each application can be input into the third sub-model in the feature extraction model to obtain the compressed feature vector corresponding to the application. The dimension of the compressed feature vector corresponding to each application is smaller than the multi-modal feature vector corresponding to the application. Exemplarily, the third sub-model can be a fully connected neural network, which can have an appropriate number of layers (e.g., three layers, four layers, five layers or more layers). Among them, exemplarily, the compressed feature vector corresponding to each application can be the feature vector output from the middle layer of the third sub-model, and its dimension is smaller than the multi-modal feature vector corresponding to the application.

[0058] Then, in step 250, clustering processing can be performed on the fused feature vector corresponding to each application, and applications similar to the target application among the multiple applications can be determined according to the multiple clustering subspaces obtained by the clustering processing, where the fused feature vector corresponding to each application is at least obtained based on the compressed feature vector corresponding to the application. As an example, the fused feature vector corresponding to each application can be the same as the compressed feature vector corresponding to the application.

[0059] Exemplarily, the multiple clustering subspaces obtained by the clustering process can be regarded as the classification results of the multiple applications. For example, it can be considered that different applications corresponding to different feature vectors in different clustering subspaces belong to different types, while different applications corresponding to different feature vectors in the same clustering subspace can be considered to be of the same type. In this example, the clustering subspace to which the fusion feature vector corresponding to the target application belongs can be used as the target clustering subspace, and then the applications corresponding to all other feature vectors in the target clustering subspace can be used as the applications similar to the target application. Alternatively, the clustering subspace to which the fusion feature vector corresponding to the target application belongs and other clustering subspaces related to this clustering subspace can be jointly used as the target clustering subspace, and then the applications corresponding to all other feature vectors in the target clustering subspace can be used as the applications similar to the target application. Exemplarily, it can be determined whether these clustering subspaces are related according to the distance between different clustering subspaces. For example, when the distance between two clustering subspaces is less than a certain threshold, it is determined that these two clustering subspaces are related. Various methods can be used to define the distance between two clustering subspaces. Exemplarily, the maximum distance between each feature vector in one clustering subspace and each feature vector in another clustering subspace can be used as the distance between these two clustering subspaces. As another example, the average value of the distances between each feature vector in one clustering subspace and each feature vector in another clustering subspace can be used as the distance between these two clustering subspaces.

[0060] Through Figure 2 The similar application detection method 200 shown in the figure can extract corresponding multimodal features (such as the above multimodal feature vectors) according to the multimodal data (such as image and text data) generated during the operation of the application, and then perform compression processing on these multimodal features to obtain corresponding compressed feature vectors, and obtain the fusion feature vector corresponding to the application based on these compressed feature vectors. By using the multiple clustering subspaces obtained by clustering these fusion feature vectors, the applications similar to the target application are determined. In this process, the multimodal features of the application can be fully utilized and the redundancy therein can be removed, and it is possible to significantly reduce or even eliminate the need for manual participation in the similar application detection process, while helping to obtain a more accurate similar application detection result.

[0061] As described above, the clustering process can be a partitioning clustering process (i.e., clustering the corresponding vectors using a partitioning clustering algorithm), which is particularly suitable for situations where the number of applications is large. However, in some cases, for example, for some non-convex feature data with large inter-class overlap, it may be difficult to separate data of different categories well using a partitioning clustering process. In this case, in addition or alternatively, other types of clustering processes may be considered, such as density clustering processes (i.e., clustering the corresponding vectors using a density clustering algorithm). Accordingly, in some embodiments, the multiple clustering subspaces can be obtained by the following steps: performing a partitioning clustering process on the compressed feature vector corresponding to each application to obtain multiple partitioning clustering subspaces; performing a density clustering process on each of the multiple partitioning clustering subspaces to obtain multiple density clustering subspaces corresponding to the partitioning clustering subspace; and determining the multiple clustering subspaces based on the multiple density clustering subspaces corresponding to each partitioning clustering subspace. For example, all density clustering subspaces corresponding to each partitioning clustering subspace can be used as the multiple clustering subspaces. This embodiment can at least alleviate the problems that may exist when using partition clustering alone, and help to obtain more refined categories of applications. For more exemplary descriptions of partition clustering processing and density clustering processing, please refer to the following about Figure 3A The description is not repeated here.

[0062] As described above, the use of density clustering processing helps to obtain more refined categories of applications, but some unexpected phenomena may occur. For example, some originally similar applications (of the same category) may be divided into different density clustering subspaces after density clustering processing. To solve this problem, in some embodiments, the two density clustering subspaces with the smallest distance in the multiple density clustering subspaces corresponding to each partition clustering subspace can be merged to update the multiple density clustering subspaces corresponding to the partition clustering subspace until the preset termination merging condition is reached. For example, Figure 3A As shown, the multiple fusion feature vectors obtained by the multiple applications through the above steps constitute a feature vector set, and a number of partition clustering subspaces (in Figure 3A In the example of , the partition clustering subspaces K0, K1 and K2 are divided into clustering subspaces K0, K1 and K2), and then the partition clustering subspaces K0, K1 and K2 are subjected to density clustering processing to obtain multiple density clustering subspaces corresponding to the partition clustering subspaces K0, K1 and K2 respectively. Figure 3A The density clustering subspaces K1_D0, K1_D1, K1_D2, K1_D3 and K1_D4 corresponding to the partition clustering subspace K1 are schematically shown in FIG. Those skilled in the art should understand that the number of partition clustering subspaces and density clustering subspaces is not limited to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46Figure 3A The situation shown in. Exemplarily, the preset termination merging condition can be determined according to the number of corresponding density clustering subspaces. In Figure 3A the example of, for the five density clustering subspaces corresponding to the partition clustering subspace K1, the merging can be terminated when the number of density clustering subspaces after merging is 3. That is, the density clustering subspaces corresponding to the final partition clustering subspace K1 include CLS0, CLS1, and CLS2. Similar processing can be performed for other partition clustering subspaces (K0 and K2 in this example) to obtain the corresponding updated density clustering subspaces, and then these density clustering subspaces as a whole are used as the classification results of the final said multiple applications.

[0063] Figure 3B Schematically shows in a more detailed manner Figure 3A Example flowchart 320 of the similar class merging process in. As Figure 3B shown, in step 321, subspace numbering can be performed to distinguish different density clustering subspaces. In fact, in Figure 3A the example of, an example of numbering is provided, that is, the five different density clustering subspaces corresponding to the partition clustering subspace K1 are represented by K1_D0, K1_D1, K1_D2, K1_D3, and K1_D4. Of course, this is not restrictive, and other numbering means can be adopted as long as different density clustering subspaces can be distinguished. In step 322, subspace distance can be defined. Exemplarily, the distance between the two farthest data points in two density clustering subspaces (for example, Euclidean distance) can be defined as the distance between these two density clustering subspaces, as shown in the following formula:

[0064] Id pair (X,Y) = argmax{D(x,y): x ∈ X, y ∈ Y}.

[0065] Among them, Id pair (X,Y) represents the distance between density clustering subspace X and density clustering subspace Y, and D(x,y) represents the distance between element x (in this application, the corresponding feature vector) in density clustering subspace X and element y (in this application, the corresponding feature vector) in density clustering subspace Y.

[0066] In step 323, the distance between each two density clustering subspaces in the multiple density clustering subspaces can be calculated using the above formula, and then in step 324, the two density clustering subspaces with the smallest distance are merged to obtain the updated multiple density clustering subspaces. In step 325, it is judged whether the preset termination merging condition is satisfied. In the above description about Figure 3AAn example of a preset termination merging condition is provided in the description. Alternatively, another example is given here: Obtain the current number of density clustering subspaces and calculate a reference value. For example, calculate the sum of squared errors (SSE) of each element of these density clustering subspaces. Update the reference value (the value of SSE in this example) after each similarity class merging is completed, and observe the change of the reference value to find the inflection point where the reference value changes from increasing to decreasing. Stop the similarity class merging when reaching this inflection point. Alternatively, another example is given here: To determine the optimal number of density clustering subspaces, in other words, to determine the number of times to execute steps 323 and 324, in addition to the method using the inflection point method in the previous example, the idea of the gating state can also be used to make this number a learnable and adjustable parameter, that is, to determine the next value based on the historical value and the current value of the parameter. If it is determined in step 325 that the preset termination merging condition is satisfied, terminate / end the merging; otherwise, continue to execute steps 323 and 324.

[0067] The applicant notes that in some cases, due to reasons related to the model (for example, the first sub-model or the second sub-model fails to extract multiple features of a certain application or certain applications), or due to reasons related to the data (for example, the display image or display text corresponding to a certain application or certain applications itself contains very little or even misleading information), the obtained multi-modal feature vectors may not accurately reflect the features of the corresponding applications, thereby affecting the accuracy of similar application detection. To alleviate or even solve this problem, in some embodiments, the above detection method 200 further includes: obtaining a set of devices, where the multiple applications are installed on at least one device in the set of devices; determining a set of application relationship sequences based on the application attributes of each application installed on each device in the set of devices, where each application relationship sequence in the set of application relationship sequences is formed by arranging each application installed on the corresponding device in the set of devices according to the corresponding application attributes; determining a relationship feature vector corresponding to each application based on the set of application relationship sequences, and then obtaining a fusion feature vector corresponding to the application by fusing the compressed feature vector and the relationship feature vector corresponding to each application.

[0068] Take Figure 1 the terminal device 120 shown as an example. The application attributes of a certain application installed on it can be the developer information of the application, the operating system information on which the application is deployed, the installation time of the application, etc. These attributes only need to reflect certain aspects of the application's features, so as to help applications with the same or similar features be classified into the same category during the classification process. As Figure 4As shown, exemplarily, the device set D has three elements, namely, it includes three devices, Device_1, Device_2, and Device_3 (for example, mobile phones, tablets, computers, etc.), and the multiple applications include APP_1, APP_2, APP_3, and APP_4. Among them, applications APP_1, APP_2, and APP_3 are installed on device Device_1, applications APP_2, APP_4, and APP_5 are installed on device Device_2, and applications APP_1, APP_3, and APP_6 are installed on device Device_1. In Figure 4 the example of Figure 4 , application APP_1 is installed on devices Device_1 and Device_3, application APP_2 is installed on devices Device_1 and Device_2, application APP_3 is installed on devices Device_1 and Device_3, and application APP_4 is installed on device Device_2. It should be noted that although at least one application is installed on each device in the example of

[0069] In Figure 4In the example of [[ID=]], the application relationship sequence set 410 includes three application relationship sequences, namely, application relationship sequence 411, application relationship sequence 412, and application relationship sequence 413. Among them, the application relationship sequence 411 is formed by arranging the applications APP_1, APP_2, and APP_3 installed on the corresponding device Device_1 in the device set D according to the corresponding application attributes. Exemplarily, they can be arranged in ascending order according to the installation times of the applications APP_1, APP_2, and APP_3. Among them, the installation time of application APP_1 is the shortest, the installation time of application APP_2 is the longest, and the installation time of application APP_3 is between that of application APP_1 and application APP_2. In this case, the application relationship sequence 411 is composed of APP_1, APP_3, and APP_2 in sequence; similarly, the application relationship sequence 412 is composed of the applications APP_4, APP_5, and APP_2 installed on the corresponding device Device_2 in the device set D in sequence. Among them, the installation time of application APP_4 is the shortest, the installation time of application APP_2 is the longest, and the installation time of application APP_5 is between that of application APP_4 and application APP_2; similarly, the application relationship sequence 413 is composed of the applications APP_6, APP_1, and APP_3 installed on the corresponding device Device_3 in the device set D in sequence. Among them, the installation time of application APP_6 is the shortest, the installation time of application APP_3 is the longest, and the installation time of application APP_1 is between that of application APP_6 and application APP_3.

[0070] Regarding each application as a word, the application relationship sequence formed by them can be regarded as a sentence. Thus, based on the application relationship sequence set (which is equivalent to a paragraph or article containing multiple sentences), the relationship feature vector corresponding to each application can be determined. Furthermore, the compressed feature vector and the relationship feature vector corresponding to each application are fused to obtain the fused feature vector corresponding to the application. Exemplarily, in order to determine the relationship feature vector corresponding to each application, in some embodiments, word embedding processing can be performed on the application relationship sequence set, and the obtained word embedding vector corresponding to each application is used as the relationship feature vector corresponding to the application. Specifically, the application relationship sequence set can be used to train a word embedding model such as Word2Vec, GloVe, or FastText. After training, each word (i.e., each application) will correspond to a vector (i.e., the relationship feature vector corresponding to the application). In this way, the fused feature vector corresponding to each application containing richer features can be obtained. In Figure 4In the example, applications APP_1, APP_2, APP_3, and APP_4 respectively correspond to fusion feature vectors 421, fusion feature vector 422, fusion feature vector 423, and fusion feature vector 424, and these fusion feature vectors form a feature vector set 420.

[0071] It should be noted that although in the example described above with reference to Figure 4 the number of applications installed on each device is the same, and thus the sizes of the generated application relationship sequences are also the same, those skilled in the art should understand that in actual applications, the number of applications installed on different devices can be different. In this case, the generated application relationship sequences can be normalized to meet the requirements for the size of each training data during the training process of the corresponding word embedding model. For example, they can be unified into a preset length, and necessary deletion or padding operations may be required during this process.

[0072] In addition, in order to obtain the fusion feature vector corresponding to each application, exemplarily, the relationship feature vector and the compressed feature vector corresponding to each application can be concatenated to obtain the fusion feature vector corresponding to the application. Alternatively, the average vector of the relationship feature vector and the compressed feature vector corresponding to each application can be calculated as the fusion feature vector corresponding to the application.

[0073] Figure 5 FIG. schematically shows an example flowchart of a training method 500 (hereinafter simply referred to as the training method 500) of a similar application detection model according to some embodiments of the present application. The similar application detection model includes a first sub-model, a second sub-model, and a third sub-model. Exemplarily, the training method 500 can be implemented by Figure 1 the remote facility 150 shown, of course, this is not restrictive.

[0074] Specifically, in step 510, the display image and display text corresponding to each sample application among multiple sample applications can be obtained. The display image corresponding to each sample application includes at least a part of the image displayed during the operation of the sample application, and the display text corresponding to each sample application includes at least a part of the text in the display image corresponding to the sample application. Those skilled in the art should understand that step 510 can be performed by referring to the description of step 210 and Figure 7 above, so it will not be elaborated here. In addition, exemplarily, the relevant data generated after the preprocessing operation shown can be stored in a data source 611 such as Figure 7 shown, and the data source 611 can be the database device 152 in the remote facility 150 described above with reference to Figure 6A Figure 1

[0075] ​​In step 520, each sample can be respectively applied with the corresponding display image and display text to the first sub-model and the second sub-model to obtain the corresponding image feature vector and text feature vector for each sample. Specifically, each sample can be applied with the corresponding display image to the first sub-model to obtain the corresponding image feature vector for this sample, and each sample can be applied with the corresponding display text to the second sub-model to obtain the corresponding text feature vector for this sample.

[0076] Exemplarily, the first sub-model can be a residual network. The second sub-model can be a TextCNN model. Among them, the first sub-model can use an auto-encoder structure to complete the reconstruction process. Both the encoder and the decoder have three layers, and the convolutional encoder backbone network uses the structure of the residual network ResNet. Figure 6B Schematically shows the architecture of the residual network adopted in this example. As Figure 6B shown, the residual network 330 includes an encoder part and a decoder part. The image (X img ) is input into the encoder part. The encoder part includes three convolutional structures (Conv1, Conv2, and Conv3), and there is a bottleneck layer between them. After these three convolutional structures, there is a fully-connected layer (network) to generate the encoded feature vector. In the decoder part, after the encoded feature vector is processed by the fully-connected layer (network), three-layer feature map deconvolution (DeConv1, DeConv2, and DeConv3) is performed to finally obtain the reconstructed picture with the same size (X' img ).

[0077] As Figure 6A shown, after each sample is applied with the corresponding display image and input into the first sub-model, the corresponding image feature vector V1 for this sample is obtained. After the corresponding display text of this sample is input into the second sub-model, the corresponding text feature vector V2 for this sample is obtained. In step 530, the corresponding image feature vector V1 and text feature vector V2 for each sample can be fused to obtain the corresponding multi-modal feature vector V3 for this sample. Exemplarily, the corresponding image feature vector V1 and text feature vector V2 for each sample can be concatenated to obtain the corresponding multi-modal feature vector V3 for this sample. Alternatively, the average vector of the corresponding image feature vector V1 and text feature vector V2 for each sample can be calculated as the corresponding multi-modal feature vector V3 for this sample.

[0078] Next, in step 540, each sample's corresponding multi-modal feature vector V3 can be input into the third sub-model to obtain the corresponding compressed feature vector V4 output from the intermediate layer of the third sub-model, and the corresponding uncompressed feature vector V5 output from the output layer of the third sub-model. The dimension of each sample's corresponding compressed feature vector V4 is smaller than that of the corresponding multi-modal feature vector V3 of the sample, and the dimension of each sample's corresponding uncompressed feature vector V5 is the same as that of the corresponding multi-modal feature vector V3 of the sample.

[0079] In step 550, the text-image loss can be determined based on each sample's corresponding multi-modal feature vector V3 and uncompressed feature vector V5. Exemplarily, the text-image loss can be determined by the Euclidean distance (L2 norm) between each sample's corresponding multi-modal feature vector V3 and uncompressed feature vector V5, as shown in the following formula:

[0080] Loss data = ∑||V3 - V5||2.

[0081] It should be noted that, in addition to being constructed using the Euclidean distance in the above example, the text-image loss can also be constructed using other types of functions, including but not limited to any one of the following: cosine similarity loss function, mean squared error loss function, and logarithmic loss function.

[0082] In step 560, clustering processing can be performed on each sample's corresponding fused feature vector V6 to determine the clustering loss, where each sample's corresponding fused feature vector V6 is obtained based at least on the corresponding compressed feature vector V4 of the sample. In Figure 6A the example, each sample's corresponding fused feature vector V6 is the corresponding compressed feature vector V4 of the sample. Of course, this is only exemplary. In other examples, each sample's corresponding fused feature vector V6 can be jointly determined by the corresponding compressed feature vector V4 of the sample and other data (for example, other feature vectors corresponding to the sample). Exemplarily, the clustering processing can be partitioning clustering processing. Partitioning clustering processing clusters based on the similarity or distance between data points. It divides data points into different clusters such that data points within the same cluster are as similar as possible, while data points between different clusters are as different as possible. Common partitioning clustering algorithms include K-means, hierarchical clustering, etc. Here, taking the K-means partitioning clustering algorithm as an example, an example of the clustering loss is given as follows:

[0083] Loss cls = ∑p cls *||V6 - V6'||2.

[0084] In the above formula, p cls is the distribution weight after normalization of all class vectors generated by the clustering partition process. p cls is recursively calculated through the following formula:

[0085] p cls = p c / ∑p c , p c = 1 / ||V6 - V6'||2, V6' = (∑p cls *V6) / ∑p cls .

[0086] Among them, p c is the reciprocal of the distance between the vector and the class center and is used to represent probability, and V6' represents the mean of the calculated clustering center feature vectors. It should be noted that those skilled in the art know that when performing recursive calculation, the calculation order is V6', p c and p cls in sequence. Exemplarily, the initial value of the distribution weight p cls corresponding to each sample applying the fused feature vector V6 can be set, then the mean V6' of the clustering center feature vectors is calculated, and then p c is calculated, and then the updated p cls is calculated. The above process is continuously iterated to determine the clustering loss Loss cls .

[0087] It should also be noted that in addition to using the K-means partition clustering algorithm in the above example to construct the clustering loss, other types of partition clustering algorithms can also be used to construct it, including but not limited to any one of the following: K-modes, K-medians.

[0088] In step 570, based on the text-image loss and the clustering loss, the target loss can be determined. Referring to the above example, exemplarily, the weighted sum of the text-image loss Loss data and the clustering loss Loss cls can be used as the target loss Loss u , as shown in the following formula:

[0089] Loss u = α * Loss data + β * Loss cls .

[0090] Among them, α and β are hyperparameters used to balance the proportion of the text-image loss Loss data and the clustering loss Loss cls . Exemplarily, the values of these two hyperparameters can both be between 0 and 1 and the sum of these two hyperparameters is 1.

[0091] Finally, at step 580, the parameters of the similar application detection model can be iteratively updated such that the target loss Loss u meets a preset condition. Exemplarily, the preset condition can be that the number of training times reaches a preset number of training times, or it can be that the target loss Loss u is minimized. It should be noted that in this application, the expression "the target loss Loss u is minimized" can mean that the number of training times of the similar application detection model reaches a threshold number of times, making the target loss Loss u take the minimum value, or it can mean that the target loss Loss u takes a value less than the corresponding loss function threshold, triggering the termination of training, resulting in the target loss Loss u taking the minimum value, or it can also mean obtaining an ideal or very close to ideal global minimum value of the target loss Loss u . Additionally, an alternative expression for "the target loss Loss u is minimized" can be "the target loss Loss u converges", or it can be "the target loss Loss u reaches the extreme value", etc. The corresponding number threshold and loss function threshold can be set according to experience or flexibly adjusted according to the application scenario, and this application does not limit this. Additionally, during the training process of the similar application detection model (i.e., during the iterative update of the parameters of the similar application detection model), various parameter update methods can be used (e.g., through the backpropagation algorithm).

[0092] Through Figure 5 the training method 500 shown, corresponding multimodal features (e.g., multimodal feature vectors) can be extracted from the multimodal data (e.g., image and text data) generated during the operation of the sample application, and then an unsupervised training of the similar application detection model can be performed by establishing a target loss including a graph-text loss and a clustering loss, thereby significantly reducing or even eliminating the need for manual participation in the model training process. Meanwhile, due to the existence of the clustering loss, during the training process of the similar application detection model, multiple sample applications will also be clustered, and thus a relatively accurate application classification result can be obtained. Of course, in addition to using the application classification result obtained by clustering, since the trained similar application detection model can obtain the multimodal features of the sample application, the corresponding sample applications can also be classified based on these multimodal features. For example, sample applications with a relatively high similarity (e.g., cosine similarity) of multimodal features (e.g., greater than 0.8, 0.9, or other preset thresholds) can be grouped into one category.

[0093] In some embodiments, the above training method 500 further includes: performing an image augmentation operation on each sample using a corresponding display image to obtain an augmented image corresponding to the sample; based on the augmented images corresponding to each sample, constructing a plurality of positive sample image pairs and a plurality of negative sample image pairs, where each positive sample image pair includes two augmented images corresponding to the same sample, and each negative sample image pair includes two augmented images corresponding to two different samples respectively, and wherein, before inputting the display image corresponding to each sample into the first sub-model, the first sub-model is trained using the plurality of positive sample image pairs and the plurality of negative sample image pairs. Exemplarily, taking sample applications App_1 and App_2 as an example, the augmented images corresponding to sample application App_1 obtained by performing an image augmentation operation on the display image Img_1 corresponding to sample application App_1 include Img_11, Img_12, and Img_13, and the augmented images corresponding to sample application App_2 obtained by performing an image augmentation operation on the display image Img_2 corresponding to sample application App_2 include Img_21 and Img_22. Then, a plurality of positive sample image pairs and a plurality of negative sample image pairs shown in the following table can be constructed:

[0094]

[0095]

[0096] As can be seen from the above table, in this example, a total of four positive sample image pairs and six negative sample image pairs are obtained. Each negative sample image pair includes two augmented images corresponding to two different samples respectively. For example, the second negative sample image pair includes the augmented image Img_11 corresponding to sample application App_1 and the augmented image Img_22 corresponding to sample application App_2, the fourth negative sample image pair includes the augmented image Img_12 corresponding to sample application App_1 and the augmented image Img_22 corresponding to sample application App_2, and the sixth negative sample image pair includes the augmented image Img_13 corresponding to sample application App_1 and the augmented image Img_22 corresponding to sample application App_2. Each positive sample image pair includes two augmented images corresponding to the same sample. For example, the first positive sample image pair includes the augmented image Img_11 and the augmented image Img_12 corresponding to sample application App_1, and the fourth positive sample image pair includes the augmented image Img_21 and the augmented image Img_22 corresponding to sample application App_2. In practical applications, the number of sample applications can be much more than the above example (e.g., hundreds, thousands, tens of thousands, or even more), and the number of augmented images obtained by augmenting the display image corresponding to each sample application can be more than the above example (e.g., four, five, six, or more). Correspondingly, the number of positive sample image pairs and negative sample image pairs can be more than the above example.

[0097] As described above, in this example, before applying each sample's corresponding display image to the first sub-model, the first sub-model is trained using the plurality of positive sample image pairs and the plurality of negative sample image pairs. Exemplarily, the first sub-model may adopt the Figure 6B architecture shown. Such a residual network often loads pre-trained parameters. By training it (i.e., fine-tuning) using the plurality of positive sample image pairs and the plurality of negative sample image pairs, the model can have a stronger recognition ability for image data.

[0098] In some embodiments, the image augmentation operation includes at least one of the following operations: changing the color of each sample's corresponding display image; adding noise to each sample's corresponding display image; scaling each sample's corresponding display image by a preset ratio (hereinafter simply referred to as the scaling operation); performing a mirror transformation on each sample's corresponding display image (hereinafter simply referred to as the mirror operation); and covering a part of each sample's corresponding display image with a solid color area of a preset size (hereinafter simply referred to as the masking operation). Figure 9 The exemplary principles of some of these operations are shown in Figure 9 As shown, reference numeral 910 represents a display image corresponding to a certain sample, that is, the original image without the image augmentation operation; reference numeral 920 represents the image obtained by performing a mirror operation on the display image 910; reference numeral 930 represents the image obtained by performing a masking operation (in this example, covering with a white area) on the display image 910; reference numeral 940 represents the image obtained by performing a scaling operation (in this example, magnifying by 50%) on the display image 910.

[0099] In the above operations, the masking operation can simulate situations such as animations, movable cursors, and screenshot delays triggered by screenshot operations during application running. The scaling operation can simulate the situation where the screenshot content distribution is unfocused due to the different resolutions of different applications themselves. The mirror operation can simulate the situation where the image screenshot appears reversed when running certain applications. Adding noise (e.g., Gaussian noise) to each sample's corresponding display image can simulate the situation where the display image is damaged. Changing the color of each sample's corresponding display image can be performed for color images, for example, performing an RGB color channel transformation on a color display image to simulate the situation where the application displays different colors on different devices. These operations can enable the similar application detection model to have a stronger recognition ability for various types of image data.

[0100] It should be noted that for different sample applications among the multiple sample applications, the applied image augmentation operations can be exactly the same, completely different, or at least partially different. Still taking the sample applications App_1 and App_2 as an example, the image augmentation operation performed on the display image Img_1 corresponding to the sample application App_1 can be a masking operation, while the image augmentation operation performed on the display image Img_2 corresponding to the sample application App_2 can be a mirroring operation. Alternatively, the image augmentation operations performed on the display image Img_1 corresponding to the sample application App_1 and the display image Img_2 corresponding to the sample application App_2 can both be masking operations. Alternatively, the image augmentation operation performed on the display image Img_1 corresponding to the sample application App_1 can be a masking operation and a scaling operation, while the image augmentation operation performed on the display image Img_2 corresponding to the sample application App_2 can be a mirroring operation and a scaling operation.

[0101] In some embodiments, the above training method 500 further includes: performing a text augmentation operation on the display text corresponding to each sample application to obtain the augmented text corresponding to the sample application; constructing a plurality of positive sample text pairs and a plurality of negative sample text pairs based on the augmented text corresponding to each sample application, where each positive sample text pair includes two augmented texts corresponding to the same sample application, each negative sample text pair includes two augmented texts respectively corresponding to two different sample applications, and wherein, before inputting the display text corresponding to each sample application into the second sub-model, using the plurality of positive sample text pairs and the plurality of negative sample text pairs to train the second sub-model. Still taking the sample applications App_1 and App_2 as an example, the augmented text corresponding to App_1 obtained by performing a text augmentation operation on the display text Text_1 corresponding to the sample application App_1 includes Text_11 and Text_12, and the augmented text corresponding to App_2 obtained by performing a text augmentation operation on the display text Text_2 corresponding to the sample application App_2 includes Text_21, Text_22, and Text_23. Then, a plurality of positive sample text pairs and a plurality of negative sample text pairs as shown in the following table can be constructed:

[0102]

[0103] As can be seen from the above table, a total of four positive sample text pairs and six negative sample text pairs are obtained in this example. Each negative sample text pair includes two augmented texts corresponding to two different sample applications respectively. Taking the fifth negative sample text pair as an example, it includes the augmented text Text_12 corresponding to the sample application App_1 and the augmented text Text_22 corresponding to the sample application App_2. Each positive sample text pair includes two augmented texts corresponding to the same sample application. Taking the second positive sample text pair as an example, it includes the augmented text Text_21 and the augmented text Text_22 corresponding to the sample application App_1. In practical applications, the number of sample applications can be much larger than that in the above example (for example, hundreds, thousands, tens of thousands or even more), and the number of augmented texts obtained by augmenting the display texts corresponding to each sample application can be larger than that in the above example (for example, four, five, six or more). Correspondingly, the number of positive sample text pairs and negative sample text pairs can be larger than that in the above example. In addition, this application does not limit the sizes of the augmented images and augmented texts. Exemplarily, their sizes can be processed to be the same as the sizes of the corresponding display images and display texts respectively.

[0104] As described above, in this example, before inputting the display text corresponding to each sample application into the second sub-model, the second sub-model is trained using the multiple positive sample text pairs and the multiple negative sample text pairs. Exemplarily, the second sub-model can be a TextCNN model loaded with pre-trained parameters. By training it (i.e., fine-tuning) using the multiple positive sample text pairs and the multiple negative sample text pairs, the model can have a stronger ability to identify text data.

[0105] In some embodiments, the text augmentation operation includes at least one of the following operations: deleting words in the display text corresponding to each sample application with a preset probability (hereinafter simply referred to as the deletion operation); selecting a preset proportion or a preset number of words in the display text corresponding to each sample application for synonym replacement (hereinafter simply referred to as the replacement operation). Exemplarily, regarding the deletion operation, a certain proportion (for example, 5%, 10% or other appropriate values) of words can be randomly selected from the display text corresponding to each sample application for deletion; regarding the replacement operation, a preset proportion (for example, 5%, 15% or other appropriate values) or a preset number (for example, 1, 2, 3 or more) of words can be selected from the display text corresponding to each sample application for synonym replacement. The deletion operation can simulate the situation of text information loss caused by running failures or other reasons (such as display reasons) during the operation of the application. The replacement operation can simulate the situation where the text semantics output by different applications of the same type are the same or similar during operation. These operations can enable the similar application detection model to have a stronger ability to identify various types of text data.

[0106] Similarly, it should be noted that for different sample applications among the multiple sample applications, the applied text augmentation operations can be exactly the same, completely different, or at least partially different. Still taking the sample applications App_1 and App_2 as an example, the text augmentation operation for the display text Text_1 corresponding to the sample application App_1 can be a deletion operation, while the text augmentation operation for the display text Text_2 corresponding to the sample application App_2 can be a replacement operation. Alternatively, the text augmentation operations for the display text Text_1 corresponding to the sample application App_1 and the display text Text_2 corresponding to the sample application App_2 can both be replacement operations. Alternatively, the text augmentation operation for the display text Text_1 corresponding to the sample application App_1 can be a deletion operation and a replacement operation, while the text augmentation operation for the display text Text_2 corresponding to the sample application App_2 can be a replacement operation.

[0107] In some embodiments, the clustering process for the fusion feature vector corresponding to each sample application to determine the clustering loss includes: based on the fusion feature vector corresponding to each sample application, determining the partition clustering loss through partition clustering processing or determining the density clustering loss through density clustering processing; determining the clustering loss based on at least one of the partition clustering loss and the density clustering loss. Exemplarily, the partition clustering loss can be determined as the clustering loss; alternatively, the density clustering loss can be determined as the clustering loss; alternatively, the partition clustering loss and the density clustering loss can be determined, and the weighted sum of the two can be used as the clustering loss. In the above description of step 560, the method of determining the partition clustering loss as the clustering loss has been described with an example of the partition clustering algorithm K-means. The following gives a more detailed description of determining the density clustering loss.

[0108] Taking the density clustering algorithm DBSCAN as an example, in the DBSCAN algorithm, clusters are formed according to the density reaching a certain threshold. Exemplarily, the density clustering loss can be defined as:

[0109] Loss cls = ∑(P[i]!= i) * max(0, log(N)).

[0110] The density clustering loss listed in the above formula aims to minimize the number of all sample points belonging to incorrect clusters, where an incorrect cluster refers to a cluster to which the sample point should not be assigned, and where P[i] is the cluster to which sample point i (in this application, it can represent the fusion feature vector corresponding to the i-th sample application) belongs, and N is the number of neighbors of sample point i (within the radius Eps and not less than the minimum number of points MinPts). The values of the radius Eps and the minimum number of points MinPts can be preset as needed. In the DBSCAN algorithm, starting from any sample point, other sample points within its neighborhood are checked. If the number of sample points within the neighborhood is greater than or equal to minPts, then this sample point is regarded as a core point, and then all points that are connected to this core point and whose own neighborhood also has a number of sample points greater than or equal to minPts are found, thus forming the cluster to which this core point (sample point) belongs. The above process is continuously executed until all sample points are assigned to a cluster or no new cluster can be formed.

[0111] In particular, in the case of a large number of sample applications, the partition clustering loss can be determined as the clustering loss through partition clustering processing to accelerate the clustering and model training speed. In some cases, for example, for some non-convex feature data with a large overlap between classes, it may be difficult to separate data of different classes well using the partition clustering algorithm. In this case, additionally or alternatively, other types of clustering algorithms, such as density clustering algorithms, can be considered.

[0112] In some embodiments, the clustering processing includes partition clustering processing, and the above training method 500 further includes: merging two density clustering subspaces with the smallest distance among multiple density clustering subspaces to update the multiple density clustering subspaces until a preset termination merging condition is reached, and each density clustering subspace among the multiple density clustering subspaces is obtained by performing density clustering processing on the corresponding partition clustering subspace among multiple partition clustering subspaces, and the multiple partition clustering subspaces are obtained through the partition clustering processing. As described above, the partition clustering algorithm may be difficult to separate data of different classes well in some cases, that is, it affects the classification of sample applications. To at least alleviate or even solve this problem, in this embodiment, density clustering processing can be performed on the basis of partition clustering processing. Figure 8 Schematically shows the overall process 800 of this embodiment. In Figure 8 which, the part of data acquisition and preprocessing can refer to the above regarding steps 510 and Figure 7Description; In multi-modal feature extraction 810, reference may be made to the description above regarding steps 520-540 to obtain the corresponding multi-modal feature vectors, uncompressed feature vectors, and fused feature vectors for each sample application, and then the construction process of the unsupervised loss is carried out. Specifically, the text-image loss may be determined in combination with the description above regarding step 550, and the clustering loss (partition clustering loss in this example) may be determined in combination with the description above regarding step 560, and then the target loss may be determined in combination with the description above regarding step 570. In this process, density clustering processing is performed on the multiple partition clustering subspaces obtained by the partition clustering process to obtain the corresponding (multiple) density clustering subspaces for each partition clustering subspace. It should be noted that in the construction process of the unsupervised loss, in addition to the multi-modal feature extraction process described in step 810, the process of applying the relationship sequence feature extraction (step 820) may also be combined, and this process can more fully extract the features of each sample application to make the classification result obtained by the subsequent clustering operation more accurate.

[0113] In some embodiments, the above training method 500 further includes: obtaining a set of devices, where the multiple sample applications are installed on at least one device in the set of devices; determining a set of application relationship sequences based on the application attributes of each application installed on each device in the set of devices, where each application relationship sequence in the set of application relationship sequences is formed by arranging each application installed on the corresponding device in the set of devices according to the corresponding application attributes; determining a relationship feature vector corresponding to each sample application based on the set of application relationship sequences; and where, the fused feature vector corresponding to each sample application is obtained by fusing the compressed feature vector and the relationship feature vector corresponding to the sample application. These steps are characterized by step 820 in Figure 8 and specifically, reference may be made to the process described above regarding Figure 4 to execute these steps, which will not be elaborated here.

[0114] In Figure 10 the training architecture 1000 of the training method according to some other embodiments of the present application schematically shown, exemplarily, the display image and display text corresponding to each sample application and the set of application relationship sequences may be stored in, such as Figure 10 the data source 1010 shown, and the data source 1010 may be the database device 152 in the remote facility 150 described above regarding Figure 1 the description. As Figure 10As shown, after each sample's corresponding display image is input into the first sub-model, the image feature vector V1 corresponding to the sample is obtained. After the display text corresponding to the sample is input into the second sub-model, the text feature vector V2 corresponding to the sample is obtained. The image feature vector V1 and the text feature vector V2 corresponding to each sample are fused to obtain the multi-modal feature vector V3 corresponding to the sample. The multi-modal feature vector V3 corresponding to each sample is input into the third sub-model to obtain the compressed feature vector V4 corresponding to the sample output from the middle layer of the third sub-model, and the uncompressed feature vector V5 corresponding to the sample output from the output layer of the third sub-model is obtained. For the application relationship sequence set, after word embedding processing, the relationship feature vector V7 corresponding to each sample is obtained. Then, the relationship feature vector V7 and the compressed feature vector V4 are fused to obtain the fused feature vector V6 corresponding to the sample. The process of separately determining the image-text loss and the clustering loss to determine the target loss has been described in detail above and will not be elaborated here.

[0115] Figure 11 FIG. schematically shows an example block diagram of a similar application detection device 1100 according to some embodiments of the present application. Exemplarily, the similar application detection device 1100 can be deployed on Figure 1 the remote facility 150 shown (e.g., on the server 151). As Figure 11 shown, the similar application detection device 1100 includes a data acquisition module 1110, a feature extraction module 1120, a feature fusion module 1130, a feature compression module 1140, and an application detection module 1150.

[0116] Specifically, the data acquisition module 1110 can be configured to acquire the display images and display texts corresponding to each of multiple applications including the target application. The display image corresponding to each application includes at least a part of the image displayed during the operation of the application, and the display text corresponding to each application includes at least a part of the text in the display image corresponding to the application. The feature extraction module 1120 can be configured to input the display image and display text corresponding to each application into the first sub-model and the second sub-model in the feature extraction model respectively to obtain the image feature vector and text feature vector corresponding to each application. The feature fusion module 1130 can be configured to fuse the image feature vector and text feature vector corresponding to each application to obtain the multi-modal feature vector corresponding to the application. The feature compression module 1140 can be configured to input the multi-modal feature vector corresponding to each application into the third sub-model in the feature extraction model to obtain the compressed feature vector corresponding to the application, and the dimension of the compressed feature vector corresponding to each application is smaller than that of the multi-modal feature vector corresponding to the application. And the application detection module 1150 can be configured to perform clustering processing on the fusion feature vector corresponding to each application, and determine the application similar to the target application among the multiple applications according to the multiple clustering sub-spaces obtained by the clustering processing, where the fusion feature vector corresponding to each application is obtained at least based on the compressed feature vector corresponding to the application.

[0117] Figure 12 FIG. schematically shows an exemplary block diagram of a device 1200 for training a similar application detection model according to some embodiments of the present application (for simplicity, hereinafter simply referred to as the model training device 1200). The similar application detection model includes a first sub-model, a second sub-model, and a third sub-model. Exemplarily, the model training device 1200 can be deployed on Figure 1 the remote facility 150 shown (for example, on the server 151). As Figure 12 shown, the model training device 1200 includes a sample data acquisition module 1210, a sample feature extraction module 1220, a sample feature fusion module 1230, a sample feature compression module 1240, a graphic and text loss determination module 1250, a clustering loss determination module 1260, a target loss determination module 1270, and a model training module 1280.

[0118] Specifically, the sample data acquisition module 1210 can be configured to acquire the display images and display texts corresponding to each of the multiple sample applications. The display image corresponding to each sample application includes at least a part of the images displayed during the running of the sample application, and the display text corresponding to each sample application includes at least a part of the texts in the display image corresponding to the sample application. The sample feature extraction module 1220 can be configured to input the display image and display text corresponding to each sample application into the first sub-model and the second sub-model respectively to obtain the image feature vector and text feature vector corresponding to each sample application. The sample feature fusion module 1230 can be configured to fuse the image feature vector and text feature vector corresponding to each sample application to obtain the multi-modal feature vector corresponding to the sample application. The sample feature compression module 1240 can be configured to input the multi-modal feature vector corresponding to each sample application into the third sub-model to obtain the compressed feature vector corresponding to the sample application output from the middle layer of the third sub-model, and obtain the uncompressed feature vector corresponding to the sample application output from the output layer of the third sub-model. The dimension of the compressed feature vector corresponding to each sample application is smaller than that of the multi-modal feature vector corresponding to the sample application, and the dimension of the uncompressed feature vector corresponding to each sample application is the same as that of the multi-modal feature vector corresponding to the sample application. The graphic-text loss determination module 1250 can be configured to determine the graphic-text loss based on the multi-modal feature vector and uncompressed feature vector corresponding to each sample application. The clustering loss determination module 1260 can be configured to perform clustering processing on the fusion feature vector corresponding to each sample application to determine the clustering loss, where the fusion feature vector corresponding to each sample application is obtained based at least on the compressed feature vector corresponding to the sample application. The target loss determination module 1270 can be configured to determine the target loss based on the graphic-text loss and the clustering loss. And the model training module 1280 can be configured to iteratively update the parameters of the similar application detection model so that the target loss meets the preset conditions.

[0119] It should be understood that both the similar application detection device 1100 and the model training device 1200 can be implemented in a software, hardware, or a combination of software and hardware manner. Multiple different modules in the similar application detection device 1100 and the model training device 1200 can be implemented in the same software or hardware structure, or one module can be implemented by multiple different software or hardware structures.

[0120] In addition, the similar application detection device 1100 and the model training device 1200 can be respectively used to implement the detection method 200 and the training method 500 described above. The relevant details have been described in detail above. For the sake of brevity, they will not be repeated here. Additionally, these devices can have the same features and advantages as described in the corresponding methods.

[0121] Figure 13The figure illustrates an example system that includes an example computing device 1300 representative of one or more systems and / or devices that may implement the various techniques described herein. The computing device 1300 may be, for example, a server used by a node in a blockchain, a device associated with the server, a system-on-chip, and / or any other suitable computing device or computing system. The similar application detection device 1100 and the model training device 1200 described above with reference to Figure 11 and Figure 12 respectively may each take the form of the computing device 1300. Alternatively, either the similar application detection device 1100 or the model training device 1200 may be implemented as a computer program in the form of an application 1316.

[0122] As Figure 13 shown, the example computing device 1300 includes a processing system 1311, one or more computer-readable media 1312, and one or more I / O interfaces 1313 that are communicatively coupled to each other. Although not shown, the computing device 1300 may also include a system bus or other data and command transfer system that couples the various components to each other. The system bus may include any one or combination of different bus structures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus utilizing any one of a variety of bus architectures. Also contemplated are various other examples, such as control and data lines.

[0123] The processing system 1311 represents the functionality to perform one or more operations using hardware. Thus, the processing system 1311 is illustrated as including hardware elements 1314 that may be configured as a processor, functional blocks, etc. This may include being implemented in hardware as an application-specific integrated circuit or other logic device formed using one or more semiconductors. The hardware elements 1314 are not limited by the materials from which they are formed or the processing mechanism employed therein. For example, a processor may be composed of (multiple) semiconductors and / or transistors (e.g., an electronic integrated circuit (IC)). In such a context, the executable instructions of the processor may be electronically executable instructions.

[0124] The computer-readable medium 1312 is illustrated as including a memory / storage 1315. The memory / storage 1315 represents the memory / storage capacity associated with one or more computer-readable media. The memory / storage 1315 may include volatile media (such as random access memory (RAM)) and / or non-volatile media (such as read-only memory (ROM), flash memory, optical discs, magnetic discs, etc.). The memory / storage 1315 may include fixed media (e.g., RAM, ROM, fixed hard disk drive, etc.) and removable media (e.g., flash memory, removable hard disk drive, optical disc, etc.). The computer-readable medium 1312 may be configured in various other ways as further described below.

[0125] One or more I / O interfaces 1313 represent the functionality that allows a user to input commands and information into the computing device 1300 using various input devices and optionally also allows information to be presented to the user and / or other components or devices using various output devices. Examples of input devices include a keyboard, a cursor control device (e.g., a mouse), a microphone (e.g., for voice input), a scanner, a touch functionality (e.g., a capacitive or other sensor configured to detect physical touch), a camera (e.g., that can detect motion not involving touch as a gesture using visible or non-visible wavelengths such as infrared frequencies), and so on. Examples of output devices include a display device (e.g., a projector), speakers, a printer, a network card, a haptic response device, etc. Thus, the computing device 1300 may be configured in various ways as further described below to support user interaction.

[0126] The computing device 1300 also includes an application 1316. The application 1316 may be, for example, a software instance of the similarity application detection device 1100 or the model training device 1200, and implements the techniques described herein in combination with other elements in the computing device 1300.

[0127] Various techniques may be described herein in the general context of software hardware elements or program modules. Generally, these modules include routines, programs, elements, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The terms "module", "function", and "component" as used herein generally represent software, firmware, hardware, or a combination thereof. The features of the techniques described herein are platform-independent, meaning that these techniques may be implemented on various computing platforms having various processors.

[0128] The implementation of the described modules and techniques may be stored on or transmitted across some form of computer-readable medium. The computer-readable medium may include various media accessible by the computing device 1300. By way of example and not limitation, the computer-readable medium may include "computer-readable storage media" and "computer-readable signal media".

[0129] Contrary to mere signal transmission, carrier, or the signal itself, a "computer-readable storage medium" refers to a medium and / or device that can persistently store information, and / or a tangible storage device. Thus, a computer-readable storage medium refers to a non-signal-bearing medium. Computer-readable storage media include hardware such as volatile and non-volatile, removable and non-removable media and / or storage devices implemented in a method or technology suitable for storing information such as computer-readable instructions, data structures, program modules, logic elements / circuits, or other data. Examples of computer-readable storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical storage devices, hard disks, cassette tapes, magnetic tapes, magnetic disk storage devices, or other magnetic storage devices, or other storage devices, tangible media, or articles of manufacture suitable for storing the desired information and accessible by a computer.

[0130] A "computer-readable signal medium" refers to a signal-bearing medium configured to send instructions to a computing device 1300, such as via a network. Signal media typically can embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave, data signal, or other transmission mechanism. Signal media also include any information delivery medium. The term "modulated data signal" refers to a signal in which one or more of the characteristics are set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media.

[0131] As previously described, hardware elements 1314 and computer-readable media 1312 represent instructions, modules, programmable device logic, and / or fixed device logic implemented in hardware, which in some embodiments can be used to implement at least some aspects of the techniques described herein. Hardware elements can include integrated circuits or systems-on-a-chip, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), and components of other hardware devices implemented in silicon or other hardware. In this context, a hardware element can serve as a processing device that executes program tasks defined by the instructions, modules, and / or logic embodied by the hardware element, and as a hardware device for storing instructions for execution, e.g., the previously described computer-readable storage media.

[0132] The foregoing combinations can also be used to implement the various techniques and modules described herein. Accordingly, software, hardware, or program modules and other program modules can be implemented as one or more instructions and / or logic on a computer-readable storage medium of a certain form and / or embodied by one or more hardware elements 1314. The computing device 1300 can be configured to implement specific instructions and / or functions corresponding to the software and / or hardware modules. Thus, for example, by using the computer-readable storage medium of the processing system and / or the hardware element 1314, the module can be implemented at least in part in hardware as a module executable by the computing device 1300 as software. The instructions and / or functions can be executable / operable by one or more articles of manufacture (e.g., one or more computing devices 1300 and / or processing systems 1311) to implement the techniques, modules, and examples described herein.

[0133] In various embodiments, the computing device 1300 can assume various different configurations. For example, the computing device 1300 can be implemented as a computer-like device including a personal computer, a desktop computer, a multi-screen computer, a laptop computer, a netbook, etc. The computing device 1300 can also be implemented as a mobile device-like device including mobile devices such as mobile phones, portable music players, portable game devices, tablet computers, multi-screen computers, etc. The computing device 1300 can also be implemented as a television-like device, which includes a device having or connected to a generally larger screen in a leisure viewing environment. These devices include televisions, set-top boxes, game consoles, etc.

[0134] The techniques described herein can be supported by these various configurations of the computing device 1300 and are not limited to the specific examples of the techniques described herein. The functionality can also be implemented in whole or in part on the "cloud" 1320 by using a distributed system, such as via the platform 1322 described below.

[0135] The cloud 1320 includes and / or represents a platform 1322 for resources 1324. The platform 1322 abstracts the underlying functionality of the hardware (e.g., servers) and software resources of the cloud 1320. The resources 1324 can include applications and / or data that can be used when performing computer processing on servers remote from the computing device 1300. The resources 1324 can also include services provided via the Internet and / or via a subscriber network such as a cellular or Wi-Fi network.

[0136] Platform 1322 can abstract resources and functions to connect computing device 1300 with other computing devices. Platform 1322 can also be used to abstract a hierarchy of resources to provide a corresponding level of hierarchy for the demands encountered for resources 1324 implemented via platform 1322. Thus, in an interconnected device embodiment, the implementation of the functions described herein can be distributed throughout system 1300. For example, the functions can be implemented partially on computing device 1300 and via platform 1322 that abstracts the functions of cloud 1320.

[0137] It should be understood that, for clarity, embodiments of the present application have been described with reference to different functional units. However, it will be apparent that, without departing from the present application, the functionality of each functional unit can be implemented in a single unit, implemented in multiple units, or implemented as part of other functional units. For example, functionality illustrated as being performed by a single unit can be performed by multiple different units. Thus, reference to a particular functional unit is only considered as a reference to an appropriate unit for providing the described functionality, rather than indicating a strict logical or physical structure or organization. Thus, the present application can be implemented in a single unit, or can be physically and functionally distributed among different units and circuits.

[0138] In embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0139] It will be understood that although the terms first, second, third, etc. may be used herein to describe various devices, elements, components, or parts, these devices, elements, components, or parts should not be limited by these terms. These terms are only used to distinguish one device, element, component, or part from another device, element, component, or part.

[0140] Although the present application has been described in connection with some embodiments, it is not intended to be limited to the specific forms set forth herein. Rather, the scope of the present application is limited only by the appended claims. Additionally, although individual features may be included in different claims, these may possibly be advantageously combined, and the inclusion in different claims does not imply that a combination of features is not feasible and / or advantageous. The order of features in the claims does not imply that the features must work in any specific order. Further, in the claims, the word "comprising" does not exclude other elements, and the terms "a" or "an" do not exclude a plurality. The reference signs in the claims are provided only as illustrative examples and should not be construed as limiting the scope of the claims in any way.

[0141] It should be understood that, for clarity, embodiments of the present application have been described with reference to different functional units. However, it will be apparent that, without departing from the present application, the functionality of each functional unit can be implemented in a single unit, implemented in multiple units, or implemented as part of other functional units. For example, the functionality described as being performed by a single unit can be performed by multiple different units. Thus, the reference to a particular functional unit is only considered as a reference to the appropriate unit for providing the described functionality, rather than indicating a strict logical or physical structure or organization. Accordingly, the present application can be implemented in a single unit, or can be physically and functionally distributed among different units and circuits.

[0142] The present application provides a computer-readable storage medium having stored thereon computer-readable instructions which, when executed, implement the above-described method for detecting similar applications or the method for training a similar application detection model.

[0143] The present application provides a computer program product or a computer program, the computer program product or the computer program including computer instructions which are stored in a computer-readable storage medium. The processor of a computing device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions such that the computing device executes the method for detecting similar applications or the method for training a similar application detection model provided in the above various alternative implementations.

[0144] By studying the drawings, the disclosure, and the appended claims, those skilled in the art will be able to understand and implement variations of the disclosed embodiments when practicing the claimed subject matter. In the claims, the word "comprising" does not exclude other elements or steps, and "a" or "an" does not exclude a plurality. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used advantageously.

[0145] It is understandable that in the specific embodiments of the present application, data related to similar application detection and the training process of the similar application detection model are involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

Claims

1. A method for detecting similar applications, comprising: Obtaining the display image and display text corresponding to each application among a plurality of applications including a target application, where the display image corresponding to each application includes at least a part of the image displayed during the running of the application, and the display text corresponding to each application includes at least a part of the text in the display image corresponding to the application; Respectively inputting the display image and display text corresponding to each application into a first sub-model and a second sub-model in a feature extraction model to obtain an image feature vector and a text feature vector corresponding to each application; Fusing the image feature vector and the text feature vector corresponding to each application to obtain a multi-modal feature vector corresponding to the application; Inputting the multi-modal feature vector corresponding to each application into a third sub-model in the feature extraction model to obtain a compressed feature vector corresponding to the application, where the dimension of the compressed feature vector corresponding to each application is smaller than that of the multi-modal feature vector corresponding to the application; And Performing clustering processing on the fused feature vector corresponding to each application, and determining the applications similar to the target application among the plurality of applications according to the plurality of clustering sub-spaces obtained by the clustering processing, where the fused feature vector corresponding to each application is at least obtained based on the compressed feature vector corresponding to the application.

2. The method according to claim 1, wherein the plurality of applications run in a target system, and wherein obtaining the display image and display text corresponding to each application among the plurality of applications includes: Using a target simulator corresponding to the target system to simulate the running of each application among the plurality of applications to obtain the display image and display text corresponding to each application.

3. The method according to claim 2, wherein the display image corresponding to each application is obtained through the following steps: Traversing all child nodes of the root node of the window component corresponding to each application, and in response to triggering a specified event, performing a screenshot operation to obtain the display image corresponding to the application.

4. The method according to claim 2, wherein the display text corresponding to each application is obtained through the following steps: Traversing all child nodes of the root node of the window component corresponding to each application, and in response to triggering a specified event, extracting the container component information of the target window to obtain the display text corresponding to the application; or Extracting text from the display image corresponding to each application as the display text corresponding to the application.

5. The method according to claim 1, wherein the plurality of clustering sub-spaces are obtained through the following steps: Performing partition clustering processing on the compressed feature vector corresponding to each application, and using the plurality of partition clustering sub-spaces obtained by the partition clustering processing as the plurality of clustering sub-spaces.

6. The method according to claim 1, wherein the plurality of clustering sub-spaces are obtained through the following steps: Performing partition clustering processing on the compressed feature vector corresponding to each application to obtain a plurality of partition clustering sub-spaces; Performing density clustering processing on each partition clustering sub-space among the plurality of partition clustering sub-spaces to obtain a plurality of density clustering sub-spaces corresponding to the partition clustering sub-space; And Based on the plurality of density clustering sub-spaces corresponding to each partition clustering sub-space, determining the plurality of clustering sub-spaces.

7. The method according to claim 6 further includes: merging two density clustering subspaces with the smallest distance among the multiple density clustering subspaces corresponding to each partition clustering subspace to update the multiple density clustering subspaces corresponding to the partition clustering subspace until a preset termination merging condition is reached.

8. The method according to claim 1, wherein the clustering the fusion feature vectors corresponding to each application and determining the applications similar to the target application among the multiple applications according to the multiple clustering subspaces obtained by the clustering includes: determining a target clustering subspace including the fusion feature vector corresponding to the target application from the multiple clustering subspaces; using the applications corresponding to the target feature vectors in the target clustering subspace among the multiple applications as the applications similar to the target application, where the target feature vectors are different from the fusion feature vectors corresponding to the target application.

9. The method according to claim 1 further includes: obtaining a device set, where the multiple applications are installed on at least one device in the device set; determining an application relationship sequence set based on the application attributes of each application installed on each device in the device set, where each application relationship sequence in the application relationship sequence set is formed by arranging each application installed on the corresponding device in the device set according to the corresponding application attributes; determining a relationship feature vector corresponding to each application based on the application relationship sequence set; and wherein, the fusion feature vector corresponding to each application is obtained by fusing the compressed feature vector and the relationship feature vector corresponding to the application.

10. The method according to claim 9, wherein the determining a relationship feature vector corresponding to each application based on the application relationship sequence set includes: performing word embedding processing on the application relationship sequence set and using the obtained word embedding vector corresponding to each application as the relationship feature vector corresponding to the application.

11. A training method for a similar application detection model, the similar application detection model includes a first sub-model, a second sub-model and a third sub-model, and the method includes: obtaining a display image and display text corresponding to each sample application among multiple sample applications, the display image corresponding to each sample application includes at least a part of the image displayed during the running of the sample application, and the display text corresponding to each sample application includes at least a part of the text in the display image corresponding to the sample application; inputting the display image and display text corresponding to each sample application into the first sub-model and the second sub-model respectively to obtain an image feature vector and a text feature vector corresponding to each sample application; fusing the image feature vector and the text feature vector corresponding to each sample application to obtain a multi-modal feature vector corresponding to the sample application; Apply the corresponding multi-modal feature vector of each sample to the third sub-model to obtain the compressed feature vector corresponding to the sample output from the intermediate layer of the third sub-model, and obtain the uncompressed feature vector corresponding to the sample output from the output layer of the third sub-model. The dimension of the compressed feature vector corresponding to each sample is less than that of the multi-modal feature vector corresponding to the sample, and the dimension of the uncompressed feature vector corresponding to each sample is the same as that of the multi-modal feature vector corresponding to the sample; Determine the image-text loss based on the multi-modal feature vector and the uncompressed feature vector corresponding to each sample; Perform clustering processing on the fused feature vector corresponding to each sample to determine the clustering loss, where the fused feature vector corresponding to each sample is obtained based at least on the compressed feature vector corresponding to the sample; Determine the target loss based on the image-text loss and the clustering loss; And Iteratively update the parameters of the similar application detection model so that the target loss meets the preset conditions.

12. The method according to claim 11, further comprising: Perform image augmentation operations on the display image corresponding to each sample to obtain the augmented image corresponding to the sample; Based on the augmented image corresponding to each sample, construct a plurality of positive sample image pairs and a plurality of negative sample image pairs, where each positive sample image pair includes two augmented images corresponding to the same sample, and each negative sample image pair includes two augmented images corresponding to two different samples respectively, and where, Before inputting the display image corresponding to each sample into the first sub-model, use the plurality of positive sample image pairs and the plurality of negative sample image pairs to train the first sub-model.

13. The method according to claim 11, further comprising: Perform text augmentation operations on the display text corresponding to each sample to obtain the augmented text corresponding to the sample; Based on the augmented text corresponding to each sample, construct a plurality of positive sample text pairs and a plurality of negative sample text pairs, where each positive sample text pair includes two augmented texts corresponding to the same sample, and each negative sample text pair includes two augmented texts corresponding to two different samples respectively, and where, Before inputting the display text corresponding to each sample into the second sub-model, use the plurality of positive sample text pairs and the plurality of negative sample text pairs to train the second sub-model.

14. The method according to claim 11, further comprising: Obtain a set of devices, where the plurality of sample applications are installed on at least one device in the set of devices; Based on the application attributes of each application installed on each device in the set of devices, determine a set of application relationship sequences, where each application relationship sequence in the set of application relationship sequences is formed by arranging each application installed on the corresponding device in the set of devices according to the corresponding application attributes; Based on the set of application relationship sequences, determine the relationship feature vector corresponding to each sample application; and where, The fused feature vector corresponding to each sample application is obtained by fusing the compressed feature vector and the relationship feature vector corresponding to the sample application.

15. A similar application detection device, comprising: A data acquisition module configured to acquire a display image and display text corresponding to each of a plurality of applications including a target application, the display image corresponding to each application including at least a part of the image displayed during the operation of the application, and the display text corresponding to each application including at least a part of the text in the display image corresponding to the application; A feature extraction module configured to input the display image and display text corresponding to each application into a first sub-model and a second sub-model in a feature extraction model respectively to obtain an image feature vector and a text feature vector corresponding to each application; A feature fusion module configured to fuse the image feature vector and the text feature vector corresponding to each application to obtain a multi-modal feature vector corresponding to the application; A feature compression module configured to input the multi-modal feature vector corresponding to each application into a third sub-model in the feature extraction model to obtain a compressed feature vector corresponding to the application, the dimension of the compressed feature vector corresponding to each application being smaller than that of the multi-modal feature vector corresponding to the application; And An application detection module configured to perform clustering processing on the fused feature vector corresponding to each application, and determine an application similar to the target application among the plurality of applications according to a plurality of clustering sub-spaces obtained by the clustering processing, wherein the fused feature vector corresponding to each application is obtained based at least on the compressed feature vector corresponding to the application.

16. A computing device, comprising: A memory configured to store computer-executable instructions; A processor configured to execute the method according to any one of claims 1 to 14 when the computer-executable instructions are executed by the processor.

17. A computer-readable storage medium storing computer-executable instructions, which when executed, execute the method according to any one of claims 1 to 14.

18. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, execute the method according to any one of claims 1 to 14.