APP and domain name association relationship acquisition method and device, and medium

By capturing packets on the initial device within a preset time period, the number of specified domains and apps in the traffic packets is obtained. The association between apps and domains is established using the method of equal and unique quantities, which solves the coverage blind spots that rely on manual analysis in existing technologies and achieves high accuracy and high reliability in obtaining association relationships.

CN121940376APending Publication Date: 2026-04-28HANGZHOU YUNSHEN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU YUNSHEN TECH CO LTD
Filing Date
2026-01-26
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, identifying the relationship between an app and its communication domain name mainly relies on manual analysis, which has blind spots and cannot fully obtain the relationship between the app and the domain name.

Method used

By capturing packets on the initial device within a preset time period, the number of specified domains and apps in the traffic packets is obtained. The association is determined by using the equality and uniqueness of the number of apps, and the association relationship between apps and domains is determined by combining feature values ​​and tag similarity.

Benefits of technology

It enables more comprehensive discovery of the relationship between apps and domains, improving accuracy and reliability, and reducing reliance on manual analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940376A_ABST
    Figure CN121940376A_ABST
Patent Text Reader

Abstract

The invention provides an APP and domain name incidence relation obtaining method and device and a medium, and relates to the technical field of data processing, and the method comprises the steps: carrying out the package capturing of a plurality of initial devices in a preset time period, obtaining an initial flow package, obtaining a designated domain name contained in the initial flow package, and carrying out the collection of the designated domain name; obtaining a first number of initial devices corresponding to the specified domain names, obtaining a second number of initial devices installed with the specified APPs, and if the first number of the target domain names is equal to the second number of the target APPs and the first number of any specified domain name except the target domain names is not equal to the first number of the target domain names, sending the specified APPs to the target domain names; and the second number of any specified APP except the target APP is not equal to the second number of the target APP, and the target APP and the target domain name are associated, so that the association relationship between the domain name and the APP is found more comprehensively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method, device and medium for obtaining the association between an app and a domain name. Background Technology

[0002] With the widespread use of mobile applications (APPs) and their domain names, the supervision of APPs and their domain names is becoming increasingly strict. In existing technologies, APPs are tagged, and then the domain name is tagged in the same way based on the relationship between the APP and the domain name, so dual supervision of APP and domain name can be achieved. Alternatively, the domain name's tags can be obtained first, and then dual supervision of APP and domain name can be achieved. Therefore, obtaining the relationship between APP and domain name is particularly important.

[0003] In existing technologies, the main method for identifying the association between an app and its communication domain name relies on security researchers or analysts manually establishing rules to identify the association. However, this method is heavily dependent on manual intervention and has blind spots in its coverage. Therefore, there is a need for a more comprehensive method to obtain the association between apps and domain names without relying heavily on manual intervention. Summary of the Invention

[0004] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows: According to a first aspect of the present invention, a method for obtaining the association relationship between an APP and a domain name is provided, the method comprising the following steps: S100 captures packets from several initial devices within a preset time period to obtain initial traffic packets; S200, obtain the specified domain name contained in the initial traffic package, and obtain the first number of initial devices corresponding to the specified domain name, wherein the specified domain name is a domain name not associated with the APP; S300, based on the APP installed on the initial device, obtain the second number of initial devices that have installed the specified APP, wherein the specified APP is an APP not associated with a domain name; S400, if the first number of target domains is equal to the second number of target apps, and the first number of any specified domain other than the target domain is not equal to the first number of target domains, and the second number of any specified app other than the target app is not equal to the second number of target apps, then associate the target app with the target domain. The target app is one of several designated apps, and the target domain name is one of several designated domain names.

[0005] According to a second aspect of the present invention, a non-transitory computer-readable storage medium is provided, wherein a computer program is stored in the storage medium, and the computer program is loaded and executed by a processor to implement the aforementioned method.

[0006] According to a third aspect of the present invention, an electronic device is provided, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned method.

[0007] The present invention has at least the following beneficial effects: In summary, within a preset time period, packet capture is performed on several initial devices to obtain initial traffic packets, the specified domain names contained in the initial traffic packets are obtained, and the first number of initial devices corresponding to the specified domain names is obtained. Based on the APP installed on the initial devices, the second number of initial devices with the specified APP installed is obtained. If the first number of target domain names is equal to the second number of target APPs, and the first number of any specified domain name other than the target domain name is not equal to the first number of target domain names, and the second number of any specified APP other than the target APP is not equal to the second number of target APPs, the target APP and the target domain name are associated. The present invention, based on the specified domain name issued by the initial devices within a preset time period and the specified APP installed on the initial devices, associates the domain name and APP by means of equal quantity, which can more comprehensively discover the association relationship between the domain name and the APP. It solves the problem that the traditional technology can only obtain the association relationship by the correspondence between the domain name and the APP name, and the accuracy and reliability are very high by means of equal quantity and uniqueness. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a flowchart illustrating a method for obtaining the association between an app and a domain name, as provided in an embodiment of the present invention. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0011] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar tasks and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0012] This invention provides a method for obtaining the association between an app and a domain name, such as... Figure 1 As shown, the method includes the following steps: S100 captures packets from several initial devices within a preset time period to obtain initial traffic packets.

[0013] Specifically, packet capture is performed on the initial device, that is, a "checkpoint" is set on the network communication path of the initial device, and a copy of the network data packets passing through the "checkpoint" is made as the initial traffic packet.

[0014] S200: Obtain the specified domain name contained in the initial traffic package, and obtain the first number of initial devices corresponding to the specified domain name. The specified domain name is a domain name not associated with the APP. Specifically, the first number is the number of initial devices containing the initial traffic package of the specified domain name. It can be understood that if two initial traffic packages of a specified domain name correspond to the same device, then the first number is 1.

[0015] In one embodiment of the present invention, the designated domain name is a domain name in the relational database that is not associated with the APP, and the relational database stores the association relationships between several domain names and the APP.

[0016] S300: Based on the apps installed on the initial devices within a preset time period, obtain a second number of initial devices that have installed a specified app, wherein the specified app is an app not associated with a domain name.

[0017] Specifically, the initial devices corresponding to the second number all have the specified APP installed, and the initial devices corresponding to the second number are different.

[0018] Those skilled in the art will understand that any method for obtaining an initially installed APP in the prior art falls within the protection scope of this invention, and will not be elaborated further here.

[0019] In one embodiment of the present invention, the designated APP is an APP in the relational database that is not associated with a domain name.

[0020] Specifically, as those skilled in the art will know, S200 and S300 can be executed in interchangeable order.

[0021] S400, if the first number of target domains is equal to the second number of target apps, and the first number of any specified domain other than the target domain is not equal to the first number of target domains, and the second number of any specified apps other than the target apps is not equal to the second number of target apps, then associate the target apps with the target domains; wherein, the target app is one of several specified apps, and the target domain is one of several specified domains.

[0022] It can be understood that if the first number of target domains is equal to the second number of target apps, and the first number of any specified domain other than the target domain is not equal to the first number of target domains, and the second number of any specified app other than the target app is not equal to the second number of target apps, that is: if the first number of target domains is equal to the second number of target apps, and the first number of target domains is unique among all first numbers, and the second number of target apps is unique among all second numbers.

[0023] In one embodiment of the present invention, after associating the target APP and the target domain name, the association relationship between the target APP and the target domain name is stored in a relational database.

[0024] In summary, within a preset time period, packet capture is performed on several initial devices to obtain initial traffic packets, the specified domain names contained in the initial traffic packets are obtained, and the first number of initial devices corresponding to the specified domain names is obtained. Based on the apps installed on the initial devices, the second number of initial devices with the specified apps installed is obtained. If the first number of target domain names equals the second number of target apps, and the first number of any specified domain name other than the target domain name does not equal the first number of target domain names, and the second number of any specified apps other than the target app does not equal the second number of target apps, the target apps and target domain names are associated. This invention, based on the specified domain names emitted by the initial devices within a preset time period and the specified apps installed on the initial devices, associates domain names and apps through the method of equal quantity, which can more comprehensively discover the association relationship between domain names and apps. It solves the problem that traditional technologies can only obtain the association relationship through the correspondence between domain names and app names, and the accuracy and reliability are very high through the equality and uniqueness of the quantity.

[0025] Following the S400, it also includes: Obtain the target tag of the target APP and use the target tag of the target APP as the tag of the target domain name, wherein the target tag is one of several preset tags.

[0026] In an exemplary description of the present invention, if the target tag of the target APP is an abnormal tag, then the tag of the target domain name is also an abnormal tag.

[0027] Following S400, if there exist several target apps whose second quantity equals the target domain's first quantity, and the first quantity of any specified domain other than the target domain is not equal to the target domain's first quantity, then the following steps are performed: S410: Obtain the initial device on which the target APP is installed as the target device, and obtain a list of intermediate APPs installed on the target device with the target tag as the target tag. The list of intermediate APPs includes the IDs of several intermediate APPs.

[0028] Specifically, if there exist several target apps whose second quantity is equal to the target domain name's first quantity, and the first quantity of any specified domain name other than the target domain name is not equal to the target domain name's first quantity, that is: the target app's second quantity is equal to the target domain name's first quantity, the target domain name's first quantity is unique among all first quantities, and the target app's second quantity is not unique among all second quantities.

[0029] Specifically, the ID of the intermediate APP is the first identifier of the intermediate APP.

[0030] S420, obtain intermediate traffic packets captured from the intermediate APP within a preset time period, and obtain an intermediate feature value list of a preset feature list corresponding to the intermediate traffic packets and an initial feature value list of a preset feature list corresponding to the initial traffic packets of the target domain. The preset feature list includes several preset features, the initial feature value list includes several initial feature values, and the intermediate feature values ​​include several intermediate feature values. The preset features include at least the time from packet capture to traffic packet.

[0031] Furthermore, the preset feature also includes: the sending time of the data packet.

[0032] Specifically, there may be one or more intermediate traffic packages corresponding to the intermediate APP, and there may be one or more initial traffic packages corresponding to the target APP.

[0033] S430, obtain the first similarity between each intermediate feature value list and each initial feature value list.

[0034] Those skilled in the art will understand that any method for obtaining intermediate feature value lists and initial feature value lists in the prior art falls within the protection scope of this invention. For example, by obtaining the Euclidean distance between intermediate feature values ​​and initial feature values, the smaller the Euclidean distance, the greater the first similarity.

[0035] S440: For a target label, obtain the average of the first similarity scores of several intermediate apps corresponding to the target label as the label similarity score.

[0036] S450, if the tag similarity of a target tag meets the first preset similarity condition, the target tag is used as the tag of the target domain name, wherein the first preset similarity condition includes at least: the tag similarity is greater than the first preset similarity threshold.

[0037] In another embodiment of the present invention, the preset similarity condition further includes: the tag similarity degree of the target tag divided by the entropy value of the specified similarity degree is greater than the preset entropy value threshold, wherein the specified similarity degree is the largest tag similarity degree other than the tag similarity degree of the target tag.

[0038] In summary, if there exist several target apps whose second quantity equals the first quantity of target domains, and the first quantity of any specified domain other than the target domain is not equal to the first quantity of the target domain, the initial device that installs the target app is obtained as the target device, and a list of intermediate apps installed on the target device with the tag being the target tag is obtained. Intermediate traffic packets captured from the intermediate apps within a preset time period are obtained, and a list of intermediate feature values ​​corresponding to the intermediate traffic packets and a list of initial feature values ​​corresponding to the preset feature list of the initial traffic packets are obtained. The first similarity between each list of intermediate feature values ​​and each list of initial feature values ​​is obtained. For a target tag, the average of the first similarity of several intermediate apps corresponding to the target tag is obtained as the tag similarity. If the tag similarity of a target tag meets the first preset similarity condition, the target tag is used as the tag of the target domain. It can be understood that when the second quantity of target apps is not unique, the association between the target app and the target domain cannot be established. This invention uses the target device as a bridge to obtain the similarity of the features of the intermediate traffic packets and the initial traffic packets of the intermediate apps of the target tag on the target device, and determines the tag of the target domain, thereby increasing the number of target domains with the target tag obtained.

[0039] In another embodiment of the invention, the method further includes the following after S420: S421, for a target device, obtain the initial feature value list of the target domain name corresponding to the target device as the first feature value list, and obtain the intermediate feature value list of the intermediate APP corresponding to the target device as the second feature value list. The first feature value list includes several first feature values, and the second feature value list includes several second feature values.

[0040] S422, obtain the similarity between each first feature value list and each second feature value list as the second similarity; and obtain the sum of the second similarity as the device similarity. By calculating the similarity between the initial feature value list and the intermediate feature value list corresponding to the target device, we can focus more on a single device and understand that the usage habits of the device are fixed. For intermediate apps under the same target tag, if the target domain name's tag belongs to that target tag, then the first feature value list of the target domain name under the target device and the second feature value list of the intermediate apps under the target device are similar.

[0041] S422, For a target label, obtain the average device similarity of all target devices of all target apps corresponding to the target label as the median similarity.

[0042] S423, if the intermediate similarity corresponding to a target tag meets the second preset similarity condition, the target tag is used as the tag of the target domain name, wherein the second preset similarity condition includes at least: the tag similarity is greater than the second preset similarity threshold.

[0043] In another embodiment of the present invention, the preset similarity condition further includes: the entropy value of the intermediate similarity of the target label divided by the candidate similarity is greater than a preset entropy threshold, wherein the candidate similarity is the largest intermediate similarity other than the intermediate similarity of the target label.

[0044] In summary, for a target device, an initial feature value list of the target domain corresponding to the target device is obtained as a first feature value list, and an intermediate feature value list of the intermediate APP corresponding to the target device is obtained as a second feature value list. The similarity between each first feature value list and each second feature value list is obtained as a second similarity. The sum of the second similarities is obtained as the device similarity. For a target tag, the average of the device similarities of all target devices of all target APPs corresponding to the target tag is obtained as the intermediate similarity. If the intermediate similarity of a target tag meets a second preset similarity condition, the target tag is used as a tag of the target domain. This invention improves the accuracy of obtaining the association between target domains and target APPs by calculating the similarity between the initial feature value list and the intermediate feature value list corresponding to the target device.

[0045] Specifically, the target tags for the target app are obtained through the following steps: S001, Obtain the training dataset, which includes several training data sets, including: training APP name, training APP package name, training APP description, and initial label of the training APP; the initial label is one of several preset labels.

[0046] Specifically, the training dataset includes training data for several training apps corresponding to each preset label. Specifically, the training app description is a textual introduction of the training app, which can be obtained through an app store.

[0047] S002, input the target instruction, training dataset, and target APP data into the large language model to obtain the initial output result of the large language model. The target APP data includes: target APP name, target APP package name, and target APP description. The target instruction is: the target label of the target APP is determined according to the training dataset, and the target label is one of the preset labels. The initial output result includes: output label and analysis text of the output label.

[0048] S003, based on the initial output of the large language model, determine the target tags for the target app.

[0049] Furthermore, S003 also includes: S031, obtain the intermediate keyword list Ai, initialize i=1, the intermediate keyword list A1 is obtained by extracting keywords from the analysis text in the initial output result, the intermediate keyword list includes several intermediate keywords.

[0050] Specifically, as those skilled in the art will know, any method for keyword extraction in the prior art falls within the scope of protection of this invention, and will not be elaborated further here.

[0051] S032, input the intermediate keyword list Ai, the initial instruction, and the target APP data into the large language model to obtain intermediate output results. The initial instruction is to determine the intermediate tags of the target APP based on the initial keywords. The intermediate output results include: intermediate tags and analysis text with the intermediate tags as the output results. The intermediate tags are one of several preset tags.

[0052] S033, if i is less than n, obtain the intermediate keyword list Ai+1 based on the analysis text whose output is an intermediate label, i=i+1, and execute S032; otherwise, execute S034, where n is the preset number of rounds. The accuracy of the target label is ensured based on the number of rounds.

[0053] S034, if there is an intermediate label in the intermediate output results of n rounds that is the same as the intermediate label in the intermediate output results of y rounds and is consistent with the output label in the initial output results, the intermediate label is used as the target label of the target APP, where y is a preset ratio threshold.

[0054] In summary, the intermediate keyword list Ai is obtained. The intermediate keyword list Ai, the initial instructions, and the target APP data are input into the large language model to obtain intermediate output results. If i is less than n, the intermediate keyword list Ai+1 is obtained based on the analysis text whose output results are intermediate labels, i=i+1, and S032 is executed; otherwise, S034 is executed. If there are intermediate labels in the intermediate output results of n rounds that are all the same and consistent with the output labels in the initial output results, the intermediate labels are used as the target labels of the target APP. The target labels of the target APP are accurately obtained through intermediate keywords and preset rounds.

[0055] Furthermore, n is determined through the following steps: S010, obtain the first sample dataset and the second sample dataset. The first sample dataset includes several first sample data, and the second sample dataset includes several second sample data.

[0056] The first sample data includes several sample keywords, and the number of sample keywords is less than a preset threshold. These sample keywords are obtained by extracting keywords from the analyzed text of the sample output results. The sample output results are input into the large language model via target instructions, training dataset, and sample APP data. The sample APP data includes: sample APP name, sample APP package name, and sample APP description. The second sample data includes several sample keywords, and the number of sample keywords is not less than a preset threshold. It can be understood that the more sample keywords there are, the more information they contain; the fewer sample keywords there are, the less information they contain. The number of sample keywords will affect the prediction of the large language model.

[0057] S020, the first training instruction is used as the target training instruction, and several first sample data are used as initial sample data. The threshold is input to obtain the model, several first thresholds are obtained, and the median of these first thresholds is used as the first median. The first training instruction is different from the second training instruction. It can be understood that the output of the threshold-obtaining model is used as the first threshold.

[0058] S030, the second training instruction is used as the target training instruction, and several first sample data are used as initial sample data. The threshold acquisition model is input to obtain several second thresholds, and the median of these second thresholds is used as the second median. The mean of the first median and the second median is used as the first mean 'a'. It can be understood that the output of the threshold acquisition model is used as the second threshold.

[0059] In one embodiment of the present invention, the first training instruction is: extract the core semantics of the target sample data based on their common features, and generate a training label. The second training instruction is: based on the target sample data, associate it with its possible extended domains, and generate a training label. It can be understood that setting the first training instruction to "focus" and "extract commonalities" and setting the second training instruction to "associate" and "extend" allows the large language model to think in different directions under the same initial sample data, thereby obtaining a more applicable preset number of rounds.

[0060] S040, the first training instruction is used as the target training instruction, and several second sample data are used as initial sample data. The threshold is input to obtain the model, several third thresholds are obtained, and the median of these third thresholds is used as the third median. It can be understood that the output of the threshold-obtaining model is used as the third threshold.

[0061] S050, the second training instruction is used as the target training instruction, and several second sample data are used as initial sample data. The threshold acquisition model is input to obtain several fourth thresholds, and the median of these fourth thresholds is used as the fourth median. The mean of the third and fourth medians is used as the second mean b. It can be understood that the output of the threshold acquisition model is used as the fourth threshold.

[0062] S060, obtain n, where n satisfies the following conditions: n is not less than the smaller of a and b, and is not greater than the larger of a and b.

[0063] In summary, the process involves obtaining a first sample dataset and a second sample dataset, using the first training instruction as the target training instruction, and using several first sample datasets as initial sample data. A threshold is input to obtain the model, several first thresholds are obtained, and the median of these first thresholds is used as the first median. The second training instruction is used as the target training instruction, and several first sample datasets are used as initial sample data. A threshold is input to obtain the model, several second thresholds are obtained, and the median of these second thresholds is used as the second median. The mean of the first and second medians is used as the first mean. The first training instruction is then used as the target... The training instructions are used, and several second sample data are used as initial sample data. The input threshold is used to obtain the model. Several third thresholds are obtained, and the median of several third thresholds is obtained as the third median. The second training instructions are used as the target training instructions, and several second sample data are used as initial sample data. The input threshold is used to obtain the model. Several fourth thresholds are obtained, and the median of several fourth thresholds is obtained as the fourth median. The mean of the third median and the fourth median is obtained as the second mean. n is obtained. By using different numbers of sample keywords / target training instructions in different directions, a more accurate preset number of rounds can be obtained.

[0064] Furthermore, the threshold acquisition model performs the following steps: S021, obtain target sample data Bj, initialize j=1, target sample data B1 is the initial sample data.

[0065] S022, input the target sample data Bj, the target training instructions and the sample APP data into the large language model, and obtain the training output result Cj+1. The training output result includes: training labels and analysis text with training labels as the training output result.

[0066] S023, if C1, C2, ..., Cj, Cj+1 meet the preset training conditions, output the threshold j+1; otherwise, obtain the analysis text of Cj+1 and use the keywords in the analysis text as the target sample data Bj+1, j=j+1, and execute S022.

[0067] Embodiments of the present invention also provide a non-transitory computer-readable storage medium that can be disposed in an electronic device to store a computer program related to implementing a method in the method embodiments, the computer program being loaded and executed by the processor to implement the method provided in the above embodiments.

[0068] Embodiments of the present invention also provide an electronic device, including: a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method provided in the above embodiments.

[0069] Embodiments of the present invention also provide a computer program product including program code, which, when the program product is run on an electronic device, causes the electronic device to perform the steps of the methods described above in various exemplary embodiments of the present invention.

[0070] While specific embodiments of the invention have been described in detail by way of examples, those skilled in the art should understand that the examples are for illustrative purposes only and are not intended to limit the scope of the invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the invention.

Claims

1. A method for obtaining the association between an app and a domain name, characterized in that, The method includes the following steps: S100 captures packets from several initial devices within a preset time period to obtain initial traffic packets; S200, obtain the specified domain name contained in the initial traffic package, and obtain the first number of initial devices corresponding to the specified domain name, wherein the specified domain name is a domain name not associated with the APP; S300, based on the APP installed on the initial device, obtain the second number of initial devices that have installed the specified APP, wherein the specified APP is an APP not associated with a domain name; S400, if the first number of target domains is equal to the second number of target apps, and the first number of any specified domain other than the target domain is not equal to the first number of target domains, and the second number of any specified app other than the target app is not equal to the second number of target apps, then associate the target app with the target domain. The target app is one of several designated apps, and the target domain name is one of several designated domain names.

2. The method for obtaining the association relationship between an APP and a domain name according to claim 1, characterized in that, Following the S400, it also includes: Obtain the target tag of the target APP and use the target tag of the target APP as the tag of the target domain name, wherein the target tag is one of several preset tags.

3. The method for obtaining the association relationship between an APP and a domain name according to claim 2, characterized in that, Following S400, if there exist several target apps whose second quantity equals the target domain's first quantity, and the first quantity of any specified domain other than the target domain is not equal to the target domain's first quantity, then the following steps are performed: S410, obtain the initial device that installs the target APP as the target device, and obtain a list of intermediate APPs installed on the target device with the target tag as the target tag. The list of intermediate APPs includes the IDs of several intermediate APPs. S420, obtain intermediate traffic packets captured from the intermediate APP within a preset time period, and obtain an intermediate feature value list of a preset feature list corresponding to the intermediate traffic packets and an initial feature value list of a preset feature list corresponding to the initial traffic packets of the target domain. The preset feature list includes several preset features, the initial feature value list includes several initial feature values, and the intermediate feature values ​​include several intermediate feature values. The preset features include at least the time when the traffic packets were captured. S430, obtain the first similarity between each intermediate feature value list and each initial feature value list; S440, For a target label, obtain the average of the first similarity of several intermediate apps corresponding to the target label as the label similarity; S450, if the tag similarity of a target tag meets the first preset similarity condition, the target tag is used as the tag of the target domain name, wherein the first preset similarity condition includes at least: the tag similarity is greater than the first preset similarity threshold.

4. The method for obtaining the association relationship between an APP and a domain name according to claim 3, characterized in that, Following S420 are: S421, for a target device, obtain the initial feature value list of the target domain name corresponding to the target device as the first feature value list, and obtain the intermediate feature value list of the intermediate APP corresponding to the target device as the second feature value list. The first feature value list includes several first feature values, and the second feature value list includes several second feature values. S422, obtain the similarity between each first feature value list and each second feature value list as a second similarity; and obtain the sum of the second similarity as the device similarity; S422, For a target label, obtain the average device similarity of all target devices of all target apps corresponding to the target label as the median similarity; S423, if the intermediate similarity corresponding to a target tag satisfies the second preset similarity condition, the target tag is used as the tag of the target domain name, wherein the second preset similarity condition includes at least: the tag similarity is greater than the second preset similarity threshold.

5. The method for obtaining the association relationship between an APP and a domain name according to claim 2, characterized in that, Obtain the target tags for the target app by following these steps: S001, Obtain the training dataset, which includes several training data sets, including: training APP name, training APP package name, training APP description, and initial label of the training APP; the initial label is one of several preset labels. S002, input the target instruction, training dataset, and target APP data into the large language model to obtain the initial output result of the large language model. The target APP data includes: target APP name, target APP package name, and target APP description. The target instruction is: the target label of the target APP is determined according to the training dataset, and the target label is one of the preset labels. The initial output result includes: output label and analysis text of the output label. S003, based on the initial output of the large language model, determine the target tags for the target app.

6. The method for obtaining the association relationship between an APP and a domain name according to claim 5, characterized in that, S003 also includes: S031, obtain the intermediate keyword list Ai, initialize i=1, the intermediate keyword list A1 is obtained by extracting keywords from the analysis text in the initial output result, the intermediate keyword list includes several intermediate keywords; S032, input the intermediate keyword list Ai, the initial instruction and the target APP data into the big language model, and obtain the intermediate output results. The initial instruction is: determine the intermediate tags of the target APP based on the initial keywords. The intermediate output results include: intermediate tags and analysis text with intermediate tags as the output results. S033, if i is less than n, obtain the intermediate keyword list Ai+1 based on the analysis text whose output is an intermediate label, i=i+1, and execute S032; otherwise, execute S034, where n is the preset number of rounds; S034, if there is an intermediate label in the intermediate output results of n rounds that is the same as the intermediate label in the intermediate output results of y rounds and is consistent with the output label in the initial output results, the intermediate label is used as the target label of the target APP, where y is a preset ratio threshold.

7. The method for obtaining the association relationship between an APP and a domain name according to claim 6, characterized in that, n is determined by the following steps: S010, Obtain the first sample dataset and the second sample dataset. The first sample dataset includes several first sample data, and the second sample dataset includes several second sample data. The first sample data includes several sample keywords and the number of sample keywords is less than a preset threshold. The sample keywords are obtained by extracting keywords from the analysis text of the sample output results. The sample output results are input into the large language model through the target instruction, training dataset and sample APP data. The sample APP data includes: sample APP name, sample APP package name and sample APP description. The second sample data includes several sample keywords and the number of sample keywords is not less than a preset threshold. S020, take the first training instruction as the target training instruction, and take several first sample data as initial sample data respectively, input threshold to obtain the model, obtain several first thresholds and obtain the median of several first thresholds as the first median, the first training instruction is different from the second training instruction; S030, take the second training instruction as the target training instruction, and take several first sample data as initial sample data respectively, input the threshold to obtain the model, obtain several second thresholds and obtain the median of several second thresholds as the second median; obtain the mean of the first median and the second median as the first mean a; S040, take the first training instruction as the target training instruction, and take several second sample data as initial sample data respectively, input the threshold to obtain the model, obtain several third thresholds and obtain the median of several third thresholds as the third median; S050, take the second training instruction as the target training instruction, and take several second sample data as initial sample data respectively, input the threshold to obtain the model, obtain several fourth thresholds and obtain the median of several fourth thresholds as the fourth median; obtain the mean of the third median and the fourth median as the second mean b; S060, obtain n, where n satisfies the following conditions: n is not less than the smaller of a and b, and is not greater than the larger of a and b.

8. The method for obtaining the association relationship between an APP and a domain name according to claim 7, characterized in that, The threshold acquisition model performs the following steps: S021, Obtain target sample data Bj, initialize j=1, target sample data B1 is the initial sample data; S022, Input the target sample data Bj, the target training instructions and the sample APP data into the large language model to obtain the training output result Cj+1. The training output result includes: training labels and analysis text in which the training output result is the training labels. S023, if C1, C2, ..., Cj, Cj+1 meet the preset training conditions, output the threshold j+1; otherwise, obtain the analysis text of Cj+1 and use the keywords in the analysis text as the target sample data Bj+1, j=j+1, and execute S022.

9. A non-transitory computer-readable storage medium, characterized in that, The storage medium stores a computer program, which is loaded and executed by a processor to implement the method for obtaining the association relationship between an APP and a domain name as described in any one of claims 1-8.

10. An electronic device, comprising: A processor, a memory, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the method for obtaining the association relationship between an APP and a domain name as described in any one of claims 1-8.