A Dark Web Similar Commodity Determination Method, Storage Medium and Computing Device Based on Multi-Mode Fusion Features

By using multi-mode fusion feature method in the dark web market to calculate the feature values ​​of product pictures and text, the difficulty in determining similar products caused by blurred pictures and simple text in dark web product products is solved, and the judgment accuracy and dynamic monitoring capabilities are improved.

CN114090901BActive Publication Date: 2025-06-13NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111367617.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-18
Publication Date
2025-06-13
Estimated Expiration
2041-11-18

AI Technical Summary

Technical Problem

The pictures of dark web products are blurred and the text information is simple, which makes it difficult to determine similar products.

Method used

Using a method based on multimode fusion feature, the product similarity is calculated by calculating the perceived hash fingerprint value of the product picture and the Word2Vec sentence vector value of the product text, and combined with the md5 value, the product similarity is calculated.

Benefits of technology

It improves the accuracy of judging similarity of products in the dark web market, simplifies the process, and enhances the monitoring and real-time early warning capabilities of dark web market dynamics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114090901B_ABST
    Figure CN114090901B_ABST
Patent Text Reader

Abstract

The present invention provides a method, a storage medium, and a computing device for determining similar dark web products based on multi-modal fusion features. The method includes: Step 10, collecting dark web product data and classifying the products; among them, the collected dark web product data includes product text and product pictures, and an md5 value is generated for the collected product pictures; Step 20, based on the dark web product data collected in Step 1 and the product classification results, calculating the perceptual hash fingerprint value of the product pictures and the Word2Vec sentence vector value of the product text; Step 30, calculating the product similarity based on the md5 value of the product pictures, the perceptual hash fingerprint value of the product pictures, and the Word2Vec sentence vector value of the product text. By constructing a similarity calculation method that fuses product pictures and product texts, the present invention can solve the problem that it is difficult to determine similar products due to the blurred product pictures and simple text information on the dark web.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dark web similar commodity determination, and more specifically, to a method, storage medium, and computing device for determining dark web similar commodities based on multi-modal fusion features. Background Art

[0002] Dark web markets (or "dark webs") are commercial websites that specialize in illegal commodity trading. They can be accessed through the dark web (such as Tor) and differ from clear web e-commerce websites in terms of specialization, technology, and main support. Most markets are for facilitating transactions between buyers and sellers of illegal commodities, but there are a large number of sellers in dark web markets, and many of the posted commodities are extremely similar or even exactly the same. To better monitor the dynamics of dark web market transactions and obtain timely information on various commodities, in addition to collecting as much information as possible on the commodities and sites of dark web markets, more work is to classify and count various commodities, filter similar commodities, and promptly discover new categories of commodities and issue warnings. Therefore, the determination of similar commodities is extremely necessary.

[0003] Currently, for commodities on clear web e-commerce websites, due to the high clarity of pictures and detailed text descriptions, the determination of similar commodities can basically be completed by relying on single picture similarity or text similarity. In addition, the determination of similar commodities is more for using collaborative filtering algorithms to make recommendations for buyer users. In contrast, for dark web commodities, the pictures are blurred and the text information is simple, making it difficult to determine similar commodities, and there are few methods in this regard on the market. Summary of the Invention

[0004] The present invention aims to provide a method, storage medium, and computing device for determining dark web similar commodities based on multi-modal fusion features to solve the problem of difficulty in determining similar dark web commodities due to blurred pictures and simple text information.

[0005] A method for determining dark web similar commodities based on multi-modal fusion features provided by the present invention includes the following steps:

[0006] Step 10: Collect dark web commodity data and classify the commodities; among them, the collected dark web commodity data includes commodity text and commodity pictures, and generate an md5 value for the collected commodity pictures.

[0007] Step 20: Calculate the perceptual hash fingerprint value of the commodity pictures and the Word2Vec sentence vector value of the commodity text based on the dark web commodity data collected in Step 1 and the commodity classification results.

[0008] Step 30: Calculate the commodity similarity based on the md5 value of the commodity pictures, the perceptual hash fingerprint value of the commodity pictures, and the Word2Vec sentence vector value of the commodity text.

[0009] Further, step 10 includes the following sub-steps:

[0010] Step 11, implement the collection of dark web commodity data. The collected dark web commodity data includes structured commodity text and commodity pictures on the commodity detail page. Among them, the commodity text includes commodity id, commodity name, and commodity description.

[0011] Step 12, for commodities with commodity pictures, while collecting the commodity pictures, use the general md5 calculation method to obtain the md5 value of the commodity pictures, and use this md5 value as the name of the commodity pictures, store them in the Seaweed database according to the set storage location, and generate the corresponding storage address string.

[0012] Step 13, classify the commodities, add first- and second-level commodity labels, and store the first- and second-level commodity labels together with the collected commodity text, the md5 value and storage address string of the commodity pictures in the ES database.

[0013] Further, step 20 includes the following sub-steps:

[0014] Step 21, read from the ES database the commodity text including commodity id, commodity name, and commodity description, and the md5 value of the commodity pictures. For commodities with non-empty md5 values, obtain the commodity pictures from the Seaweed database according to the corresponding storage address string.

[0015] Step 22, merge the commodity id, commodity name, and commodity description as the complete commodity text.

[0016] Step 23, calculate the perceptual hash fingerprint value of the commodity pictures.

[0017] Step 24, calculate the Word2Vec sentence vector value of the commodity text.

[0018] Step 25, store the data obtained in steps 21, 23, and 24 into the MySQL commodity feature vector table together.

[0019] Further, step 30 includes the following sub-steps:

[0020] Step 31, read the commodity id of the new commodity, denoted as id1, read the md5 value of the picture of the new commodity, denoted as md5_1, read the perceptual hash fingerprint value of the commodity picture of the new commodity, denoted as h1, read the Word2Vec sentence vector value of the commodity text of the new commodity, denoted as v1, and read the first- and second-level commodity labels of the new commodity, denoted as c, t.

[0021] Step 32: Read the attribute values of the products that have the same first- and second-level product tags as the new product from the MySQL product feature vector table, which also includes the product ID, the MD5 value of the product image, the perceptual hash fingerprint value of the product image, and the Word2Vec sentence vector value of the product text:

[0022] Step 33: From the products corresponding to the attribute values of each product obtained in Step 32, sequentially select a product that has not participated in the comparison yet, and set its product ID as id2, the MD5 value of the product image as md5_2, the perceptual hash fingerprint value of the product image as h2, and the Word2Vec sentence vector value of the product text as v2;

[0023] Step 34: If md5_1 == md5_2, then set the similarity s of the two products to 1;

[0024] Step 35: If the condition in Step 34 is not satisfied, then:

[0025] (1) Calculate the Hamming distance d = hamming_dist(h1, h2) of the perceptual hash fingerprint values of the product images of the two products;

[0026] (2) Calculate the cosine similarity c = cos_similarity(v1, v2) of the Word2Vec sentence vector values of the product texts of the two products;

[0027] (3) Set the similarity of the two products as s = (1 / ln(e + d / 10) + c) / 2, where e is the natural logarithm;

[0028] Step 36: Compare the similarity s of the two products solved in Step 34 or 35 with the pre-set similarity threshold lambda, and screen out the two products with the similarity s ≥ lambda;

[0029] Step 37: Combine the product IDs and similarity s of the two products that meet the conditions in Step 36 into a triple (id1, id2, s) and store it in the MySQL product similarity table;

[0030] Step 38: Return to Step 33 until all products have been compared.

[0031] The present invention also provides a computer terminal storage medium storing computer terminal executable instructions, characterized in that the computer terminal executable instructions are used to execute the above-mentioned method for determining similar dark web products based on multi-modal fusion features.

[0032] The present invention also provides a computing device, characterized by including:

[0033] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-described method for determining similar dark web goods based on multi-modal fusion features.

[0034] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are as follows:

[0035] 1. By constructing a similarity calculation method that fuses product images and product texts, the present invention can calculate the similarity of products in the dark web market, obtain similar products under various categories, facilitate better classification of products in the dark web market, improve the accuracy of judging product similarity, with a simple method and strong interpretability, thus solving the problem of difficult judgment of similar products due to blurred product images and simple text information in the dark web.

[0036] 2. Through dark web data collection, product image feature calculation, product text feature calculation, and similarity calculation, the present invention can effectively monitor the buying and selling of new products, provide real-time warnings, and better track the dynamics of the dark web market. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 It is the overall flowchart of the method for determining similar dark web goods based on multi-modal fusion features according to the embodiments of the present invention.

[0039] Figure 2 It is the flowchart of step 10 in the method for determining similar dark web goods based on multi-modal fusion features according to the embodiments of the present invention.

[0040] Figure 3 It is the flowchart of step 20 in the method for determining similar dark web goods based on multi-modal fusion features according to the embodiments of the present invention.

[0041] Figure 4 It is the flowchart of step 30 in the method for determining similar dark web goods based on multi-modal fusion features according to the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated herein generally can be arranged and designed in a variety of different configurations.

[0043] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0044] Embodiment

[0045] As Figure 1 shown, this embodiment proposes a method for determining similar dark web products based on multi-modal fusion features, including the following steps:

[0046] Step 10, collect dark web product data and classify the products; among them, the collected dark web product data includes product text and product pictures, and generate an md5 value for the collected product pictures; this step mainly uses dark web data collection technology to collect structured data such as product id, product name, and product description of product text. For products with product pictures, the corresponding product pictures are also collected at the same time. In addition, generate the md5 value of the product picture, classify the product, store the structured data in the ES database, and store the product pictures in the Seaweed database. As Figure 2 shown, step 10 includes the following sub-steps:

[0047] Step 11, implement the collection of dark web product data, and the collected dark web product data includes structured product text and product pictures on the product details page; among them, the product text includes product id, product name, and product description;

[0048] Step 12, for products with product pictures, while collecting the product pictures, use a general md5 calculation method to obtain the md5 value of the product picture, and use this md5 value as the name of the product picture, store it in the Seaweed database according to the set storage location, and generate the corresponding storage address string;

[0049] Step 13, classify the products, add first- and second-level product tags, and store the first- and second-level product tags together with the collected product text, the md5 value of the product picture, and the storage address string in the ES database.

[0050] Step 20: Calculate the perceptual hashing fingerprint value of the product image and the Word2Vec sentence vector value of the product text based on the dark web product data and product classification results collected in Step 1. This step mainly calculates the perceptual hashing fingerprint value of the product image and the Word2Vec sentence vector value of the product text respectively based on the product text including product ID, product name and product description, product image, md5 value of the product image and product classification results collected in Step 1. Finally, these feature values together with the basic information of the product are stored in the MySQL product feature vector table. As Figure 3 shown, Step 20 includes the following steps:

[0051] Step 21: Read the product text including product ID, product name, product description and the md5 value of the product image from the ES database. For products with non-empty md5 values, obtain the product images from the Seaweed database according to the corresponding storage address string;

[0052] Step 22: Combine the product ID, product name and product description as the complete product text;

[0053] Step 23: Calculate the perceptual hashing fingerprint value of the product image; The method for calculating the perceptual hashing fingerprint value is a prior art and will not be elaborated here.

[0054] Step 24: Calculate the Word2Vec sentence vector value of the product text; The method for calculating the Word2Vec sentence vector value is a prior art and will not be elaborated here.

[0055] Step 25: Store the data obtained in Steps 21, 23 and 24 into the MySQL product feature vector table together.

[0056] Step 30: Calculate the product similarity based on the md5 value of the product image, the perceptual hashing fingerprint value of the product image and the Word2Vec sentence vector value of the product text. This step mainly calculates the product similarity using the similarity algorithms based on Hamming distance and cosine similarity respectively based on the data obtained in Step 20, such as the md5 value of the product image, the perceptual hashing fingerprint value of the product image and the Word2Vec sentence vector value of the product text, and store the product IDs and similarities of two products with similarities greater than the pre-set similarity threshold in the MySQL product similarity table. As Figure 4 shown, Step 30 includes the following sub-steps:

[0057] Step 31, read the product ID of the new product (i.e., the newly collected product), denoted as id1, read the MD5 value of the product picture, denoted as md5_1, read the perceptual hash fingerprint value of the product picture of the new product, denoted as h1, read the Word2Vec sentence vector value of the product text of the new product, denoted as v1, and read the first- and second-level product tags of the new product, denoted as c, t;

[0058] Step 32, read the attribute values of the products with the same first- and second-level product tags as the new product from the MySQL product feature vector table, which also includes the product ID, the MD5 value of the product picture, the perceptual hash fingerprint value of the product picture, and the Word2Vec sentence vector value of the product text:

[0059] Step 33, for the products corresponding to the attribute values obtained in Step 32, sequentially select a product that has not participated in the comparison yet, and set its product ID as id2, the MD5 value of the product picture as md5_2, the perceptual hash fingerprint value of the product picture as h2, and the Word2Vec sentence vector value of the product text as v2;

[0060] Step 34, if md5_1 == md5_2, then set the similarity s of the two products to 1;

[0061] Step 35, if the condition in Step 34 is not satisfied, then:

[0062] (1) Calculate the Hamming distance d = hamming_dist(h1, h2) of the perceptual hash fingerprint values of the product pictures of the two products;

[0063] (2) Calculate the cosine similarity c = cos_similarity(v1, v2) of the Word2Vec sentence vector values of the product texts of the two products;

[0064] (3) Set the similarity of the two products as s = (1 / ln(e + d / 10) + c) / 2, where e is the natural logarithm;

[0065] Step 36, compare the similarity s of the two products solved in Step 34 or 35 with the pre-set similarity threshold lambda, and screen out the two products with similarity s ≥ lambda;

[0066] Step 37, combine the product IDs and similarity s of the two products that meet the conditions in Step 36 into a triple (id1, id2, s) and store it in the MySQL product similarity table;

[0067] Step 38, return to Step 33 until all products have been compared.

[0068] So far, the determination of similar dark web products has been completed through the above-mentioned method for determining similar dark web products based on multi-modal fusion features. In this method:

[0069] (1) By constructing a similarity calculation method that fuses product images and product texts, the similarity calculation of products in the dark web market can be realized, similar products in various categories can be obtained, which is beneficial to better classify the products in the dark web market, can improve the accuracy of judging product similarity, has a simple method and strong interpretability, thus solving the problem that it is difficult to judge similar products due to the blurred product images and simple text information in the dark web.

[0070] (2) Through dark web data collection, product image feature calculation, product text feature calculation, and similarity calculation, the buying and selling of new products can be effectively monitored, real-time warnings can be issued, and the dynamics of the dark web market can be better tracked.

[0071] In addition, in some embodiments, a computer terminal storage medium is proposed, which stores computer terminal executable instructions, and is characterized in that the computer terminal executable instructions are used to execute the method for determining similar dark web products based on multi-modal fusion features as described in the previous embodiments. Examples of computer storage media include magnetic storage media (such as floppy disks, hard disks, etc.), optical recording media (such as CD-ROMs, DVDs, etc.) or memories, such as memory cards, ROMs or RAMs, etc. The computer storage medium can also be distributed on computer systems connected by a network, such as an application store.

[0072] In addition, in some embodiments, a computing device is proposed, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for determining similar dark web products based on multi-modal fusion features as described in the previous embodiments. Examples of computing devices include PCs, tablet computers, smart phones or PDAs, etc.

[0073] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for determining similar dark web products based on multi-modal fusion features, characterized in that, it includes the following steps: Step 10, collect dark web product data and classify the products; among them, the collected dark web product data includes product text and product pictures, and generate md5 values for the collected product pictures; Step 20, based on the dark web product data and product classification results collected in Step 1, calculate the perceptual hash fingerprint value of the product picture and the Word2Vec sentence vector value of the product text; Step 30, calculate the product similarity based on the md5 value of the product picture, the perceptual hash fingerprint value of the product picture, and the Word2Vec sentence vector value of the product text; Step 30 includes the following sub-steps: Step 31, read the product id of the new product, denoted as id1, read the md5 value of the product picture of the new product, denoted as md5_1, read the perceptual hash fingerprint value of the product picture of the new product, denoted as h1, read the Word2Vec sentence vector value of the product text of the new product, denoted as v1, and read the primary and secondary product labels of the new product, denoted as c, t; Step 32, read the attribute values of the products with the same primary and secondary product labels as the new product from the MySQL product feature vector table, which also includes the product id, the md5 value of the product picture, the perceptual hash fingerprint value of the product picture, and the Word2Vec sentence vector value of the product text: Step 33, among the products corresponding to the attribute values of each product obtained in Step 32, sequentially select a product that has not participated in the comparison, and set its product id as id2, the md5 value of the product picture as md5_2, the perceptual hash fingerprint value of the product picture as h2, and the Word2Vec sentence vector value of the product text as v2; Step 34, if md5_1 == md5_2, then set the similarity s of the two products to 1; Step 35, if the condition in Step 34 is not satisfied, then: (1) Calculate the Hamming distance d = hamming_dist(h1, h2) of the perceptual hash fingerprint values of the product pictures of the two products; (2) Calculate the cosine similarity c = cos_similarity(v1, v2) of the Word2Vec sentence vector values of the product texts of the two products; (3) Set the similarity of the two products as s = (1 / ln(e + d / 10) + c) / 2, where e is the natural logarithm; Step 36, compare the similarity s of the two products solved in Step 34 or 35 with the pre-set similarity threshold lambda, and screen out the two products with similarity s ≥ lambda; Step 37, combine the product ids and similarity s of the two products that meet the conditions in Step 36 into a triple (id1, id2, s) and store it in the MySQL product similarity table; Step 38, return to Step 33 until all products have been compared.

2. The method for determining similar dark web products based on multi-modal fusion features according to claim 1, characterized in that, Step 10 includes the following sub-steps: Step 11: Implement the collection of dark web commodity data. The collected dark web commodity data includes structured commodity text and commodity pictures on the commodity detail page. Among them, the commodity text includes commodity ID, commodity name, and commodity description. Step 12: For commodities with commodity pictures, while collecting the commodity pictures, use a common md5 calculation method to obtain the md5 value of the commodity pictures, and use this md5 value as the name of the commodity pictures. Store them in the Seaweed database according to the set storage location and generate the corresponding storage address string. Step 13: Classify the commodities, add first- and second-level commodity labels, and store the first- and second-level commodity labels together with the collected commodity text, the md5 value of the commodity pictures, and the storage address string in the ES database.

3. The method for determining similar dark web commodities based on multi-modal fusion features according to claim 2, wherein, Step 20 includes the following sub-steps: Step 21: Read from the ES database the commodity text including commodity ID, commodity name, and commodity description, and the md5 value of the commodity pictures. For commodities with non-empty md5 values, obtain the commodity pictures from the Seaweed database according to the corresponding storage address string. Step 22: Combine the commodity ID, commodity name, and commodity description as the complete commodity text. Step 23: Calculate the perceptual hash fingerprint value of the commodity pictures. Step 24: Calculate the Word2Vec sentence vector value of the commodity text. Step 25: Store the data obtained in Steps 21, 23, and 24 in the MySQL commodity feature vector table.

4. A computer terminal storage medium storing computer terminal executable instructions, wherein, the computer terminal executable instructions are used to execute the method for determining similar dark web commodities based on multi-modal fusion features according to any one of claims 1-3.

5. A computing device, wherein, it includes: at least one processor; and a memory communicatively connected to the at least one processor. Among them, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for determining similar dark web commodities based on multi-modal fusion features according to any one of claims 1-3.

Citation Information

Patent Citations

  • Chinese herbal medicine plant picture capturing method based on professional term vector of traditional Chinese medicine and pharmacy field

    CN105824904A

  • Similar information processing method and device

    CN112084448A