A data intelligent recognition, distribution and execution method and system
Through the combination of acquisition, segmentation, identification and distribution modules, the problem of time-consuming and low accuracy of data identification system in local distribution is solved, and fast and accurate data distribution is achieved, multi-point deployment and batch URL data acquisition is supported, and system efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202211045149.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-08-30
AI Technical Summary
The existing data identification system takes a long time to distribute the local area and has low accuracy, so it cannot automatically transmit the local area.
The data submitted by the user is collected and segmented by the evidence collection module, and the local identification module is used for identification. The data distribution module is divided according to the area and adaptively opened up cache space for distribution. Combined with the cache management unit, the database management unit and the registration query unit, the accuracy and efficiency of local identification are improved.
It realizes fast and accurate data distribution, reduces workload, improves system efficiency, supports multi-point deployment and batch URL data collection, avoids resource waste and delays, and ensures the accuracy and efficiency of data distribution.
Smart Images

Figure CN115396415B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of data processing and artificial intelligence, and particularly to a data intelligent recognition, distribution and execution method and system. Background Art
[0002] Due to the openness of the network, online public opinion can form rapidly and have a huge impact on society. Especially when negative online public opinion appears, if not controlled in time, it is very easy to form an opinion crisis, and in severe cases, it even affects public safety. For relevant departments, how to control negative content in time and effectively guide it has become a major difficulty in online public opinion management. In this case, it is very necessary to build a system that can quickly distribute public opinion data.
[0003] Currently, in traditional business systems, for the distribution of public opinion data, it is necessary for salespersons to manually judge which jurisdiction to distribute this piece of data to according to the URL of the public opinion data. This method takes too much time and has a low accuracy rate.
[0004] Therefore, there is a need for a data intelligent recognition, distribution and execution method and system that can automatically perform jurisdiction distribution, is convenient and fast, and has a high accuracy rate. Summary of the Invention
[0005] The purpose of the present invention is to solve the defects of the existing data recognition system, such as long time-consuming for distributing jurisdictions, inability to automatically distribute to jurisdictions, and low distribution accuracy rate, and provide a data intelligent recognition, distribution and execution method and system that can automatically perform jurisdiction distribution, is convenient and fast, and has a high accuracy rate.
[0006] A data intelligent recognition, distribution and execution method according to the present invention includes the following steps:
[0007] S1. Use a collection and evidence-taking module to collect data submitted by users;
[0008] S2. Segment the collected data to obtain a segmented matrix;
[0009] S3. Use a jurisdiction recognition module to recognize the segmented matrix;
[0010] S4. Divide the recognition result by region through a data distribution module to obtain a recognition result matrix;
[0011] S5. Adaptively allocate cache space according to the number of non-zero elements in each column of the recognition result matrix, and distribute the data to the receiving area management module.
[0012] Furthermore: The collection and evidence-taking module includes a monitoring unit, a collection unit, an extraction unit, a screenshot unit and a download unit; in S1, it specifically includes the following steps:
[0013] S11. Perform multi-process acquisition and evidence collection on the data submitted by the user through the acquisition unit, where the data is URL data; use the monitoring unit to monitor the acquisition process in real time;
[0014] S12. Use the screenshot unit to take screenshots of the URL data page;
[0015] S13. During the acquisition process, use the extraction unit to extract the URL data submitted by the user in real time, and at the same time use the download unit to download the extracted data.
[0016] Further: In S1, the acquisition unit, the screenshot unit, and the extraction unit all adopt the restful service method.
[0017] Further: The territorial identification module includes a domain name extraction unit, a policy management unit, and a territorial identification unit; in S3, it specifically includes the following steps:
[0018] S31. Transmit the segmented matrix to the domain name extraction unit;
[0019] S32. Extract valid data from the elements in the segmented matrix, and use the domain name extraction unit to extract the data in the order from fine to coarse granularity, and put the extracted domain names into the domain name pool;
[0020] S33. Statistically analyze the domain names in the domain name pool, set a threshold through the policy management unit, if the total amount of data uploaded by the user reaches the threshold, batch call the territorial identification unit, if not, call the territorial identification unit individually;
[0021] S34. The territorial identification unit obtains the corresponding territorial information according to the extracted domain names.
[0022] Further: The domain name extraction unit includes a cache management unit, a database management unit, a record-filing location query unit, and a regional display unit; in S32, it specifically includes the following steps:
[0023] S321. Call the cache management unit to identify whether there is regional information corresponding to the data in the current cache library. If no regional information is found, call the database management unit to identify whether there is regional information corresponding to the data in the current database; if still not found, call the record-filing location query unit to identify the regional information corresponding to the data through the record-filing location query website, and newly create the queried regional information in the database and the cache library;
[0024] S322. Send the identified regional information and the regional information manually added by the user to the regional display unit for display, and the regional information includes domain names, territorial areas, and / or websites.
[0025] Furthermore, the data distribution module includes a data distribution unit, an anomaly detection unit, and a duplication detection unit. In S4, the following steps are specifically included:
[0026] S41. The distribution unit distributes the recognition result.
[0027] S42. The anomaly detection unit detects the distribution process in real time. If it detects that the territorial information is empty or the territory is not in the receiving location list, it marks the recognition result and prevents the distribution unit from distributing.
[0028] S43. The duplication detection unit puts the already distributed recognition results into the distribution pool in real time, and compares the next recognition result to be distributed with the distribution pool. If the distribution pool already contains this recognition result, it prevents the distribution unit from distributing.
[0029] Furthermore, in S5, a reshaped matrix is obtained based on the differences of the vector elements in the columns of the recognition result matrix. According to the number of non-zero elements in each column of the reshaped matrix, the receiving location management module adaptively allocates cache space for data distribution. At the same time, the receiving location list in the receiving location management module is called to compare with the identified territory to ensure that the name of the receiving location is consistent with the territory, and then data distribution to the receiving location is allowed.
[0030] A data intelligent recognition and distribution system according to the present invention includes a collection and evidence-taking module, a territorial recognition module, a data distribution module, a domain name extraction unit, and a receiving location management module. The output end of the collection and evidence-taking module is communicatively connected to the input end of the territorial recognition module. The output end of the territorial recognition module is communicatively connected to the input end of the data distribution module. The output end of the domain name extraction unit is communicatively connected to the input end of the territorial recognition module. The output end of the receiving location management module is communicatively connected to the input end of the data distribution module.
[0031] The collection and evidence-taking module is used to collect and take evidence of the data submitted by the user.
[0032] The territorial recognition module is used to segment and recognize the collected data.
[0033] The data distribution module is used to distribute the recognized data to the corresponding receiving locations.
[0034] The domain name extraction unit is used to provide the territorial information of the domain name and provide data support for the data recognition of the territorial recognition module.
[0035] The receiving location management module is used to manage all receiving locations and provide data support according to the distribution tasks of the data distribution module.
[0036] Furthermore, the acquisition and evidence collection module includes a monitoring unit, a collection unit, an extraction unit, a screenshot unit, and a download unit; the collection unit is used to collect the data submitted by the user, the monitoring unit is used to monitor the collection process of the collection unit, the extraction unit is used to extract the data submitted by the user, the screenshot unit is used to take screenshots of the data page, and the download unit is used to download the extracted data;
[0037] The territorial identification module includes a domain name extraction unit, a policy management unit, and a territorial identification unit; the domain name extraction unit is used to extract the data and put the extracted domain name into the domain name pool, the policy management unit is used to set a threshold and call the territorial identification unit according to the comparison result between the total data volume of the current domain name and the threshold according to a preset rule, and the territorial identification unit is used to obtain the territorial information of the enemy camp according to the extracted domain name;
[0038] The domain name extraction unit includes a cache management unit, a database management unit, a filing location query unit, and a region display unit; the cache management unit is used to identify whether there is territorial information corresponding to the data in the current cache library, the database management unit is used to identify whether there is territorial information corresponding to the data in the current database, and the filing location query unit is used to identify the territorial information corresponding to the data through the filing location query website and create the queried territorial information in the database and the cache library;
[0039] The data distribution module includes a data distribution unit, an anomaly detection unit, and a duplication detection unit; the distribution unit is used to distribute the recognition result, the anomaly detection unit is used to detect the distribution process, and the duplication detection unit is used to compare the recognition result to be distributed with the distribution pool.
[0040] The beneficial effects of the present invention are:
[0041] The present invention can effectively solve the technical problems of URL (Uniform Resource Locator) data distribution being too time-consuming and having low accuracy, and through a series of effect investigations, by introducing a spatial modulation algorithm in a monitoring unit and extracting and detecting some threads, the workload is reduced, the running time is reduced, and the efficiency of the system is further improved. At the same time, multi-process operation is adopted to support multi-point deployment and batch URL data collection and evidence collection; by introducing a domain name extraction unit including a cache management unit, a database management unit, and a filing location query unit for calling by a territory identification module, the accuracy and efficiency of territory identification are improved, and delays caused by frequent access to the database and the filing location query website are avoided; when the territory identification module is running, the cache management unit, the database management unit, and the filing location query unit are called in sequence to improve the efficiency of territory identification, and delays caused by frequent access to the database and the filing location query website are avoided. Identifying the territory according to the granularity from fine to coarse can accurately locate the territory corresponding to the domain name, and solve the problem of different territories issued by different websites of large manufacturers; the present invention obtains a territory identification matrix by dividing the batch identification results into regions, and opens up cache space according to the number of non-zero elements in each column of the territory identification matrix, reasonably allocates cache resources, avoids resource waste, and further improves the system distribution efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 Provide a block diagram of the data intelligent identification and distribution system;
[0043] Figure 2 Provide a flow chart for territory identification;
[0044] Figure 3 Flowchart for data distribution. DETAILED DESCRIPTION
[0045] The following are only preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a technician familiar with the technical field within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. The embodiments described below are only used to explain the present invention and cannot be interpreted as limitations on the present invention. The protection scope of the present invention should be based on the protection scope of the claims. The embodiments of the present invention are described in detail below. In order to facilitate the description of the present invention and simplify the description, the technical terms used in the specification of the present invention should be interpreted in a broad sense, including but not limited to conventional replacement schemes not mentioned in this application, and also including direct implementation methods and indirect implementation methods.
[0046] Example 1
[0047] Combination Figure 1 This embodiment describes a method for intelligently identifying and distributing data, including the following steps:
[0048] S1. Use the collection and evidence collection module to collect the data submitted by the user;
[0049] S2. Segment the collected data to obtain a segmentation matrix; segment the domain name of the URL to prepare for further identification; segment the collected URL to obtain a segmentation matrix, transmit the segmentation matrix to the domain name extraction unit, extract valid URL data from the elements in the segmentation matrix, extract the data in the order of fine to coarse granularity according to the URL protocol, and put the extracted domain name into a domain name pool, the domain name pool is composed of domain names, and the domain names in the domain name pool are counted to obtain a statistic C, and the policy management unit sets a threshold , if the statistics C of the URL data uploaded by the user reaches the threshold, the location identification unit is called in batches; if it does not reach the threshold, the location identification unit is called individually;
[0050] S3. Use the territory identification module to identify the segmented matrix; identify the URL data encoded by the user according to the fine and coarse criteria; the territory identification unit is responsible for obtaining the corresponding territory and website according to the domain name, and first calls the cache management unit of the domain name extraction unit; if the cache management unit does not find the data, it calls the database management unit; if the database management unit also does not find the data, it calls the filing location query unit, and further, the domain name identified according to the URL submitted by the user and the domain name manually added by the user are displayed by the domain name display unit, and the displayed domain name information includes the domain name, territory, website and other information;
[0051] S4, dividing the recognition results by region through the data distribution module to obtain a recognition result matrix; transmitting the local recognition results to the data distribution unit in the data distribution module, and dividing the batch recognition results by region in the distribution unit to obtain a recognition result matrix address;
[0052] S5, adaptively open up cache space according to the number of non-zero elements in each column of the recognition result matrix, and distribute the data to the receiving location management module. The recognition result is divided by region through the data distribution module to obtain the recognition result matrix, and adaptively open up cache space according to the number of non-zero elements in each column of the recognition result matrix to distribute the URL data and finally distribute it to the receiving location. Adaptively open up cache space according to the number of non-zero elements in each column of the matrix address to distribute the URL data.
[0053] Example 2
[0054] This embodiment is described in conjunction with Example 1. This embodiment discloses a method for intelligent data identification and distribution execution. The acquisition and evidence collection module includes a monitoring unit, an acquisition unit, an extraction unit, a screenshot unit, and a download unit. In S1, the following steps are specifically included:
[0055] S11. Use the collection unit to collect and obtain evidence of the data submitted by the user in multiple processes. The data is URL data. Use the monitoring unit to monitor the collection process in real time. Use the collection unit and the screenshot unit in the collection unit to collect the data submitted by the user. In this embodiment, Haproxy (a free and open-source software written in C language) is used to implement multi-process and multi-machine deployment, which is different from the traditional Slimer JS component (Slimer JS is a server-side JavaScript API tool) deployment that can only perform single-process deployment. Haproxy is used in the system to implement multi-process and multi-machine deployment, improving the concurrency of collection and enabling the collection and evidence obtaining of more URL data simultaneously.
[0056] S12. Use the screenshot unit to take screenshots of the URL data page. Use the screenshot unit (the existing SlimerJS component technology) to collect and take screenshots of the URL page. The collection unit and the screenshot unit consume a large amount of system resources when collecting URLs, and the program may terminate unexpectedly. In this embodiment, a monitoring program is developed to set the threshold number of active processes. When it is detected that the number of processes is , automatically restart all program nodes.
[0057] S13. During the collection process, use the extraction unit to extract the URL data submitted by the user in real time, and at the same time use the download unit to download the extracted data. The collection unit, the screenshot unit, and the extraction unit all adopt the restful service (RESTFUL is a design style and development method of network application programs) method, support JPG and PNG format screenshots, and support JavaScript (abbreviation "JS", a lightweight, interpreted or just-in-time compiled programming language with function priority) parsing. The extraction unit adopts the restful service method, starts multiple service processes, extracts the title, author, publication time, source, and content of the web page, and at the same time provides services externally by the unified httpd (the main program of the Apache Hypertext Transfer Protocol server). When the service process terminates unexpectedly, it will automatically restart; use the download unit to perform one-key download on the URL-related content (collection content and screenshots) processed by the extraction unit. The download is in the form of a compressed package, and the downloaded content includes an Excel table and screenshot pictures. The Excel table contains the title, text, publication date, etc. of the URL. The generation of the compressed package consumes a lot of CPU resources, so the system limits that a user can only download one at a time, and the compressed space of a single server cannot exceed 30G.
[0058] Embodiment 3
[0059] This embodiment is described in combination with Embodiment 1. A data intelligent recognition, distribution and execution method disclosed in this embodiment. In S1, a monitoring unit is used to monitor the acquisition process in real time; the acquisition unit, the screenshot unit and the extraction unit all adopt the restful service (RESTFUL is a design style and development method of network applications, based on HTTP, and can be defined in XML format or JSON format). At the same time, a monitoring program introducing the spatial modulation algorithm is developed to monitor the acquisition process; some threads are extracted and detected, reducing the workload, shortening the running time, further improving the system efficiency, and at the same time adopting multi-process operation, which supports multi-point deployment and the acquisition and evidence collection of batch URLs. The running speed and accuracy of the monitoring program are effectively improved, and the activity of the process is monitored more accurately. The specific process of calculating the thread activity is as follows:
[0060] Encode the created processes, and denote all process sets as , , where represents the total number of processes, represents the th process, . When the monitoring program runs, the spatial modulation algorithm is used to selectively detect the processes in the program. The specific detection process is as follows:
[0061] For any processes, perform times of status detection, where , represents rounding, and denote the set of encoded active processes detected as ;
[0062] First time: The set of encoded active processes detected is: , represents the number of elements in the first detection set. Let , The corresponding number of detected active processes is , where represents the number of elements in the set;
[0063] Second time: The set of encoded active processes detected is: , represents the number of elements in the second detection set. Let , , and the corresponding number of detected active processes is . If , then re-select processes for detection, where is obtained according to multiple experimental simulations, Represents the union of two sets;
[0064] Third time: The detected active process code set is: , Represents the number of elements in the third detection set. Let , , and the corresponding number of detected active processes is . If , then reselect processes for detection;
[0065]
[0066] The th time: The detected active process code set is: , Represents the th detection set's number of elements. Let , , and the corresponding number of detected active processes is . If , then reselect processes for detection, where ;
[0067]
[0068] The th time: The detected active process code set is: , Represents the th detection set's number of elements. Let , , and the corresponding number of detected active processes is . If , then reselect processes for detection;
[0069] In particular, during detection, when , when reselecting M processes for detection, if the number of reselecting times is greater than , then automatically restart all program nodes, where is obtained based on multiple experiment simulations.
[0070] Preferably, according to the detection randomness of the status of any processes in the detection program, a certain degree of error in the number of active processes is allowed. Denote the error as , The value is obtained from multiple experiment simulations.
[0071] The number of detected active processes calculated in this embodiment is:
[0072]
[0073] Finally, for the number of detected active processes a determination is made, and the determination threshold for the number of active processes is set to , , in this embodiment, the determination threshold for the number of active processes is defined as:
[0074]
[0075]
[0076] wherein, represents taking the remainder. If the number of detected active processes when, then the monitoring program continues to execute. If the number of detected active processes when, all program nodes are automatically restarted. If the number of detected active processes when, the monitoring is repeated times to obtain value still falls within the interval then all program nodes are automatically restarted, obtained based on multiple experimental simulations.
[0077] Define the number of detected active processes , set the determination threshold for the number of active processes , and make a determination on the number of detected active processes . If the number of detected active processes when, then the monitoring program continues to execute. If the number of detected active processes when, all program nodes are automatically restarted. If the number of detected active processes when, the monitoring is repeated times to obtain value still falls within the interval then all program nodes are automatically restarted, obtained based on multiple experimental simulations.
[0078] Embodiment 4
[0079] Combined with Figure 2 and Embodiment 1 to illustrate this embodiment. A data intelligent recognition and distribution execution method disclosed in this embodiment, the territorial recognition module includes a domain name extraction unit, a policy management unit, and a territorial recognition unit; in S3, it specifically includes the following steps:
[0080] S31. Transmit the segmented matrix to the domain name extraction unit; the specific recognition process is as follows:
[0081] First, the URL submitted by the user is segmented using Java according to the URL syntax and protocol, and Data is recorded as the set of all submitted URLs. , Indicates the number of submitted URLs. Indicates URL data, , for each Do segmentation processing, , Indicates URL Segments, ; You can get the segment matrix representation of the URL submitted by the user:
[0082]
[0083] S32, extracting valid data from the elements in the segmented matrix, using the domain name extraction unit to extract the data in order of granularity from fine to coarse, and putting the extracted domain names into the domain name pool; Transmitted to the domain name extraction unit, the matrix The element extracts valid URL data, extracts the data in order of granularity from fine to coarse according to the URL protocol, and puts the extracted domain name into a domain name pool, the domain name pool is composed of domain names, and the domain names in the domain name pool are counted to obtain a statistic C;
[0084] S33, collect statistics on the domain names in the domain name pool, set a threshold through the policy management unit, if the total amount of data uploaded by the user reaches the threshold, call the location identification unit in batches, if it does not reach the threshold, call the location identification unit individually; the policy management unit sets the threshold , if the amount C of URLs uploaded by the user reaches a threshold, the location identification unit is called in batches; if the threshold is not reached, the location identification unit is called individually;
[0085] S34. The territorial identification unit obtains the corresponding territorial information according to the extracted domain name. The territorial identification unit is responsible for obtaining the corresponding territory and website according to the domain name. Taking the following URL data as an example, the user inputs the URL: https: / / developers.weixin.qq.com / doc / oplatform / service_market / intro.html. The domain name is extracted as developers.weixin.qq.com, weixin.qq.com, qq.com in the order of decreasing granularity, and then sent to the territorial identification module for territorial identification. When identifying, the cache management unit of the domain name extraction unit is first called. If the corresponding domain name data is found in the cache management unit, the identification is successful, and the identification result is region X. The corresponding URL is sent to region X.
[0086] Example 5
[0087] Combined with Figure 3 and Example 4 to illustrate this example. A data intelligent identification and distribution execution method disclosed in this example, the domain name extraction unit includes a cache management unit, a database management unit, a record-filing place query unit, and a region display unit; in S32, it specifically includes the following steps:
[0088] S321, calling the cache management unit to identify whether there is regional information corresponding to the data in the current cache library. If no regional information is found, calling the database management unit to identify whether there is regional information corresponding to the data in the current database; if still not found, calling the registration location query unit to identify the regional information corresponding to the data through the registration location query website, and newly adding the queried regional information to the database and the cache library; first calling the cache management unit of the domain name extraction unit, the cache management unit is responsible for constructing the domain name pool in the cache from scratch, and providing an interface for adding, deleting, modifying and querying the domain name in the cache management unit; if the cache management unit does not find the data, calling the database management unit, the database management unit is responsible for constructing the domain name pool in the database from scratch The structure provides an interface for adding, deleting, modifying and checking domain names in the database management unit, and is responsible for calling the adding, deleting, modifying and checking interface of the cache management unit to synchronously update the cached data; if the database management unit also does not find the data, the registration site query unit is called, and the registration site query unit is responsible for remotely calling the registration site query website for domain names that are not in the cache management unit and the database management unit, and at the same time calling the adding, deleting, modifying and checking interface in the database management unit and the adding, deleting, modifying and checking interface in the cache management unit to update the domain name information in the database management unit and the cache management unit; the location identification module in this embodiment sequentially calls the cache management unit, the database management unit, and the registration site query unit during operation in order to improve the efficiency of location identification and avoid delays caused by frequent access to the database and the registration site query website. The location identification according to the granularity from fine to coarse can accurately locate the location corresponding to the domain name, solving the problem of different locations issued by different websites of large manufacturers.
[0089] Preferably, when the database management unit is called and a hit occurs, the domain name information needs to be written into the cache, and when the record location query unit is called and a hit occurs, the domain name information needs to be written into the cache and the database respectively, so that the next time the information of the same domain name is queried, it can be directly hit from the cache. If the cache, database, and record location query website all fail to hit the domain name information, the local website of the domain name is set to empty, and the user can manually edit the local website information.
[0090] Preferably, if the domain name automatically identified by the system is incorrect, the user can modify the domain name. After the user adds, deletes, or modifies the domain name, the domain name information in the cache and database will be modified synchronously.
[0091] S322: Send the identified region information and the region information manually added by the user to the region display unit for display, wherein the region information includes domain name, location and / or website. The domain name automatically identified by the system based on the URL submitted by the user and the domain name manually added by the user are displayed by the domain name display unit, and the displayed domain name information includes domain name, location, website and other information.
[0092] Example 6
[0093] This example is described in combination with Example 1. A data intelligent recognition distribution execution method disclosed in this example, the data distribution module includes a data distribution unit, an anomaly detection unit, and a duplication detection unit; in S4, it specifically includes the following steps:
[0094] S41. The distribution unit distributes the recognition result; the territorial recognition result is transmitted to the data distribution unit in the data distribution module, and the batch recognition result is divided into regions in the distribution unit to obtain the recognition result matrix address,
[0095]
[0096] where X represents the total number of territories, and Y represents the number of URLs, represents the y-th URL of the x-th territory, ;
[0097] Specifically, when the number of territorial URL recognition results is less than Y, the row vector corresponding to the territory contains 0 elements, that is, if the x1-th territory has y1 URLs, and , then the value of the x1-th column corresponding to the matrix address is ;
[0098] The cache space is adaptively allocated for URL data distribution according to the number of non-zero elements in each column of the matrix address; at the same time, the receiving area management module is called to compare with the recognized territory to ensure that the name of the receiving area is consistent with the territory. The recognized territory of the URL must be in the receiving area list to allow data distribution to determine that each piece of data has a corresponding organization to receive. In this example, the batch recognition result is divided into regions to obtain the territorial recognition matrix, and the cache space is allocated according to the number of non-zero elements in each column of the territorial recognition matrix, rationally allocating cache resources, avoiding resource waste, and further improving the system distribution efficiency.
[0099] S42. The anomaly detection unit continuously monitors the distribution process. If the territorial information is found to be empty or the territory is not in the receiving location list, it marks the recognition result and prevents the distribution unit from distributing. During the distribution process, it is detected by the anomaly monitoring unit. The anomaly detection unit mainly detects the territorial characteristics of the URL. If the territory of the URL is found to be empty or the territory is not in the receiving location list, the URL is highlighted in red and cannot be distributed. Specifically, the user must manually edit the territory of the URL to ensure that the territory is in the receiving location list before data can be distributed. Preferably, after the user edits the territory of the URL, the cache management unit and the database management unit in the domain name extraction unit are synchronously updated. Further, when the user uploads a URL with the same domain name next time, it will be recognized as the territory edited by the user.
[0100] S43. The repeatability detection unit continuously puts the distributed recognition results into the distribution pool and compares the next recognition result to be distributed with the distribution pool. If the distribution pool already contains this recognition result, it prevents the distribution unit from distributing. Preferably, the repeatability detection unit puts the MD5 values generated from the URLs already distributed by the user into the distribution pool and compares the MD5 value generated from the URL to be distributed by the user with the distribution pool to prompt the user whether the URL has been distributed before.
[0101] Example 7
[0102] Combined with Example 1 to illustrate this example. A data intelligent recognition and distribution execution method disclosed in this example. In S5, a reshaped matrix is obtained based on the differences of the vector elements in the columns of the recognition result matrix. According to the number of non-zero elements in each column of the reshaped matrix, the receiving location management module adaptively allocates cache space for data distribution. At the same time, the receiving location list in the receiving location management module is called to compare with the recognized territory to ensure that the name of the receiving location is consistent with the territory, and then data is allowed to be distributed to the receiving location.
[0103] The recognition results are divided by region through the data distribution module to obtain a recognition result matrix. A reshaped matrix is obtained based on the differences of the vector elements in the columns of the recognition result matrix. Cache space is adaptively allocated for URL data distribution according to the number of non-zero elements in each column of the reshaped matrix and finally distributed to the receiving location; Cache space is adaptively allocated for URL data distribution according to the number of non-zero elements in each column of the recognition result matrix and finally distributed to the receiving location;
[0104] Example 8
[0105] This embodiment is described in combination with Embodiment 1. A data intelligent recognition and distribution system disclosed in this embodiment includes a collection and evidence-taking module, a territorial identification module, a data distribution module, a domain name extraction unit, and a receiving location management module. The output end of the collection and evidence-taking module is communicatively connected to the input end of the territorial identification module. The output end of the territorial identification module is communicatively connected to the input end of the data distribution module. The output end of the domain name extraction unit is communicatively connected to the input end of the territorial identification module. The output end of the receiving location management module is communicatively connected to the input end of the data distribution module;
[0106] The collection and evidence-taking module is used to collect and take evidence on the data submitted by users; it provides the function of collecting and taking evidence for URLs, facilitating subsequent viewing of the original content of URLs;
[0107] The territorial identification module is used to segment and identify the collected data; it provides the function of identifying the territory and website of the domain name of the URL; it distributes to the corresponding receiving locations according to the territory identified by the territorial identification module;
[0108] The data distribution module is used to distribute the identified data to the corresponding receiving locations; it distributes to the corresponding receiving locations according to the territory identified by the territorial identification module;
[0109] The domain name extraction unit is used to provide the territorial information of the domain name and provide data support for the data identification of the territorial identification module; it manages the territorial and website information of the domain name and provides data support for the territorial identification module;
[0110] The receiving location management module is used to manage all receiving locations and provide data support according to the distribution tasks of the data distribution module. It manages all receiving locations and provides data support for data distribution.
[0111] Embodiment 9
[0112] This embodiment is described in combination with Embodiment 1. A data intelligent recognition and distribution system disclosed in this embodiment, the collection and evidence-taking module includes a collection unit, an extraction unit, a screenshot unit, and a download unit; the collection unit is used to collect the data submitted by users, the extraction unit is used to extract the data submitted by users, the screenshot unit is used to take screenshots of the data page, and the download unit is used to download the extracted data;
[0113] The territorial identification module includes a domain name extraction unit, a policy management unit, and a territorial identification unit; the domain name extraction unit is used to extract data and put the extracted domain name into a domain name pool, the policy management unit is used to set a threshold and call the territorial identification unit according to a preset rule based on the comparison result between the total amount of current domain name data and the threshold, and the territorial identification unit is used to obtain the territorial information of the enemy camp according to the extracted domain name;
[0114] The domain name extraction unit includes a cache management unit, a database management unit, a filing location query unit, and a regional display unit; the cache management unit is used to identify whether there is regional information corresponding to the data in the current cache library, the database management unit is used to identify whether there is regional information corresponding to the data in the current database, the filing location query unit is used to identify the regional information corresponding to the data through a filing location query website, and create the queried regional information in the database and the cache library;
[0115] The data distribution module includes a data distribution unit, an anomaly detection unit, and a duplication detection unit; the distribution unit is used to distribute the recognition result, the anomaly detection unit is used to detect the distribution process, and the duplication detection unit is used to compare the recognition result to be distributed with the distribution pool.
Claims
1. A data intelligent recognition, distribution and execution method, characterized in that It includes the following steps: S1. Use the acquisition and evidence collection module to collect the data submitted by the user; S2. Segment the collected data to obtain a segmented matrix; S3. Use the territorial identification module to identify the segmented matrix; the territorial identification module includes a domain name extraction unit, a policy management unit, and a territorial identification unit; in S3, it specifically includes the following steps: S31. Transmit the segmented matrix to the domain name extraction unit; S32. Extract valid data from the elements in the segmented matrix, use the domain name extraction unit to extract the data in the order from fine to coarse granularity, and put the extracted domain names into the domain name pool; S33. Statistically analyze the domain names in the domain name pool, set a threshold through the policy management unit, if the total amount of data uploaded by the user reaches the threshold, batch call the territorial identification unit, if not, individually call the territorial identification unit; S34. The territorial identification unit obtains the corresponding territorial information according to the extracted domain names; S4. Divide the recognition result by region through the data distribution module to obtain a recognition result matrix; S5. Adaptively allocate cache space according to the number of non-zero elements in each column of the recognition result matrix, and distribute the data to the receiving area management module.
2. The data intelligent recognition, distribution and execution method according to claim 1, wherein The acquisition and evidence collection module includes a monitoring unit, an acquisition unit, an extraction unit, a screenshot unit, and a download unit; in S1, it specifically includes the following steps: S11. Use the acquisition unit to perform multi-process acquisition and evidence collection on the data submitted by the user, and the data is URL data; use the monitoring unit to perform real-time monitoring on the acquisition process; S12. Use the screenshot unit to take screenshots of the URL data page; S13. During the acquisition process, use the extraction unit to perform real-time extraction on the URL data submitted by the user, and at the same time use the download unit to download the extracted data.
3. The data intelligent recognition, distribution and execution method according to claim 2, characterized in that, In S1, the acquisition unit, the screenshot unit, and the extraction unit all adopt the restful service method.
4. The data intelligent recognition, distribution and execution method according to claim 1, wherein The domain name extraction unit includes a cache management unit, a database management unit, a record-filing location query unit, and a region display unit; in S32, it specifically includes the following steps: S321. Call the cache management unit to identify whether there is territorial information corresponding to the data in the current cache library. If no territorial information is found, call the database management unit to identify whether there is territorial information corresponding to the data in the current database; if still not found, call the record-filing location query unit to identify the territorial information corresponding to the data through the record-filing location query website, and create the queried territorial information in the database and the cache library; S322. Send the identified territorial information and the territorial information manually added by the user to the region display unit for display, and the territorial information includes domain names, territorial areas, and / or websites.
5. The data intelligent recognition, distribution and execution method according to claim 1, characterized in that The data distribution module includes a data distribution unit, an anomaly detection unit, and a repeatability detection unit; in S4, it specifically includes the following steps: S41. The distribution unit distributes the recognition result; S42. The anomaly detection unit detects the distribution process in real time. If the territorial information is found to be empty or the territory is not in the receiving location list, it marks the recognition result and prevents the distribution unit from distributing. S43. The duplicate detection unit puts the already distributed recognition results into the distribution pool in real time and compares the next recognition result to be distributed with the distribution pool. If the distribution pool already contains this recognition result, it prevents the distribution unit from distributing.
6. The data intelligent recognition, distribution and execution method according to claim 1, wherein In S5, a reshaped matrix is obtained based on the differences of the vector elements in the columns of the recognition result matrix. According to the number of non-zero elements in each column of the reshaped matrix, the receiving location management module adaptively allocates cache space for data distribution. At the same time, the receiving location list in the receiving location management module is called to compare with the identified territory to ensure that the name of the receiving location is consistent with the territory, and then data distribution to the receiving location is allowed.
7. A data intelligent recognition and distribution system for implementing a data intelligent recognition, distribution, and execution method according to any one of claims 1-6, characterized in that, It includes a collection and evidence-taking module, a territorial recognition module, a data distribution module, a domain name extraction unit, and a receiving location management module. The output end of the collection and evidence-taking module is communicatively connected to the input end of the territorial recognition module. The output end of the territorial recognition module is communicatively connected to the input end of the data distribution module. The output end of the domain name extraction unit is communicatively connected to the input end of the territorial recognition module. The output end of the receiving location management module is communicatively connected to the input end of the data distribution module. The collection and evidence-taking module is used to collect and take evidence of the data submitted by the user. The territorial recognition module is used to segment and recognize the collected data. The data distribution module is used to distribute the recognized data to the corresponding receiving locations. The domain name extraction unit is used to provide the territorial information of the domain name and provide data support for the data recognition of the territorial recognition module. The receiving location management module is used to manage all receiving locations and provide data support according to the distribution tasks of the data distribution module.
8. A data intelligent recognition and distribution system according to claim 7, characterized in that, The collection and evidence-taking module includes a monitoring unit, a collection unit, an extraction unit, a screenshot unit, and a download unit. The collection unit is used to collect the data submitted by the user. The monitoring unit is used to monitor the collection process of the collection unit. The extraction unit is used to extract the data submitted by the user. The screenshot unit is used to take screenshots of the data page. The download unit is used to download the extracted data. The territorial recognition module includes a domain name extraction unit, a policy management unit, and a territorial recognition unit. The domain name extraction unit is used to extract the data and put the extracted domain names into the domain name pool. The policy management unit is used to set thresholds and call the territorial recognition unit according to the comparison result between the total data volume of the current domain name and the threshold according to preset rules. The territorial recognition unit is used to obtain the territorial information of the enemy camp based on the extracted domain names. The domain name extraction unit includes a cache management unit, a database management unit, a filing location query unit, and a region display unit; the cache management unit is used to identify whether there is region information corresponding to the data in the current cache library, the database management unit is used to identify whether there is region information corresponding to the data in the current database, the filing location query unit is used to identify the region information corresponding to the data through the filing location query website, and newly create the queried region information into the database and the cache library; The data distribution module includes a data distribution unit, an anomaly detection unit, and a duplication detection unit; the distribution unit is used to distribute the recognition result, the anomaly detection unit is used to detect the distribution process, and the duplication detection unit is used to compare the recognition result to be distributed with the distribution pool.
Citation Information
Patent Citations
Website administration home location identification method based on web spider
CN107590265A