Distributed data management system and method based on artificial intelligence

Through a distributed data management system based on artificial intelligence, using feature recognition and differentiation storage technology, the problem of time-consuming and insufficient retrieval in massive network data management is solved, and efficient and accurate data management and storage is achieved, diverse needs are met, and data security and retrieval efficiency are ensured.

CN120386809AActive Publication Date: 2025-07-29SHANXI ZHONGYUN ZHIGU DATA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510470338.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-29
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

In the management of massive network data, the management accuracy of simple distinction based on data update time and data format is insufficient, resulting in a long search time and a large load on the search engine, which affects the accuracy and efficiency of data matching.

Method used

A distributed data management system based on artificial intelligence is adopted, and through upload modules, feature recognition modules, analysis modules, management modules and cloud databases, the feature recognition, cleaning, correlation analysis and distinction storage of data content are realized. The picking unit and the extraction unit are used to identify text and image features, and the semantic similarity and object name repetition rate are combined for finely distinguishing storage.

Benefits of technology

It improves the utilization rate of storage space, accurately identify data characteristics, improves the accuracy and efficiency of data management, meets the needs of diverse users, ensures the security and efficiency of data management, reduces redundant storage, and improves the accuracy and efficiency of data retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386809A_ABST
    Figure CN120386809A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed data management system and method based on artificial intelligence, and relates to the field of data processing, and the system comprises an uploading module which is used for uploading data contents needing to be stored and managed, and placing the uploaded data contents in a storage medium for temporary storage; the feature recognition module is used for traversing the data content temporarily stored in the storage medium and performing feature recognition on the data content; and the analysis module is used for acquiring the feature recognition result of the data content in the feature recognition module and analyzing the feature relevance of the data content based on the feature recognition result. According to the invention, uploaded data can be automatically cleaned, repeated items can be removed, and the storage utilization rate can be improved. Data features of different formats can be accurately identified, association between data is mined, and valuable information is provided for users. In terms of storage management, storage is finely distinguished according to data features, and the storage management precision is improved. And the user can customize the key preset value, so that diversified requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and specifically provides a distributed data management system and method based on artificial intelligence. Background Art

[0002] Data storage management is the process of effectively storing, organizing, and maintaining data. It involves selecting appropriate storage media and technologies to ensure the security, integrity, and efficient access of data, while also performing tasks such as data backup, recovery, and reasonable allocation and management of storage space.

[0003] The invention patent application with the application number 202411118702.X discloses a distributed data matching method based on artificial intelligence. The method includes: constructing a data set, collecting historical matching data in the data set, and dividing the historical matching data; obtaining feature data in the historical matching data, analyzing the association rules of each feature data in the historical matching data, and generating a data model at the same time; performing data traceability in the historical matching data, collecting the procurement or manufacturing process corresponding to the matching data, uploading the data traceability information to the data set, and updating the data set; obtaining the data with the most historical matching times in the data set, collecting the association features between the data and the procurement or manufacturing process, and performing data annotation on the association features. This application aims to solve the problem that "when the existing distributed data matching technology matches and searches for data in the production or procurement process, it is difficult to connect the data information in the upstream, midstream, and downstream of intelligent manufacturing, sometimes resulting in incomplete matching data information, thereby affecting the accuracy of data matching, reducing the data processing and matching efficiency, and causing delays in the procurement or production of intelligent manufacturing".

[0004] However, for a large amount of network data content, when the existing technology performs storage management, it mostly conducts differential management based on simple conditions such as data update time and data format. The accuracy of this differential management is limited, so that when configuring a search engine to retrieve target data content, the engine load is large and the retrieval time is long.

[0005] Therefore, a distributed data management system and method based on artificial intelligence are proposed. Summary of the Invention

[0006] In view of the above-mentioned drawbacks of the existing technology, the present invention provides a distributed data management system and method based on artificial intelligence, which can effectively solve the problems of the existing technology.

[0007] To achieve the above objectives, the present invention is realized through the following technical solutions:

[0008] The present invention discloses a distributed data management system based on artificial intelligence, including:

[0009] An upload module for uploading data content that needs to be stored and managed, and temporarily storing the uploaded data content in a storage medium; a feature recognition module for traversing the data content temporarily stored in the storage medium and performing feature recognition on the data content; an analysis module for obtaining the feature recognition results of the data content in the feature recognition module, analyzing the feature relevance of the data content based on the feature recognition results, and outputting associated data content by applying the analysis results of the feature relevance of the data content; a management module for receiving the feature recognition results of each data content in the feature recognition module, differentiating the data content based on the feature recognition results, and forwarding the data content to the cloud database based on the differentiation results; a cloud database for obtaining the data content forwarded from the management module and performing differentiated storage operations on the data content based on its differentiation results in the management module; a control module for monitoring whether new data content is uploaded in the upload module, and when the monitoring result is yes, controlling the system to refresh and run

[0010] Furthermore, the data content uploaded in the upload module is from any authorized system-side user, the storage space size of the storage medium is manually configured by any system-side user, when the data content is temporarily stored in the storage medium, it is sent to the storage medium through medium transmission, and after each new data content is temporarily stored in the storage medium, the storage medium performs a data cleaning operation through a preset built-in program;

[0011] Among them, when the data content performs a data cleaning operation in the storage medium based on a preset built-in program, it follows: differentiating the data content based on the data format, and performing a data cleaning operation once in each differentiation interval. In the data cleaning operation stage, comparing the similarity of each data content in the same differentiation interval, and taking any one data in a group of data with a similarity of 100% as the deletion target and performing a deletion operation.

[0012] Furthermore, a picking unit and an extraction unit are arranged at a lower level of the feature recognition module. The picking unit is used for data content containing text information and picks feature text information in the data content. The extraction module is used for data content in image format and extracts feature information in the data content;

[0013] Among them, the feature recognition module runs continuously. Each time the feature recognition module runs, it takes the data content stored in a differentiation interval in the storage medium as the processing target, identifies the processing target format. When the processing target format is data content containing text information, each data in the processing target is sequentially sent to the picking unit. When the processing target format is image format, each data in the processing target is sequentially sent to the extraction unit.

[0014] Furthermore, during the operation phase of the picking unit, after receiving the data content, the picking logic for the feature text information in the data content is expressed as:

[0015] Perform word segmentation on the data content to obtain a set of words, denoted as W = {w1, w2,..., w n}, and calculate the word frequency score, part-of-speech score, and position score of each word respectively;

[0016]

[0017] Total word score: S(ω i ) = S f (ω i ) × α + pos(w i ) × β + S l × γ;

[0018] In the formula: S(ω i ) is the total word score; S f (ω i ), pos(w i ), S l are the word frequency score, part-of-speech score, and position score; α, β, γ are weights; n is the total number of all words in the text; f(ω i ) is the number of times the word (w i ) appears in the text; f(ω j ) is the sum of the number of times all words in the text appear;

[0019] Among them, indicates that when the word is a noun, verb, adjective, or other part of speech, its corresponding value is 0.4, 0.3, 0.2, 0.1, indicates that when the word appears in the first 10% or the last 10% of the text or is the first word of a paragraph, the value is 0.3, and the value is 0.1 in other positions. After calculating the total score for each word, the words are sorted in descending order based on the score, and the top 5% of the words in the descending order are denoted as the feature text information.

[0020] Furthermore, during the operation phase of the picking unit, after receiving the data content, the extraction logic for the feature information in the data content is expressed as:

[0021] Logic1: Obtain the data content, and based on the target recognition algorithm, identify the objects and object names in the data content;

[0022] Logic2: Traverse the data content, identify the color values of each pixel in the data content, and based on the color value recognition results, count the number of pixels corresponding to each color value;

[0023] Among them, when Logic2 is executed, after Logic1 identifies the object and the object name, the pixels in the area image outside the area image representing the object are used as the recognition and statistical targets, and the statistical results of the object name and the number of pixels corresponding to each color value are recorded as feature information.

[0024] Furthermore, when the analysis module analyzes the feature relevance of data content, the recognition results of the feature text information and the feature information are recorded as Dataset 1 and Dataset 2. Taking each group of feature information in Dataset 2 as the analysis target, the feature relevance analysis is carried out with each feature text information in Dataset 1;

[0025] In the feature relevance analysis stage, taking the object name in the feature information as the comparison content, comparing it with the feature text information, and recording the data content from which the feature information and the feature text information with the highest comparison matching degree are derived as the associated data content.

[0026] Furthermore, the matching degree comparison logic between the comparison content and the feature text information is expressed as:

[0027]

[0028] In the formula: K is the matching degree between the comparison content and the feature text information; m(a x ∩b x ) is the intersection number of the object name included in the comparison content a x in the feature information and the words included in the feature text information b x in Dataset 1; is the total amount of object names in the comparison content a x and the total amount of words in the feature text information b x in Dataset 1; d(r, t) is the semantic similarity between the r-th object name and the t-th word; is the averaging symbol;

[0029] Among them, the larger K is, the higher the matching degree between the comparison content and the feature text information. Based on the above formula, the similarity between each comparison content and each feature text information is calculated for comparison to obtain the comparison content with the highest matching degree for each feature text information. The data content from which the comparison content is derived and the data content from which the feature text information is derived are recorded as the associated data content.

[0030] Further, when the management module differentiates data content, for data content containing text information, it differentiates based on the semantic similarity of each characteristic text information, so that the semantic similarity between the data content containing text information stored in each differentiation interval is not less than a first preset value. For data content in image format, it differentiates based on the repetition rate of object names in each characteristic information, so that the repetition rate of object names in the characteristic information to which the data content in image format stored in each differentiation interval belongs is not less than a second preset value;

[0031] After the data content is differentiated based on the management module, each differentiation storage interval storing data content containing text information is used as a processing target, the associated data content of each data content in the processing target is obtained, the associated data content is retrieved from its own differentiation storage interval, and is transmitted to the processing target;

[0032] Among them, both the first preset value and the second preset value are user-defined by the system-side user. The first preset value and the second preset value are initially default set to 80% and 50%, and the data content stored in the cloud database is synchronized and deleted in the upload module.

[0033] Further, the upload module is wirelessly interactively connected to a feature recognition module. The lower level of the feature module is wirelessly interactively connected to a picking unit and an extraction unit. The feature recognition module is wirelessly interactively connected to an analysis module and a management module. The management module is wirelessly interactively connected to a cloud database. The cloud database is wirelessly interactively connected to a control module.

[0034] On the other hand, a distributed data management method based on artificial intelligence includes:

[0035] Step 1: Upload data content, identify the basic format of the data content, differentiate the data content based on the basic format of the data content, and temporarily store it using a storage medium;

[0036] Step 2: Clean the data content stored in the storage medium, perform feature recognition on the cleaned data content, analyze the relevance of each data content based on the feature recognition result to obtain associated data content;

[0037] Step 3: Obtain the feature recognition results of each data content, analyze the matching degree (i.e., semantic similarity) of each data content according to the comparison of the feature recognition results of each data content, and differentiate the data content in combination with the semantic similarity and the data content matching degree;

[0038] Step 4: Create a cloud database, use the cloud database to store the data content. During the storage stage of the data content by the cloud database, differentiated storage is performed based on the differentiation result of the data content in the previous step;

[0039] Step 5: When new data content is uploaded, refresh the step execution.

[0040] Adopting the technical solution provided by the present invention, compared with the known prior art, it has the following beneficial effects:

[0041] The present invention can efficiently clean the uploaded data, automatically identify and delete duplicate data, reduce redundant storage, improve the utilization rate of storage space. For data in different formats, whether it is data containing text information or image format data, it can accurately identify its features, find associated data content by analyzing the feature correlation of the data content, and provide users with more valuable data association information. In terms of storage management, the system can finely distinguish and store data according to the features of the data, such as the semantic similarity of text information and the repetition rate of object names in image information, greatly improving the storage management accuracy. When uploading to cloud storage, through preset values for customization, it can meet the diverse needs of different users, and the original uploaded data will be synchronously deleted after cloud storage, further ensuring the security and efficiency of data management. Based on this, it effectively improves the efficiency and accuracy of later data search. Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0043] Figure 1 It is a schematic structural diagram of a distributed data management system based on artificial intelligence;

[0044] Figure 2 It is a schematic flow diagram of a distributed data management method based on artificial intelligence. Detailed Embodiments

[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] The following further describes the present invention with reference to the embodiments.

[0047] Embodiment 1:

[0048] A distributed data management system based on artificial intelligence in this embodiment, as Figure 1 shown, includes:

[0049] An upload module, used to upload the data content that needs to be stored and managed, and place the uploaded data content in a storage medium for temporary storage;

[0050] The data content uploaded in the upload module comes from any authorized system-side user. The storage space size of the storage medium is manually configured by any system-side user. When the data content is placed in the storage medium for temporary storage, it is sent to the storage medium through medium transmission. After each new temporary storage data content is added to the storage medium, the storage medium executes a data cleaning operation through a preset built-in program;

[0051] Among them, when the data content executes a data cleaning operation in the storage medium based on a preset built-in program, it follows: distinguish the data content based on the data format, and execute a data cleaning operation in each distinguished interval. During the data cleaning operation stage, compare the similarity of each data content in the same distinguished interval, and use any one data in a group of data with a similarity of 100% as the deletion target and execute the deletion operation;

[0052] A feature recognition module, used to traverse the data content temporarily stored in the storage medium and perform feature recognition on the data content;

[0053] The data content uploaded in the upload module comes from any authorized system-side user. The storage space size of the storage medium is manually configured by any system-side user. When the data content is placed in the storage medium for temporary storage, it is sent to the storage medium through medium transmission. After each new temporary storage data content is added to the storage medium, the storage medium executes a data cleaning operation through a preset built-in program;

[0054] Among them, when the data content executes a data cleaning operation in the storage medium based on a preset built-in program, it follows: distinguish the data content based on the data format, and execute a data cleaning operation in each distinguished interval. During the data cleaning operation stage, compare the similarity of each data content in the same distinguished interval, and use any one data in a group of data with a similarity of 100% as the deletion target and execute the deletion operation;

[0055] During the operation stage of the pickup unit, after receiving the data content, the pickup logic of the feature text information in the data content is expressed as:

[0056] Perform word segmentation on the data content to obtain a word set, denoted as W = {w1, w2,..., w n}, and calculate the word frequency score, part-of-speech score, and position score of each word respectively;

[0057]

[0058] Total score of words: S(ω i ) = S f (ω i ) × α + pos(w i ) × β + S l × γ;

[0059] In the formula: S(ω i ) is the total score of words; S f (ω i ), pos(w i ), S l are the word frequency score, the part-of-speech score, and the position score; α, β, γ are weights; n is the total number of all words in the text; f(ω i ) is the number of times the word (w i ) appears in the text; f(ω j ) is the sum of the number of times all words in the text appear;

[0060] Among them, indicates that when the word is a noun, verb, adjective, or other part of speech, its corresponding value is 0.4, 0.3, 0.2, 0.1, indicates that when the word appears in the first 10% of the text, the last 10% of the text, or is the first word of a paragraph, the value is 0.3, and the value is 0.1 in other positions. After calculating the total score of each word, the words are sorted in descending order based on the score, and the top 5% of the words in the descending order are recorded as the characteristic text information;

[0061] Through the above logical formula calculation, the words from the data content are scored, and thus this is used as the pickup logic of the characteristic text information.

[0062] During the operation stage of the pickup unit, after receiving the data content, the extraction logic of the characteristic information in the data content is expressed as:

[0063] Logic1: Obtain the data content, and based on the target recognition algorithm, identify the objects and object names in the data content;

[0064] Logic2: Traverse the data content, identify the color values of each pixel in the data content, and based on the color value recognition result, count the number of pixels corresponding to each color value;

[0065] Among them, when Logic2 is executed, the pixels in the area image outside the area image representing the object after Logic1 identifies the objects and object names are used as the recognition and statistical targets, and the statistical results of the object names and the number of pixels corresponding to each color value are recorded as the characteristic information;

[0066] An analysis module, configured to obtain the feature recognition results of the data content in the feature recognition module, analyze the feature correlation of the data content based on the feature recognition results, and output associated data content by applying the analysis results of the feature correlation of the data content;

[0067] When analyzing the feature correlation of the data content, the analysis module records the recognition results of the feature text information and the feature information as Dataset One and Dataset Two. Taking each group of feature information in Dataset Two as the analysis target, it conducts feature correlation analysis with each feature text information in Dataset One;

[0068] In the feature correlation analysis stage, taking the object name in the feature information as the comparison content, comparing it with the feature text information, and recording the data content from which the feature information and the feature text information with the highest comparison matching degree are derived as the associated data content;

[0069] The matching degree comparison logic between the comparison content and the feature text information is expressed as:

[0070]

[0071] In the formula: K is the matching degree between the comparison content and the feature text information; m(a x ∩b x ) is the intersection quantity of the object name included in the comparison content a x in the feature information and the words included in the feature text information b x in Dataset One; is the total quantity of object names in the comparison content a x and the total quantity of words in the feature text information b x in Dataset One; d(r, t) is the semantic similarity between the r-th object name and the t-th word; is the symbol for taking the average;

[0072] Among them, the larger K is, the higher the matching degree between the comparison content and the feature text information. Based on the above formula, the similarity between each comparison content and each feature text information is calculated and compared to obtain the comparison content with the highest matching degree for each feature text information. The data content from which the comparison content is derived and the data content from which the feature text information is derived are recorded as the associated data content;

[0073] Through the above logical formula calculation, the matching degree between the comparison content and the feature text information is analyzed to provide support for the further operation of the management module in this embodiment of the system.

[0074] A management module, configured to receive the feature recognition results of each data content in the feature recognition module, distinguish the data content based on the feature recognition results, and forward the data content to the cloud database based on the distinction results;

[0075] When the management module differentiates data content, for the data content containing text information, it differentiates based on the semantic similarity of each characteristic text information, so that the semantic similarity between the data content containing text information stored in each differentiation interval is not less than a preset value one. For the data content in image format, it differentiates based on the repetition rate of object names in each characteristic information, so that the repetition rate of object names in the characteristic information to which the data content in image format stored in each differentiation interval belongs to each other is not less than a preset value two;

[0076] After the data content is differentiated based on the management module, each differentiation storage interval storing the data content containing text information is used as a processing target, the associated data content of each data content in the processing target is obtained, the associated data content is retrieved from its own differentiation storage interval, and is transmitted to the processing target;

[0077] Among them, both the preset value one and the preset value two are user-defined by the system-end user. The preset value one and the preset value two are initially default set to 80% and 50%, and the data content stored in the cloud database is synchronized and deleted in the upload module;

[0078] The cloud database is used to obtain the data content forwarded from the management module, and perform a differentiated storage operation on the data content based on its differentiation result in the management module;

[0079] The control module is used to monitor whether new data content is uploaded in the upload module, and when the monitoring result is yes, control the system to refresh and run;

[0080] The upload module is wirelessly interactively connected to a feature recognition module. The lower level of the feature module is wirelessly interactively connected to a pickup unit and an extraction unit. The feature recognition module is wirelessly interactively connected to an analysis module and a management module. The management module is wirelessly interactively connected to a cloud database. The cloud database is wirelessly interactively connected to the control module.

[0081] In this embodiment, the upload module runs to upload the data content that needs to be stored and managed, places the uploaded data content in the storage medium for temporary storage. The feature recognition module further traverses the data content temporarily stored in the storage medium, performs feature recognition on the data content, the picking unit synchronously picks up the data content containing text information, picks up the feature text information in the data content, the extraction module extracts the data content in real-time image format, extracts the feature information in the data content, the analysis module runs later to obtain the feature recognition results of the data content in the feature recognition module, analyzes the feature relevance of the data content based on the feature recognition results, applies the analysis results of the feature relevance of the data content to output the associated data content, and then the management module receives the feature recognition results of each data content in the feature recognition module, can distinguish the data content based on the feature recognition results, and forwards the data content to the cloud database based on the differentiation results, and obtains the data content forwarded from the management module through the cloud database, performs differentiated storage operations on the data content based on its differentiation results in the management module, and finally the control module monitors whether new data content is uploaded in the upload module. When the monitoring result is yes, the control system refreshes and runs.

[0082] Through the system operation in the above embodiment, it effectively provides a new, more refined technology for data storage management that is conducive to later data retrieval and query, and ensures the stability of massive data in the storage scenario.

[0083] Embodiment 2:

[0084] At the specific implementation level, on the basis of Embodiment 1, this embodiment refers to Figure 2 to further specifically describe a distributed data management system based on artificial intelligence in Embodiment 1:

[0085] A distributed data management method based on artificial intelligence includes:

[0086] Step 1: Upload the data content, identify the basic format of the data content, distinguish the data content based on the basic format of the data content and apply temporary storage in the storage medium;

[0087] Step 2: Clean the data content stored in the storage medium, perform feature recognition on the cleaned data content, and analyze the relevance of each data content based on the feature recognition results to obtain the associated data content;

[0088] Step 3: Obtain the feature recognition results of each data content, analyze the matching degree (i.e., semantic similarity) of each data content according to the comparison of the feature recognition results of each data content, and distinguish the data content by combining the semantic similarity and the data content matching degree;

[0089] Step 4: Create a cloud database and use the cloud database to store the data content. During the data storage phase, the cloud database stores the data content based on the differentiation results of the data content in the previous step.

[0090] Step 5: When new data content is uploaded, the refresh step is executed.

[0091] In summary, the system in the above embodiment can automatically clean uploaded data, remove duplicates, and improve storage utilization. It can accurately identify the characteristics of data in different formats, explore relationships between data, and provide users with valuable information. In terms of storage management, storage is finely differentiated by data characteristics, improving storage management accuracy. Users can also customize key preset values to meet diverse needs. After data is stored in the cloud, the uploader automatically deletes the data, ensuring the security of data management and achieving efficient, accurate, and secure data management.

[0092] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A distributed data management system based on artificial intelligence, characterized in that Including: An upload module, which is used to upload the data content that needs to be stored and managed, and place the uploaded data content in the storage medium for temporary storage; A feature recognition module, which is used to traverse the data content temporarily stored in the storage medium and perform feature recognition on the data content; An analysis module, which is used to obtain the feature recognition results of the data content in the feature recognition module, analyze the feature correlation of the data content based on the feature recognition results, and output associated data content by applying the analysis results of the feature correlation of the data content; A management module, which is used to receive the feature recognition results of each data content in the feature recognition module, distinguish the data content based on the feature recognition results, and forward the data content to the cloud database based on the distinction results; A cloud database, which is used to obtain the data content forwarded from the management module and perform a distinction storage operation on the data content based on its distinction results in the management module; A control module, which is used to monitor whether new data content is uploaded in the upload module, and when the monitoring result is yes, the control system refreshes and runs.

2. The distributed data management system based on artificial intelligence according to claim 1, characterized in that The data content uploaded in the upload module comes from any authorized system-side user. The storage space size of the storage medium is manually configured by any system-side user. When the data content is placed in the storage medium for temporary storage, it is sent to the storage medium through medium transmission. After each new temporary storage data content is added to the storage medium, the storage medium performs a data cleaning operation through a preset built-in program; Among them, when the data content performs a data cleaning operation in the storage medium based on a preset built-in program, it follows: distinguish the data content based on the data format, and perform a data cleaning operation once in each distinction interval. During the data cleaning operation stage, compare the similarity of each data content in the same distinction interval, and use any one data in a group of data with a similarity of 100% as the deletion target and perform a deletion operation.

3. An artificial intelligence-based distributed data management system according to claim 1, characterized in that A pickup unit and an extraction unit are set at a lower level of the feature recognition module. The pickup unit is used for data content containing text information and picks up feature text information in the data content. The extraction module is used for data content in image format and extracts feature information in the data content; Among them, the feature recognition module runs continuously. Each time the feature recognition module runs, it uses the data content stored in a distinction interval in the storage medium as the processing target, identifies the processing target format. When the processing target format is data content containing text information, each data in the processing target is sequentially sent to the pickup unit. When the processing target format is image format, each data in the processing target is sequentially sent to the extraction unit.

4. An artificial intelligence-based distributed data management system according to claim 3, characterized in that, During the operation stage of the pickup unit, after receiving the data content, the pickup logic for the feature text information in the data content is expressed as: Tokenize the data content to obtain a set of words, denoted as W = {w1, w2,..., w n}, and calculate the word frequency score, part-of-speech score, and position score of each word respectively; Total word score: S(ω i ) = S f (ω i ) × α + pos(w i ) × β + S l × γ; Where: S(ω i ) is the total word score; S f (ω i ), pos(w i ), S l are the word frequency score, the part-of-speech score, and the position score; α, β, γ are weights; n is the total number of all words in the text; f(ω i ) is the number of times the word (w i ) appears in the text; f(ω j ) is the sum of the number of times all words in the text appear; Among them, It means that when the word is a noun, verb, adjective, or other part of speech, its corresponding value is 0.4, 0.3, 0.2, 0.

1. It means that when the word appears in the first 10% of the text, the last 10% of the text, or is the first word of a paragraph, the value is 0.3, and the value in other positions is 0.

1. After obtaining the total score for each word, the words are sorted in descending order based on the score, and the top 5% of the words in the descending order are recorded as the characteristic text information.

5. An artificial intelligence-based distributed data management system according to claim 3, characterized in that, During the operation stage of the pickup unit, after receiving the data content, the extraction logic for the feature information in the data content is expressed as: Logic1: Obtain the data content and identify the objects and object names in the data content based on the target recognition algorithm; Logic2: Traverse the data content, identify the color values of each pixel in the data content, and count the number of pixels corresponding to each color value based on the color value recognition results; Among them, when Logic2 is executed, after Logic1 identifies the object and the object name, the pixels in the area image outside the area image representing the object are used as the recognition and statistical targets, and the statistical results of the object name and the pixel quantities corresponding to the respective color values are recorded as feature information.

6. The distributed data management system based on artificial intelligence according to claim 1, characterized in that When the analysis module analyzes the feature relevance of the data content, the recognition results of the feature text information and the feature information are recorded as Dataset 1 and Dataset 2. Using each group of feature information in Dataset 2 as the analysis target, the feature relevance analysis is performed with each feature text information in Dataset 1; In the feature relevance analysis stage, using the object name in the feature information as the comparison content, comparing it with the feature text information, and recording the data content from which the feature information and the feature text information with the highest comparison matching degree are derived as the associated data content.

7. An artificial intelligence-based distributed data management system according to claim 6, characterized in that, The matching degree comparison logic between the comparison content and the feature text information is expressed as: Where: K is the matching degree between the comparison content and the characteristic text information; m(a x ∩b x ) is the number of intersections of the comparison content a x containing the object name and the characteristic text information b x in the feature information; is the total amount of object names in the comparison content a x and the total amount of words in the characteristic text information b x ; d(r,t) is the semantic similarity between the r-th object name and the t-th word; is the symbol for taking the average; Among them, the larger K is, the higher the matching degree between the comparison content and the feature text information. Based on the above formula, the similarity between each comparison content and each feature text information is obtained for comparison to obtain the comparison content with the highest matching degree for each feature text information. The data content from which the comparison content is derived and the data content from which the feature text information is derived are recorded as the associated data content.

8. An artificial intelligence-based distributed data management system according to claim 1, characterized in that, When the management module differentiates the data content, for the data content containing text information, it differentiates based on the semantic similarity of each feature text information, so that the semantic similarity between the data content containing text information stored in each differentiation interval is not less than a preset value 1. For the data content in image format, it differentiates based on the repetition rate of the object name in each feature information, so that the repetition rate of the object name in the feature information to which the data content in image format stored in each differentiation interval belongs is not less than a preset value 2; After the data content is differentiated based on the management module, taking each differentiation storage interval storing the data content containing text information as the processing target, obtaining the associated data content of each data content in the processing target, retrieving the associated data content from its own differentiation storage interval, and transmitting it to the processing target; Among them, both the preset value 1 and the preset value 2 are user-defined by the system terminal user. The preset value 1 and the preset value 2 are initially default set to 80% and 50%, and the data content stored in the cloud database is synchronized and deleted in the upload module.

9. An artificial intelligence-based distributed data management system according to claim 1, wherein, The upload module is wirelessly interconnected with a feature recognition module. The lower level of the feature module is wirelessly interconnected with a picking unit and an extraction unit. The feature recognition module is wirelessly interconnected with the analysis module and the management module. The management module is wirelessly interconnected with the cloud database. The cloud database is wirelessly interconnected with the control module.

10. A distributed data management method based on artificial intelligence, which is an implementation method of a distributed data management system based on artificial intelligence as described in any one of claims 1-9, characterized in that, Including: Step 1: Upload the data content, identify the basic format of the data content, differentiate the data content based on the basic format of the data content, and temporarily store it using a storage medium; Step 2: Clean the data content stored in the storage medium, perform feature recognition on the data content after cleaning, analyze the relevance of each data content based on the feature recognition results to obtain the associated data content; Step 3: Obtain the recognition results of each data content feature, analyze the matching degree of each data content, i.e., the semantic similarity, according to the comparison of the recognition results of each data content feature, and distinguish the data content in combination with the semantic similarity and the data content matching degree; Step 4: Create a cloud database, and use the cloud database to store the data content. During the storage stage of the data content by the cloud database, perform differential storage based on the differential results of the data content in the previous step; Step 5: When new data content is uploaded, refresh the step execution.

Citation Information

Patent Citations

  • Distributed data matching method and system based on artificial intelligence

    CN119004134A

  • Data and text processing system and method based on machine learning

    CN115577698A

  • Data storage system and data writing method

    CN119311922A