A data processing method, apparatus and device
By obtaining the pinyin and word segmentation results of the site name, and combining weight coefficients and text features, a text classification model is used to automatically identify fuzzy resources, solving the problems of low efficiency and high cost in existing technologies, and achieving efficient site name correction.
Patent Information
- Application Number
- CN202111341557.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-12
- Publication Date
- 2026-05-15
- Estimated Expiration
- 2041-11-12
AI Technical Summary
Existing data processing solutions for determining fuzzy resources are inefficient and costly, especially when the names of spatial resource sites are entered incorrectly, manual error correction is both inefficient and costly.
By acquiring the pinyin information and word segmentation results of the site name, and combining weight coefficients and text features, a text classification model is used to automatically analyze whether the site name belongs to fuzzy resources, including indicators such as pinyin overlap, character differences, word segmentation position, and latitude and longitude distance, to achieve intelligent fuzzy resource identification.
It improves the efficiency of fuzzy resource identification, reduces labor costs, and achieves automated and intelligent site name correction processing.
Smart Images

Figure CN116127052B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a data processing method, apparatus, and device. Background Technology
[0002] With the development of the network, there are more and more spatial resource sites (a site refers to a location that includes a computer room, generally a building containing one or more computer rooms). These sites are recorded in the spatial resource table. However, since these resource site names are all manually entered, there are often cases where the data entry personnel accidentally make typos, resulting in these incorrectly entered resource sites becoming ambiguous resources that require manual correction later. This process is inefficient and costly. In addition, it is almost impossible to manually correct the massive number of site names.
[0003] As can be seen from the above, the existing data processing solutions for determining fuzzy resources have problems such as low efficiency and high labor costs. Summary of the Invention
[0004] The purpose of this invention is to provide a data processing method, apparatus, and device to solve the problems of low efficiency and high labor costs in existing data processing schemes for determining fuzzy resources.
[0005] To address the aforementioned technical problems, embodiments of the present invention provide a data processing method, comprising:
[0006] Obtain the first pinyin to be matched corresponding to the name of the site to be processed, and the first target pinyin of the first information in the first target site name corresponding to the name of the site to be processed;
[0007] Based on the first pinyin to be matched and the first target pinyin, a first processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information;
[0008] The name of the site to be processed is segmented into words to obtain at least one first segmentation result;
[0009] Based on the at least one first word segmentation result and the reference site name, a second processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information;
[0010] Based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, the final result of whether the site name to be processed belongs to fuzzy resource information is obtained.
[0011] The first information refers to the text information in the name of the first target site, excluding the text used for level classification.
[0012] Optionally, the step of obtaining a first processing result regarding whether the site name to be processed belongs to fuzzy resource information based on the first pinyin to be matched and the first target pinyin includes:
[0013] Determine whether there is any overlapping pinyin between the first pinyin to be matched and the first target pinyin;
[0014] In the absence of overlapping pinyin, the first processing result is that the name of the site to be processed does not belong to fuzzy resource information;
[0015] In the case of overlapping pinyin, obtain the first text corresponding to the overlapping pinyin in the name of the site to be processed and the second text corresponding to the first target site name;
[0016] If the first text and the second text are the same, the first processing result is that the site name to be processed does not belong to the fuzzy resource information;
[0017] If the first text and the second text are different, the number of characters in the first text is obtained, and based on the number of characters, a first processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information.
[0018] Optionally, the step of obtaining a first processing result based on the number of characters in the text to determine whether the site name to be processed belongs to fuzzy resource information includes:
[0019] When the number of characters in the text is equal to 1, the first processing result is that the name of the site to be processed does not belong to the fuzzy resource information.
[0020] If the number of characters in the text is greater than 1, the first processing result is obtained that the name of the site to be processed belongs to fuzzy resource information.
[0021] Optionally, the step of obtaining a second processing result—whether the site name to be processed belongs to fuzzy resource information—based on the at least one first word segmentation result and the reference site name includes:
[0022] If the second information is not stored in a preset location, obtain the word segmentation pinyin corresponding to the second information, and use the second information as the third text corresponding to the word segmentation pinyin;
[0023] If the total number of all texts corresponding to the segmented pinyin is greater than 1, then the segmented pinyin corresponds to the third text and at least one fourth text.
[0024] Obtain the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts;
[0025] If the distance is less than the first threshold, a second processing result is obtained based on the fourth text, the third text, and the reference site name corresponding to the distance, to determine whether the site name to be processed belongs to fuzzy resource information.
[0026] The second information is either the first word segmentation result or the aggregated result of the first word segmentation result after being aggregated according to a preset slider length.
[0027] Optionally, obtaining the distance between the latitude and longitude location corresponding to the third text and the latitude and longitude location corresponding to each of the fourth texts includes:
[0028] If the number of characters corresponding to the word segmentation pinyin is greater than the second threshold, obtain the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts;
[0029] Wherein, the second threshold is equal to the preset slider length.
[0030] Optionally, the second processing result of determining whether the site name to be processed belongs to fuzzy resource information based on the fourth text, the third text, and the reference site name corresponding to the distance includes:
[0031] If the fourth text corresponding to the distance is contained in the reference site name, a second processing result is obtained in which the site name to be processed belongs to fuzzy resource information;
[0032] If the third text is contained in the reference site name, a second processing result is obtained in which the site name to be processed does not belong to the fuzzy resource information.
[0033] Optional, also includes:
[0034] Based on the set of reference site names, obtain text features;
[0035] Based on the text features, a text classification model is constructed; the text classification model can determine whether the input site name belongs to fuzzy resource information.
[0036] Before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, the process further includes:
[0037] Using the text classification model, a third processing result is obtained to determine whether the name of the site to be processed belongs to fuzzy resource information;
[0038] The step of obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result includes:
[0039] Based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result, the final result of whether the site name to be processed belongs to fuzzy resource information is obtained.
[0040] Optionally, obtaining text features based on the set of reference site names includes:
[0041] Based on the set of reference site names, obtain the forward and reverse datasets;
[0042] The forward and reverse datasets are randomly shuffled to form a complete dataset, and the training set in the complete dataset is obtained.
[0043] Using a pre-trained language model, text features are obtained from the training set;
[0044] The forward dataset is the set of reference site names, and the reverse dataset includes a first reverse dataset and a second reverse dataset. The first reverse dataset is obtained by swapping any two characters in each reference site name in the set of reference site names. The second reverse dataset is obtained by randomly replacing any character in the forward dataset with any character from a preset set of characters.
[0045] Optionally, before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, the method further includes:
[0046] Based on the name of the first tagged site, obtain the error rate corresponding to the preset operation;
[0047] Based on the error rate, obtain the corresponding weighting coefficient;
[0048] Wherein, the preset operation is a first operation and the weight coefficient is a first weight coefficient; or, the preset operation is a second operation and the weight coefficient is a second weight coefficient; or, the preset operation is a third operation and the weight coefficient is a third weight coefficient.
[0049] The first operation includes: obtaining the second pinyin to be matched corresponding to the first site name, and the second target pinyin of the second information in the second target site name corresponding to the first site name; obtaining the result of whether the first site name belongs to fuzzy resource information based on the second pinyin to be matched and the second target pinyin; the second information is the text information in the second target site name excluding the text used for level classification;
[0050] The second operation includes: segmenting the first site name into words to obtain at least one second segmentation result; and determining whether the first site name belongs to fuzzy resource information based on the at least one second segmentation result and the reference site name.
[0051] The third operation includes: using the text classification model to obtain the result of whether the first site name belongs to fuzzy resource information.
[0052] Optionally, obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result includes:
[0053] Using the first formula, based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result, the final result of whether the site name to be processed belongs to fuzzy resource information is obtained;
[0054] The first formula is:
[0055] ;
[0056] Indicates the final result. This represents the weight coefficient corresponding to the k-th processing result. This represents the k-th processing result, x represents the name of the site to be processed, k is greater than or equal to 1, and k is less than or equal to the total number of processing results.
[0057] This invention also provides a data processing apparatus, comprising:
[0058] The first acquisition module is used to acquire the first pinyin to be matched corresponding to the site name to be processed, and the first target pinyin of the first information in the first target site name corresponding to the site name to be processed;
[0059] The first processing module is used to obtain a first processing result as to whether the site name to be processed belongs to fuzzy resource information based on the first pinyin to be matched and the first target pinyin;
[0060] The second processing module is used to segment the site name to be processed into words to obtain at least one first word segmentation result;
[0061] The third processing module is used to obtain a second processing result as to whether the site name to be processed belongs to fuzzy resource information based on the at least one first word segmentation result and the reference site name;
[0062] The fourth processing module is used to obtain the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result.
[0063] The first information refers to the text information in the name of the first target site, excluding the text used for level classification.
[0064] Optionally, the step of obtaining a first processing result regarding whether the site name to be processed belongs to fuzzy resource information based on the first pinyin to be matched and the first target pinyin includes:
[0065] Determine whether there is any overlapping pinyin between the first pinyin to be matched and the first target pinyin;
[0066] In the absence of overlapping pinyin, the first processing result is that the name of the site to be processed does not belong to fuzzy resource information;
[0067] In the case of overlapping pinyin, obtain the first text corresponding to the overlapping pinyin in the name of the site to be processed and the second text corresponding to the first target site name;
[0068] If the first text and the second text are the same, the first processing result is that the site name to be processed does not belong to the fuzzy resource information;
[0069] If the first text and the second text are different, the number of characters in the first text is obtained, and based on the number of characters, a first processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information.
[0070] Optionally, the step of obtaining a first processing result based on the number of characters in the text to determine whether the site name to be processed belongs to fuzzy resource information includes:
[0071] When the number of characters in the text is equal to 1, the first processing result is that the name of the site to be processed does not belong to the fuzzy resource information.
[0072] If the number of characters in the text is greater than 1, the first processing result is obtained that the name of the site to be processed belongs to fuzzy resource information.
[0073] Optionally, the step of obtaining a second processing result—whether the site name to be processed belongs to fuzzy resource information—based on the at least one first word segmentation result and the reference site name includes:
[0074] If the second information is not stored in a preset location, obtain the word segmentation pinyin corresponding to the second information, and use the second information as the third text corresponding to the word segmentation pinyin;
[0075] If the total number of all texts corresponding to the segmented pinyin is greater than 1, then the segmented pinyin corresponds to the third text and at least one fourth text.
[0076] Obtain the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts;
[0077] If the distance is less than the first threshold, a second processing result is obtained based on the fourth text, the third text, and the reference site name corresponding to the distance, to determine whether the site name to be processed belongs to fuzzy resource information.
[0078] The second information is either the first word segmentation result or the aggregated result of the first word segmentation result after being aggregated according to a preset slider length.
[0079] Optionally, obtaining the distance between the latitude and longitude location corresponding to the third text and the latitude and longitude location corresponding to each of the fourth texts includes:
[0080] If the number of characters corresponding to the word segmentation pinyin is greater than the second threshold, obtain the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts;
[0081] Wherein, the second threshold is equal to the preset slider length.
[0082] Optionally, the second processing result of determining whether the site name to be processed belongs to fuzzy resource information based on the fourth text, the third text, and the reference site name corresponding to the distance includes:
[0083] If the fourth text corresponding to the distance is contained in the reference site name, a second processing result is obtained in which the site name to be processed belongs to fuzzy resource information;
[0084] If the third text is contained in the reference site name, a second processing result is obtained in which the site name to be processed does not belong to the fuzzy resource information.
[0085] Optional, also includes:
[0086] The second acquisition module is used to obtain text features based on the set of reference site names;
[0087] The first construction module is used to construct a text classification model based on the text features; the text classification model can determine whether the input site name belongs to fuzzy resource information.
[0088] The device further includes:
[0089] The third acquisition module is used to obtain a third processing result of whether the site name to be processed belongs to fuzzy resource information by using the text classification model before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result.
[0090] The step of obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result includes:
[0091] Based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result, the final result of whether the site name to be processed belongs to fuzzy resource information is obtained.
[0092] Optionally, obtaining text features based on the set of reference site names includes:
[0093] Based on the set of reference site names, obtain the forward and reverse datasets;
[0094] The forward and reverse datasets are randomly shuffled to form a complete dataset, and the training set in the complete dataset is obtained.
[0095] Using a pre-trained language model, text features are obtained from the training set;
[0096] The forward dataset is the set of reference site names, and the reverse dataset includes a first reverse dataset and a second reverse dataset. The first reverse dataset is obtained by swapping any two characters in each reference site name in the set of reference site names. The second reverse dataset is obtained by randomly replacing any character in the forward dataset with any character from a preset set of characters.
[0097] Optional, also includes:
[0098] The fourth acquisition module is used to obtain the error rate corresponding to the preset operation based on the first site name that has been labeled, before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result and the second weight coefficient corresponding to the second processing result.
[0099] The fifth acquisition module is used to acquire the corresponding weight coefficient based on the error rate;
[0100] Wherein, the preset operation is a first operation and the weight coefficient is a first weight coefficient; or, the preset operation is a second operation and the weight coefficient is a second weight coefficient; or, the preset operation is a third operation and the weight coefficient is a third weight coefficient.
[0101] The first operation includes: obtaining the second pinyin to be matched corresponding to the first site name, and the second target pinyin of the second information in the second target site name corresponding to the first site name; obtaining the result of whether the first site name belongs to fuzzy resource information based on the second pinyin to be matched and the second target pinyin; the second information is the text information in the second target site name excluding the text used for level classification;
[0102] The second operation includes: segmenting the first site name into words to obtain at least one second segmentation result; and determining whether the first site name belongs to fuzzy resource information based on the at least one second segmentation result and the reference site name.
[0103] The third operation includes: using the text classification model to obtain the result of whether the first site name belongs to fuzzy resource information.
[0104] Optionally, obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result includes:
[0105] Using the first formula, based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result, the final result of whether the site name to be processed belongs to fuzzy resource information is obtained;
[0106] The first formula is:
[0107] ;
[0108] Indicates the final result. This represents the weight coefficient corresponding to the k-th processing result. This represents the k-th processing result, x represents the name of the site to be processed, k is greater than or equal to 1, and k is less than or equal to the total number of processing results.
[0109] This invention also provides a data processing device, including: a processor and a transceiver;
[0110] The processor is used to obtain the first pinyin to be matched corresponding to the name of the site to be processed, and the first target pinyin of the first information in the first target site name corresponding to the name of the site to be processed;
[0111] Based on the first pinyin to be matched and the first target pinyin, a first processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information;
[0112] The name of the site to be processed is segmented into words to obtain at least one first segmentation result;
[0113] Based on the at least one first word segmentation result and the reference site name, a second processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information;
[0114] Based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, the final result of whether the site name to be processed belongs to fuzzy resource information is obtained.
[0115] The first information refers to the text information in the name of the first target site, excluding the text used for level classification.
[0116] Optionally, the step of obtaining a first processing result regarding whether the site name to be processed belongs to fuzzy resource information based on the first pinyin to be matched and the first target pinyin includes:
[0117] Determine whether there is any overlapping pinyin between the first pinyin to be matched and the first target pinyin;
[0118] In the absence of overlapping pinyin, the first processing result is that the name of the site to be processed does not belong to fuzzy resource information;
[0119] In the case of overlapping pinyin, obtain the first text corresponding to the overlapping pinyin in the name of the site to be processed and the second text corresponding to the first target site name;
[0120] If the first text and the second text are the same, the first processing result is that the site name to be processed does not belong to the fuzzy resource information;
[0121] If the first text and the second text are different, the number of characters in the first text is obtained, and based on the number of characters, a first processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information.
[0122] Optionally, the step of obtaining a first processing result based on the number of characters in the text to determine whether the site name to be processed belongs to fuzzy resource information includes:
[0123] When the number of characters in the text is equal to 1, the first processing result is that the name of the site to be processed does not belong to the fuzzy resource information.
[0124] If the number of characters in the text is greater than 1, the first processing result is obtained that the name of the site to be processed belongs to fuzzy resource information.
[0125] Optionally, the step of obtaining a second processing result—whether the site name to be processed belongs to fuzzy resource information—based on the at least one first word segmentation result and the reference site name includes:
[0126] If the second information is not stored in a preset location, obtain the word segmentation pinyin corresponding to the second information, and use the second information as the third text corresponding to the word segmentation pinyin;
[0127] If the total number of all texts corresponding to the segmented pinyin is greater than 1, then the segmented pinyin corresponds to the third text and at least one fourth text.
[0128] Obtain the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts;
[0129] If the distance is less than the first threshold, a second processing result is obtained based on the fourth text, the third text, and the reference site name corresponding to the distance, to determine whether the site name to be processed belongs to fuzzy resource information.
[0130] The second information is either the first word segmentation result or the aggregated result of the first word segmentation result after being aggregated according to a preset slider length.
[0131] Optionally, obtaining the distance between the latitude and longitude location corresponding to the third text and the latitude and longitude location corresponding to each of the fourth texts includes:
[0132] If the number of characters corresponding to the word segmentation pinyin is greater than the second threshold, obtain the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts;
[0133] Wherein, the second threshold is equal to the preset slider length.
[0134] Optionally, the second processing result of determining whether the site name to be processed belongs to fuzzy resource information based on the fourth text, the third text, and the reference site name corresponding to the distance includes:
[0135] If the fourth text corresponding to the distance is contained in the reference site name, a second processing result is obtained in which the site name to be processed belongs to fuzzy resource information;
[0136] If the third text is contained in the reference site name, a second processing result is obtained in which the site name to be processed does not belong to the fuzzy resource information.
[0137] Optionally, the processor is further configured to:
[0138] Based on the set of reference site names, obtain text features;
[0139] Based on the text features, a text classification model is constructed; the text classification model can determine whether the input site name belongs to fuzzy resource information.
[0140] The processor is also used for:
[0141] Before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, the text classification model is used to obtain the third processing result of whether the site name to be processed belongs to fuzzy resource information.
[0142] The step of obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result includes:
[0143] Based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result, the final result of whether the site name to be processed belongs to fuzzy resource information is obtained.
[0144] Optionally, obtaining text features based on the set of reference site names includes:
[0145] Based on the set of reference site names, obtain the forward and reverse datasets;
[0146] The forward and reverse datasets are randomly shuffled to form a complete dataset, and the training set in the complete dataset is obtained.
[0147] Using a pre-trained language model, text features are obtained from the training set;
[0148] The forward dataset is the set of reference site names, and the reverse dataset includes a first reverse dataset and a second reverse dataset. The first reverse dataset is obtained by swapping any two characters in each reference site name in the set of reference site names. The second reverse dataset is obtained by randomly replacing any character in the forward dataset with any character from a preset set of characters.
[0149] Optionally, the processor is further configured to:
[0150] Before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, the error rate corresponding to the preset operation is obtained based on the first site name that has been labeled.
[0151] Based on the error rate, obtain the corresponding weighting coefficient;
[0152] Wherein, the preset operation is a first operation and the weight coefficient is a first weight coefficient; or, the preset operation is a second operation and the weight coefficient is a second weight coefficient; or, the preset operation is a third operation and the weight coefficient is a third weight coefficient.
[0153] The first operation includes: obtaining the second pinyin to be matched corresponding to the first site name, and the second target pinyin of the second information in the second target site name corresponding to the first site name; obtaining the result of whether the first site name belongs to fuzzy resource information based on the second pinyin to be matched and the second target pinyin; the second information is the text information in the second target site name excluding the text used for level classification;
[0154] The second operation includes: segmenting the first site name into words to obtain at least one second segmentation result; and determining whether the first site name belongs to fuzzy resource information based on the at least one second segmentation result and the reference site name.
[0155] The third operation includes: using the text classification model to obtain the result of whether the first site name belongs to fuzzy resource information.
[0156] Optionally, obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result includes:
[0157] Using the first formula, based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result, the final result of whether the site name to be processed belongs to fuzzy resource information is obtained;
[0158] The first formula is:
[0159] ;
[0160] Indicates the final result. This represents the weight coefficient corresponding to the k-th processing result. This represents the k-th processing result, x represents the name of the site to be processed, k is greater than or equal to 1, and k is less than or equal to the total number of processing results.
[0161] This invention also provides a data processing device, including a memory, a processor, and a program stored in the memory and executable on the processor; when the processor executes the program, it implements the above-described data processing method.
[0162] This invention also provides a readable storage medium storing a program that, when executed by a processor, implements the steps in the data processing method described above.
[0163] The beneficial effects of the above-described technical solution of the present invention are as follows:
[0164] In the above scheme, the data processing method obtains the first pinyin to be matched corresponding to the site name to be processed, and the first target pinyin of the first information in the first target site name corresponding to the site name to be processed; based on the first pinyin to be matched and the first target pinyin, it obtains a first processing result of whether the site name to be processed belongs to fuzzy resource information; it segments the site name to be processed into words to obtain at least one first segmentation result; based on the at least one first segmentation result and the reference site name, it obtains a second processing result of whether the site name to be processed belongs to fuzzy resource information; based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, it obtains the final result of whether the site name to be processed belongs to fuzzy resource information; wherein, the first information is the text information in the first target site name excluding the text used for level classification; it can realize automated and intelligent analysis of whether the site name belongs to fuzzy resource information (such as whether there are text errors), improves processing efficiency, and reduces labor costs, and effectively solves the problems of low efficiency and high labor costs in the existing data processing schemes for determining fuzzy resources (corresponding to the above-mentioned fuzzy resource information). Attached Figure Description
[0165] Figure 1 This is a schematic diagram of the data processing method according to an embodiment of the present invention;
[0166] Figure 2 This is a schematic diagram illustrating the specific implementation flow of the data processing method in this embodiment of the invention. Figure 1 ;
[0167] Figure 3 This is a schematic diagram illustrating the specific implementation flow of the data processing method in this embodiment of the invention. Figure 2 ;
[0168] Figure 4 This is a schematic diagram illustrating the specific implementation flow of the data processing method in this embodiment of the invention. Figure 3 ;
[0169] Figure 5 This is a schematic diagram illustrating the specific implementation flow of the data processing method in this embodiment of the invention. Figure 4 ;
[0170] Figure 6 This is a schematic diagram of the data processing device structure according to an embodiment of the present invention;
[0171] Figure 7 This is a schematic diagram of the data processing device structure according to an embodiment of the present invention. Detailed Implementation
[0172] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0173] This invention addresses the problems of low efficiency and high labor costs in existing data processing schemes for determining fuzzy resources by providing a data processing method (also known as an information processing method), such as... Figure 1 As shown, it includes:
[0174] Step 11: Obtain the first pinyin to be matched corresponding to the site name to be processed, and the first target pinyin of the first information in the first target site name corresponding to the site name to be processed;
[0175] Step 12: Based on the first pinyin to be matched and the first target pinyin, obtain the first processing result of whether the site name to be processed belongs to fuzzy resource information;
[0176] Step 13: Segment the site name to be processed into words to obtain at least one first segmentation result;
[0177] Step 14: Based on the at least one first word segmentation result and the reference site name, obtain a second processing result as to whether the site name to be processed belongs to fuzzy resource information;
[0178] Step 15: Based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, obtain the final result of whether the name of the site to be processed belongs to fuzzy resource information; wherein, the first information is the text information in the first target site name excluding the text used for level classification.
[0179] The "text used for hierarchical classification" can be text used for address hierarchical classification, such as city, county, district, village, etc. Steps 11 and 13 are not performed in any particular order.
[0180] The data processing method provided by the embodiment of the present invention obtains the first to-be-matched pinyin corresponding to the to-be-processed site name and the first target pinyin of the first information in the first target site name corresponding to the to-be-processed site name; obtains the first processing result of whether the to-be-processed site name belongs to fuzzy resource information according to the first to-be-matched pinyin and the first target pinyin; segments the to-be-processed site name to obtain at least one first segmentation result; obtains the second processing result of whether the to-be-processed site name belongs to fuzzy resource information according to the at least one first segmentation result and the reference site name; obtains the final result of whether the to-be-processed site name belongs to fuzzy resource information according to the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result; wherein the first information is the text information in the first target site name except the text for level division; it can realize automatic and intelligent analysis of whether the site name belongs to fuzzy resource information (such as whether there are text errors), improves the processing efficiency, and reduces the labor cost, and well solves the problems of low efficiency and high labor cost of the data processing scheme for determining fuzzy resources (corresponding to the above fuzzy resource information) in the prior art.
[0181] In the embodiment of the present invention, the obtaining the first processing result of whether the to-be-processed site name belongs to fuzzy resource information according to the first to-be-matched pinyin and the first target pinyin includes: determining whether there is a coincident pinyin between the first to-be-matched pinyin and the first target pinyin; in the case of no coincident pinyin, obtaining the first processing result that the to-be-processed site name does not belong to fuzzy resource information; in the case of having a coincident pinyin, obtaining the first text corresponding to the coincident pinyin in the to-be-processed site name and the second text corresponding to the coincident pinyin in the first target site name; in the case of the first text being the same as the second text, obtaining the first processing result that the to-be-processed site name does not belong to fuzzy resource information; in the case of the first text being different from the second text, obtaining the number of characters of the first text, and obtaining the first processing result of whether the to-be-processed site name belongs to fuzzy resource information according to the number of characters.
[0182] In this way, the first processing result can be accurately obtained. The "determining whether there is a coincident pinyin between the first to-be-matched pinyin and the first target pinyin" can specifically include: determining whether there is a coincident pinyin between the pinyin corresponding to each address level in the first to-be-matched pinyin and the first target pinyin; for example, whether there is a coincident pinyin between the pinyin of XX (the content before the level division character "city") and the first target pinyin.
[0183] The step of obtaining a first processing result of whether the site name to be processed belongs to fuzzy resource information based on the number of characters in the text includes: obtaining a first processing result that the site name to be processed does not belong to fuzzy resource information when the number of characters in the text is equal to 1; and obtaining a first processing result that the site name to be processed belongs to fuzzy resource information when the number of characters in the text is greater than 1.
[0184] This can further ensure the accuracy of the first processing result.
[0185] In this embodiment of the invention, the step of obtaining a second processing result regarding whether the site name to be processed belongs to fuzzy resource information based on at least one first word segmentation result and a reference site name includes: when the second information is not stored in a preset location, obtaining the word segmentation pinyin corresponding to the second information and using the second information as the third text corresponding to the word segmentation pinyin; when the total number of all texts corresponding to the word segmentation pinyin is greater than 1, determining that the word segmentation pinyin corresponds to the third text and at least one fourth text; obtaining the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts; when the distance is less than a first threshold, obtaining a second processing result regarding whether the site name to be processed belongs to fuzzy resource information based on the fourth text, the third text, and the reference site name corresponding to the distance; wherein, the second information is the first word segmentation result, or the aggregation result of the first word segmentation result after aggregation according to a preset slider length.
[0186] This allows for accurate acquisition of the second processing result. "Not stored in the preset location" can be understood as not having appeared in the preset location. The "preset slider length" can specifically be 2, for example, to combine (aggregate) two word segmentation pinyin entries together. In this embodiment of the invention, when the first word segmentation result has been stored in a preset position (i.e., it has already appeared), a skip operation can be performed on the current site name to be processed, which can also be understood as updating the site name to be processed; that is, the operation of obtaining the above final result is performed on the next site name to be processed, which can also be understood as returning to the updated site name to be processed and performing "obtaining the first pinyin to be matched corresponding to the site name to be processed, and the first target pinyin of the first information in the first target site name corresponding to the site name to be processed; obtaining a first processing result of whether the site name to be processed belongs to fuzzy resource information based on the first pinyin to be matched and the first target pinyin; segmenting the site name to be processed to obtain at least one first word segmentation result; obtaining a second processing result of whether the site name to be processed belongs to fuzzy resource information based on the at least one first word segmentation result and the reference site name; obtaining a final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result; wherein, the first information is the text information in the first target site name excluding the text used for level classification".
[0187] Furthermore, in this embodiment of the invention, when the number of third texts corresponding to the segmented pinyin is equal to 1, a skip operation can be performed on the current site name to be processed, which can also be understood as updating the site name to be processed; that is, performing the operation on the next site name to be processed. When the distance is greater than or equal to the first threshold, a skip operation can be performed on the current site name to be processed, which can also be understood as updating the site name to be processed; that is, performing the operation on the next site name to be processed.
[0188] The first threshold can be an empirical value, but it is not a limitation.
[0189] In this embodiment of the invention, obtaining the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts includes: when the number of characters corresponding to the word segmentation pinyin is greater than a second threshold, obtaining the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts; wherein, the second threshold is equal to a preset slider length.
[0190] This avoids excessive execution of the distance acquisition operation, saving processing resources. Specifically, if the number of characters corresponding to the segmented pinyin is less than or equal to the second threshold, a skip operation can be performed on the current site name to be processed, which can also be understood as updating the site name to be processed; that is, performing the operation on the next site name to be processed.
[0191] The step of obtaining a second processing result regarding whether the site name to be processed belongs to fuzzy resource information based on the fourth text, the third text, and the reference site name corresponding to the distance includes: obtaining a second processing result that the site name to be processed belongs to fuzzy resource information when the fourth text corresponding to the distance is included in the reference site name; and obtaining a second processing result that the site name to be processed does not belong to fuzzy resource information when the third text is included in the reference site name.
[0192] This allows for a quick and accurate acquisition of the second processing result.
[0193] Furthermore, the data processing method further includes: obtaining text features based on a set of reference site names; constructing a text classification model based on the text features; the text classification model being able to determine whether the input site name belongs to fuzzy resource information; before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, the method further includes: obtaining a third processing result of whether the site name to be processed belongs to fuzzy resource information using the text classification model; the step of obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result includes: obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result.
[0194] This ensures a more comprehensive and accurate final result.
[0195] The step of obtaining text features based on the set of reference site names includes: obtaining a forward dataset and a reverse dataset based on the set of reference site names; randomly shuffling the forward dataset and the reverse dataset to form a complete dataset, and obtaining a training set from the complete dataset; and using a pre-trained language model to obtain text features based on the training set. The forward dataset is the set of reference site names, and the reverse dataset includes a first reverse dataset and a second reverse dataset. The first reverse dataset is obtained by swapping any two characters within each reference site name in the set of reference site names. The second reverse dataset is obtained by randomly replacing any character in the forward dataset with any character from a preset set of characters.
[0196] This allows for accurate acquisition of text features, ensuring the precision of the text classification model. The complete dataset also includes a test set, which can be used to test the text classification model and further optimize it.
[0197] Furthermore, before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, the method further includes: obtaining the error rate corresponding to a preset operation based on the first site name that has been labeled; obtaining the corresponding weight coefficient based on the error rate; wherein, the preset operation is a first operation and the weight coefficient is a first weight coefficient; or, the preset operation is a second operation and the weight coefficient is a second weight coefficient; or, the preset operation is a third operation and the weight coefficient is a third weight coefficient; the first operation includes: obtaining the first site name corresponding to... The second operation includes: the second pinyin to be matched, and the second target pinyin of the second information in the second target site name corresponding to the first site name; based on the second pinyin to be matched and the second target pinyin, the result of whether the first site name belongs to fuzzy resource information is obtained; the second information is the text information in the second target site name excluding the text used for level classification; the second operation includes: segmenting the first site name into words to obtain at least one second word segmentation result; based on the at least one second word segmentation result and the reference site name, the result of whether the first site name belongs to fuzzy resource information is obtained; the third operation includes: using the text classification model to obtain the result of whether the first site name belongs to fuzzy resource information.
[0198] This further ensures the final result obtained subsequently. The "text used for hierarchical classification" can be text similar to that used for address hierarchical classification, such as city, county, district, village, etc.
[0199] In this embodiment of the invention, obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result includes: using a first formula, obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result; wherein, the first formula is:
[0200] ; Indicates the final result. This represents the weight coefficient corresponding to the k-th processing result. This represents the k-th processing result, x represents the name of the site to be processed, k is greater than or equal to 1, and k is less than or equal to the total number of processing results.
[0201] In this embodiment of the invention, the method may further include: indicating the location of the ambiguous resource information based on the first text, the third text, and / or the fourth text. This allows the user to directly know the specific location of the ambiguous resource information to perform related operations, such as deletion or correction.
[0202] In addition, it may include: providing online reminders based on the final result; this enables online reminders of errors when manually entering site names.
[0203] The data processing method provided in the embodiments of the present invention will be illustrated below, wherein the site name can be abbreviated as site name.
[0204] To address the aforementioned technical problems, and considering that a site name is composed of several of the following: the city, district / county, township, village, or community to which the site belongs, along with its specific location name (which can also be understood as a cascading of place names), this invention provides a data processing method. Specifically, it can be implemented as a site name verification method for fuzzy resource matching (or a comprehensive fuzzy resource discrimination method), which can automatically and intelligently locate fuzzy resources (corresponding to the aforementioned fuzzy resource information) in a resource table (also called a spatial resource table). This involves:
[0205] This solution is based on natural language processing technology in AI (artificial intelligence). It makes full use of the characteristic that the cascading of place names (the above-mentioned site names to be processed) can find standard naming (corresponding to the above-mentioned first target site name) references, and the characteristic that most site names in this cascading form will have overlapping place names. It can comprehensively analyze which resources in the resource table are fuzzy resources, and can also indicate the location of errors. It can achieve a high accuracy rate in judging fuzzy resources, and can also provide online reminders of errors when manually entering site resource names.
[0206] Specifically, the solution provided in this embodiment of the invention involves three sub-models (sub-model 1, sub-model 2, and sub-model 3) and the fusion of the three sub-models (to obtain the overall model), as detailed in the following four parts:
[0207] Part 1, Regarding Sub-model 1;
[0208] Sub-model 1 is designed to enhance the overall model's ability to identify errors in place names (i.e., site names) such as "Xinyang City Gushi Chenlinzi Agricultural Bank." The township to which this site belongs should be Chenlinzi Town; clearly, the character "Lin" was misspelled. Although the site name column (corresponding to the aforementioned site name) in the resource table is manually entered, the city, district / county, township, village, or community are dropdown options (corresponding to the first target site name), serving as a reference. The execution flow of sub-model 1 is as follows: Figure 2 As shown, it includes:
[0209] Operation 1. Remove the suffixes (corresponding to the characters used for level division) of the city, district, township, and village or community names (corresponding to the first target site name) of the site (corresponding to the site name to be processed above). After removal, the pinyin M1-Mn corresponding to each level of division is obtained (corresponding to the first target pinyin above, with each level corresponding to one M). In addition, obtain the pinyin N of the site name (corresponding to the site name to be processed above) (corresponding to the first pinyin to be matched above).
[0210] Operation 2. Determine whether there is any overlap between M1-Mn and N (corresponding to the above determination of whether there is any overlap between the first pinyin to be matched and the first target pinyin). If not (i.e., if there is no overlap), then this site name (i.e., this resource) is a non-fuzzy resource in sub-model 1 (corresponding to the first processing result of obtaining that the site name to be processed does not belong to the fuzzy resource information), and sub-model 1 ends; if yes (i.e., if there is an overlap), then obtain the text w (corresponding to the first text above) corresponding to the overlapping pinyin in N and the corresponding text w' (corresponding to the second text above) in M1-Mn.
[0211] Operation 3. Determine whether w and w' are the same (i.e., determine whether w=w' is true). If they are the same, then this site name (i.e. this resource) is a non-fuzzy resource in sub-model 1, and sub-model 1 ends; otherwise, proceed to operation 4, i.e., if they are different, proceed to the next operation.
[0212] Operation 4. Determine whether the length len(w) of w is greater than 1. If not (i.e., if not greater than 1), then this site name (i.e., this resource) is a non-fuzzy resource in sub-model 1, and sub-model 1 ends; if yes (i.e., if greater than 1), then this site name (i.e., this resource) is a fuzzy resource in sub-model 1, and sub-model 1 ends.
[0213] Part Two, Regarding Sub-Model 2;
[0214] Sub-model 2 aims to enhance the overall model by identifying errors such as "Luoyang Jianxi Wahaha Group Special Line" and "Luoyang Jianxi District Luoyang Wahaha Hengfan Beverage Co., Ltd." Clearly, one of the "Wahaha" or "Wahaha" in these two site names is incorrect. While these errors may not be explicitly listed in the city, district, township, or village (or community) names in the resource table, many correct site names can be compared, providing a useful comparative tool. The execution flow of sub-model 2 is as follows: Figure 3 As shown, it includes:
[0215] Step 1. Obtain a pre-provided list of site names (corresponding to the reference site names above) that is accurate.
[0216] Operation 2. Perform Chinese word segmentation on each site name (corresponding to the site name to be processed mentioned above) in the resource table to obtain each word s (corresponding to the first word segmentation result mentioned above). Then, aggregate s from left to right with a slider length of 2 (corresponding to the preset slider length mentioned above) to obtain s'. After that, obtain the pinyin of s and s'.
[0217] Operation 3. Determine if s or s' (corresponding to the second information above) has appeared before. If not (i.e., if it has not appeared), add L=[s or s', corresponding station name, longitude, latitude] to the list of the corresponding pinyin (corresponding to the word segmentation pinyin above), so that the number of different characters for the same pinyin corresponds to the number of Ls in the pinyin list; if yes (i.e., if it has appeared), do not process it (i.e., skip it), and repeat operations 1 and 2 until the resource table is traversed to obtain the complete pinyin list. Specifically, the above judgment operation and subsequent operations are performed on s and s' respectively.
[0218] Operation 4. Determine if the length of the (pinyin) list (i.e. the number of different character shapes) is greater than 1. If not, skip it. That is, if it is not greater than 1, it means that there is only one character shape and no processing is required. If it is greater than 1, then determine if the number of characters corresponding to this pinyin (corresponding to the above word segmentation pinyin) is greater than 2 (corresponding to the above second threshold). If not, skip it. That is, if it is not greater than 2, no processing is required. If it is greater than 2, then calculate the straight-line distance dis between each pair of stations in the list (corresponding to each of the two character shapes) based on latitude and longitude (corresponding to the above distance).
[0219] Operation 5. (Can be manually set) Given a distance threshold k (corresponding to the first threshold mentioned above), determine whether dis is less than k. If not, skip it; if not less than k, do not process it. If so (i.e., if less than k), classify these two resources into an incorrect group, i.e., one of them is a fuzzy resource. Specifically, it can be determined which site name in the list is in the correct site name list obtained in (Operation 1), and the rest are fuzzy resources (corresponding to the second processing result that the site name to be processed belongs to fuzzy resource information when the fourth text corresponding to the distance is included in the reference site name; and the second processing result that the site name to be processed does not belong to fuzzy resource information when the third text is included in the reference site name).
[0220] The execution flow of sub-model 2 will be illustrated in detail below.
[0221] For example, the resource table contains both the site names 'Luoyang City Jianxi Wahaha Group Special Line' and 'Luoyang City Jianxi District Luoyang Wahaha Hengfan Beverage Co., Ltd.'. During the traversal, the former (corresponding to the first target site name mentioned above) is read first. Performing Chinese word segmentation on it yields s: ['Luoyang City', 'Jian', 'Xi', 'Wahaha', 'Group', 'Special Line']. This allows us to obtain the pinyin for 'Wahaha': 'wa-ha-ha' (included in the first target pinyin). Since 'Wahaha' has not appeared in the previous traversal, and the pinyin 'wa-ha-ha' has not appeared either, L=['Wahaha', 'Luoyang City Jianxi Wahaha Group Special Line', this site's longitude, this site's latitude] is added to the 'wa-ha-ha' list.
[0222] During the traversal, 'Luoyang City, Jianxi District, Luoyang Wahaha Hengfan Beverage Co., Ltd.' (corresponding to the aforementioned site name) was encountered. Chinese word segmentation was performed on it, resulting in s: ['Luoyang City', 'Jianxi District', 'Luoyang', 'Wa', 'Haha', 'Hengfan', 'Beverage', 'Co., Ltd.']. Since 'Wa' and 'Haha' are separate and cannot be matched, a slider aggregation of length 2 was used to combine them into 'Wahaha' (corresponding to the second information above, which is the aggregation result of the first word segmentation result according to the preset slider length), resulting in the pinyin 'wa-ha-ha'. It was determined that 'Wahaha' had not appeared in the previous traversal, but the pinyin 'wa-ha-ha' had already appeared, so L=['Wahaha', 'Luoyang City, Jianxi District, Luoyang Wahaha Hengfan Beverage Co., Ltd.', this site's longitude, this site's latitude] was added to the 'wa-ha-ha' list.
[0223] At this point, the 'wa-ha-ha' list should be [['Wahaha', 'Luoyang Jianxi Wahaha Group Special Line', longitude of this station, latitude of this station], ['Wahaha', 'Luoyang Jianxi District Luoyang Wahaha Hengfan Beverage Co., Ltd.', longitude of this station, latitude of this station]]. We determine that the length of this list (corresponding to the total number of texts corresponding to the above word segmentation pinyin) is greater than 1 and the number of characters in 'Wahaha' is 3 (greater than 2, corresponding to the number of characters corresponding to the above word segmentation pinyin being greater than the second threshold). Based on the straight-line distance dis between these two stations in the latitude and longitude list (assumed to be 320 meters), and given a distance threshold k (assumed to be 500 meters), we determine that dis is less than k. Furthermore, based on the above list of correct station names, we know that 'Luoyang Jianxi Wahaha Group Special Line' is a correct station name. Therefore, 'Luoyang Jianxi District Luoyang Wahaha Hengfan Beverage Co., Ltd.' is a fuzzy resource.
[0224] The execution flow of sub-model 2 in this embodiment of the invention can be understood as obtaining characters with the same pinyin; and determining which are fuzzy resources based on the known correct characters.
[0225] Part Three, Regarding Sub-Model 3;
[0226] Sub-model 3 is designed to enhance the generality of the overall model, aiming to find all fuzzy resources not covered by the previous two sub-models. The specific process for obtaining sub-model 3 is as follows: Figure 4 As shown, it includes:
[0227] Operation 1. Construct the dataset. Obtain accurate site names (corresponding to the above set of reference site names) and label them as the forward dataset. Randomly swap the positions of Chinese characters in the forward dataset and label them as the reverse dataset (specifically, this may include randomly swapping the positions of two Chinese characters (belonging to one site name) in the forward dataset and labeling them as the (first) reverse dataset). Then, randomly replace a Chinese character in the forward dataset with one of the 3500 commonly used Chinese characters and label it as the (second) reverse dataset. After merging the forward and reverse datasets, randomly shuffle them to form the complete dataset. Divide the complete dataset into a training set and a test set.
[0228] Step 2. Obtain text features using a pre-trained language model. Any pre-trained language model can be used in this step. Here, we take the Chinese BERT model as an example. Replace the Chinese characters in the site name with the corresponding numerical sequence number in the BERT dictionary, and vectorize them to obtain the site text vector. Use the pre-trained Chinese BERT model to input the obtained site text vector into the Chinese BERT model to obtain the (text) features.
[0229] Step 3. Construct a classification model. After obtaining the (text) features, a text classification model can be constructed based on these features to perform binary classification (the basic models used here include, but are not limited to, fully connected neural networks, Text-CNN, Multi-LSTM, etc.). Obtain the probability that the text classification model classifies the input site name as a fuzzy resource. Finally, calculate the cross-entropy between the classification result and the actual label as the loss function. Minimize this loss function to train the text classification model, obtaining the fuzzy resource classification model (i.e., the final text classification model). Save the trained model. The model can then be tested using a test set for further optimization.
[0230] As shown above, all three sub-models have systemic significance, and choosing a good combination strategy to combine them for optimization is crucial. During the optimization process, the combination strategy can draw inspiration from the Adaboost weak learner combination concept in ensemble learning, enabling the overall model to adapt and automatically adjust the weights of each sub-model based on the data. For details on sub-model fusion, please refer to Section Four below.
[0231] Part Four, regarding the fusion of the three sub-models;
[0232] Part Four mainly involves combining strategy models to obtain the overall model; that is, the sub-models are merged into the overall model. Specifically, it can be seen that... Figure 5 As shown, more specifically, the following operations are involved:
[0233] Operation 1. First, use -1 to represent the fuzzy resource output results of each sub-model, and use 1 to represent the non-fuzzy resource output results of each sub-model. Then, take another batch of site names that have been manually labeled with fuzzy resource tags y (corresponding to the first batch of labeled site names mentioned above).
[0234] Operation 2. Retrieve the output results of each sub-model (That is, the input result of the kth sub-model is) ), x represents the input site name (corresponding to the first site name mentioned above, and can also correspond to the site name to be processed when using the overall model for data processing). (Corresponding to the output results of the first sub-model, the second sub-model, and the third sub-model above), calculate the output error rate of the k-th sub-model (corresponding to the error rate mentioned above), which can be achieved using the following formula:
[0235] ;
[0236] in, This represents the output error rate of the k-th sub-model. This represents the output of the k-th sub-model for the i-th input site name, where m is the total number of outputs from the k-th sub-model. This indicates the output result of the k-th sub-model. With the corresponding tags If the values are the same, the result is 1; otherwise, it is 0.
[0237] Step 3. Next, calculate the weight coefficients for each model, which can be done using the following formula:
[0238] ;
[0239] As can be seen from the above formula, the larger the error rate, the smaller the weight coefficient, and the smaller the error rate, the larger the weight coefficient.
[0240] in, This represents the weight coefficient of the k-th sub-model; This represents the output error rate of the k-th sub-model. Specifically, it can be that the weight coefficient of the 1st sub-model corresponds to the first weight coefficient mentioned above, the weight coefficient of the 2nd sub-model corresponds to the second weight coefficient mentioned above, and the weight coefficient of the 3rd sub-model corresponds to the third weight coefficient mentioned above.
[0241] Step 4. With the weight coefficients of each anomaly detection model (corresponding to the three sub-models mentioned above), the final combination strategy can be as follows (corresponding to the implementation using the first formula mentioned above):
[0242] ;
[0243] in, This represents the final result described above, where sign is the sign function. This represents the weight coefficient corresponding to the k-th processing result (i.e., the output result of the k-th sub-model). This represents the k-th processing result, x represents the name of the site to be processed, k is greater than or equal to 1, and k is less than or equal to the total number of processing results (for example, if there are three sub-models above, corresponding to a total of 3 processing results, then k is less than or equal to 3; more specifically, k can be equal to the total number of processing results, such as 3).
[0244] At this point, all the sub-models are combined together to form the final overall model output.
[0245] Furthermore, the solution provided in this embodiment of the invention can also indicate the location of the error based on sub-model 1 and / or sub-model 2, and can also provide online error reminders when manually entering site resource names (e.g., reminding you if there is an error after entering the name).
[0246] As described above, the solution provided in the embodiments of the invention involves: when determining fuzzy resources, three sub-models are first proposed, and then the three sub-models are merged to obtain the final overall model; so as to accurately determine which site names are fuzzy resources.
[0247] This invention also provides a data processing apparatus, such as... Figure 6 As shown, it includes:
[0248] The first acquisition module 61 is used to acquire the first pinyin to be matched corresponding to the site name to be processed, and the first target pinyin of the first information in the first target site name corresponding to the site name to be processed;
[0249] The first processing module 62 is used to obtain a first processing result as to whether the site name to be processed belongs to fuzzy resource information based on the first pinyin to be matched and the first target pinyin;
[0250] The second processing module 63 is used to segment the site name to be processed into words to obtain at least one first word segmentation result;
[0251] The third processing module 64 is used to obtain a second processing result as to whether the site name to be processed belongs to fuzzy resource information based on the at least one first word segmentation result and the reference site name;
[0252] The fourth processing module 65 is used to obtain the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result.
[0253] The first information refers to the text information in the name of the first target site, excluding the text used for level classification.
[0254] The data processing device provided in this embodiment of the invention obtains the first pinyin to be matched corresponding to the name of the site to be processed, and the first target pinyin of the first information in the first target site name corresponding to the name of the site to be processed; obtains a first processing result of whether the name of the site to be processed belongs to fuzzy resource information based on the first pinyin to be matched and the first target pinyin; segments the name of the site to be processed into words to obtain at least one first segmentation result; obtains a second processing result of whether the name of the site to be processed belongs to fuzzy resource information based on the at least one first segmentation result and a reference site name; and obtains a final result of whether the name of the site to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result; wherein, the first information is the text information in the first target site name excluding the text used for level classification; it can realize automated and intelligent analysis of whether the site name belongs to fuzzy resource information (such as whether there are text errors), improves processing efficiency, and reduces labor costs, and effectively solves the problems of low efficiency and high labor costs in the data processing scheme for fuzzy resources (corresponding to the above-mentioned fuzzy resource information) in the prior art.
[0255] The step of obtaining a first processing result regarding whether the site name to be processed belongs to fuzzy resource information based on the first pinyin to be matched and the first target pinyin includes: determining whether there is an overlap between the first pinyin to be matched and the first target pinyin; if there is no overlap, obtaining a first processing result indicating that the site name to be processed does not belong to fuzzy resource information; if there is an overlap, obtaining a first text corresponding to the overlap in the site name to be processed and a second text corresponding to the overlap in the first target site name; if the first text and the second text are the same, obtaining a first processing result indicating that the site name to be processed does not belong to fuzzy resource information; if the first text and the second text are different, obtaining the number of characters in the first text, and obtaining a first processing result indicating whether the site name to be processed belongs to fuzzy resource information based on the number of characters.
[0256] In this embodiment of the invention, the step of obtaining a first processing result of whether the site name to be processed belongs to fuzzy resource information based on the number of characters in the text includes: when the number of characters in the text is equal to 1, obtaining a first processing result that the site name to be processed does not belong to fuzzy resource information; and when the number of characters in the text is greater than 1, obtaining a first processing result that the site name to be processed belongs to fuzzy resource information.
[0257] The step of obtaining a second processing result regarding whether the site name to be processed belongs to fuzzy resource information based on at least one first word segmentation result and a reference site name includes: when the second information is not stored in a preset location, obtaining the word segmentation pinyin corresponding to the second information and using the second information as the third text corresponding to the word segmentation pinyin; when the total number of all texts corresponding to the word segmentation pinyin is greater than 1, determining that the word segmentation pinyin corresponds to the third text and at least one fourth text; obtaining the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts; when the distance is less than a first threshold, obtaining a second processing result regarding whether the site name to be processed belongs to fuzzy resource information based on the fourth text, the third text, and the reference site name corresponding to the distance; wherein the second information is the first word segmentation result, or the aggregation result of the first word segmentation result after aggregation according to a preset slider length.
[0258] In this embodiment of the invention, obtaining the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts includes: when the number of characters corresponding to the word segmentation pinyin is greater than a second threshold, obtaining the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts; wherein, the second threshold is equal to a preset slider length.
[0259] The step of obtaining a second processing result regarding whether the site name to be processed belongs to fuzzy resource information based on the fourth text, the third text, and the reference site name corresponding to the distance includes: obtaining a second processing result that the site name to be processed belongs to fuzzy resource information when the fourth text corresponding to the distance is included in the reference site name; and obtaining a second processing result that the site name to be processed does not belong to fuzzy resource information when the third text is included in the reference site name.
[0260] Furthermore, the data processing device further includes: a second acquisition module, used to acquire text features based on a set of reference site names; a first construction module, used to construct a text classification model based on the text features; the text classification model is capable of determining whether the input site name belongs to fuzzy resource information; the device further includes: a third acquisition module, used to acquire a third processing result of whether the site name to be processed belongs to fuzzy resource information using the text classification model before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result; obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result includes: obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result.
[0261] The step of obtaining text features based on the set of reference site names includes: obtaining a forward dataset and a reverse dataset based on the set of reference site names; randomly shuffling the forward dataset and the reverse dataset to form a complete dataset, and obtaining a training set from the complete dataset; and using a pre-trained language model to obtain text features based on the training set. The forward dataset is the set of reference site names, and the reverse dataset includes a first reverse dataset and a second reverse dataset. The first reverse dataset is obtained by swapping any two characters within each reference site name in the set of reference site names. The second reverse dataset is obtained by randomly replacing any character in the forward dataset with any character from a preset set of characters.
[0262] Furthermore, the data processing device further includes: a fourth acquisition module, configured to, before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processed result, the first weight coefficient corresponding to the first processed result, the second processed result, and the second weight coefficient corresponding to the second processed result, acquire an error rate corresponding to a preset operation based on the first site name with labels; a fifth acquisition module, configured to acquire a corresponding weight coefficient based on the error rate; wherein the preset operation is a first operation and the weight coefficient is a first weight coefficient; or, the preset operation is a second operation and the weight coefficient is a second weight coefficient; or, the preset operation is a third operation and the weight coefficient is a third weight coefficient; the first operation The operation includes: obtaining the second pinyin to be matched corresponding to the first site name, and the second target pinyin of the second information in the second target site name corresponding to the first site name; obtaining the result of whether the first site name belongs to fuzzy resource information based on the second pinyin to be matched and the second target pinyin; the second information is the text information in the second target site name excluding the text used for level classification; the second operation includes: segmenting the first site name into words to obtain at least one second word segmentation result; obtaining the result of whether the first site name belongs to fuzzy resource information based on the at least one second word segmentation result and the reference site name; the third operation includes: using the text classification model to obtain the result of whether the first site name belongs to fuzzy resource information.
[0263] In this embodiment of the invention, obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result includes: using a first formula, obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result; wherein, the first formula is:
[0264] ; Indicates the final result. This represents the weight coefficient corresponding to the k-th processing result. This represents the k-th processing result, x represents the name of the site to be processed, k is greater than or equal to 1, and k is less than or equal to the total number of processing results.
[0265] The implementation embodiments of the above data processing method are all applicable to the embodiments of the data processing device and can achieve the same technical effect.
[0266] This invention also provides a data processing device, such as... Figure 7 As shown, it includes: a processor 71 and a transceiver 72;
[0267] The processor 71 is used to obtain the first pinyin to be matched corresponding to the name of the site to be processed, and the first target pinyin of the first information in the first target site name corresponding to the name of the site to be processed;
[0268] Based on the first pinyin to be matched and the first target pinyin, a first processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information;
[0269] The name of the site to be processed is segmented into words to obtain at least one first segmentation result;
[0270] Based on the at least one first word segmentation result and the reference site name, a second processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information;
[0271] Based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, the final result of whether the site name to be processed belongs to fuzzy resource information is obtained.
[0272] The first information refers to the text information in the name of the first target site, excluding the text used for level classification.
[0273] In this embodiment of the invention, the transceiver is capable of communicating with the processor.
[0274] The data processing device provided in this embodiment of the invention obtains the first pinyin to be matched corresponding to the name of the site to be processed, and the first target pinyin of the first information in the first target site name corresponding to the name of the site to be processed; obtains a first processing result of whether the name of the site to be processed belongs to fuzzy resource information based on the first pinyin to be matched and the first target pinyin; segments the name of the site to be processed into words to obtain at least one first word segmentation result; obtains a second processing result of whether the name of the site to be processed belongs to fuzzy resource information based on the at least one first word segmentation result and a reference site name; and obtains a final result of whether the name of the site to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result; wherein, the first information is the text information in the first target site name excluding the text used for level classification; it can realize automated and intelligent analysis of whether the site name belongs to fuzzy resource information (such as whether there are text errors), improves processing efficiency, and reduces labor costs, and effectively solves the problems of low efficiency and high labor costs in the data processing scheme for fuzzy resources (corresponding to the above-mentioned fuzzy resource information) in the prior art.
[0275] The step of obtaining a first processing result regarding whether the site name to be processed belongs to fuzzy resource information based on the first pinyin to be matched and the first target pinyin includes: determining whether there is an overlap between the first pinyin to be matched and the first target pinyin; if there is no overlap, obtaining a first processing result indicating that the site name to be processed does not belong to fuzzy resource information; if there is an overlap, obtaining a first text corresponding to the overlap in the site name to be processed and a second text corresponding to the overlap in the first target site name; if the first text and the second text are the same, obtaining a first processing result indicating that the site name to be processed does not belong to fuzzy resource information; if the first text and the second text are different, obtaining the number of characters in the first text, and obtaining a first processing result indicating whether the site name to be processed belongs to fuzzy resource information based on the number of characters.
[0276] In this embodiment of the invention, the step of obtaining a first processing result of whether the site name to be processed belongs to fuzzy resource information based on the number of characters in the text includes: when the number of characters in the text is equal to 1, obtaining a first processing result that the site name to be processed does not belong to fuzzy resource information; and when the number of characters in the text is greater than 1, obtaining a first processing result that the site name to be processed belongs to fuzzy resource information.
[0277] The step of obtaining a second processing result regarding whether the site name to be processed belongs to fuzzy resource information based on at least one first word segmentation result and a reference site name includes: when the second information is not stored in a preset location, obtaining the word segmentation pinyin corresponding to the second information and using the second information as the third text corresponding to the word segmentation pinyin; when the total number of all texts corresponding to the word segmentation pinyin is greater than 1, determining that the word segmentation pinyin corresponds to the third text and at least one fourth text; obtaining the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts; when the distance is less than a first threshold, obtaining a second processing result regarding whether the site name to be processed belongs to fuzzy resource information based on the fourth text, the third text, and the reference site name corresponding to the distance; wherein the second information is the first word segmentation result, or the aggregation result of the first word segmentation result after aggregation according to a preset slider length.
[0278] In this embodiment of the invention, obtaining the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts includes: when the number of characters corresponding to the word segmentation pinyin is greater than a second threshold, obtaining the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts; wherein, the second threshold is equal to a preset slider length.
[0279] The step of obtaining a second processing result regarding whether the site name to be processed belongs to fuzzy resource information based on the fourth text, the third text, and the reference site name corresponding to the distance includes: obtaining a second processing result that the site name to be processed belongs to fuzzy resource information when the fourth text corresponding to the distance is included in the reference site name; and obtaining a second processing result that the site name to be processed does not belong to fuzzy resource information when the third text is included in the reference site name.
[0280] Furthermore, the processor is also configured to: obtain text features based on a set of reference site names; construct a text classification model based on the text features; the text classification model is capable of determining whether the input site name belongs to fuzzy resource information; the processor is also configured to: before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, obtain a third processing result of whether the site name to be processed belongs to fuzzy resource information using the text classification model; the step of obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result includes: obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result.
[0281] In this embodiment of the invention, obtaining text features based on a set of reference site names includes: obtaining a forward dataset and a reverse dataset based on the set of reference site names; randomly shuffling the forward dataset and the reverse dataset to form a complete dataset, and obtaining a training set from the complete dataset; using a pre-trained language model to obtain text features based on the training set; wherein, the forward dataset is the set of reference site names, and the reverse dataset includes a first reverse dataset and a second reverse dataset; the first reverse dataset is obtained by swapping any two characters in each reference site name in the set of reference site names; the second reverse dataset is obtained by randomly replacing any character in the forward dataset with any character from a preset set of characters.
[0282] Furthermore, the processor is also configured to: before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, obtain an error rate corresponding to a preset operation based on the first site name that has been labeled; obtain a corresponding weight coefficient based on the error rate; wherein the preset operation is a first operation and the weight coefficient is a first weight coefficient; or, the preset operation is a second operation and the weight coefficient is a second weight coefficient; or, the preset operation is a third operation and the weight coefficient is a third weight coefficient; the first operation includes: obtaining the first site name. The second operation involves: defining the corresponding second pinyin to be matched, and the second target pinyin of the second information in the second target site name corresponding to the first site name; determining whether the first site name belongs to fuzzy resource information based on the second pinyin to be matched and the second target pinyin; the second information being the text information in the second target site name excluding the text used for level classification; the second operation includes: segmenting the first site name to obtain at least one second segmentation result; determining whether the first site name belongs to fuzzy resource information based on the at least one second segmentation result and the reference site name; the third operation includes: using the text classification model to obtain the result of whether the first site name belongs to fuzzy resource information.
[0283] In this embodiment of the invention, obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result includes: using a first formula, obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result; wherein, the first formula is:
[0284] ; Indicates the final result. This represents the weight coefficient corresponding to the k-th processing result. This represents the k-th processing result, x represents the name of the site to be processed, k is greater than or equal to 1, and k is less than or equal to the total number of processing results.
[0285] The implementation embodiments of the above data processing method are all applicable to the embodiments of the data processing device and can achieve the same technical effect.
[0286] This invention also provides a data processing device, including a memory, a processor, and a program stored in the memory and executable on the processor; when the processor executes the program, it implements the above-described data processing method.
[0287] The implementation embodiments of the above data processing method are all applicable to the embodiments of the data processing device and can achieve the same technical effect.
[0288] This invention also provides a readable storage medium storing a program that, when executed by a processor, implements the steps in the data processing method described above.
[0289] The implementation embodiments of the above data processing methods are all applicable to the embodiments of the readable storage medium and can achieve the same technical effect.
[0290] It should be noted that many of the functional components described in this specification are referred to as modules in order to more specifically emphasize the independence of their implementation.
[0291] In this embodiment of the invention, the module can be implemented in software so that it can be executed by various types of processors. For example, an identified executable code module may include one or more physical or logical blocks of computer instructions, which may be constructed as objects, procedures, or functions. Nevertheless, the executable code of the identified module does not need to be physically located together, but may include different instructions stored in different bits, which, when logically combined, constitute the module and achieve the module's intended purpose.
[0292] In practice, an executable code module can be a single instruction or many instructions, and can even be distributed across multiple different code segments, different programs, and across multiple memory devices. Similarly, operational data can be identified within the module and can be implemented in any suitable form and organized within any suitable data structure. This operational data can be collected as a single dataset or distributed across different locations (including different storage devices), and can exist, at least in part, solely as electronic signals within the system or network.
[0293] When a module can be implemented using software, considering the current level of hardware technology, modules that can be implemented in software can be implemented using hardware circuits by those skilled in the art to achieve the corresponding functions, without considering cost. These hardware circuits include conventional very-large-scale integrated circuits (VLSI) or gate arrays, as well as existing semiconductors such as logic chips and transistors, or other discrete components. Modules can also be implemented using programmable hardware devices, such as field-programmable gate arrays, programmable array logic, and programmable logic devices.
[0294] The above describes the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A data processing method, characterized in that, include: Obtain the first pinyin to be matched corresponding to the name of the site to be processed, and the first target pinyin of the first information in the first target site name corresponding to the name of the site to be processed; Based on the first pinyin to be matched and the first target pinyin, a first processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information; The name of the site to be processed is segmented into words to obtain at least one first segmentation result; Based on the at least one first word segmentation result and the reference site name, a second processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information; Based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, the final result of whether the site name to be processed belongs to fuzzy resource information is obtained. Wherein, the first information is the text information in the name of the first target site, excluding the text used for level classification; The first processing result, which determines whether the site name to be processed belongs to fuzzy resource information based on the first pinyin to be matched and the first target pinyin, includes: Determine whether there is any overlapping pinyin between the first pinyin to be matched and the first target pinyin; In the absence of overlapping pinyin, the first processing result is that the name of the site to be processed does not belong to fuzzy resource information; In the case of overlapping pinyin, obtain the first text corresponding to the overlapping pinyin in the name of the site to be processed and the second text corresponding to the first target site name; If the first text and the second text are the same, the first processing result is that the site name to be processed does not belong to the fuzzy resource information; If the first text and the second text are different, the number of characters in the first text is obtained, and based on the number of characters, a first processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information.
2. The data processing method according to claim 1, characterized in that, The first processing result, which determines whether the site name to be processed belongs to fuzzy resource information based on the number of characters in the text, includes: When the number of characters in the text is equal to 1, the first processing result is that the name of the site to be processed does not belong to the fuzzy resource information. If the number of characters in the text is greater than 1, the first processing result is obtained that the name of the site to be processed belongs to fuzzy resource information.
3. The data processing method according to claim 1, characterized in that, The second processing result, which determines whether the site name to be processed belongs to fuzzy resource information based on at least one first word segmentation result and a reference site name, includes: If the second information is not stored in a preset location, obtain the word segmentation pinyin corresponding to the second information, and use the second information as the third text corresponding to the word segmentation pinyin; If the total number of all texts corresponding to the segmented pinyin is greater than 1, then the segmented pinyin corresponds to the third text and at least one fourth text; wherein, the fourth text is a text used to indicate the location of the fuzzy resource information; Obtain the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts; If the distance is less than the first threshold, a second processing result is obtained based on the fourth text, the third text, and the reference site name corresponding to the distance, to determine whether the site name to be processed belongs to fuzzy resource information. The second information is either the first word segmentation result or the aggregated result of the first word segmentation result after being aggregated according to a preset slider length.
4. The data processing method according to claim 3, characterized in that, The step of obtaining the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts includes: If the number of characters corresponding to the word segmentation pinyin is greater than the second threshold, obtain the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts; Wherein, the second threshold is equal to the preset slider length.
5. The data processing method according to claim 3, characterized in that, The second processing result, which determines whether the site name to be processed belongs to fuzzy resource information based on the fourth text, the third text, and the reference site name corresponding to the distance, includes: If the fourth text corresponding to the distance is contained in the reference site name, a second processing result is obtained in which the site name to be processed belongs to fuzzy resource information; If the third text is contained in the reference site name, a second processing result is obtained in which the site name to be processed does not belong to the fuzzy resource information.
6. The data processing method according to claim 1, characterized in that, Also includes: Based on the set of reference site names, obtain text features; Based on the text features, construct a text classification model; The text classification model can determine whether the input site name belongs to fuzzy resource information; Before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, the process further includes: Using the text classification model, a third processing result is obtained to determine whether the name of the site to be processed belongs to fuzzy resource information; The step of obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result includes: Based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result, the final result of whether the site name to be processed belongs to fuzzy resource information is obtained.
7. The data processing method according to claim 6, characterized in that, The step of obtaining text features based on the set of reference site names includes: Based on the set of reference site names, obtain the forward and reverse datasets; The forward and reverse datasets are randomly shuffled to form a complete dataset, and the training set in the complete dataset is obtained. Using a pre-trained language model, text features are obtained from the training set; The forward dataset is the set of reference site names, and the reverse dataset includes a first reverse dataset and a second reverse dataset. The first reverse dataset is obtained by swapping any two characters in each reference site name in the set of reference site names. The second reverse dataset is obtained by randomly replacing any character in the forward dataset with any character from a preset set of characters.
8. The data processing method according to claim 6, characterized in that, Before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, the process further includes: Based on the name of the first tagged site, obtain the error rate corresponding to the preset operation; Based on the error rate, obtain the corresponding weighting coefficient; Wherein, the preset operation is a first operation and the weight coefficient is a first weight coefficient; or, the preset operation is a second operation and the weight coefficient is a second weight coefficient; or, the preset operation is a third operation and the weight coefficient is a third weight coefficient. The first operation includes: obtaining the second pinyin to be matched corresponding to the first site name, and the second target pinyin of the second information in the second target site name corresponding to the first site name; obtaining the result of whether the first site name belongs to fuzzy resource information based on the second pinyin to be matched and the second target pinyin; the second information is the text information in the second target site name excluding the text used for level classification; The second operation includes: segmenting the first site name into words to obtain at least one second segmentation result; and determining whether the first site name belongs to fuzzy resource information based on the at least one second segmentation result and the reference site name. The third operation includes: using the text classification model to obtain the result of whether the first site name belongs to fuzzy resource information.
9. The data processing method according to claim 6, characterized in that, The step of obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result includes: Using the first formula, based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result, the final result of whether the site name to be processed belongs to fuzzy resource information is obtained; The first formula is: ; Indicates the final result. This represents the weight coefficient corresponding to the k-th processing result. This represents the k-th processing result, x represents the name of the site to be processed, k is greater than or equal to 1, and k is less than or equal to the total number of processing results.
10. A data processing apparatus, characterized in that, include: The first acquisition module is used to acquire the first pinyin to be matched corresponding to the site name to be processed, and the first target pinyin of the first information in the first target site name corresponding to the site name to be processed; The first processing module is used to obtain a first processing result as to whether the site name to be processed belongs to fuzzy resource information based on the first pinyin to be matched and the first target pinyin; The second processing module is used to segment the site name to be processed into words to obtain at least one first word segmentation result; The third processing module is used to obtain a second processing result as to whether the site name to be processed belongs to fuzzy resource information based on the at least one first word segmentation result and the reference site name; The fourth processing module is used to obtain the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result. Wherein, the first information is the text information in the name of the first target site, excluding the text used for level classification; The first processing result, which determines whether the site name to be processed belongs to fuzzy resource information based on the first pinyin to be matched and the first target pinyin, includes: Determine whether there is any overlapping pinyin between the first pinyin to be matched and the first target pinyin; In the absence of overlapping pinyin, the first processing result is that the name of the site to be processed does not belong to fuzzy resource information; In the case of overlapping pinyin, obtain the first text corresponding to the overlapping pinyin in the name of the site to be processed and the second text corresponding to the first target site name; If the first text and the second text are the same, the first processing result is that the site name to be processed does not belong to the fuzzy resource information; If the first text and the second text are different, the number of characters in the first text is obtained, and based on the number of characters, a first processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information.
11. The data processing apparatus according to claim 10, characterized in that, The first processing result, which determines whether the site name to be processed belongs to fuzzy resource information based on the number of characters in the text, includes: When the number of characters in the text is equal to 1, the first processing result is that the name of the site to be processed does not belong to the fuzzy resource information. If the number of characters in the text is greater than 1, the first processing result is obtained that the name of the site to be processed belongs to fuzzy resource information.
12. The data processing apparatus according to claim 10, characterized in that, The second processing result, which determines whether the site name to be processed belongs to fuzzy resource information based on at least one first word segmentation result and a reference site name, includes: If the second information is not stored in a preset location, obtain the word segmentation pinyin corresponding to the second information, and use the second information as the third text corresponding to the word segmentation pinyin; If the total number of all texts corresponding to the segmented pinyin is greater than 1, then the segmented pinyin corresponds to the third text and at least one fourth text; wherein, the fourth text is a text used to indicate the location of the fuzzy resource information; Obtain the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts; If the distance is less than the first threshold, a second processing result is obtained based on the fourth text, the third text, and the reference site name corresponding to the distance, to determine whether the site name to be processed belongs to fuzzy resource information. The second information is either the first word segmentation result or the aggregated result of the first word segmentation result after being aggregated according to a preset slider length.
13. The data processing apparatus according to claim 12, characterized in that, The step of obtaining the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts includes: If the number of characters corresponding to the word segmentation pinyin is greater than the second threshold, obtain the distance between the latitude and longitude position corresponding to the third text and the latitude and longitude position corresponding to each of the fourth texts; Wherein, the second threshold is equal to the preset slider length.
14. The data processing apparatus according to claim 13, characterized in that, The second processing result, which determines whether the site name to be processed belongs to fuzzy resource information based on the fourth text, the third text, and the reference site name corresponding to the distance, includes: If the fourth text corresponding to the distance is contained in the reference site name, a second processing result is obtained in which the site name to be processed belongs to fuzzy resource information; If the third text is contained in the reference site name, a second processing result is obtained in which the site name to be processed does not belong to the fuzzy resource information.
15. The data processing apparatus according to claim 10, characterized in that, Also includes: The second acquisition module is used to obtain text features based on the set of reference site names; The first construction module is used to construct a text classification model based on the text features; The text classification model can determine whether the input site name belongs to fuzzy resource information; The device further includes: The third acquisition module is used to obtain a third processing result of whether the site name to be processed belongs to fuzzy resource information by using the text classification model before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result. The step of obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result includes: Based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result, the final result of whether the site name to be processed belongs to fuzzy resource information is obtained.
16. The data processing apparatus according to claim 15, characterized in that, The step of obtaining text features based on the set of reference site names includes: Based on the set of reference site names, obtain the forward and reverse datasets; The forward and reverse datasets are randomly shuffled to form a complete dataset, and the training set in the complete dataset is obtained. Using a pre-trained language model, text features are obtained from the training set; The forward dataset is the set of reference site names, and the reverse dataset includes a first reverse dataset and a second reverse dataset. The first reverse dataset is obtained by swapping any two characters in each reference site name in the set of reference site names. The second reverse dataset is obtained by randomly replacing any character in the forward dataset with any character from a preset set of characters.
17. The data processing apparatus according to claim 15, characterized in that, Also includes: The fourth acquisition module is used to obtain the error rate corresponding to the preset operation based on the first site name that has been labeled, before obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result and the second weight coefficient corresponding to the second processing result. The fifth acquisition module is used to acquire the corresponding weight coefficient based on the error rate; Wherein, the preset operation is a first operation and the weight coefficient is a first weight coefficient; or, the preset operation is a second operation and the weight coefficient is a second weight coefficient; or, the preset operation is a third operation and the weight coefficient is a third weight coefficient. The first operation includes: obtaining the second pinyin to be matched corresponding to the first site name, and the second target pinyin of the second information in the second target site name corresponding to the first site name; obtaining the result of whether the first site name belongs to fuzzy resource information based on the second pinyin to be matched and the second target pinyin; the second information is the text information in the second target site name excluding the text used for level classification; The second operation includes: segmenting the first site name into words to obtain at least one second segmentation result; and determining whether the first site name belongs to fuzzy resource information based on the at least one second segmentation result and the reference site name. The third operation includes: using the text classification model to obtain the result of whether the first site name belongs to fuzzy resource information.
18. The data processing apparatus according to claim 15, characterized in that, The step of obtaining the final result of whether the site name to be processed belongs to fuzzy resource information based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result includes: Using the first formula, based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, the second weight coefficient corresponding to the second processing result, the third processing result, and the third weight coefficient corresponding to the third processing result, the final result of whether the site name to be processed belongs to fuzzy resource information is obtained; The first formula is: ; Indicates the final result. This represents the weight coefficient corresponding to the k-th processing result. This represents the k-th processing result, x represents the name of the site to be processed, k is greater than or equal to 1, and k is less than or equal to the total number of processing results.
19. A data processing device, characterized in that, include: Processor and transceiver; The processor is used to obtain the first pinyin to be matched corresponding to the name of the site to be processed, and the first target pinyin of the first information in the first target site name corresponding to the name of the site to be processed; Based on the first pinyin to be matched and the first target pinyin, a first processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information; The name of the site to be processed is segmented into words to obtain at least one first segmentation result; Based on the at least one first word segmentation result and the reference site name, a second processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information; Based on the first processing result, the first weight coefficient corresponding to the first processing result, the second processing result, and the second weight coefficient corresponding to the second processing result, the final result of whether the site name to be processed belongs to fuzzy resource information is obtained. Wherein, the first information is the text information in the name of the first target site, excluding the text used for level classification; The first processing result, which determines whether the site name to be processed belongs to fuzzy resource information based on the first pinyin to be matched and the first target pinyin, includes: Determine whether there is any overlapping pinyin between the first pinyin to be matched and the first target pinyin; In the absence of overlapping pinyin, the first processing result is that the name of the site to be processed does not belong to fuzzy resource information; In the case of overlapping pinyin, obtain the first text corresponding to the overlapping pinyin in the name of the site to be processed and the second text corresponding to the first target site name; If the first text and the second text are the same, the first processing result is that the site name to be processed does not belong to the fuzzy resource information. If the first text and the second text are different, the number of characters in the first text is obtained, and based on the number of characters, a first processing result is obtained as to whether the site name to be processed belongs to fuzzy resource information.
20. A data processing apparatus, comprising a memory, a processor, and a program stored in the memory and executable on the processor; characterized in that, When the processor executes the program, it implements the data processing method as described in any one of claims 1 to 9.
21. A readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the data processing method as described in any one of claims 1 to 9.