A big data-based sampling error dynamic calibration method and system

By constructing a sampling keyword library and a text library, evaluating the semantic connection strength and separation, and dynamically adjusting the clustering scale, the problem of semantic community division distortion in traditional methods is solved, and dynamic calibration and fault risk assessment of big data text sampling are realized.

CN120851033BActive Publication Date: 2025-11-25SUZHOU NOVOSENSE MICROELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511334126.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-11-25
Estimated Expiration
2045-09-18

AI Technical Summary

Technical Problem

Traditional text sampling methods in natural language processing cannot dynamically adjust cluster boundaries when faced with large-scale unstructured data, leading to distorted semantic community divisions and a lack of systematic constraints on semantic clustering results, resulting in inaccurate risk assessment of fault lines.

Method used

Establish a sampling keyword library and a text library, construct sampling dimension combination pairs with directed attributes to form a hierarchical sample set, evaluate the semantic connection strength, construct a two-dimensional coordinate system for semantic communities, dynamically adjust the clustering scale through distribution density and separation evaluation, generate the optimal semantic community classification result, and conduct a fault risk assessment.

Benefits of technology

It enables timely assessment of fault risk when semantic distribution changes, generates optimal semantic community classification results, balances classification precision with error calibration efficiency, and dynamically calibrates sampling error.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120851033B_ABST
    Figure CN120851033B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on big data's sampling error dynamic calibration method and system, belong to natural language processing technical field.Keyword library and sampling text library are established, sampling dimension combination pair with directed attribute is constructed, and sampling text is stratified according to this, to form stratified sample set, to evaluate the semantic connection strength of sampling dimension combination pair between sampling text, and then construct semantic community two-dimensional coordinate system;By setting clustering scale, calculating distribution density index, combining separation degree evaluation and semantic clustering constraint condition, dynamically adjust clustering scale, to balance classification precision and error calibration efficiency, generate optimal semantic community classification result, to separate the connection keyword between semantic community and carry out fault risk assessment and atlas display, realize the dynamic calibration of sampling error;Further capable of evaluating the fault risk between semantic community lagging behind semantic change in time when semantic distribution changes with data update.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology, specifically to a method and system for dynamic calibration of sampling errors based on big data. Background Technology

[0002] In text sampling and analysis scenarios within the field of natural language processing, traditional sampling methods and error calibration techniques generally suffer from insufficient assessment of semantic fault risk when dealing with large-scale unstructured data. This deficiency manifests itself in the following ways:

[0003] Traditional clustering algorithms typically use fixed scales or preset parameters for classification, which cannot dynamically adjust cluster boundaries according to data distribution characteristics. When the semantic distribution of sampled text changes with data updates, static clustering can lead to distortion of semantic community division. For example, in social media sentiment analysis, the evolution of semantic features may cause the original cluster boundaries to overlap or break. However, existing technologies lack a dynamic evaluation mechanism based on distribution density and separation, making it difficult to calibrate the cluster scale in real time, which in turn leads to the risk assessment of faults lagging behind semantic changes.

[0004] Meanwhile, traditional methods lack systematic constraints on semantic clustering results when assessing sampling errors, and do not establish quantitative indicators such as distribution density threshold and separation threshold. For example, when there are low-density semantic clusters or overlapping communities in the clustering results, existing technologies cannot eliminate invalid classifications through explicit constraints, resulting in inaccurate fault risk assessment results and difficulty in balancing classification precision and error calibration efficiency. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for dynamic calibration of sampling errors based on big data, so as to solve the problems mentioned in the background art.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] A sampling error dynamic calibration system based on big data, comprising: a database module, a sampling dimension processing module, a semantic analysis module, a sample visualization module, and a sampling error dynamic calibration module;

[0008] The database module is used to establish a sampling keyword library and a sampling text library;

[0009] The sampling dimension processing module constructs a pair of sampling dimensions with directed attributes based on the sampling keyword library, and thereby layers the sampled text to form a layered sample set.

[0010] The semantic analysis module evaluates the semantic connection strength between sampled texts based on the hierarchical sample set, so as to construct a two-dimensional coordinate system for semantic communities.

[0011] The visualization sample module is used to perform semantic clustering of the index in the two-dimensional coordinate system of the semantic community, and to depict a visualized semantic clustering circle to form a set of circular domains containing the index.

[0012] The sampling error dynamic calibration module is used to quantify the distribution density index and separation degree assessment, and dynamically adjust the clustering scale through semantic clustering constraints to generate the optimal semantic community classification result. It outputs and displays the sampled keywords within the semantic clustering circle and performs fault risk assessment.

[0013] Furthermore, the sampling dimension processing module includes a keyword connection unit and a sample stratification unit;

[0014] The keyword contact unit constructs sampling dimension combination pairs based on the sampled keyword library;

[0015] The sample stratification unit stratifies the sampled text based on the sampling dimension combination pairs, and identifies all the sampling keywords that exist between the sampling keywords in the sampled text to form a stratified sample set of the sampled text.

[0016] Furthermore, the semantic analysis module includes a semantic connection strength analysis unit and a scatter plotting unit;

[0017] The semantic connectivity strength analysis unit evaluates the semantic connectivity strength of sampled dimension combination pairs between sampled texts based on a hierarchical sample set.

[0018] The scattered characterization unit uses the keyword number as the index of the point coordinates to form a two-dimensional coordinate system for the semantic community, and the point value corresponding to the index is the semantic connection strength.

[0019] Furthermore, the visualization sample module includes a semantic clustering unit and a semantic clustering circle visualization unit;

[0020] The semantic clustering unit is used to perform semantic clustering of the index in the two-dimensional coordinate system of the semantic community, and to set the initial clustering scale;

[0021] The semantic clustering circle visualization unit uses the clustering scale as the radius of the semantic clustering circle, and depicts several semantic clustering circles in the two-dimensional coordinate system of the semantic community, and captures all the indices contained in the semantic clustering circle to form a circle domain set.

[0022] Furthermore, the sampling error dynamic calibration module includes a distribution density analysis unit, a sampling error dynamic calibration model unit, and an optimization display unit;

[0023] The distribution density analysis unit calculates the distribution density index of point values ​​within each circular domain set based on the circular domain set.

[0024] The sampling error dynamic calibration model unit is used to construct a sampling error dynamic calibration model, evaluate the separation degree between semantic clustering circles, and introduce a semantic clustering constraint model to generate the optimal semantic community classification result by iteratively optimizing the clustering scale.

[0025] The optimized display unit is used to display each sampled keyword within the optimized semantic clustering circle on the display terminal and to perform a fault risk assessment.

[0026] A dynamic calibration method for sampling errors based on big data, comprising the following steps:

[0027] Step S1: Establish a sampling keyword library and a sampling text library;

[0028] Step S2: Construct sampling dimension combination pairs with directed attributes based on the sampling keyword library, and stratify the sampled text accordingly to form a stratified sample set;

[0029] Step S3: Based on the hierarchical sample set, evaluate the semantic connection strength between sampled texts by combining sampling dimensions to construct a two-dimensional coordinate system for semantic communities;

[0030] Step S4: Perform semantic clustering in the two-dimensional coordinate system of semantic communities. By setting the clustering scale and calculating the distribution density index, and combining the separation degree evaluation and semantic clustering constraints, dynamically adjust the clustering scale to generate the optimal semantic community classification result. Based on the optimized semantic clustering circle, separate the connecting keywords between semantic communities and conduct a fault risk assessment to achieve dynamic calibration and display of sampling error.

[0031] Furthermore, the specific implementation process of step S1 includes:

[0032] Establish a sampling keyword library, which records several sampling keywords. Let any i-th sampling keyword be denoted as... The sampled keyword library is then represented as Where I represents the total number of sampled keywords;

[0033] Establish a sampled text library, which records a number of sampled texts. Let any x-th sampled text be denoted as... The sampled text library is then represented as , where Y represents the total number of sampled texts.

[0034] Furthermore, the specific implementation process of step S2 includes:

[0035] Based on the sampling keyword library, sampling dimension combination pairs are constructed, and the r-th sampling dimension combination pair is denoted as . Where i ≠ j, Let j represent the j-th sampled keyword, and let the combination of sampling dimensions have a directed attribute, i.e. ;

[0036] Based on the combination of sampling dimensions, the sampled text Perform stratification and in the sampled text Identifying sampling keywords With sampling keywords All existing sampling keywords are used to construct the sampling text. The hierarchical sample set, denoted as ,in, Indicates a combination of sampling dimensions The r-th hierarchical sample cluster obtained by identification, and .

[0037] Furthermore, the specific implementation process of step S3 includes:

[0038] Based on a hierarchical sample set, the semantic connectivity strength of sampled dimension combination pairs is evaluated among sampled texts:

[0039] ;

[0040] In the formula, Indicates the combination of sampling dimensions semantic connection strength, This represents the y-th sampled text. Indicates belonging to the hierarchical sample set Hierarchical sample clusters, Indicates belonging to the hierarchical sample set Hierarchical sample clusters, This represents the number of sampled keywords contained in the intersection set of the hierarchical sample clusters. This represents the number of sampled keywords contained in the union set of the hierarchical sample clusters. This represents the number of permutations and combinations among the sampled texts in the sampled text library, and , ! represents the factorial symbol;

[0041] Using (i, j) as the index of the point coordinates, a two-dimensional coordinate system for the semantic community is constructed, and the point value corresponding to the index (i, j) in the two-dimensional coordinate system of the semantic community is... .

[0042] Furthermore, the specific implementation process of step S4 includes:

[0043] In the semantic community two-dimensional coordinate system, semantic clustering is performed on the index, and the initial clustering scale is set, denoted as . ;

[0044] Clustering scale Let be the radius of the semantic clustering circle, and denote several semantic clustering circles in the two-dimensional coordinate system of the semantic community. Let the h-th semantic clustering circle be denoted as . Capture semantic clustering circles All indices (i, j) contained therein form a circular domain set, denoted as . ;

[0045] Based on circular region set Calculate the distribution density index of point values ​​within each circular domain set. ;

[0046] In the formula, Represents the circular field set The number of index points included;

[0047] Construct a dynamic calibration model for sampling errors:

[0048] Evaluate the h-th semantic clustering circle With the s-th semantic cluster circle The degree of separation between them, h≠s:

[0049] In the formula, m and n represent the coding sequence numbers of the sampling keywords, and (m, n) represents the coordinates of the points formed by the sampling keywords. Represents semantic clustering circles Distribution density index;

[0050] In the above method, the separation degree is used to characterize the classification clarity between circular domain sets. The numerator of the separation degree formula represents the minimum geometric distance between the index points in the two circular domain sets in the two-dimensional coordinate system of the semantic community. The denominator of the separation degree formula is the larger value of the distribution density of the two circular domain sets. By normalizing the distance and density, the interference of density difference on the separation degree is eliminated. The larger the distance value and the smaller the distribution density, the more obvious the separation of the two semantic categories is.

[0051] Introducing semantic clustering constraints:

[0052] ;

[0053] In the formula, The standard deviation of the distribution density index. Let H be the standard deviation of the separation, and H be the number of semantic cluster circles. The preset number of semantic community categories;

[0054] In the above method, the standard deviation of the distribution density index is used as a critical value to measure the density of semantic connection strength points within the circular domain set, ensuring that the samples within each classification cluster have sufficient relevance and cohesion, and avoiding invalid classification due to excessive clustering; the standard deviation of the separation degree is used as a critical value to measure the clarity of classification boundaries between different circular domain sets, avoiding overlap between different semantic community categories, and ensuring the uniqueness and distinguishability of the classification results; the preset number of semantic community classifications reflects the expected value of the number of semantic community classifications by users or business scenarios, reflecting the fineness of classification;

[0055] Dynamically adjust clustering scale The value of is used to iteratively optimize the clustering scale and generate the optimal semantic community classification result. , This represents the h-th semantic clustering circle after iterative optimization;

[0056] For the optimized semantic clustering circle, the connecting keywords between semantic communities are separated. In the formula, This represents the optimized s-th semantic clustering circle. Represents the optimized semantic clustering circle The formed circular domain set, Represents the optimized semantic clustering circle The resulting circular domain set;

[0057] Randomly select sampling keywords from the sampling keyword library Based on the separation of connecting keywords between semantic communities, the risk of line breaks is assessed when sampled keywords are used as connecting keywords between semantic communities:

[0058] ;

[0059] In the formula, Indicates sampling keywords Risk value of discontinuity when used as a connecting keyword between semantic communities , Denotes a counting function, if it satisfies Then let If satisfied Then let ;

[0060] Output the connection keywords between semantic communities and the fault risk values ​​corresponding to the connection keywords, and display the sampled map on the display terminal.

[0061] Compared with existing technologies, the beneficial effects achieved by this invention are as follows: This invention provides a dynamic calibration method and system for sampling errors based on big data. It establishes a sampling keyword library and a sampling text library, constructs sampling dimension combination pairs with directed attributes, and stratifies the sampling texts accordingly to form a stratified sample set. This assesses the semantic connection strength between sampling dimension combination pairs and the sampling texts, thereby constructing a two-dimensional coordinate system for semantic communities. By setting the clustering scale, calculating the distribution density index, and combining separation evaluation and semantic clustering constraints, the clustering scale is dynamically adjusted to balance classification precision and error calibration efficiency, generating optimal semantic community classification results. This allows for the separation of connecting keywords between semantic communities and the assessment of discontinuity risk and the visualization of the graph, achieving dynamic calibration of sampling errors. Furthermore, it enables timely assessment of discontinuity risk between semantic communities that lag behind semantic changes when the semantic distribution changes with data updates. Attached Figure Description

[0062] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0063] Figure 1 This is a schematic diagram illustrating the steps of a dynamic calibration method for sampling errors based on big data according to the present invention.

[0064] Figure 2 This is the execution logic diagram of a sampling error dynamic calibration method based on big data according to the present invention. Detailed Implementation

[0065] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0066] In this first embodiment: a sampling error dynamic calibration system based on big data is provided. The system includes: a database module, a sampling dimension processing module, a semantic analysis module, a visualization sample module, and a sampling error dynamic calibration module.

[0067] The database module is used to build a sampling keyword library and a sampling text library;

[0068] The sampling dimension processing module constructs a combination of sampling dimensions with directed attributes based on the sampling keyword library, and then layers the sampled text to form a layered sample set.

[0069] The sampling dimension processing module includes a keyword connection unit and a sample stratification unit.

[0070] Keyword contact unit, based on a sampled keyword library, constructs sampling dimension combination pairs;

[0071] The sample stratification unit, based on the combination of sampling dimensions, stratifies the sampled text and identifies all the sampling keywords that exist between the sampling keywords in the sampled text to form a stratified sample set of the sampled text.

[0072] The semantic analysis module evaluates the semantic connection strength between sampled texts based on the hierarchical sample set, in order to construct a two-dimensional coordinate system for semantic communities.

[0073] The semantic analysis module includes a semantic connection strength analysis unit and a scatter plotting unit.

[0074] The semantic connectivity strength analysis unit evaluates the semantic connectivity strength of sampled dimension combination pairs between sampled texts based on a hierarchical sample set.

[0075] The scattered characterization unit uses the keyword number as the index of the point coordinates to form a two-dimensional coordinate system for the semantic community, and the point value corresponding to the index is the semantic connection strength.

[0076] The visualization sample module is used to perform semantic clustering of the index in the two-dimensional coordinate system of the semantic community, and to depict the visualized semantic clustering circle, forming a set of circular domains containing the index;

[0077] The visualization sample module includes a semantic clustering unit and a semantic clustering circle visualization unit.

[0078] Semantic clustering unit, used to perform semantic clustering of indexes in a two-dimensional coordinate system of semantic community, and to set the initial clustering scale;

[0079] The semantic clustering circle visualization unit uses the clustering scale as the radius of the semantic clustering circle, and depicts several semantic clustering circles in the two-dimensional coordinate system of the semantic community. It also captures all the indices contained in the semantic clustering circle to form a circle domain set.

[0080] The sampling error dynamic calibration module is used to quantify the distribution density index and separation evaluation, and dynamically adjust the clustering scale through semantic clustering constraints to generate the optimal semantic community classification results. It outputs and displays the sampled keywords within the semantic clustering circle and performs fault risk assessment.

[0081] The sampling error dynamic calibration module includes a distribution density analysis unit, a sampling error dynamic calibration model unit, and an optimization display unit.

[0082] The distribution density analysis unit calculates the distribution density index of point values ​​within each circular domain set based on the circular domain set.

[0083] The sampling error dynamic calibration model unit is used to construct a sampling error dynamic calibration model, evaluate the separation between semantic clustering circles, and introduce a semantic clustering constraint model to iteratively optimize the clustering scale to generate the optimal semantic community classification result.

[0084] The optimized display unit is used to display each sampled keyword within the optimized semantic clustering circle on the display terminal and to conduct a fault risk assessment.

[0085] Please see Figures 1-2 In this second embodiment, a dynamic calibration method for sampling errors based on big data is provided to be applicable to the first embodiment described above. The specific implementation of the method of the present invention may include the following steps:

[0086] Step S1: Establish a sampling keyword library and a sampling text library;

[0087] For example, a sampling keyword library is established, which records several sampling keywords. Any i-th sampling keyword is denoted as... The sampled keyword library is then represented as Where I represents the total number of sampled keywords;

[0088] Establish a sampled text library containing several sampled texts. Let the x-th sampled text be denoted as . The sampled text library is then represented as Where Y represents the total number of sampled texts;

[0089] Step S2: Construct sampling dimension combination pairs with directed attributes based on the sampling keyword library, and stratify the sampled text accordingly to form a stratified sample set;

[0090] For example, based on the sampling keyword library, sampling dimension combination pairs are constructed, and the r-th sampling dimension combination pair is denoted as . Where i ≠ j, Let j represent the j-th sampled keyword, and let the combination of sampling dimensions have a directed attribute, i.e. ;

[0091] Based on the combination of sampling dimensions, the sampled text Perform stratification and in the sampled text Identifying sampling keywords With sampling keywords All existing sampling keywords are used to construct the sampling text. The hierarchical sample set, denoted as ,in, Indicates a combination of sampling dimensions The r-th hierarchical sample cluster obtained by identification, and ;

[0092] Involves semantic space construction:

[0093] For example, in e-commerce review analysis, firstly, sampling keywords such as "price," "quality," and "logistics" are extracted from massive product reviews to build a thesaurus, while reviews of different products are collected as a text library; by constructing dimensional combinations of directed attributes (such as "price → quality" and "quality → price" being considered different dimensions), reviews are layered according to keyword associations to form a hierarchical sample set containing semantic associations; for example, when analyzing mobile phone reviews, the dimensional combination of "battery life → charging speed" will identify review paragraphs that involve both keywords.

[0094] Step S3: Based on the hierarchical sample set, evaluate the semantic connection strength between sampled texts by combining sampling dimensions to construct a two-dimensional coordinate system for semantic communities;

[0095] For example, based on a hierarchical sample set, the semantic connectivity strength of sampled dimension combination pairs is evaluated among sampled texts:

[0096] ;

[0097] In the formula, Indicates the combination of sampling dimensions semantic connection strength, This represents the y-th sampled text. Indicates belonging to the hierarchical sample set Hierarchical sample clusters, Indicates belonging to the hierarchical sample set Hierarchical sample clusters, This represents the number of sampled keywords contained in the intersection set of the hierarchical sample clusters. This represents the number of sampled keywords contained in the union set of the hierarchical sample clusters. This represents the number of permutations and combinations among the sampled texts in the sampled text library, and , ! represents the factorial symbol;

[0098] Using (i, j) as the index of the point coordinates, a two-dimensional coordinate system for the semantic community is constructed, and the point value corresponding to the index (i, j) in the two-dimensional coordinate system of the semantic community is... ;

[0099] Involves semantic connectivity quantization:

[0100] Referring to the Jaccard similarity principle, the semantic connection strength between texts is calculated. Taking social media sentiment analysis as an example, for two articles discussing "environmental protection policies", the semantic relevance is quantified by calculating the ratio of the common keyword clusters to the total keyword clusters. The number of intersections in the formula represents common semantic features, and the number of unions represents differences. Finally, the semantic strength point value in the two-dimensional coordinate system is obtained by normalizing the number of permutations and combinations.

[0101] For example, in news text sampling, assuming Y = 100 technology news articles, when calculating the connection strength along the "artificial intelligence → machine learning" dimension, If the keyword clusters of two articles have an intersection of 3 and a union of 5, then .

[0102] Step S4: Perform semantic clustering in the two-dimensional coordinate system of semantic communities. By setting the clustering scale and calculating the distribution density index, and combining the separation degree evaluation and semantic clustering constraints, dynamically adjust the clustering scale to generate the optimal semantic community classification result. Based on the optimized semantic clustering circle, separate the connecting keywords between semantic communities and conduct a fault risk assessment to achieve dynamic calibration and display of sampling error.

[0103] For example, in the two-dimensional coordinate system of the semantic community, semantic clustering is performed on the index, and an initial clustering scale is set, denoted as . ;

[0104] Clustering scale Let be the radius of the semantic clustering circle, and denote several semantic clustering circles in the two-dimensional coordinate system of the semantic community. Let the h-th semantic clustering circle be denoted as . Capture semantic clustering circles All indices (i, j) contained therein form a circular domain set, denoted as . ;

[0105] Based on circular region set Calculate the distribution density index of point values ​​within each circular domain set. ;

[0106] In the formula, Represents the circular field set The number of index points included;

[0107] Construct a dynamic calibration model for sampling errors:

[0108] Evaluate the h-th semantic clustering circle With the s-th semantic cluster circle The degree of separation between them, h≠s:

[0109] In the formula, m and n represent the coding sequence numbers of the sampling keywords, and (m, n) represents the coordinates of the points formed by the sampling keywords. Represents semantic clustering circles Distribution density index;

[0110] Introducing semantic clustering constraints:

[0111] ;

[0112] In the formula, The standard deviation of the distribution density index. Let H be the standard deviation of the separation, and H be the number of semantic cluster circles. The preset number of semantic community categories;

[0113] For example, if a cluster circle contains 10 index points and the total semantic strength is 8, then the density K = 0.8; if another cluster circle has a density of 0.6 and the minimum distance between the two circles is 2, then the separation E = 2 / 0.8 = 2.5. At that time, it was considered that the classification was clear;

[0114] Dynamically adjust clustering scale The value of is used to iteratively optimize the clustering scale and generate the optimal semantic community classification result. , This represents the h-th semantic clustering circle after iterative optimization;

[0115] For the optimized semantic clustering circle, the connecting keywords between semantic communities are separated. In the formula, This represents the optimized s-th semantic clustering circle. Represents the optimized semantic clustering circle The formed circular domain set, Represents the optimized semantic clustering circle The resulting circular domain set;

[0116] Randomly select sampling keywords from the sampling keyword library Based on the separation of connecting keywords between semantic communities, the risk of line breaks is assessed when sampled keywords are used as connecting keywords between semantic communities:

[0117] ;

[0118] In the formula, Indicates sampling keywords Risk value of discontinuity when used as a connecting keyword between semantic communities , Denotes a counting function, if it satisfies Then let If satisfied Then let ;

[0119] Output the connection keywords between semantic communities and the fault risk values ​​corresponding to the connection keywords, and display the sampled map on the display terminal;

[0120] Involves dynamic clustering calibration:

[0121] For example, in a financial text risk assessment scenario, clustering circles are drawn in a semantic coordinate system with an initial clustering scale (e.g., radius 5), and the distribution density of semantic points within the circle (e.g., the total semantic intensity per unit area) is calculated. The separation degree between clustering circles (the ratio of minimum geometric distance to maximum density) is evaluated, combined with a preset density threshold (e.g., ... ) and resolution threshold (e.g. =0.5), dynamically adjust the clustering scale; for example, when overlapping categories appear in the financial news cluster, reduce the radius to improve classification accuracy, and finally locate the source of sampling bias by displaying the graph connecting keywords (such as "interest rate" and "stock market");

[0122] Involves an error feedback mechanism:

[0123] For example, in a medical literature retrieval scenario, if the initial sampling results do not cover enough literature related to "cancer treatment", the system identifies semantic gaps by connecting keywords between cluster circles (such as "targeted drugs → chemotherapy"), automatically adjusts the combination of sampling dimensions, supplements the missing keyword associations, and forms a closed-loop calibration.

[0124] For example, when analyzing 500,000 comments on a certain brand of mobile phone on Weibo, traditional sampling methods, because they do not consider the semantic relationship between "battery life" and "charging", mistakenly categorize some related comments as "performance". This invention constructs 120 directed dimensions such as "battery life → charging" to form 25 semantic communities. By evaluating the separation degree, "fast charging technology" is identified as a connecting keyword. After supplementing the sampling dimensions, the accuracy of related comment classification is improved.

[0125] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0126] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A dynamic calibration method for sampling errors based on big data, characterized in that, The method includes the following steps: Step S1: Establish a sampling keyword library and a sampling text library; Step S2: Construct sampling dimension combination pairs with directed attributes based on the sampling keyword library, and stratify the sampled text accordingly to form a stratified sample set; Step S3: Based on the hierarchical sample set, evaluate the semantic connection strength between sampled texts by combining sampling dimensions to construct a two-dimensional coordinate system for semantic communities; Step S4: Perform semantic clustering in the two-dimensional coordinate system of semantic communities. By setting the clustering scale and calculating the distribution density index, and combining the separation degree evaluation and semantic clustering constraints, dynamically adjust the clustering scale to generate the optimal semantic community classification result. Based on the optimized semantic clustering circle, separate the connecting keywords between semantic communities and conduct a fault risk assessment to achieve dynamic calibration and display of sampling error.

2. The method for dynamic calibration of sampling errors based on big data according to claim 1, characterized in that, The specific implementation process of step S1 includes: Establish a sampling keyword library, which records several sampling keywords. Let any i-th sampling keyword be denoted as... The sampled keyword library is then represented as Where I represents the total number of sampled keywords; Establish a sampled text library, which records a number of sampled texts. Let any x-th sampled text be denoted as... The sampled text library is then represented as , where Y represents the total number of sampled texts.

3. The method for dynamic calibration of sampling errors based on big data according to claim 2, characterized in that, The specific implementation process of step S2 includes: Based on the sampling keyword library, sampling dimension combination pairs are constructed, and the r-th sampling dimension combination pair is denoted as . Where i ≠ j, Let j represent the j-th sampled keyword, and let the combination of sampling dimensions have a directed attribute, i.e. ; Based on the combination of sampling dimensions, the sampled text Perform stratification and in the sampled text Identifying sampling keywords With sampling keywords All existing sampling keywords are used to construct the sampling text. The hierarchical sample set, denoted as ,in, Indicates a combination of sampling dimensions The r-th hierarchical sample cluster obtained by identification, and .

4. The method for dynamic calibration of sampling errors based on big data according to claim 3, characterized in that, The specific implementation process of step S3 includes: Based on a hierarchical sample set, the semantic connectivity strength of sampled dimension combination pairs is evaluated among sampled texts: ; In the formula, Indicates the combination of sampling dimensions semantic connection strength, This represents the y-th sampled text. Indicates belonging to the hierarchical sample set Hierarchical sample clusters, Indicates belonging to the hierarchical sample set Hierarchical sample clusters, This represents the number of sampled keywords contained in the intersection set of the hierarchical sample clusters. This represents the number of sampled keywords contained in the union set of the hierarchical sample clusters. This represents the number of permutations and combinations among the sampled texts in the sampled text library, and , ! represents the factorial symbol; Using (i, j) as the index of the point coordinates, a two-dimensional coordinate system for the semantic community is constructed, and the point value corresponding to the index (i, j) in the two-dimensional coordinate system of the semantic community is... .

5. The method for dynamic calibration of sampling errors based on big data according to claim 4, characterized in that, The specific implementation process of step S4 includes: In the semantic community two-dimensional coordinate system, semantic clustering is performed on the index, and the initial clustering scale is set, denoted as . ; Clustering scale Let be the radius of the semantic clustering circle, and denote several semantic clustering circles in the two-dimensional coordinate system of the semantic community. Let the h-th semantic clustering circle be denoted as . Capture semantic clustering circles All indices (i, j) contained therein form a circular domain set, denoted as . ; Based on circular region set Calculate the distribution density index of point values ​​within each circular domain set. ; In the formula, Represents the circular field set The number of index points included; Construct a dynamic calibration model for sampling errors: Evaluate the h-th semantic cluster circle With the s-th semantic cluster circle The degree of separation between them, h≠s: In the formula, m and n represent the coding sequence numbers of the sampling keywords, and (m, n) represents the coordinates of the points formed by the sampling keywords. Represents semantic clustering circles Distribution density index; Introducing semantic clustering constraints: ; In the formula, The standard deviation of the distribution density index Let H be the standard deviation of the separation, and H be the number of semantic cluster circles. The preset number of semantic community categories; Dynamically adjust clustering scale The value of is used to iteratively optimize the clustering scale and generate the optimal semantic community classification result. , This represents the h-th semantic clustering circle after iterative optimization; For the optimized semantic clustering circle, the connecting keywords between semantic communities are separated. In the formula, This represents the optimized s-th semantic clustering circle. Represents the optimized semantic clustering circle The formed circular domain set, Represents the optimized semantic clustering circle The resulting circular domain set; Randomly select sampling keywords from the sampling keyword library Based on the separation of connecting keywords between semantic communities, the risk of line breaks is assessed when sampled keywords are used as connecting keywords between semantic communities: ; In the formula, Indicates sampling keywords Risk value of discontinuity when used as a connecting keyword between semantic communities , Denotes a counting function, if it satisfies Then let If satisfied Then let ; Output the connection keywords between semantic communities and the fault risk values ​​corresponding to the connection keywords, and display the sampled map on the display terminal.

6. A sampling error dynamic calibration system based on big data, executing the sampling error dynamic calibration method according to any one of claims 1-5, characterized in that, The system includes: a database module, a sampling dimension processing module, a semantic analysis module, a sample visualization module, and a sampling error dynamic calibration module; The database module is used to establish a sampling keyword library and a sampling text library; The sampling dimension processing module constructs a pair of sampling dimensions with directed attributes based on the sampling keyword library, and thereby layers the sampled text to form a layered sample set. The semantic analysis module evaluates the semantic connection strength between sampled texts based on the hierarchical sample set, so as to construct a two-dimensional coordinate system for semantic communities. The visualization sample module is used to perform semantic clustering of the index in the two-dimensional coordinate system of the semantic community, and to depict a visualized semantic clustering circle to form a set of circular domains containing the index. The sampling error dynamic calibration module is used to quantify the distribution density index and separation degree assessment, and dynamically adjust the clustering scale through semantic clustering constraints to generate the optimal semantic community classification result. It outputs and displays the sampled keywords within the semantic clustering circle and performs fault risk assessment.

7. The sampling error dynamic calibration system based on big data according to claim 6, characterized in that, The sampling dimension processing module includes a keyword connection unit and a sample stratification unit; The keyword contact unit constructs sampling dimension combination pairs based on the sampled keyword library; The sample stratification unit stratifies the sampled text based on the sampling dimension combination pairs, and identifies all the sampling keywords that exist between the sampling keywords in the sampled text to form a stratified sample set of the sampled text.

8. The sampling error dynamic calibration system based on big data according to claim 7, characterized in that, The semantic analysis module includes a semantic connection strength analysis unit and a scatter plotting unit; The semantic connectivity strength analysis unit evaluates the semantic connectivity strength of sampled dimension combination pairs between sampled texts based on a hierarchical sample set. The scattered characterization unit uses the keyword number as the index of the point coordinates to form a two-dimensional coordinate system for the semantic community, and the point value corresponding to the index is the semantic connection strength.

9. A sampling error dynamic calibration system based on big data according to claim 8, characterized in that, The visualization sample module includes a semantic clustering unit and a semantic clustering circle visualization unit; The semantic clustering unit is used to perform semantic clustering of the index in the two-dimensional coordinate system of the semantic community, and to set the initial clustering scale; The semantic clustering circle visualization unit uses the clustering scale as the radius of the semantic clustering circle, and depicts several semantic clustering circles in the two-dimensional coordinate system of the semantic community, and captures all the indices contained in the semantic clustering circle to form a circle domain set.

10. A sampling error dynamic calibration system based on big data according to claim 9, characterized in that, The sampling error dynamic calibration module includes a distribution density analysis unit, a sampling error dynamic calibration model unit, and an optimization display unit. The distribution density analysis unit calculates the distribution density index of point values ​​within each circular domain set based on the circular domain set. The sampling error dynamic calibration model unit is used to construct a sampling error dynamic calibration model, evaluate the separation degree between semantic clustering circles, and introduce a semantic clustering constraint model to generate the optimal semantic community classification result by iteratively optimizing the clustering scale. The optimized display unit is used to display each sampled keyword within the optimized semantic clustering circle on the display terminal and to perform a fault risk assessment.

Citation Information

Patent Citations

  • Text sampling method and device, equipment and storage medium

    CN117149942A

  • Weak supervision indoor point cloud semantic segmentation method and device based on clustering thought and medium

    CN118135225A