System and device for identifying names of other persons selected by elected handwritten Chinese characters
By using multi-scale convolutional neural networks and generative adversarial networks to repair the stroke features of handwritten Chinese characters, combined with recurrent neural networks and weighted sorting of language models, high-precision recognition of handwritten names in election scenarios is achieved, solving the problems of complex writing styles and lack of contextual information, and improving recognition accuracy and efficiency.
Patent Information
- Application Number
- CN202510707971.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-12
AI Technical Summary
Existing handwritten Chinese character name recognition systems have difficulty achieving high-precision recognition when faced with complex writing styles and a lack of contextual information, resulting in a high recognition error rate and affecting the reliability of election results.
A multi-scale convolutional neural network is used to extract initial stroke features, a generative adversarial network is used to repair broken areas, a recurrent neural network is combined to analyze stroke sequences, an attention mechanism and a language model are used for weighted sorting, and a multimodal verification algorithm is used to determine the final set of name candidates.
It improves the accuracy and efficiency of handwritten name recognition in election scenarios, ensuring the reliability and accuracy of election results.
Smart Images

Figure CN120635916A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of recognition, and in particular relates to a system and device for recognizing the names of others selected by handwritten Chinese characters in elections. Background Art
[0002] During the election process, the recognition of voters' handwritten alternative names is a key technology that is directly related to the fairness and accuracy of the election. The importance of this field lies in ensuring that the true wishes of voters can be correctly recorded and reflected, especially in complex scenarios involving handwritten Chinese characters. The reliability of the recognition system becomes the core of ensuring the credibility of the democratic process. However, current recognition methods still have significant limitations when dealing with handwritten Chinese characters. Traditional solutions often rely on optical character recognition technology or simple template matching, but these methods often have difficulty adapting to the diversity of voters' writing styles, such as sloppy handwriting, deformed strokes, or differences in writing habits, resulting in high recognition error rates, especially when dealing with uncommon names. In addition, existing methods generally lack comprehensive analysis of contextual information or the frequency of name use, making it difficult to achieve high-precision recognition in complex election scenarios.
[0003] First, the extraction and quantification of stroke features are key difficulties. Due to the complex structure of Chinese characters and the different writing styles, it is difficult for the system to uniformly extract stable feature vectors. Secondly, the accuracy of name prediction is limited by the model's ability to integrate historical data and context. Especially in cases where handwriting is blurred or strokes are missing, the difficulty of inferring the correct name increases significantly. Finally, the reasonable allocation of feature weights in the quantitative model is also a major problem. The adaptability of weight settings in different writing scenarios directly affects the recognition effect. These unresolved technical factors result in the system often being unable to accurately distinguish similar Chinese characters or infer uncommon names when processing the names of other people, which in turn affects the reliability of the election results. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a system and device for recognizing the names of others selected from handwritten Chinese characters in elections, comprising:
[0005] The initial stroke extraction module is used to analyze the handwritten Chinese character image through a pre-established Chinese character stroke database, extract the initial stroke sequence, and use a multi-scale convolutional neural network to extract features from the initial stroke sequence to obtain a standardized stroke feature vector;
[0006] The stroke repair module is used to detect broken or discontinuous areas of the strokes based on the standardized stroke feature vector. If the number of broken points exceeds a preset threshold, the broken area is repaired using a generative adversarial network to obtain the repaired stroke feature vector;
[0007] The preliminary candidate generation module is used to analyze the time dependency of the stroke sequence based on the restored stroke feature vector using a recurrent neural network and combine it with the predefined stroke order rules of Chinese characters to obtain a preliminary candidate set of Chinese characters;
[0008] The weighted sorting module is used to calculate the frequency distribution of candidate Chinese characters in names by combining the preliminary candidate set of Chinese characters with the pre-established election scenario name database, and then use the attention mechanism to perform weighted sorting on the candidate Chinese characters to obtain a weighted Chinese character candidate sequence;
[0009] The semantic optimization module is used to analyze the semantic collocation rationality of the weighted Chinese character candidate sequence through the pre-trained language model if the highest confidence level is lower than the preset threshold, and obtain the optimized Chinese character candidate sequence;
[0010] A similarity calculation module is used to calculate the similarity between the optimized Chinese character candidate sequence and the names in the name database using a sequence matching algorithm based on the optimized Chinese character candidate sequence to obtain a preliminary name candidate set;
[0011] A weight adjustment module is used to obtain a final set of name candidates by combining the preliminary set of name candidates with the pre-collected frequency distribution of names in the election scenario area;
[0012] The multimodal verification module is used to use a multimodal verification algorithm for the final name candidate set, integrate the repaired stroke feature vector and the rationality of semantic matching, determine the confidence of the final name candidate set, and obtain the confirmed name record.
[0013] Preferably, the initial stroke extraction module includes:
[0014] A data acquisition submodule is used to acquire a handwritten Chinese character image, match the handwritten Chinese character image with a preset Chinese character stroke database, and obtain an initial stroke sequence;
[0015] An extraction submodule, configured to process the initial stroke sequence using a multi-scale convolutional neural network, extract multi-level features, and obtain a preliminary feature set;
[0016] a vector generation module, configured to generate a standardized feature vector by normalizing the preliminary feature set;
[0017] a judgment submodule, configured to determine if the dimension of the standardized feature vector meets a preset threshold, then reduce the dimension of the standardized feature vector using principal component analysis to obtain an optimized feature vector; if the dimension of the standardized feature vector does not meet the preset threshold, then adjust the structure of the multi-scale convolutional neural network to regenerate the standardized feature vector;
[0018] Preferably, the stroke repair module includes:
[0019] The initial judgment submodule is used to obtain the standardized stroke feature vector in a continuous manner, determine the regional continuity of the stroke through feature vector analysis, and obtain the initial continuity judgment result;
[0020] The number acquisition submodule is used to determine if the initial continuity judgment result shows that there is a discontinuous area, then perform breakpoint detection on the discontinuous area to obtain the number of breakpoints;
[0021] The processing submodule is used to compare the number of breakpoints with a preset threshold. If the number of breakpoints exceeds the preset threshold, a generative adversarial network is used to process the discontinuous area to obtain a feature vector after repairing the broken area.
[0022] The detection submodule is used to detect the regional continuity of the repaired vector through feature vector analysis and determine the integrity judgment result of the repaired vector;
[0023] The optimization submodule is used to obtain the integrity judgment result of the repaired vector, perform local optimization on the vector in the area with insufficient integrity, and obtain the optimized feature vector;
[0024] The repair submodule is used to judge the continuity and integrity of the final stroke feature vector based on the optimized feature vector and adopt a vector integrity verification algorithm to obtain the repaired stroke feature vector.
[0025] Preferably, the preliminary candidate generation module includes:
[0026] The feature sequence acquisition submodule is used to obtain sequence data from the repaired stroke feature vector using feature extraction technology to obtain a stroke feature sequence;
[0027] The dependency feature acquisition submodule is used to analyze the stroke feature sequence through a recurrent neural network, determine the time dependency relationship, and obtain the sequence dependency feature;
[0028] The preliminary candidate set acquisition submodule is used to match candidate Chinese characters based on sequence dependency features and predefined Chinese character stroke order rules to generate a preliminary candidate set;
[0029] The sorting submodule is used to determine if the number of Chinese characters in the preliminary candidate set exceeds a preset threshold, and then use the rule matching method to sort the candidate set and determine the priority sorting result.
[0030] Preferably, the weighted ranking module includes:
[0031] The frequency analysis submodule is used to obtain Chinese character data in a preset name database through a preliminary Chinese character candidate set, calculate the frequency distribution of the candidate Chinese characters in the name, and obtain the Chinese character frequency analysis results;
[0032] The data matching submodule is used to extract name data related to the candidate Chinese characters from a preset name database using database query processing based on the results of the Chinese character frequency analysis to obtain name data matching results;
[0033] The Chinese character feature extraction submodule is used to extract the context features of candidate Chinese characters based on the name data matching results and generate Chinese character feature extraction results;
[0034] The sorting weight calculation submodule is used to calculate the sorting weight of candidate Chinese characters through the Chinese character feature extraction results and apply the attention mechanism to obtain the sorting weight calculation results;
[0035] The candidate Chinese character sorting submodule is used to perform weighted sorting on the candidate Chinese characters if the weight value in the sorting weight calculation result meets the preset threshold value to obtain a candidate Chinese character sorting sequence;
[0036] The candidate set optimization submodule is used to optimize the preliminary Chinese character candidate set according to the candidate Chinese character sorting sequence and generate the candidate set optimization result;
[0037] The generation submodule is used to generate a weighted Chinese character candidate sequence through the candidate set optimization results and determine the final weighted Chinese character sequence.
[0038] Preferably, the semantic optimization module includes:
[0039] The context information acquisition submodule is used to determine if the highest confidence level of the weighted Chinese character sequence is lower than a preset threshold, and then obtain the context information of each Chinese character in the sequence to obtain an initial semantic feature representation;
[0040] The embedding vector generation submodule is used to encode the initial semantic feature representation through the pre-trained language model to obtain the semantic embedding vector;
[0041] The semantic collocation submodule is used to calculate the matching degree between semantic embedding vectors using the cosine similarity algorithm to determine the rationality of semantic collocation;
[0042] The candidate replacement submodule is used to determine if the semantic collocation rationality is lower than a preset threshold, and then obtain high-frequency synonymous Chinese characters from the vocabulary of the pre-trained language model to obtain a candidate replacement sequence;
[0043] The Chinese character candidate sequence optimization submodule is used to reorder the candidate replacement sequences through the sequence-to-sequence model to obtain the optimized Chinese character candidate sequence;
[0044] The final output submodule is used to optimize the candidate sequence of Chinese characters, calculate the semantic coherence of the entire sequence using a dynamic programming algorithm, and determine the final output sequence.
[0045] Preferably, the similarity calculation module includes:
[0046] A similarity submodule is used to calculate the similarity between the optimized Chinese character candidate sequence and the names in the name database using a sequence matching algorithm to obtain an initial name candidate set;
[0047] The name set submodule is used to determine if the similarity is higher than a preset threshold, and then include the corresponding name in the initial name candidate set to obtain a filtered name set;
[0048] The pinyin feature vector submodule is used to obtain the pinyin features of each name based on the filtered name set and generate a pinyin feature vector;
[0049] The clustering submodule is used to group the name set using the pinyin feature vector and a clustering algorithm to obtain the classified name subsets;
[0050] The subset calculation submodule is used to calculate the semantic feature vector of the name in each subset after classification and generate a semantic feature set;
[0051] The name candidate set submodule is used to determine if the semantic feature set matches the preset semantic template, and then retain the corresponding name subset to obtain the final preliminary name candidate set.
[0052] Preferably, the weight adjustment module includes:
[0053] The distribution comparison submodule is used to obtain the preliminary name candidate set and the regional name frequency distribution from the data source, and use statistical methods to calculate the distribution comparison between the two to obtain the preliminary matching degree;
[0054] The distribution comparison and extraction submodule is used to determine if the preliminary matching degree is lower than the preset threshold, and then extract the difference features through the distribution comparison results to obtain a feature set;
[0055] The probability distribution submodule is used to calculate the conditional probability of each candidate name based on the feature set using the Bayesian inference model to obtain the probability distribution;
[0056] The adjustment submodule is used to adjust the weight of each name in the preliminary candidate set through probability distribution to obtain an optimized weight set;
[0057] The sorting and screening submodule is used to sort and screen the preliminary candidate set according to the optimized weight set to obtain the adjusted candidate set;
[0058] The verification result calculation submodule is used to recalculate the matching degree with the regional frequency using the distribution comparison method for the adjusted candidate set, determine whether it meets the preset threshold, and obtain the verification result;
[0059] The verification submodule is used to determine if the verification result meets the preset threshold and output the adjusted candidate set as the final name candidate set. If not, the feature set is updated according to the verification result, and the Bayesian inference and weight adjustment are repeated to obtain the final name candidate set.
[0060] Preferably, the multimodal verification module includes:
[0061] A filtering submodule is used to obtain a final set of name candidates, filter the set using the candidate set screening rules to obtain a selected name list, and obtain a comprehensive feature vector based on the selected name list;
[0062] The confidence calculation submodule is used to calculate the confidence of the comprehensive feature vector through a multimodal verification algorithm to obtain a confidence calculation result;
[0063] The record confirmation submodule is used to select the name with the highest confidence based on the confidence calculation result and confirm the name record using the record confirmation mechanism;
[0064] The descending sorting submodule is used to use a sorting algorithm to sort the names in descending order of confidence if there are multiple names with high confidence in the confidence calculation results, so as to obtain the final confirmed name record.
[0065] On the other hand, the present invention further provides an electronic device, comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the system is implemented when the processor executes the computing program.
[0066] Compared with the prior art, the present invention has the following advantages and technical effects:
[0067] The present invention discloses a method for recognizing handwritten names in election scenarios. By pre-building a Chinese character stroke database and an election scenario name database, combined with deep learning technologies such as multi-scale convolutional neural networks, generative adversarial networks, and recurrent neural networks, it realizes stroke feature extraction, break repair, and sequence analysis of handwritten Chinese character images. At the same time, it integrates algorithms such as attention mechanisms, language models, and Bayesian reasoning to perform weighted sorting and optimization on candidate Chinese characters and names, and verifies them based on regional name frequency distribution. Finally, through a multimodal verification algorithm, the confidence level of the final name candidate set is determined, achieving accurate recognition of handwritten names in election scenarios, improving recognition accuracy and efficiency, and providing strong technical support for election work. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0069] Figure 1Schematic diagram of the system structure of an embodiment of the present invention;
[0070] Figure 2 is a first flow chart of an embodiment of the present invention;
[0071] Figure 3 This is a second flow chart of an embodiment of the present invention. DETAILED DESCRIPTION
[0072] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0073] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0074] Example 1
[0075] like Figure 1-3 As shown, this embodiment provides a system and device for recognizing the names of other people selected by handwritten Chinese characters in an election, including:
[0076] The initial stroke extraction module is used to analyze the handwritten Chinese character image through a pre-established Chinese character stroke database, extract the initial stroke sequence, and use a multi-scale convolutional neural network to extract features from the initial stroke sequence to obtain a standardized stroke feature vector;
[0077] The stroke repair module is used to detect broken or discontinuous areas of the strokes based on the standardized stroke feature vector. If the number of broken points exceeds a preset threshold, the broken area is repaired using a generative adversarial network to obtain the repaired stroke feature vector;
[0078] The preliminary candidate generation module is used to analyze the time dependency of the stroke sequence based on the restored stroke feature vector using a recurrent neural network and combine it with the predefined stroke order rules of Chinese characters to obtain a preliminary candidate set of Chinese characters;
[0079] The weighted sorting module is used to calculate the frequency distribution of candidate Chinese characters in names by combining the preliminary candidate set of Chinese characters with the pre-established election scenario name database, and then use the attention mechanism to perform weighted sorting on the candidate Chinese characters to obtain a weighted Chinese character candidate sequence;
[0080] The semantic optimization module is used to analyze the semantic collocation rationality of the weighted Chinese character candidate sequence through the pre-trained language model if the highest confidence level is lower than the preset threshold, and obtain the optimized Chinese character candidate sequence;
[0081] A similarity calculation module is used to calculate the similarity between the optimized Chinese character candidate sequence and the names in the name database using a sequence matching algorithm based on the optimized Chinese character candidate sequence to obtain a preliminary name candidate set;
[0082] A weight adjustment module is used to obtain a final set of name candidates by combining the preliminary set of name candidates with the pre-collected frequency distribution of names in the election scenario area;
[0083] The multimodal verification module is used to use a multimodal verification algorithm for the final name candidate set, integrate the repaired stroke feature vector and the rationality of semantic matching, determine the confidence of the final name candidate set, and obtain the confirmed name record.
[0084] Furthermore, the initial stroke extraction module includes:
[0085] A data acquisition submodule is used to acquire a handwritten Chinese character image, match the handwritten Chinese character image with a preset Chinese character stroke database, and obtain an initial stroke sequence;
[0086] An extraction submodule, configured to process the initial stroke sequence using a multi-scale convolutional neural network, extract multi-level features, and obtain a preliminary feature set;
[0087] a vector generation module, configured to generate a standardized feature vector by normalizing the preliminary feature set;
[0088] a judgment submodule, configured to determine if the dimension of the standardized feature vector meets a preset threshold, then reduce the dimension of the standardized feature vector using principal component analysis to obtain an optimized feature vector; if the dimension of the standardized feature vector does not meet the preset threshold, then adjust the structure of the multi-scale convolutional neural network to regenerate the standardized feature vector;
[0089] Obtain a handwritten Chinese character image, match the handwritten Chinese character image from a preset Chinese character stroke database, and obtain an initial stroke sequence. If the initial stroke sequence is complete, use image analysis technology to segment the region of the handwritten Chinese character image to obtain stroke distribution features; if the initial stroke sequence is incomplete, complete the missing strokes through the Chinese character stroke database to obtain the stroke distribution features. Process the stroke distribution features using a multi-scale convolutional neural network, extract multi-level features, and obtain a preliminary feature set. Generate a standardized feature vector by standardizing the preliminary feature set. If the dimension of the standardized feature vector meets the preset threshold, use principal component analysis to reduce the dimension of the standardized feature vector to obtain an optimized feature vector; if the dimension of the standardized feature vector does not meet the preset threshold, adjust the structure of the multi-scale convolutional neural network and regenerate the standardized feature vector. Match the optimized feature vector with the Chinese character stroke database to determine the Chinese character recognition result. Update the Chinese character stroke database based on the Chinese character recognition result to obtain a dynamically adjusted Chinese character stroke database.
[0090] Specifically, the handwritten Chinese character recognition technology involves image processing and machine learning. The core is to extract features from the handwritten image and match the stroke information in the database. The following analyzes and exemplifies each technical topic, focusing on the handwritten Chinese character recognition scenario. Obtaining a handwritten Chinese character image is to collect an image from a user input device such as a touch screen or a scanner.
[0091] For example, when a user writes the character "yong" on a tablet computer, the system captures the image with a resolution of 1024x1024 pixels in grayscale format for subsequent processing. This step ensures that the input data is clear and provides a basis for stroke matching. Matching the initial stroke sequence from the Chinese character stroke database requires a preset database containing common Chinese character strokes, such as "horizontal, vertical, left-falling stroke, right-falling stroke", etc.
[0092] Exemplarily, for the image of the character "yong", the system extracts the stroke contour through edge detection, compares it with the database, and initially identifies the sequence of "dot, horizontal, vertical, left-falling stroke, right-falling stroke". This step relies on the completeness of the database and helps to quickly locate the Chinese character structure. If the initial stroke sequence is complete, segment the image region through image analysis technology to obtain stroke distribution features.
[0093] In a possible implementation, for the character "yong", the system divides the image into an 8x8 grid, analyzes the stroke density and direction within each grid, generates a feature map, and records information such as "high dot density in the upper left corner" and "wide right-falling stroke distribution in the lower right corner". This segmentation can refine feature extraction and improve the subsequent recognition accuracy. If the initial stroke sequence is incomplete, for example, the "folding" stroke of the character "yong" is not recognized, it is completed through the database.
[0094] It should be noted that the system can infer the missing parts according to the context. For example, by combining the common structures of the character "yong", the character "zhe" is completed and the feature map is regenerated. This step ensures the integrity of the features and avoids misjudgment caused by ambiguous input. A multi-scale convolutional neural network is used to process the stroke distribution features and extract multi-level features. Specifically, the network contains 3 convolutional kernels with sizes of 3x3, 5x5, and 7x7 respectively, which capture fine strokes, local structures, and overall layouts.
[0095] Furthermore, the stroke repair module includes:
[0096] An initial judgment sub-module, which is used to continuously obtain the standardized stroke feature vector, determine the regional continuity of the strokes through feature vector analysis, and obtain the initial continuity judgment result;
[0097] A quantity acquisition sub-module, which is used to judge that if the initial continuity judgment result shows that there are discontinuous regions, then detect the break points for the discontinuous regions to obtain the number of break points;
[0098] A processing sub-module, which is used to compare the number of break points with a preset threshold. If the number of break points exceeds the preset threshold, then use a generative adversarial network to process the discontinuous regions to obtain the feature vector after repairing the broken regions;
[0099] A detection sub-module, which is used to detect the regional continuity of the repaired vector through feature vector analysis and determine the integrity judgment result of the repaired vector;
[0100] An optimization sub-module, which is used to obtain the integrity judgment result of the repaired vector, perform local vector optimization for the regions with insufficient integrity, and obtain the optimized feature vector;
[0101] A repair sub-module, which is used to judge the continuity and integrity of the final stroke feature vector according to the optimized feature vector by using a vector integrity verification algorithm, and obtain the repaired stroke feature vector.
[0102] Obtain the standardized stroke feature vector, determine the regional continuity of the stroke through feature vector analysis, and obtain the initial continuity judgment result. If the initial continuity judgment result shows the existence of a discontinuous area, perform breakpoint detection on the discontinuous area to obtain the number of breakpoints. Compare the number of breakpoints with the preset threshold. If the number of breakpoints exceeds the preset threshold, use a generative adversarial network to process the discontinuous area to obtain the feature vector after repairing the broken area. Detect the regional continuity of the repaired vector through feature vector analysis and determine the integrity judgment result of the repaired vector. Obtain the integrity judgment result of the repaired vector, perform local vector optimization on the area with insufficient integrity, and obtain the optimized feature vector. Based on the optimized feature vector, use a vector integrity verification algorithm to judge the continuity and integrity of the final stroke feature vector to obtain a verified feature vector. Generate the final stroke feature representation through the verified feature vector to determine the integrity of the stroke structure.
[0103] For example, after obtaining the standardized stroke feature vector, when analyzing the continuity of the stroke area, the coherence of the stroke can be determined by detecting the change in the distance between each point in the feature vector. Assuming that the feature vector represents the stroke trajectory of a Chinese character, the continuity analysis will check whether there are mutation points in the vector, such as the sudden interruption of the vector value of a certain trajectory. In one possible implementation, the distance threshold can be set to 0.05. If the distance between adjacent points exceeds this value, it is marked as a potential discontinuous area. This method can effectively identify the breaks in the strokes and provide a basis for subsequent processing.
[0104] In one embodiment, if the initial continuity determination result indicates the presence of discontinuous regions, the breakpoint detection further analyzes these regions. The number of breakpoints can be determined by scanning the feature vector and counting the number of sudden changes in the vector value.
[0105] For example, if three mutation points are detected in a stroke area, it can be considered that there are three break points.
[0106] It should be noted that the threshold of the number of breakpoints can be set according to the complexity of the Chinese characters, such as 2 for simple Chinese characters and 4 for complex Chinese characters, to adapt to different situations.
[0107] Specifically, if the number of breakpoints exceeds the threshold, for example, if 5 breakpoints are detected and the threshold is 3, you can use:
[0108] R(I)=α·M(I)+(1-α)·S(I);
[0109] R represents the restoration result, I represents the input image, M represents the main restoration content, S represents the structural restoration content, and α represents the fusion weight coefficient;
[0110]
[0111] L_GAN represents the loss function of the generative adversarial network, D represents the discriminator network, G represents the generator network, p_data represents the real data distribution, and p_z represents the random noise distribution. The generative adversarial network repairs discontinuous areas. The generative adversarial network learns the continuity characteristics of strokes and generates smooth trajectories to fill in the broken areas.
[0112] For example, for a broken horizontal stroke, the network can generate a smooth vector sequence based on the trends of the previous and subsequent strokes, making the repaired area look natural. This method can effectively restore the integrity of the stroke.
[0113] Furthermore, the continuity check of the repaired vector will reanalyze the vector to confirm the repair effect. This can be determined by comparing the smoothness of the vector before and after the repair.
[0114] For example, if there is a significant jump in the vector at the fracture before repair and the jump disappears after repair, and the trajectory is continuous, then the repair is considered successful. This verification ensures the reliability of the feature vector.
[0115] In a possible implementation, if there are still areas of insufficient integrity in the repaired vector, local optimization may be performed.
[0116] For example, if a stroke trajectory isn't smooth enough, the coordinates of key points in the vector can be adjusted to better align with the stroke's trajectory. If the vector's smoothness improves by 20% after optimization, the optimization is considered effective. This approach can further improve the quality of feature vectors.
[0117] It is understandable that when the vector integrity verification algorithm is finally adopted, it can be judged by checking the global continuity and local details of the vector.
[0118] For example, the verification algorithm analyzes the entire stroke trajectory to see if it conforms to the Chinese character structure, while also checking for minor breaks in local areas. If the verification results show 95% of the strokes are continuous, the vector is considered qualified. This verification ensures the accuracy of the feature vector.
[0119] Furthermore, the preliminary candidate generation module includes:
[0120] The feature sequence acquisition submodule is used to obtain sequence data from the repaired stroke feature vector using feature extraction technology to obtain a stroke feature sequence;
[0121] The dependency feature acquisition submodule is used to analyze the stroke feature sequence through a recurrent neural network, determine the time dependency relationship, and obtain the sequence dependency feature;
[0122] The preliminary candidate set acquisition submodule is used to match candidate Chinese characters based on sequence dependency features and predefined Chinese character stroke order rules to generate a preliminary candidate set;
[0123] Using feature extraction technology, sequence data is obtained from the repaired stroke feature vectors to get a stroke feature sequence. The stroke feature sequence is analyzed by a recurrent neural network to judge the time dependence relationship, and sequence dependence features are obtained. According to the sequence dependence features and combined with the predefined Chinese character stroke order rules, candidate Chinese characters are matched to generate a preliminary candidate set.
[0124] Specifically, obtaining sequence data from the repaired stroke feature vectors by feature extraction technology is a key link in Chinese character recognition.
[0125] Exemplarily, feature extraction can be based on the spatial distribution and direction information of strokes to decompose the vector into sequence data representing the starting and ending points, length, and angle of strokes.
[0126] For example, a certain stroke vector may be parsed into a sequence containing the starting coordinates x1, y1, the ending coordinates x2, y2, and the direction angle of 30 degrees. This method provides structured input for subsequent analysis by capturing the geometric characteristics of strokes.
[0127] In a possible implementation manner, the stroke feature sequence is analyzed by a recurrent neural network to judge the time dependence relationship. The recurrent neural network is good at processing sequence data and can capture the order and dynamic changes between strokes.
[0128] Specifically, for a group of sequence data, such as "horizontal, vertical, left-falling stroke", the network can analyze the probability of "horizontal" followed by "vertical" and generate dependence features reflecting the stroke order.
[0129] For example, inputting the stroke sequence of the character "field", the network may output the dependence relationship of "horizontal - vertical - horizontal - vertical - horizontal", indicating the timing logic between strokes.
[0130] It should be noted that the sequence dependence features are combined with the predefined Chinese character stroke order rules to be used for matching candidate Chinese characters.
[0131] Furthermore, the rule library contains the stroke orders of common Chinese characters, such as "dot, horizontal, vertical, left-falling stroke, right-falling stroke" of the character "eternal". By comparing the dependence features with the rule library, a preliminary candidate set is generated. <00002A data matching sub-module, which is used to extract name data related to candidate Chinese characters from a preset name database through database query processing according to the Chinese character frequency analysis result, and obtain a name data matching result;
[0135] A Chinese character feature extraction sub-module, which is used to extract the context features of candidate Chinese characters for the name data matching result and generate a Chinese character feature extraction result;
[0136] A sorting weight calculation sub-module, which is used to calculate the sorting weight of candidate Chinese characters through the Chinese character feature extraction result by applying the attention mechanism and obtain a sorting weight calculation result;
[0137] A candidate Chinese character sorting sub-module, which is used to perform weighted sorting on candidate Chinese characters if the weight value in the sorting weight calculation result meets a preset threshold, and obtain a candidate Chinese character sorting sequence;
[0138] A candidate set optimization sub-module, which is used to optimize the preliminary Chinese character candidate set according to the candidate Chinese character sorting sequence and generate a candidate set optimization result;
[0139] A generation sub-module, which is used to generate a weighted Chinese character candidate sequence through the candidate set optimization result and determine the final weighted Chinese character sequence.
[0140] Obtain Chinese character data in the preset name database through the preliminary Chinese character candidate set, count the frequency distribution of candidate Chinese characters in names, and obtain a Chinese character frequency analysis result. According to the Chinese character frequency analysis result, extract name data related to candidate Chinese characters from the preset name database through database query processing, and obtain a name data matching result. For the name data matching result, extract the context features of candidate Chinese characters and generate a Chinese character feature extraction result. Through the Chinese character feature extraction result, apply the attention mechanism to calculate the sorting weight of candidate Chinese characters and obtain a sorting weight calculation result. If the weight value in the sorting weight calculation result meets a preset threshold, perform weighted sorting on candidate Chinese characters to obtain a candidate Chinese character sorting sequence. Optimize the preliminary Chinese character candidate set according to the candidate Chinese character sorting sequence and generate a candidate set optimization result. Generate a weighted Chinese character candidate sequence through the candidate set optimization result and determine the final weighted Chinese character sequence.
[0141] Specifically, obtain Chinese character data in the preset name database through the preliminary Chinese character candidate set, count the frequency distribution of candidate Chinese characters in names, and obtain a Chinese character frequency analysis result.
[0142] For example, assume that the preliminary candidate set contains three Chinese characters: "Lin", "Lin", and "Lin". Query in the name database and find that the frequency of "Lin" in names is 60%, "Lin" is 30%, and "Lin" is 10%. This frequency distribution reflects the commonness of Chinese characters in names and provides a basis for subsequent screening.
[0143] Furthermore, the semantic optimization module includes:
[0144] The context information acquisition submodule is used to determine if the highest confidence level of the weighted Chinese character sequence is lower than a preset threshold, and then obtain the context information of each Chinese character in the sequence to obtain an initial semantic feature representation;
[0145] The embedding vector generation submodule is used to encode the initial semantic feature representation through the pre-trained language model to obtain the semantic embedding vector;
[0146] The semantic collocation submodule is used to calculate the matching degree between semantic embedding vectors using the cosine similarity algorithm to determine the rationality of semantic collocation;
[0147] The candidate replacement submodule is used to determine if the semantic collocation rationality is lower than a preset threshold, and then obtain high-frequency synonymous Chinese characters from the vocabulary of the pre-trained language model to obtain a candidate replacement sequence;
[0148] The Chinese character candidate sequence optimization submodule is used to reorder the candidate replacement sequences through the sequence-to-sequence model to obtain the optimized Chinese character candidate sequence;
[0149] The final output submodule is used to optimize the candidate sequence of Chinese characters, calculate the semantic coherence of the entire sequence using a dynamic programming algorithm, and determine the final output sequence.
[0150] If the highest confidence of the weighted Chinese character sequence is lower than the preset threshold, the context information of each Chinese character in the sequence is obtained to obtain the initial semantic feature representation. The initial semantic feature representation is encoded by the pre-trained language model to obtain a semantic embedding vector. The cosine similarity algorithm is used to calculate the matching degree between the semantic embedding vectors to judge the rationality of the semantic collocation. If the rationality of the semantic collocation is lower than the preset threshold, high-frequency synonymous Chinese characters are obtained from the vocabulary of the pre-trained language model to obtain a candidate replacement sequence. The candidate replacement sequence is reordered by the sequence-to-sequence model to obtain an optimized Chinese character candidate sequence. For the optimized Chinese character candidate sequence, a dynamic programming algorithm is used to calculate the semantic coherence of the entire sequence to determine the final output sequence. If the semantic coherence of the final output sequence is higher than the preset threshold, the result is output through the text generation interface to obtain the target text.
[0151] For example, when the highest confidence level of the weighted Chinese character sequence is lower than a preset threshold, the initial semantic feature representation can be obtained by analyzing the context information of the Chinese characters. The context information refers to the collocation relationship between the Chinese characters in the name.
[0152] For example, in a name database, assume that a certain candidate sequence contains "Li" and "Fang". The context information may include that "Li" often appears in combinations such as "beautiful" and "Lijuan", while "Fang" is more common in "Fanghua" and "Fangting". By counting the frequencies of these collocations, an initial semantic feature representation is generated to reflect the semantic tendencies of Chinese characters in the name scenario.
[0153] In one possible implementation, a pre-trained language model is used to encode the initial semantic feature representation to obtain a semantic embedding vector. The pre-trained language model has learned the semantic relationships of Chinese characters through large-scale texts.
[0154] For example, after inputting the context features of "Li" and "Fang", the model may generate high-dimensional vectors. Among them, the vector of "Li" is closer to semantics such as "elegant" and "beautiful", while "Fang" tends to "fragrant" and "gentle". These vectors capture the deep semantic associations of Chinese characters in names.
[0155] Specifically, the cosine similarity algorithm is used to calculate the matching degree between semantic embedding vectors to judge the rationality of semantic collocations. The cosine similarity judges whether two Chinese characters are semantically compatible by comparing the included angle of vectors.
[0156] For example, assume that the included angle between the vectors of "Li" and "Fang" is small, indicating that they are naturally collocated in the name. For example, "Lifang" sounds smooth and harmonious. If the included angle is large, such as between "Li" and "Gang", then they may be semantically不协调, and the matching degree is lower than the threshold.
[0157] Furthermore, when the rationality of semantic collocations is insufficient, high-frequency synonymous Chinese characters are obtained from the vocabulary of the pre-trained language model to generate candidate replacement sequences.
[0158] For example, the synonymous Chinese characters of "Li" may be "Mei" and "Yan", while the replacement options for "Fang" may include "Fen" and "Xiang". By analyzing the usage frequencies of these Chinese characters in the name database, more common combinations are screened out, such as "Meifang" or "Yanfen", to form candidate replacement sequences.
[0159] Furthermore, the similarity calculation module includes:
[0160] A similarity sub-module for calculating the similarity between the optimized Chinese character candidate sequence and the names in the name database using a sequence matching algorithm to obtain an initial name candidate set;
[0161] A name set sub-module for judging that if the similarity is higher than a preset threshold, the corresponding name is included in the initial name candidate set to obtain a screened name set;
[0162] A pinyin feature vector sub-module for obtaining the pinyin features of each name according to the screened name set and generating pinyin feature vectors;
[0163] A clustering sub-module, which is used to group a set of names through a pinyin feature vector by using a clustering algorithm to obtain a classified name subset;
[0164] A subset calculation sub-module, which is used to calculate the semantic feature vectors of the names in each subset for the classified name subsets and generate a semantic feature set;
[0165] A name candidate set sub-module, which is used to determine that if the semantic feature set matches a preset semantic template, the corresponding name subset is retained to obtain a final preliminary name candidate set.
[0166] Adopt a sequence matching algorithm to calculate the similarity between the optimized Chinese character candidate sequence and the names in the name database to obtain a preliminary name candidate set. If the similarity is higher than a preset threshold, the corresponding name is included in the preliminary candidate set to obtain a filtered name set. According to the filtered name set, obtain the pinyin features of each name and generate a pinyin feature vector. Through the pinyin feature vector, use a clustering algorithm to group the name set to obtain a classified name subset. For the classified name subsets, calculate the semantic feature vectors of the names in each subset to generate a semantic feature set. If the semantic feature set matches a preset semantic template, the corresponding name subset is retained to obtain an optimized name candidate set. According to the optimized name candidate set, use a sorting algorithm to sort in descending order of similarity and output the final name candidate sequence.
[0167] Specifically, the sequence matching algorithm is used to calculate the similarity between the optimized Chinese character candidate sequence and the names in the name database, aiming to screen out the names that best meet the conditions.
[0168] Exemplarily, assume that there is an optimized Chinese character candidate sequence containing Chinese character combinations such as "Zhang Wei" and "Li Qiang", and a large number of common names such as "Zhang Wei", "Li Qiang", and "Wang Fang" are stored in the name database. The sequence matching algorithm can calculate the similarity between the candidate sequence and the names in the database by comparing the character structures and combination patterns of the Chinese characters.
[0169] For example, adopt the edit distance algorithm to measure the number of characters that need to be replaced, inserted, or deleted between two sequences. If "Zhang Wei" is exactly the same as "Zhang Wei" in the database, the similarity is 100%, while compared with "Zhang Wei", the similarity may drop to 90% due to the difference between "Wei" and "Wei".
[0170] Furthermore, set the threshold to 85%, and the names with a similarity higher than this value are included in the preliminary candidate set. This method can quickly filter out the names that obviously do not match.
[0171] In a possible implementation, after screening the preliminary candidate set, it is necessary to obtain the pinyin features of each name to generate a pinyin feature vector. <0,
[0172] Specifically, for the name "Zhang Wei", its pinyin is "zhāngwěi", and features such as initials, finals, and tones can be extracted to form a multi-dimensional vector.
[0173] For example, the initial of "zhāng" is "zh", the final is "ang", and the tone is the first tone; the initial of "wěi" is "w", the final is "ei", and the tone is the third tone. These features are encoded into a vector, such as [zh,ang,1,w,ei,3]. In this way, the pinyin feature vector not only retains the pronunciation information but also facilitates subsequent calculations.
[0174] It can be understood that this vectorized representation provides a basis for the clustering algorithm. Using the clustering algorithm to group the name set can group names with similar pronunciations into one category.
[0175] For example, based on the K-means clustering algorithm, set the number of clusters to 3, and divide names with similar pinyin feature vectors into the same subset.
[0176] In one embodiment, the names "Zhang Wei" and "Zhang Wei" are assigned to the same subset because their pinyins "zhāngwěi" and "zhāngwěi" are highly similar; while "Li Qiang" (lǐqiáng) is assigned to another subset. This grouping method helps to distinguish names with similar pronunciations but different Chinese characters, facilitating subsequent semantic analysis. For the classified name subsets, calculate semantic feature vectors to generate a semantic feature set.
[0177] It should be noted that the semantic feature vector reflects the meaning of the name in a specific context.
[0178] For example, through a pre-trained word vector model, map "Zhang Wei" to a high-dimensional vector to capture its semantic characteristics as a common name. Assuming that the vectors of "Zhang Wei" and "Zhang Wei" are close in the high-dimensional space, it indicates that they are semantically similar.
[0179] Furthermore, if the preset semantic template requires the name to have the characteristics of "common and concise", the subset that matches the template is retained.
[0180] For example, the subset where "Zhang Wei" is located is retained because it meets the template requirements, while the subset containing rare names may be excluded.
[0181] In one embodiment, the optimized name candidate set is sorted in descending order of similarity by a sorting algorithm, and the final sequence is output.
[0182] Specifically, assume that in the candidate set, the similarity of "Zhang Wei" is 95%, the similarity of "Li Qiang" is 90%, and the similarity of "Zhang Wei" is 88%. After sorting, the output sequence is "Zhang Wei, Li Qiang, Zhang Wei". This arrangement ensures that the most matching names are presented first.
[0183] For example, in practical applications, the system may preferentially recommend "Zhang Wei" because it is highly consistent with the input sequence and is common. The advantage of this method is that it can intuitively present the optimal option.
[0184] Furthermore, the weight adjustment module includes:
[0185] A distribution comparison sub-module, which is used to obtain a preliminary name candidate set and a regional name frequency distribution from a data source, calculate the distribution comparison between the two using statistical methods, and obtain a preliminary matching degree;
[0186] A distribution comparison extraction sub-module, which is used to determine that if the preliminary matching degree is lower than a preset threshold, extract differential features through the distribution comparison result to obtain a feature set;
[0187] A probability distribution sub-module, which is used to calculate the conditional probability of each candidate name using a Bayesian inference model according to the feature set to obtain a probability distribution;
[0188] An adjustment sub-module, which is used to adjust the weight of each name in the preliminary candidate set through the probability distribution to obtain an optimized weight set;
[0189] A sorting and screening sub-module, which is used to sort and screen the preliminary candidate set according to the optimized weight set to obtain an adjusted candidate set;
[0190] A verification result calculation sub-module, which is used to calculate the matching degree with the regional frequency again using the distribution comparison method for the adjusted candidate set, determine whether it meets the preset threshold, and obtain a verification result;
[0191] A verification sub-module, which is used to determine that if the verification result meets the preset threshold, output the adjusted candidate set as the final name candidate set, and if it does not meet, update the feature set according to the verification result, repeat the Bayesian inference and weight adjustment to obtain the final name candidate set.
[0192] In a possible implementation manner, when obtaining the preliminary name candidate set and the regional name frequency distribution from a data source, it usually involves a large-scale name database and regional population statistical data.
[0193] For example, the data source may include a list of names in the household register information of a certain city and a regional name frequency distribution table based on the census statistics. The preliminary candidate set may include 1,000 common names such as "Zhang Wei" and "Li Na", and the regional frequency distribution reflects the occurrence probability of these names in a specific region. For example, "Zhang Wei" accounts for 10% in northern cities and "Li Na" accounts for 5%. When obtaining these data, it is necessary to ensure the integrity and representativeness of the data to avoid affecting subsequent analysis due to sample bias.
[0194] Specifically, when using statistical methods to calculate distribution comparison, the degree of match can be quantified by comparing the difference between the name frequency and the region frequency of the preliminary candidate set.
[0195] For example, if "Wang Qiang" accounts for 8% of the initial candidate set, while the regional frequency shows it to be only 3%, the name is a low match. Statistical methods may generate a preliminary match score based on the sum of the absolute values of the frequency differences. If the score falls below a preset threshold, such as 60%, differential feature extraction is triggered. Differential features may include surname preference, name length, or common character combinations, such as whether the surname "Wang" is frequent in the region.
[0196] In one embodiment, after extracting the difference features, a feature set is generated, which may include the frequency of surname occurrence, the number of characters in the name, the pinyin tone, etc.
[0197] For example, the feature set might show that two-word names account for 70% of the candidate set, while the region favors single-word names. Based on this, the Bayesian inference model can calculate the conditional probability of each name.
[0198] Furthermore, the multimodal verification module includes:
[0199] A filtering submodule is used to obtain a final set of name candidates, filter the set using the candidate set screening rules to obtain a selected name list, and obtain a comprehensive feature vector based on the selected name list;
[0200] The confidence calculation submodule is used to calculate the confidence of the comprehensive feature vector through a multimodal verification algorithm to obtain a confidence calculation result;
[0201] The record confirmation submodule is used to select the name with the highest confidence based on the confidence calculation result and confirm the name record using the record confirmation mechanism;
[0202] The descending sorting submodule is used to use a sorting algorithm to sort the names in descending order of confidence if there are multiple names with high confidence in the confidence calculation results, so as to obtain the final confirmed name record.
[0203] For example, filtering rules can be based on conditions such as name length and regional preferences. For example, if the candidate set contains 100 names, 50 names with a length of 2-3 characters and that are consistent with regional culture will be retained after filtering to form a selected list of names. This approach ensures that the list is more in line with actual needs.
[0204] In one possible implementation, a stroke feature extraction tool is used to extract stroke feature vectors from a selected name list.
[0205] Specifically, the number of strokes of each Chinese character can be quantified as a numerical vector. For example, "Zhang Wei" can be extracted as a vector of [11, 11]. To ensure accuracy, the visual features of the name can be supplemented by combining the glyph structure, such as the radical. This extraction method provides a reliable data basis for subsequent analysis.
[0206] It should be noted that the denoising algorithm is used to repair abnormal data in the stroke feature vector.
[0207] For example, if the number of strokes of a Chinese character is incorrect due to an incomplete font library during the extraction process, such as "Wei" being mislabeled as 10 strokes, the denoising algorithm can correct it to 11 strokes by comparing with the standard font library.
[0208] Furthermore, smoothing processing can be introduced to eliminate the微小偏差 caused by manual input and ensure the integrity of the feature vector.
[0209] Specifically, the semantic analysis model is used to calculate the rationality of name semantic collocations.
[0210] In one embodiment, the cultural meaning of the name combination can be analyzed through a pre-trained language model. For example, "Zhang Wei" may be interpreted as an image of simplicity and stability with a relatively high score, while "Zhang Long" may have a slightly lower score because the character "Long" is too strong. Assuming the threshold is 0.8 and "Zhang Wei" gets 0.85, it meets the condition. This analysis ensures that the name is more coordinated semantically. <000043On the other hand, this embodiment further provides an electronic device, including a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the system is implemented when the processor executes the computing program.
[0217] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A handwritten Chinese character name recognition system for elections, characterized in that: include: The initial stroke extraction module is used to analyze the handwritten Chinese character image through a pre-established Chinese character stroke database, extract the initial stroke sequence, and use a multi-scale convolutional neural network to extract features from the initial stroke sequence to obtain a standardized stroke feature vector; The stroke repair module is used to detect broken or discontinuous areas of the strokes based on the standardized stroke feature vector. If the number of broken points exceeds a preset threshold, the broken area is repaired using a generative adversarial network to obtain the repaired stroke feature vector; The preliminary candidate generation module is used to analyze the time dependency of the stroke sequence based on the restored stroke feature vector using a recurrent neural network and combine it with the predefined stroke order rules of Chinese characters to obtain a preliminary candidate set of Chinese characters; The weighted sorting module is used to calculate the frequency distribution of candidate Chinese characters in names by combining the preliminary candidate set of Chinese characters with the pre-established election scenario name database, and then use the attention mechanism to perform weighted sorting on the candidate Chinese characters to obtain a weighted Chinese character candidate sequence; The semantic optimization module is used to analyze the semantic collocation rationality of the weighted Chinese character candidate sequence through the pre-trained language model if the highest confidence level is lower than the preset threshold, and obtain the optimized Chinese character candidate sequence; A similarity calculation module is used to calculate the similarity between the optimized Chinese character candidate sequence and the names in the name database using a sequence matching algorithm based on the optimized Chinese character candidate sequence to obtain a preliminary name candidate set; A weight adjustment module is used to obtain a final set of name candidates by combining the preliminary set of name candidates with the pre-collected frequency distribution of names in the election scenario area; The multimodal verification module is used to use a multimodal verification algorithm for the final name candidate set, integrate the repaired stroke feature vector and the rationality of semantic matching, determine the confidence of the final name candidate set, and obtain the confirmed name record.
2. The system according to claim 1, wherein: The initial stroke extraction module includes: A data acquisition submodule is used to acquire a handwritten Chinese character image, match the handwritten Chinese character image with a preset Chinese character stroke database, and obtain an initial stroke sequence; An extraction submodule, configured to process the initial stroke sequence using a multi-scale convolutional neural network, extract multi-level features, and obtain a preliminary feature set; a vector generation module, configured to generate a standardized feature vector by normalizing the preliminary feature set; The judgment submodule is used to determine whether the dimension of the standardized feature vector meets the preset threshold, and then use principal component analysis to reduce the dimension of the standardized feature vector to obtain an optimized feature vector; if the dimension of the standardized feature vector does not meet the preset threshold, adjust the structure of the multi-scale convolutional neural network and regenerate the standardized feature vector.
3. The system according to claim 1, wherein: The stroke repair module includes: The initial judgment submodule is used to obtain the standardized stroke feature vector in a continuous manner, determine the regional continuity of the stroke through feature vector analysis, and obtain the initial continuity judgment result; The number acquisition submodule is used to determine if the initial continuity judgment result shows that there is a discontinuous area, then perform breakpoint detection on the discontinuous area to obtain the number of breakpoints; The processing submodule is used to compare the number of breakpoints with a preset threshold. If the number of breakpoints exceeds the preset threshold, a generative adversarial network is used to process the discontinuous area to obtain a feature vector after repairing the broken area. The detection submodule is used to detect the regional continuity of the repaired vector through feature vector analysis and determine the integrity judgment result of the repaired vector; The optimization submodule is used to obtain the integrity judgment result of the repaired vector, perform local optimization on the vector in the area with insufficient integrity, and obtain the optimized feature vector; The repair submodule is used to judge the continuity and integrity of the final stroke feature vector based on the optimized feature vector and adopt a vector integrity verification algorithm to obtain the repaired stroke feature vector.
4. The system according to claim 1, wherein: The preliminary candidate generation module includes: The feature sequence acquisition submodule is used to obtain sequence data from the repaired stroke feature vector using feature extraction technology to obtain a stroke feature sequence; The dependency feature acquisition submodule is used to analyze the stroke feature sequence through a recurrent neural network, determine the time dependency relationship, and obtain the sequence dependency feature; The preliminary candidate set acquisition submodule is used to match candidate Chinese characters based on sequence dependency features and predefined Chinese character stroke order rules to generate a preliminary candidate set; The sorting submodule is used to determine if the number of Chinese characters in the preliminary candidate set exceeds a preset threshold, and then use the rule matching method to sort the candidate set and determine the priority sorting result.
5. The system according to claim 1, wherein: The weighted sorting module includes: The frequency analysis submodule is used to obtain Chinese character data in a preset name database through a preliminary Chinese character candidate set, calculate the frequency distribution of the candidate Chinese characters in the name, and obtain the Chinese character frequency analysis results; The data matching submodule is used to extract name data related to the candidate Chinese characters from a preset name database using database query processing based on the results of the Chinese character frequency analysis to obtain name data matching results; The Chinese character feature extraction submodule is used to extract the context features of candidate Chinese characters based on the name data matching results and generate Chinese character feature extraction results; The sorting weight calculation submodule is used to calculate the sorting weight of candidate Chinese characters through the Chinese character feature extraction results and apply the attention mechanism to obtain the sorting weight calculation results; The candidate Chinese character sorting submodule is used to perform weighted sorting on the candidate Chinese characters if the weight value in the sorting weight calculation result meets the preset threshold value to obtain a candidate Chinese character sorting sequence; The candidate set optimization submodule is used to optimize the preliminary Chinese character candidate set according to the candidate Chinese character sorting sequence and generate the candidate set optimization result; The generation submodule is used to generate a weighted Chinese character candidate sequence through the candidate set optimization results and determine the final weighted Chinese character sequence.
6. The system according to claim 1, wherein: The semantic optimization module includes: The context information acquisition submodule is used to determine if the highest confidence level of the weighted Chinese character sequence is lower than a preset threshold, and then obtain the context information of each Chinese character in the sequence to obtain an initial semantic feature representation; The embedding vector generation submodule is used to encode the initial semantic feature representation through the pre-trained language model to obtain the semantic embedding vector; The semantic collocation submodule is used to calculate the matching degree between semantic embedding vectors using the cosine similarity algorithm to determine the rationality of semantic collocation; The candidate replacement submodule is used to determine if the semantic collocation rationality is lower than a preset threshold, and then obtain high-frequency synonymous Chinese characters from the vocabulary of the pre-trained language model to obtain a candidate replacement sequence; The Chinese character candidate sequence optimization submodule is used to reorder the candidate replacement sequences through the sequence-to-sequence model to obtain the optimized Chinese character candidate sequence; The final output submodule is used to optimize the candidate sequence of Chinese characters, calculate the semantic coherence of the entire sequence using a dynamic programming algorithm, and determine the final output sequence.
7. The system according to claim 1, wherein: The similarity calculation module includes: A similarity submodule is used to calculate the similarity between the optimized Chinese character candidate sequence and the names in the name database using a sequence matching algorithm to obtain an initial name candidate set; The name set submodule is used to determine if the similarity is higher than a preset threshold, and then include the corresponding name in the initial name candidate set to obtain a filtered name set; The pinyin feature vector submodule is used to obtain the pinyin features of each name based on the filtered name set and generate a pinyin feature vector; The clustering submodule is used to group the name set using the pinyin feature vector and a clustering algorithm to obtain the classified name subsets; The subset calculation submodule is used to calculate the semantic feature vector of the name in each subset after classification and generate a semantic feature set; The name candidate set submodule is used to determine if the semantic feature set matches the preset semantic template, and then retain the corresponding name subset to obtain the final preliminary name candidate set.
8. The system according to claim 1, wherein: The weight adjustment module includes: The distribution comparison submodule is used to obtain the preliminary name candidate set and the regional name frequency distribution from the data source, and use statistical methods to calculate the distribution comparison between the two to obtain the preliminary matching degree; The distribution comparison and extraction submodule is used to determine if the preliminary matching degree is lower than the preset threshold, and then extract the difference features through the distribution comparison results to obtain a feature set; The probability distribution submodule is used to calculate the conditional probability of each candidate name based on the feature set using the Bayesian inference model to obtain the probability distribution; The adjustment submodule is used to adjust the weight of each name in the preliminary candidate set through probability distribution to obtain an optimized weight set; The sorting and screening submodule is used to sort and screen the preliminary candidate set according to the optimized weight set to obtain the adjusted candidate set; The verification result calculation submodule is used to recalculate the matching degree with the regional frequency using the distribution comparison method for the adjusted candidate set, determine whether it meets the preset threshold, and obtain the verification result; The verification submodule is used to determine if the verification result meets the preset threshold and output the adjusted candidate set as the final name candidate set. If not, the feature set is updated according to the verification result, and the Bayesian inference and weight adjustment are repeated to obtain the final name candidate set.
9. The system according to claim 1, wherein: The multimodal verification module includes: A filtering submodule is used to obtain a final set of name candidates, filter the set using the candidate set screening rules to obtain a selected name list, and obtain a comprehensive feature vector based on the selected name list; The confidence calculation submodule is used to calculate the confidence of the comprehensive feature vector through a multimodal verification algorithm to obtain a confidence calculation result; The record confirmation submodule is used to select the name with the highest confidence based on the confidence calculation result and confirm the name record using the record confirmation mechanism; The descending sorting submodule is used to use a sorting algorithm to sort the names in descending order of confidence if there are multiple names with high confidence in the confidence calculation results, so as to obtain the final confirmed name record.
10. An electronic device comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein: When the processor executes the computing program, the system according to any one of claims 1 to 9 is implemented.
Citation Information
Cited By
Multi-model signature matching method and device, storage medium and program product
CN121096031A