Similar Chinese character determination method and related device

By combining the shape, pronunciation and different characters of Chinese characters in scenarios such as trademark registration and domain name registration, the similarity of Chinese characters is solved, and the recognition effect of brand protection and other scenarios is improved.

CN120180146APending Publication Date: 2025-06-20CHINA INTERNET NETWORK INFORMATION CENTER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510249522.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the current technology, in the scenarios such as trademark registration and domain name registration, the method used to determine the similarity of Chinese characters has not been comprehensively analyzed from the shape, pronunciation and different characters of Chinese characters, resulting in the incompleteness of similar Chinese characters, which affects the effectiveness of scenarios such as brand protection.

Method used

By obtaining the shape feature data and pronunciation feature data of the Chinese characters to be compared and the target Chinese characters in the target character set, calculate their shape similarity and pronunciation similarity, and combining the collection of different characters, generate Chinese character recognition results and determine the list of similar Chinese characters.

Benefits of technology

It improves the comprehensiveness of similar Chinese characters, and can more effectively identify similar Chinese characters that are easily confused with Chinese characters to be compared in scenarios such as brand protection, and provides more comprehensive search results for similar Chinese characters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180146A_ABST
    Figure CN120180146A_ABST
Patent Text Reader

Abstract

The invention discloses a similar Chinese character determination method and a related device, and the method comprises the steps: obtaining a target character set, and determining the target shape similarity and target pronunciation similarity of a to-be-compared Chinese character and a target Chinese character according to the shape feature data and pronunciation feature data corresponding to the to-be-compared Chinese character and the target Chinese character in the target character set; according to the target shape similarity and the target pronunciation similarity, in combination with an allomorphic character set which is homophonous and synonymous with the to-be-compared Chinese character but different in writing mode, a Chinese character recognition result of the target Chinese character is obtained, then all Chinese characters in the target character set are executed according to the steps, and a similar Chinese character list of the to-be-compared Chinese character is generated. Therefore, the feedback information of the similar Chinese character list carrying the to-be-compared Chinese character can be generated when the query request carrying the to-be-compared Chinese character is obtained by considering the three dimensions of the shape, the pronunciation and the variant character, and the feedback information can more comprehensively provide the Chinese character which is easily confused with the to-be-compared Chinese character in the scene of brand protection and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method for determining similar Chinese characters and related devices. Background Art

[0002] In scenarios such as trademark registration and domain name registration, text is an important identifier for distinguishing different brands, different domain names, etc. For example, an enterprise submits a registration application to the trademark office and completes trademark registration for the brand field composed of Chinese characters, so as to distinguish the goods or services of this enterprise from those of other enterprises.

[0003] In the related art, image recognition, stroke comparison and other methods are usually used to determine the similarity of different Chinese characters.

[0004] However, the above methods do not analyze from the dimension of the distinguishing attributes of Chinese characters, resulting in the similar Chinese characters obtained by the related art being not comprehensive enough, thus leading to effective protection in scenarios such as brand protection. Summary of the Invention

[0005] In view of the above problems, the present application provides a method for determining similar Chinese characters and related devices, which is used to improve the comprehensiveness of determining similar Chinese characters, so as to improve effective protection in scenarios such as brand protection.

[0006] Based on this, the present application discloses the following technical solutions:

[0007] In a first aspect, an embodiment of the present application provides a method for determining similar Chinese characters, the method comprising:

[0008] Obtain a target character set;

[0009] Determine the target shape similarity between the Chinese character to be compared and the target Chinese character according to the shape feature data respectively corresponding to the Chinese character to be compared and the target Chinese character in the target character set;

[0010] Determine the target pronunciation similarity between the Chinese character to be compared and the target Chinese character according to the pronunciation feature data respectively corresponding to the Chinese character to be compared and the target Chinese character in the target character set;

[0011] Generate a Chinese character recognition result of the target Chinese character according to the target shape similarity, the target pronunciation similarity and the set of variant Chinese characters of the Chinese character to be compared, where the variant Chinese characters included in the set of variant Chinese characters have the same pronunciation, the same meaning and different writing methods as the Chinese character to be compared;

[0012] Respectively determine each Chinese character in the target character set as the target Chinese character, obtain the Chinese character recognition results of each Chinese character, and generate a list of similar Chinese characters of the Chinese character to be compared according to the Chinese character recognition results of each Chinese character;

[0013] In response to obtaining a query request carrying a Chinese character to be compared, feedback information carrying a list of similar Chinese characters of the Chinese character to be compared is generated.

[0014] Optionally, generating the Chinese character recognition result of the target Chinese character according to the target shape similarity, the target pronunciation similarity, and the variant character set of the Chinese character to be compared includes:

[0015] Obtain a first threshold corresponding to the shape similarity and a second threshold corresponding to the pronunciation similarity;

[0016] If the target shape similarity is greater than the first threshold, or the target pronunciation similarity is greater than the second threshold, or the target Chinese character belongs to the variant character set of the Chinese character to be compared, then generate a Chinese character recognition result for indicating that the target Chinese character is a similar Chinese character of the Chinese character to be compared.

[0017] Optionally, the shape feature data includes a glyph structure, a radical, and a conversion graph. Determining the target shape similarity between the Chinese character to be compared and the target Chinese character according to the shape feature data respectively corresponding to the Chinese character to be compared and the target Chinese character in the target character set includes:

[0018] According to the glyph structures respectively corresponding to the Chinese character to be compared and the target Chinese character, determine the structural similarity between the Chinese character to be compared and the target Chinese character through a structural preset rule, and the structural preset rule is used to determine the similarity degree between different glyph structures;

[0019] According to the radicals respectively corresponding to the Chinese character to be compared and the target Chinese character, determine the radical similarity between the Chinese character to be compared and the target Chinese character through a radical preset rule;

[0020] Identify the conversion graphs respectively corresponding to the Chinese character to be compared and the target Chinese character through an image recognition algorithm to obtain the graphic similarity between the Chinese character to be compared and the target Chinese character;

[0021] Determine the target shape similarity according to the structural similarity, the radical similarity, and the graphic similarity respectively corresponding to the Chinese character to be compared and the target Chinese character.

[0022] Optionally, the pronunciation feature data includes an initial, a final, and a tone. Determining the target pronunciation similarity between the Chinese character to be compared and the target Chinese character according to the pronunciation feature data respectively corresponding to the Chinese character to be compared and the target Chinese character in the target character set includes:

[0023] According to the initials respectively corresponding to the Chinese character to be compared and the target Chinese character, determine the initial similarity between the Chinese character to be compared and the target Chinese character through an initial preset rule;

[0024] According to the respective finals corresponding to the Chinese character to be compared and the target Chinese character, determine the finals similarity between the Chinese character to be compared and the target Chinese character through the preset finals rules;

[0025] According to the respective tones corresponding to the Chinese character to be compared and the target Chinese character, determine the tone similarity between the Chinese character to be compared and the target Chinese character;

[0026] According to the initial similarity, finals similarity and tone similarity corresponding to the Chinese character to be compared and the target Chinese character respectively, determine the target pronunciation similarity.

[0027] Optionally, the initial preset rules include a first initial preset rule and a second initial preset rule. The first initial rule is used to determine the similarity degree between the flat tongue sound and the retroflex sound, and the second initial rule is used to determine the similarity degree between different initials based on the initial confusion degree in the target area.

[0028] Optionally, if the field to be compared includes a first Chinese character to be compared and a second Chinese character to be compared, the method further includes:

[0029] Obtain a first list of similar Chinese characters of the first Chinese character to be compared and a second list of similar Chinese characters of the second Chinese character to be compared;

[0030] Generate a combination of similar fields of the field to be compared according to the first list of similar Chinese characters and the second list of similar Chinese characters.

[0031] Optionally, if the field to be compared is registered, the method further includes:

[0032] In response to obtaining a registration request carrying a field to be registered, if the field to be registered is one of the fields in the combination of similar fields of the field to be compared, generate a prompt message for indicating that the field to be registered is similar to the field to be compared.

[0033] In a second aspect, an embodiment of the present application provides a similar Chinese character determination device, and the device includes: an acquisition unit, a determination unit and a generation unit;

[0034] The acquisition unit is used to acquire a target character set;

[0035] The determination unit is used to determine the target shape similarity between the Chinese character to be compared and the target Chinese character according to the respective shape feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set;

[0036] The determination unit is further used to determine the target pronunciation similarity between the Chinese character to be compared and the target Chinese character according to the respective pronunciation feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set;

[0037] The generating unit is configured to generate a Chinese character recognition result of the target Chinese character according to the target shape similarity, the target pronunciation similarity, and the variant character set of the Chinese character to be compared, where the variant characters included in the variant character set have the same pronunciation, the same meaning, and different writing forms as the Chinese character to be compared;

[0038] The generating unit is further configured to respectively determine each Chinese character in the target character set as the target Chinese character, obtain the Chinese character recognition results of each Chinese character, and generate a list of similar Chinese characters of the Chinese character to be compared according to the Chinese character recognition results of each Chinese character;

[0039] The generating unit is further configured to, in response to obtaining a query request carrying the Chinese character to be compared, generate feedback information carrying the list of similar Chinese characters of the Chinese character to be compared.

[0040] In a third aspect, an embodiment of the present application provides a computer device, where the computer device includes a processor and a memory:

[0041] The memory is configured to store a computer program and transmit the computer program to the processor;

[0042] The processor is configured to execute the method described in the first aspect above according to the computer program.

[0043] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium is configured to store a computer program, and the computer program is configured to execute the method described in the first aspect above.

[0044] In a fifth aspect, an embodiment of the present application provides a computer program product including a computer program, which, when running on a computer device, causes the computer device to execute the method described in the first aspect above.

[0045] From the above technical solutions, it can be seen that the present application has at least the following beneficial effects:

[0046] The similar Chinese character determination method provided by this application determines the list of corresponding similar Chinese characters for each Chinese character in the target character set based on three distinguishable dimensions: shape, pronunciation, and variant Chinese characters, considering these three dimensions, which improves the comprehensiveness of determining similar Chinese characters. Taking the example of determining the list of similar Chinese characters for the Chinese character to be compared in the target character set, after obtaining the target character set, according to the shape feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set, the target shape similarity between the Chinese character to be compared and the target Chinese character is determined. According to the pronunciation feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set, the target pronunciation similarity between the Chinese character to be compared and the target Chinese character is determined. Based on the target shape similarity, the target pronunciation similarity, and in combination with the set of variant Chinese characters that are homophonous and synonymous with the Chinese character to be compared but have different writing forms, the Chinese character recognition result of the target Chinese character is obtained, thereby determining whether the target Chinese character is a similar Chinese character of the Chinese character to be compared. Then, each Chinese character in the target character set is executed according to the above steps respectively to generate the list of similar Chinese characters of the Chinese character to be compared. Chinese characters with similar shapes, similar pronunciations, or variant Chinese characters are likely to cause confusion among the public and reduce the distinguishability between Chinese characters. Therefore, by considering these three dimensions of shape, pronunciation, and variant Chinese characters, it is possible to more comprehensively identify the similar Chinese characters that are likely to be confused with the Chinese character to be compared. Thus, when obtaining a query request carrying the Chinese character to be compared, it is possible to generate feedback information carrying the list of similar Chinese characters of the Chinese character to be compared, and this feedback information can more comprehensively provide Chinese characters that are similar to the Chinese character to be compared and are likely to cause confusion in scenarios such as brand protection. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0048] Figure 1 It is a schematic flowchart of a method for determining similar Chinese characters provided by an embodiment of the present application;

[0049] Figure 2 It is a schematic flowchart of an application scenario of a method for determining similar Chinese characters provided by an embodiment of the present application;

[0050] Figure 3 It is a schematic structural diagram of a device for determining similar Chinese characters provided by an embodiment of the present application;

[0051] Figure 4 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] Embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.

[0053] In the related art, similar Chinese characters are usually determined by traditional image comparison methods, such as image pixel point comparison, stroke comparison, etc. It does not consider that in scenarios such as brand protection, Chinese characters, as identifiers for distinguishing from other entities, can cause confusion among the public in multiple dimensions. For example, Chinese characters with the same radical and semi-surrounded structure are likely to cause confusion at the visual level, Chinese characters with the same or similar pronunciation are likely to cause confusion at the auditory level, and variant Chinese characters with the same pronunciation and meaning but different writing styles have individual strokes different from the standard characters, which are also likely to cause confusion. These Chinese characters that are likely to cause public confusion are interspersed in trademarks, domain names, etc., making the public prone to misjudge identifiers such as brands and domain names. And single-dimensional Chinese character image recognition cannot analyze and judge the above three dimensions that are likely to cause public confusion, and the obtained similar Chinese characters are not comprehensive enough, resulting in a relatively low comprehensiveness of identifying similar fields when the similar Chinese characters obtained based on the related art are applied in scenarios such as brand protection and domain name protection.

[0054] Based on this, the embodiments of the present application provide a method and related device for determining similar Chinese characters. By considering three dimensions that are likely to cause public confusion, a list of similar Chinese characters for the Chinese characters to be compared is generated according to the shape similarity, pronunciation similarity, and the set of variant Chinese characters in the target character set, so as to be able to more comprehensively pre-calibrate and obtain the similar Chinese characters of each Chinese character in the target character set in advance, and improve the comprehensiveness of identifying similar Chinese characters in scenarios such as brand protection.

[0055] The method for determining similar Chinese characters provided by the present application can be applied to computer devices with the ability to determine similar Chinese characters, such as terminal devices and servers. Among them, the terminal device can specifically be a desktop computer, a laptop computer, a mobile phone, a tablet computer, etc.; the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, etc. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this here.

[0056] See Figure 1 , which is a schematic flowchart of the method for determining similar Chinese characters provided by the embodiments of the present application. For the convenience of description, the following embodiments will be introduced by taking the execution subject of this method for determining similar Chinese characters as a server as an example. As Figure 1As shown, the similar Chinese character determination method includes S101-S104.

[0057] S101: Obtain a target character set.

[0058] The embodiment of the present application predetermines a list of similar Chinese characters that match multiple Chinese characters in three dimensions: shape, pronunciation, and variant forms, thereby being able to identify similar trademarks, similar domain names, etc., to protect distinctive entities such as brands.

[0059] First, a target character set is obtained. The target character set is a set of Chinese characters used to determine similar Chinese characters. The target character set may be a predefined Chinese character set including multiple Chinese characters. For example, the target character set is "Information Technology Chinese Coded Character Set GB18030-2022".

[0060] S102: Determine the target shape similarity between the Chinese character to be compared and the target Chinese character according to the shape feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set.

[0061] The Chinese character to be compared is a Chinese character in the target character set used for similarity comparison with other Chinese characters. The following describes the similar Chinese character determination method provided in the embodiment of the present application by taking determining similar Chinese characters of the Chinese character to be compared as an example.

[0062] The target Chinese character is the Chinese character that is compared with the to-be-compared Chinese character in a similarity comparison process in the target character set, and the target Chinese character and the to-be-compared Chinese character are different Chinese characters. The shape feature data is the data that can provide visual information at the shape feature level of the Chinese character.

[0063] The target shape similarity is the degree of similarity between the Chinese character to be compared and the target Chinese character in the shape dimension. If the target shape similarity is high, it means that the Chinese character to be compared and the target Chinese character are easily confused visually. By determining the target shape similarity between the Chinese character to be compared and the target Chinese character, the degree to which the two Chinese characters are easily confused can be quantified. For example, the Chinese character to be compared is "康", and the target Chinese character is "康". If it is determined based on the target shape similarity that "康" and "康" are visually similar and easily confused, it means that the brand name "康康" is visually difficult to distinguish from "康康", "康康", "康康", etc. Therefore, by determining the target shape similarity between "康" and "康", we can measure whether the two characters are visually easy to confuse.

[0064] The target shape similarity can not only be obtained by treating the Chinese character as a graphic for image comparison calculation, but also can be obtained by analyzing the text features of the Chinese character itself, such as comparing the radical, strokes, pictographic features, glyph structure and other text features of the Chinese character to be compared with the target Chinese character to obtain the target shape similarity.

[0065] S103: Determine the target pronunciation similarity between the Chinese character to be compared and the target Chinese character according to the pronunciation feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set.

[0066] The target pronunciation similarity is the similarity between the Chinese character to be compared and the target Chinese character in the pronunciation dimension. If the target pronunciation similarity is high, it means that the Chinese character to be compared and the target Chinese character are easily confused in hearing. By determining the target shape similarity between the Chinese character to be compared and the target Chinese character, the degree to which the two Chinese characters are easily confused by hearing can be quantified. For example, the Chinese character to be compared is "酒" and the target Chinese character is "久". Although the two characters have different writing methods, glyph structures, radicals, shapes and other shape-related features, they have the same pronunciation and cannot be distinguished from "酒" and "久" in hearing. This will also reduce the distinction between the two, resulting in easy confusion by the public in scenarios such as brand protection.

[0067] The target pronunciation similarity can be obtained not only by analyzing and comparing the audio features of the standard audio data of the Chinese characters to be compared and the target Chinese characters, but also by analyzing and comparing the pronunciation feature data involved in the reading process of the Chinese characters to be compared and the target Chinese characters, such as pinyin, tone, etc., to determine the degree of auditory confusion between the two.

[0068] S104: Generate a Chinese character recognition result of the target Chinese character according to the target shape similarity, the target pronunciation similarity and the variant character set of the Chinese character to be compared.

[0069] Each process of determining similar Chinese characters is carried out by comparing the feature data between the two Chinese characters in three dimensions: shape, pronunciation and variants, thereby generating a Chinese character recognition result for each similar Chinese character determination process. The Chinese character recognition result is used to indicate whether the target Chinese character is a similar Chinese character to the Chinese character to be compared.

[0070] Variant characters are Chinese characters with the same pronunciation and meaning, but different only in shape. The variant characters correspond to the standard characters, that is, the difference between variant characters and standard characters is only in the way of writing, which cannot be distinguished auditorily, and express the same meaning. The Chinese characters to be compared are standard characters, and the variant character set of the Chinese characters to be compared includes at least one variant character of the Chinese characters to be compared. By searching whether the target Chinese character is included in the variant character set of the Chinese characters to be compared, it can be determined whether the target Chinese character is a variant character of the Chinese characters to be compared. Variant characters have a high probability of being confused in scenarios such as trademarks and domain names. By replacing the standard characters with variant characters at the same position in field A to obtain field B, it is easy to cause field A and field B to be confused. Therefore, the variant character set of the Chinese characters to be compared is used to identify whether the target Chinese character is a variant character of the Chinese characters to be compared. If so, it is determined that the target Chinese character is similar to the Chinese character to be compared in the variant character dimension, and more comprehensive Chinese character recognition information in the variant character dimension is provided.

[0071] As an implementation manner, the set of variant Chinese characters for each Chinese character can be obtained through a preset variant Chinese character database.

[0072] In addition to the dimension of variant Chinese characters, it is also necessary to determine the similarity degree between the target Chinese character and the Chinese character to be compared in the visual dimension through the target shape similarity, and in the auditory dimension through the target pronunciation similarity. That is, according to the three judgment bases of the target shape similarity, the target pronunciation similarity, and whether the target Chinese character is in the set of variant Chinese characters of the Chinese character to be compared, the Chinese character recognition information corresponding to the visual dimension, the auditory dimension, and the variant Chinese character dimension is generated to obtain the final Chinese character recognition result.

[0073] S105: Each Chinese character in the target character set is respectively determined as the target Chinese character to obtain the Chinese character recognition result of each Chinese character, and a list of similar Chinese characters of the Chinese character to be compared is generated according to the Chinese character recognition result of each Chinese character.

[0074] The Chinese characters included in the list of similar Chinese characters of the Chinese character to be compared are the similar Chinese characters of the Chinese character to be compared.

[0075] By respectively determining each Chinese character in the target character set except the Chinese character to be compared as the target Chinese character and executing the processes of S102 - S104, the Chinese character recognition result of each Chinese character is obtained. If the Chinese character recognition result indicates a similar Chinese character of the Chinese character to be compared, the Chinese character corresponding to the Chinese character recognition result is added to the list of similar Chinese characters of the Chinese character to be compared. After traversing each Chinese character in the target character set, a complete list of similar Chinese characters of the Chinese character to be compared is obtained.

[0076] S106: In response to obtaining a query request carrying the Chinese character to be compared, feedback information carrying the list of similar Chinese characters of the Chinese character to be compared is generated.

[0077] The query request is a request for querying whether there is a detailed Chinese character for a Chinese character. The query request can come from a client system or a management system. Each query request carries at least one Chinese character for querying the similar Chinese characters of this Chinese character. For example, the query request can come from a trademark registration system. If a field to be registered carried by the query request, a list of similar Chinese characters of each Chinese character in the field can be returned and subsequent processes can be carried out. Another example is that the query request can come from the intellectual property system of an enterprise. Based on the original fields of the brands under the enterprise, the detailed Chinese characters of each Chinese character in the original fields are queried for protective intellectual property layout or infringement evidence collection.

[0078] Taking the query request carrying the Chinese character to be compared as an example, in response to obtaining the query request, feedback information based on the query request is generated. The feedback information carries the list of similar Chinese characters of the Chinese character to be compared generated in the foregoing steps. The similar Chinese character list is generated based on three aspects: shape similarity, pronunciation similarity, and variant Chinese characters. Thus, the feedback information carrying the list of similar Chinese characters can provide a more comprehensive similar Chinese character query result.

[0079] It should be noted that the above is only an exemplary description of the process of determining similar Chinese characters using the Chinese characters to be compared as an example. The method for determining similar Chinese characters provided in the embodiments of the present application can also take each Chinese character in the target character set as the Chinese character to be compared and separately execute the steps of S102 - S105, so as to obtain a list of similar Chinese characters corresponding to each Chinese character in the target character set, thereby obtaining a similar Chinese character database. When a query request is obtained, the list of similar Chinese characters of the Chinese character carried in the query request can be obtained from the similar Chinese character database, and corresponding feedback information can be generated. The embodiments of the present application do not limit this here.

[0080] As can be seen from the above technical solution, the method for determining similar Chinese characters provided in the present application considers three distinguishable dimensions of Chinese characters: shape, pronunciation, and variant characters. Based on these three dimensions, a list of similar Chinese characters corresponding to each Chinese character in the target character set can be determined, improving the comprehensiveness of determining similar Chinese characters. Taking the example of determining the list of similar Chinese characters of the Chinese character to be compared in the target character set, after obtaining the target character set, according to the shape feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set, the target shape similarity between the Chinese character to be compared and the target Chinese character is determined. According to the pronunciation feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set, the target pronunciation similarity between the Chinese character to be compared and the target Chinese character is determined. Based on the target shape similarity, the target pronunciation similarity, and in combination with the set of variant characters that are homophonous and synonymous with the Chinese character to be compared but have different writing forms, the Chinese character recognition result of the target Chinese character is obtained, so as to determine whether the target Chinese character is a similar Chinese character of the Chinese character to be compared. Then, each Chinese character in the target character set is separately executed according to the above steps to generate a list of similar Chinese characters of the Chinese character to be compared. Chinese characters with similar shapes, similar pronunciations, or variant characters are likely to cause confusion among the public, reducing the distinguishability between Chinese characters. Therefore, by considering these three dimensions of shape, pronunciation, and variant characters, it is possible to more comprehensively identify similar Chinese characters that are likely to be confused with the Chinese character to be compared. Thus, when a query request carrying the Chinese character to be compared is obtained, feedback information carrying the list of similar Chinese characters of the Chinese character to be compared can be generated. This feedback information can more comprehensively provide Chinese characters that are similar to the Chinese character to be compared and are likely to cause confusion in scenarios such as brand protection.

[0081] In a possible implementation manner, during the process of generating the Chinese character recognition result, the first threshold corresponding to the shape similarity and the second threshold corresponding to the pronunciation similarity can be obtained first. The first threshold is used to determine whether the shape similarity is too high, and the second threshold is used to determine whether the pronunciation similarity is too high. The first threshold and the second threshold can be set by those skilled in the art according to actual needs. The embodiments of the present application do not limit this here.

[0082] If the target shape similarity is greater than the first threshold, or the target pronunciation similarity is greater than the second threshold, or the target Chinese character belongs to the set of variants of the Chinese character to be compared, a Chinese character recognition result is generated to indicate that the target Chinese character is a similar Chinese character to the Chinese character to be compared. That is, the target shape similarity is greater than the threshold as the first condition, the target pronunciation similarity is greater than the second threshold as the second condition, and the target Chinese character belongs to the set of variants of the Chinese character to be compared as the third condition. As long as at least one of the conditions is met, it means that the target Chinese character is a similar Chinese character to the Chinese character to be compared, and then the corresponding Chinese character recognition result is generated.

[0083] If the target Chinese character meets one condition, it means that the target Chinese character is easily confused with the Chinese character to be compared in the dimension corresponding to the condition. Therefore, the recognition range of similar Chinese characters can be expanded as much as possible, that is, as many Chinese characters similar to the Chinese characters to be compared as possible are added to the Chinese character similarity list, reducing the probability of missing similar Chinese characters and improving the comprehensiveness of identifying similar Chinese characters in scenarios such as brand and domain name protection.

[0084] The embodiments of the present application do not specifically limit how to determine the target shape similarity and the target pronunciation similarity, and exemplary explanations are given below.

[0085] In a possible implementation, the shape feature data includes a glyph structure, a radical, and a conversion graphic. The glyph structure is the basic construction method of Chinese characters. For example, the glyph structure of Chinese characters includes a left-right structure, a left-middle-right structure, an upper-lower structure, an upper-middle-lower structure, a semi-enclosed structure, a fully-enclosed structure, and the like. The radical is a basic component of a Chinese character and has a semantic function. "河" and "海" both have the same radical "氵", indicating that they are related to water. The conversion graphic is a graphic expression form that converts Chinese characters into a specific format, such as a vector map or a bitmap. As an implementation, the glyph structure and radical of Chinese characters can be obtained through a preset glyph structure database and a preset radical database.

[0086] The structural preset rules are used to measure the correspondence between the glyph structure relationship and the structural similarity of different Chinese characters, and are used to determine the similarity between different glyph structures. According to the glyph structures corresponding to the Chinese characters to be compared and the target Chinese characters, the structural similarity between the Chinese characters to be compared and the target Chinese characters is determined by the structural preset rules. For example, in the structural preset rules, the structural similarity of the same glyph structure is 1, the structural similarity of the semi-enclosed structure and the fully enclosed structure is 0.9, the structural similarity of the left-right structure and the left-middle-right structure is 0.9, the structural similarity of the upper-lower structure and the upper-middle-lower structure is 0.9, and the structural similarity of other glyph structure relationships is 0.5.

[0087] The radical preset rule is to measure the corresponding relationship between the radical relationships of different Chinese characters and the radical similarity. According to the radicals corresponding to the Chinese character to be compared and the target Chinese character respectively, the radical similarity between the Chinese character to be compared and the target Chinese character is determined through the radical preset rule. For example, in the radical preset rule, the radical similarity of the same radical is 1, and the radical similarity of different radicals is 0.8.

[0088] The conversion graphics corresponding to the Chinese character to be compared and the target Chinese character are recognized through an image recognition algorithm to obtain the graphic similarity between the Chinese character to be compared and the target Chinese character. First, the Chinese character to be compared and the target Chinese character can be standardized and converted into a first conversion graphic and a second conversion graphic of the same size, and then the first conversion graphic and the second conversion graphic are processed through an image recognition algorithm to calculate the graphic similarity between the Chinese character to be compared and the target Chinese character. For example, the image recognition algorithm can be the mean hash algorithm, the difference hash algorithm, the perceptual hash algorithm, the three-histogram algorithm, etc.

[0089] According to the structural similarity, radical similarity, and graphic similarity corresponding to the Chinese character to be compared and the target Chinese character respectively, the target shape similarity is determined. For example, the target shape similarity can be determined by the following formula:

[0090] LX = structural similarity × radical similarity × graphic similarity, where LX is the target shape similarity.

[0091] Thus, the graphic similarity calculated through the image recognition algorithm can reflect the similarity degree of Chinese characters in the form of images. On this basis, combined with the radical features and structural features of Chinese characters themselves, the determination process of the target shape similarity is more in line with the Chinese character formation principle, improving the comprehensiveness of measuring the confusion degree of the Chinese character to be compared and the target Chinese character at the visual level.

[0092] In a possible implementation manner, the pronunciation feature data includes initials, finals, and tones. The initial is the initial consonant part of the Chinese character pronunciation, the final is the vowel or vowel combination part following the initial, and the tones include the first tone, the second tone, the third tone, and the fourth tone.

[0093] The initial preset rule is to measure the corresponding relationship between the initial relationships of different Chinese characters and the initial similarity. According to the initials corresponding to the Chinese character to be compared and the target Chinese character respectively, the initial similarity between the Chinese character to be compared and the target Chinese character is determined through the initial preset rule. For example, the initial similarity of the same initial is 1, the initial similarity of easily confused initials is 0.9, and the initial similarity in other cases is 0.

[0094] In a possible implementation, the initial consonant preset rules include a first initial consonant preset rule and a second initial consonant preset rule. The first initial consonant rule is used to determine the similarity degree between the flat tongue sound and the retroflex tongue sound, and the second initial consonant rule is used to determine the similarity degree between different initial consonants based on the confusion degree of the initial consonants in the target area.

[0095] For example, in the first initial consonant preset rule, the initial consonant similarity between zh and z is 0.9, the initial consonant similarity between ch and c is 0.9, and the initial consonant similarity between sh and s is 0.9. The target area is an area where the pronunciation of initial consonants is confused due to mother tongue habits. For example, in the second initial consonant preset rule, the initial consonant similarity between h and f is 0.9, the initial consonant similarity between r and l is 0.9, and the initial consonant similarity between n and l is 0.9. These initial consonants are the initial consonants that are easily confused in the target area and also need to be considered in the process of identifying similar Chinese characters.

[0096] Thus, in the process of determining the initial consonant similarity, not only the situation where flat tongue sounds and retroflex tongue sounds are easily confused is considered, but also the characteristics of easily confused special initial consonants in the target area are considered, improving the comprehensiveness of the initial consonant similarity.

[0097] The final vowel preset rule is to measure the corresponding relationship between the final vowel relationships of different Chinese characters and the final vowel similarity. According to the final vowels corresponding to the Chinese character to be compared and the target Chinese character respectively, the final vowel similarity between the Chinese character to be compared and the target Chinese character is determined through the final vowel preset rule. For example, the final vowel similarity of the same final vowel is 1, and the final vowel similarity between ang and an, eng and en, ing and in, iang and ian, uang and uan is 0.9, and the final vowel similarity in other cases is 0.

[0098] According to the tones corresponding to the Chinese character to be compared and the target Chinese character respectively, the tone similarity between the Chinese character to be compared and the target Chinese character is determined. Based on the pronunciation characteristics of Chinese characters, similar initial consonants and final vowels are more likely to cause auditory confusion. Therefore, the influence degree of the tone similarity on the target pronunciation similarity can be appropriately reduced. For example, the tone similarity of the same tone is 1, and in other cases it is 0.9.

[0099] According to the initial consonant similarity, final vowel similarity and tone similarity corresponding to the Chinese character to be compared and the target Chinese character respectively, the target pronunciation similarity is determined. For example, the target pronunciation similarity can be determined by the following formula:

[0100] LY = initial consonant similarity × final vowel similarity × tone similarity, where LY is the target pronunciation similarity.

[0101] Thus, based on the pronunciation rules of Chinese characters, the pronunciation is divided into initial consonants, final vowels and tones, and the similarity degree between the Chinese character to be compared and the target Chinese character is analyzed more carefully from different angles, improving the comprehensiveness of the target pronunciation similarity.

[0102] In a possible implementation, to expand the application scenarios of the method for determining similar Chinese characters, the embodiments of the present application further provide a method for generating similar field combinations. Taking the generation of similar field combinations of the to-be-compared field as an example, the to-be-compared field includes a first to-be-compared Chinese character and a second to-be-compared Chinese character. Obtain a first list of similar Chinese characters of the first to-be-compared Chinese character and a second list of similar Chinese characters of the second to-be-compared Chinese character. That is, after the first to-be-compared Chinese character and the second to-be-compared Chinese character respectively execute the steps of S102 - S105, obtain the corresponding lists of similar Chinese characters respectively. Generate similar field combinations of the to-be-compared field according to the first list of similar Chinese characters and the second list of similar Chinese characters. Specifically, there are n + 1 (n is the number of Chinese characters in the first list of similar Chinese characters, 1 is the first to-be-compared Chinese character) possibilities at the position of the first to-be-compared Chinese character, and there are m + 1 (m is the number of Chinese characters in the second list of similar Chinese characters, 1 is the second to-be-compared Chinese character) possibilities at the position of the second to-be-compared Chinese character. Therefore, a total of (n + 1)×(m + 1) - 1 permutations and combinations are generated except for the to-be-compared field, and similar field combinations of the to-be-compared Chinese characters are obtained.

[0103] Thus, by generating similar field combinations of the to-be-compared field based on the lists of similar Chinese characters of each already generated Chinese character, it is possible to more comprehensively provide similar field results in scenarios such as enterprise protection registration and trademark registration similar field query.

[0104] Taking the to-be-registered field for which registration is requested as an example, the to-be-registered field is a field that has not completed registration. If the to-be-compared field has completed registration, and after obtaining a registration request carrying the to-be-registered field, it is found and determined that the to-be-registered field is one of the fields in the similar field combination of the to-be-compared field, then a prompt message is generated to indicate that the to-be-registered field is similar to the to-be-compared field.

[0105] Thus, by pre-generating similar field combinations of each already registered field, when a new field registration request is obtained, it is possible to quickly generate a prompt message carrying a similar comparison result.

[0106] See Figure 2 , Figure 2 which is a schematic flowchart of the application scenario of a method for determining similar Chinese characters provided by the embodiments of the present application.

[0107] First, select each Chinese character in the "Chinese Coded Character Set GB18030-2022 for Information Technology", calculate the similarity of the shape (target shape similarity) LX between each Chinese character and any other Chinese character, calculate the similarity of the pronunciation (target pronunciation similarity) LY between each Chinese character and other Chinese characters, and calculate the similarity of variant Chinese characters LZ between each Chinese character and other Chinese characters. LZ is determined by whether the Chinese character is in the set of variant Chinese characters of the Chinese character to be compared. Then, split the Chinese trademark into multiple Chinese characters. For each Chinese character, calculate the Chinese characters whose shape similarity and pronunciation similarity are greater than the preset threshold based on the above steps, and the set of variant Chinese characters corresponding to each Chinese character. The Chinese characters similar in the three dimensions of pronunciation, shape, and variant Chinese characters are combined in sequence to generate multiple similar Chinese trademarks for subsequent screening.

[0108] See Figure 3 , Figure 3 A device for determining similar Chinese characters provided by an embodiment of the present application. The device 300 includes an acquisition unit 301, a determination unit 302, and a generation unit 303;

[0109] The acquisition unit 301 is used to acquire a target character set;

[0110] The determination unit 302 is used to determine the target shape similarity between the Chinese character to be compared and the target Chinese character according to the shape feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set;

[0111] The determination unit 302 is further used to determine the target pronunciation similarity between the Chinese character to be compared and the target Chinese character according to the pronunciation feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set;

[0112] The generation unit 303 is used to generate a Chinese character recognition result of the target Chinese character according to the target shape similarity, the target pronunciation similarity, and the set of variant Chinese characters of the Chinese character to be compared. The set of variant Chinese characters includes variant Chinese characters with the same pronunciation, the same meaning, and different writing methods as the Chinese character to be compared;

[0113] The generation unit 303 is further used to respectively determine each Chinese character in the target character set as the target Chinese character, obtain the Chinese character recognition results of each Chinese character, and generate a list of similar Chinese characters of the Chinese character to be compared according to the Chinese character recognition results of each Chinese character;

[0114] The generation unit 303 is further used to generate feedback information carrying the list of similar Chinese characters of the Chinese character to be compared in response to acquiring a query request carrying the Chinese character to be compared.

[0115] As can be seen from the above technical solutions, the similar Chinese character determination device provided by this application determines the list of similar Chinese characters corresponding to each Chinese character in the target character set based on three distinguishable dimensions of Chinese characters: shape, pronunciation, and variant characters, improving the comprehensiveness of determining similar Chinese characters. Taking the example of determining the list of similar Chinese characters for the Chinese character to be compared in the target character set, after the acquisition unit obtains the target character set, the determination unit determines the target shape similarity between the Chinese character to be compared and the target Chinese character according to the shape feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set. The determination unit determines the target pronunciation similarity between the Chinese character to be compared and the target Chinese character according to the pronunciation feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set. The generation unit obtains the Chinese character recognition result of the target Chinese character based on the target shape similarity, the target pronunciation similarity, and the set of variant characters that are homophonic and synonymous with the Chinese character to be compared but have different writing forms, so as to determine whether the target Chinese character is a similar Chinese character of the Chinese character to be compared. Then, the generation unit executes the above steps for each Chinese character in the target character set respectively to generate the list of similar Chinese characters of the Chinese character to be compared. Chinese characters with similar shapes, similar pronunciations, or variant characters are likely to cause confusion among the public and reduce the distinguishability between Chinese characters. Therefore, by considering these three dimensions of shape, pronunciation, and variant characters, it is possible to more comprehensively identify the similar Chinese characters that are likely to be confused with the Chinese character to be compared. Thus, when obtaining a query request carrying the Chinese character to be compared, the generation unit can generate feedback information carrying the list of similar Chinese characters of the Chinese character to be compared, and this feedback information can more comprehensively provide Chinese characters that are similar to the Chinese character to be compared and are likely to cause confusion in scenarios such as brand protection.

[0116] As a possible implementation manner, the generation unit is specifically configured to:

[0117] Obtain a first threshold corresponding to the shape similarity and a second threshold corresponding to the pronunciation similarity;

[0118] If the target shape similarity is greater than the first threshold, or the target pronunciation similarity is greater than the second threshold, or the target Chinese character belongs to the set of variant characters of the Chinese character to be compared, then generate a Chinese character recognition result for indicating that the target Chinese character is a similar Chinese character of the Chinese character to be compared.

[0119] As a possible implementation manner, the shape feature data includes the glyph structure, radicals, and conversion graphics, and the determination unit is specifically configured to:

[0120] According to the glyph structures corresponding to the Chinese character to be compared and the target Chinese character respectively, determine the structural similarity between the Chinese character to be compared and the target Chinese character through a structure preset rule, and the structure preset rule is used to determine the similarity degree between different glyph structures;

[0121] Determine the radical similarity between the Chinese character to be compared and the target Chinese character according to the radicals corresponding to the Chinese character to be compared and the target Chinese character respectively through the preset rules of radicals.

[0122] Identify the converted graphics corresponding to the Chinese character to be compared and the target Chinese character respectively through an image recognition algorithm to obtain the graphic similarity between the Chinese character to be compared and the target Chinese character.

[0123] Determine the target shape similarity according to the structural similarity, radical similarity and graphic similarity corresponding to the Chinese character to be compared and the target Chinese character respectively.

[0124] As a possible implementation manner, the pronunciation feature data includes initials, finals and tones, and the determining unit is specifically configured to:

[0125] Determine the initial similarity between the Chinese character to be compared and the target Chinese character according to the initials corresponding to the Chinese character to be compared and the target Chinese character respectively through the preset rules of initials.

[0126] Determine the final similarity between the Chinese character to be compared and the target Chinese character according to the finals corresponding to the Chinese character to be compared and the target Chinese character respectively through the preset rules of finals.

[0127] Determine the tone similarity between the Chinese character to be compared and the target Chinese character according to the tones corresponding to the Chinese character to be compared and the target Chinese character respectively.

[0128] Determine the target pronunciation similarity according to the initial similarity, final similarity and tone similarity corresponding to the Chinese character to be compared and the target Chinese character respectively.

[0129] As a possible implementation manner, the preset rules of initials include the first preset rule of initials and the second preset rule of initials. The first initial rule is used to determine the similarity degree between the flat tongue sound and the retroflex sound, and the second initial rule is used to determine the similarity degree between different initials based on the confusion degree of initials in the target area.

[0130] As a possible implementation manner, if the field to be compared includes a first Chinese character to be compared and a second Chinese character to be compared, the device further includes a field generation unit, which is configured to:

[0131] Obtain a first list of similar Chinese characters of the first Chinese character to be compared and a second list of similar Chinese characters of the second Chinese character to be compared;

[0132] Generate a combination of similar fields of the field to be compared according to the first list of similar Chinese characters and the second list of similar Chinese characters.

[0133] As a possible implementation, if the field to be compared is registered, the device further includes a prompting unit:

[0134] In response to obtaining a registration request carrying a field to be registered, if the field to be registered is one of the similar field combinations of the field to be compared, a prompt message for indicating that the field to be registered is similar to the field to be compared is generated.

[0135] See Figure 4 , this embodiment of the present application further provides a computer device, which includes a memory 401 and a processor 402:

[0136] The memory is used to store a computer program and transmit the computer program to the processor;

[0137] The processor is used to execute the method of the above method embodiment according to the computer program.

[0138] This embodiment of the present application further provides a computer-readable storage medium, which is characterized in that the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method of the above method embodiment.

[0139] This embodiment of the present application further provides a computer program product including a computer program. When it runs on a computer device, it causes the computer device to execute the method of the above method embodiment.

[0140] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0141] The term "including" and its variants used herein are open-ended inclusions, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0142] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0143] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0144] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0145] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to these embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for determining similar Chinese characters, characterized in that: The method comprises: Get the target character set; Determine the target shape similarity between the Chinese character to be compared and the target Chinese character according to the shape feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set; Determine the target pronunciation similarity between the Chinese character to be compared and the target Chinese character according to the pronunciation feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set; Generate a Chinese character recognition result of the target Chinese character according to the target shape similarity, the target pronunciation similarity and the variant character set of the Chinese character to be compared, wherein the variant characters included in the variant character set have the same pronunciation and meaning as the Chinese character to be compared but are written in different ways; Determine each Chinese character in the target character set as the target Chinese character, obtain a Chinese character recognition result for each Chinese character, and generate a similar Chinese character list of the Chinese character to be compared according to the Chinese character recognition result for each Chinese character; In response to obtaining a query request carrying a Chinese character to be compared, feedback information of a similar Chinese character list carrying the Chinese character to be compared is generated.

2. The method according to claim 1, characterized in that The step of generating a Chinese character recognition result of the target Chinese character according to the target shape similarity, the target pronunciation similarity and the variant character set of the Chinese character to be compared comprises: Obtaining a first threshold corresponding to shape similarity and a second threshold corresponding to pronunciation similarity; If the target shape similarity is greater than the first threshold, or the target pronunciation similarity is greater than the second threshold, or the target Chinese character belongs to the variant character set of the Chinese character to be compared, a Chinese character recognition result is generated to indicate that the target Chinese character is a similar Chinese character to the Chinese character to be compared.

3. The method according to claim 1, characterized in that The shape feature data includes a glyph structure, a radical and a conversion graph, and determining the target shape similarity between the Chinese character to be compared and the target Chinese character according to the shape feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set, comprises: According to the glyph structures respectively corresponding to the Chinese character to be compared and the target Chinese character, the structural similarity between the Chinese character to be compared and the target Chinese character is determined by a structural preset rule, wherein the structural preset rule is used to determine the similarity between different glyph structures; According to the radicals corresponding to the Chinese character to be compared and the target Chinese character respectively, determining the radical similarity between the Chinese character to be compared and the target Chinese character by using a preset radical rule; Identify the conversion graphics corresponding to the Chinese character to be compared and the target Chinese character respectively by an image recognition algorithm, and obtain the graphic similarity between the Chinese character to be compared and the target Chinese character; The target shape similarity is determined according to the structural similarity, radical similarity and graphic similarity respectively corresponding to the Chinese character to be compared and the target Chinese character.

4. The method according to claim 1, characterized in that: The pronunciation feature data includes initial consonants, final vowels and tones, and determining the target pronunciation similarity between the Chinese character to be compared and the target Chinese character according to the pronunciation feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set, comprises: According to the initial consonants corresponding to the Chinese character to be compared and the target Chinese character respectively, determining the initial consonant similarity between the Chinese character to be compared and the target Chinese character by using the initial consonant preset rule; According to the finals corresponding to the Chinese character to be compared and the target Chinese character respectively, determining the final similarity between the Chinese character to be compared and the target Chinese character by using a final preset rule; Determining the tone similarity between the Chinese character to be compared and the target Chinese character according to the tones respectively corresponding to the Chinese character to be compared and the target Chinese character; The target pronunciation similarity is determined according to the initial consonant similarity, final consonant similarity and tone similarity respectively corresponding to the Chinese character to be compared and the target Chinese character.

5. The method according to claim 4, characterized in that The initial consonant preset rules include a first initial consonant preset rule and a second initial consonant preset rule, the first initial consonant rule is used to determine the similarity between flat tongue sounds and curled tongue sounds, and the second initial consonant rule is used to determine the similarity between different initial consonants based on the degree of initial consonant confusion in the target area.

6. The method according to claim 1, characterized in that If the field to be compared includes a first Chinese character to be compared and a second Chinese character to be compared, the method further includes: Obtaining a first similar Chinese character list of the first Chinese character to be compared, and a second similar Chinese character list of the second Chinese character to be compared; A similar field combination of the field to be compared is generated according to the first similar Chinese character list and the second similar Chinese character list.

7. The method according to claim 6, characterized in that If the field to be compared is registered, the method further includes: In response to obtaining a registration request carrying a field to be registered, if the field to be registered is a field in a similar field combination of the field to be compared, prompt information is generated to indicate that the field to be registered is similar to the field to be compared.

8. A similar Chinese character determination device, characterized in that: The device comprises: an acquisition unit, a determination unit and a generation unit; The acquisition unit is used to acquire the target character set; The determining unit is used to determine the target shape similarity between the Chinese character to be compared and the target Chinese character according to the shape feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set; The determination unit is further used to determine the target pronunciation similarity between the Chinese character to be compared and the target Chinese character according to the pronunciation feature data corresponding to the Chinese character to be compared and the target Chinese character in the target character set; The generating unit is used to generate a Chinese character recognition result of the target Chinese character according to the target shape similarity, the target pronunciation similarity and the variant character set of the Chinese character to be compared, wherein the variant characters included in the variant character set have the same pronunciation and meaning as the Chinese character to be compared but are written in different ways; The generating unit is further used to respectively determine each Chinese character in the target character set as the target Chinese character, obtain a Chinese character recognition result of each Chinese character, and generate a similar Chinese character list of the Chinese character to be compared according to the Chinese character recognition result of each Chinese character; The generating unit is further configured to generate feedback information of a similar Chinese character list carrying the Chinese character to be compared in response to obtaining a query request carrying the Chinese character to be compared.

9. A computer device, characterized in that: The computer device comprises a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is configured to execute the method according to any one of claims 1 to 7 according to the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method according to any one of claims 1 to 7.