Word bank data processing method and device, equipment and medium

By detecting the syllables and word counts of the lexicon data, length data and checksum data are generated, which solves the problem of data abnormality detection and repair in the lexicon, and improves the recording accuracy and input experience of entry data.

CN120257979APending Publication Date: 2025-07-04BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410007873.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The data content recorded in the existing thesaurus is difficult to detect abnormalities and repair data, resulting in the inability to accurately record the entry data, affecting the object's input experience.

Method used

By detecting the number of syllables and word numbers of word data, syllable length data and word length data are generated, combined with the total length data and word checksum data, word entry data is established, and abnormal detection and repair processing is performed on the lexicon library.

Benefits of technology

Accurate recording and abnormal detection of entry data is realized, improving the input experience of the object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257979A_ABST
    Figure CN120257979A_ABST
Patent Text Reader

Abstract

The invention discloses a lexicon data processing method and device, equipment and a medium. The method comprises the steps of obtaining word data input by a target object in response to an input selection operation of the target object, detecting the number of syllables and the number of characters of the word data, generating syllable length data and word length data, and generating total length data according to the total number of the number of the syllables and the number of the characters; the code data corresponding to each word in the word data is queried, and the code data corresponding to all the words in the word data are accumulated to obtain word check sum data, so that the word data can be conveniently checked; and according to the word data, the syllable length data, the word length data, the total length data and the word checksum data, establishing entry data, and storing the entry data in a target word bank corresponding to the target object. According to the technical scheme provided by the invention, the entry data can be accurately recorded, and the input experience of the object is improved. The technical scheme of the invention can be widely applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information technology, and particularly to a method, apparatus, device, and medium for processing lexicon data. Background Art

[0002] Currently, with the rapid development of information technology, related applications have gradually integrated into people's lives, providing various services for people. For example, in the text input scenario, there is an application of an object lexicon. The object lexicon can record the commonly used word data input by the object, which can facilitate the object to more conveniently and quickly achieve text content editing, and better meet the personalized needs of different objects.

[0003] In the related art, generally, entry data is established according to the input frequency of each word data, and all the entry data is integrated to form a lexicon corresponding to the object. Since the lexicon is often read and written when used, errors or losses may occur during data transmission. The data content recorded in the current lexicon is difficult to perform anomaly detection and data repair, resulting in inaccurate recording of entry data and easily affecting the input experience of the object.

[0004] In summary, the technical problems existing in the related art need to be improved. Summary of the Invention

[0005] Embodiments of the present application provide a method, apparatus, device, and medium for processing lexicon data, which can improve the development efficiency of the list component and reduce the development cost.

[0006] One aspect of the embodiments of the present application provides a method for processing lexicon data, the method comprising:

[0007] In response to an input selection operation of a target object, obtaining word data input by the target object; the word data includes at least one character;

[0008] Detecting the number of syllables and the number of characters in the word data, generating syllable length data according to the number of syllables, generating word length data according to the number of characters, and generating total length data according to the total number of the number of syllables and the number of characters;

[0009] Querying the encoding data corresponding to each character in the word data, and accumulating the encoding data corresponding to all characters in the word data to obtain word checksum data;

[0010] According to the word data, the syllable length data, the word length data, the total length data, and the word checksum data, establishing entry data, and storing the entry data into a target lexicon corresponding to the target object.

[0011] On the other hand, an embodiment of the present application provides a processing device for lexicon data, and the device includes:

[0012] An acquisition unit, configured to acquire word data input by the target object in response to an input selection operation of the target object; at least one character is included in the word data;

[0013] A generation unit, configured to detect the number of syllables and the number of characters of the word data, generate syllable length data according to the number of syllables, generate word length data according to the number of characters, and generate total length data according to the total number of the number of syllables and the number of characters;

[0014] A query unit, configured to query the encoding data corresponding to each character in the word data, and accumulate the encoding data corresponding to all characters in the word data to obtain word checksum data;

[0015] A creation unit, configured to create entry data according to the word data, the syllable length data, the word length data, the total length data, and the word checksum data, and store the entry data into a target lexicon corresponding to the target object.

[0016] Optionally, the device further includes a processing unit, and the processing unit is specifically configured to:

[0017] In response to a lexicon detection instruction, acquire each entry data in the target lexicon;

[0018] Perform anomaly detection on each of the entry data to obtain an anomaly detection result corresponding to each entry data; wherein, the anomaly detection result is that there is an anomaly or there is no anomaly;

[0019] When the anomaly detection result corresponding to the entry data is that there is an anomaly, perform repair processing on the entry data.

[0020] Optionally, the device further includes a generation unit; the generation unit is specifically configured to:

[0021] Query a first time period between the current time node and a first time node when the last lexicon detection was performed; when the first time period reaches a preset interval threshold, generate a lexicon detection instruction;

[0022] Or, when the storage location of the target lexicon changes, generate a lexicon detection instruction.

[0023] Optionally, the processing unit is specifically configured to:

[0024] In response to a lexicon detection instruction, perform a hash calculation on the current target lexicon to obtain a first hash value;

[0025] Obtain a second hash value, and compare the first hash value with the second hash value; the second hash value is the hash value obtained by performing hash calculation after the target thesaurus was last updated;

[0026] If the second hash value is different from the first hash value, obtain each entry data in the target thesaurus.

[0027] Optionally, the processing unit is specifically configured to:

[0028] Calculate the product of the syllable length data and 2 to obtain a first length data;

[0029] Calculate the product of the word length data and 2 to obtain a second length data;

[0030] Detect whether the first length data, the second length data, and the total length data are the same;

[0031] If any two of the first length data, the second length data, and the total length data are different, determine that the abnormal detection result corresponding to the entry data includes abnormal length data.

[0032] Optionally, the processing unit is specifically configured to:

[0033] If the first length data is the same as the second length data, and the total length data is different from the first length data, determine a new total length data according to the first length data or the second length data;

[0034] Or, if the first length data is the same as the total length data, and the second length data is different from the first length data, determine a new word length data according to the syllable length data or the total length data;

[0035] Or, if the second length data is the same as the total length data, and the first length data is different from the second length data, determine a new syllable length data according to the word length data or the total length data.

[0036] Optionally, the word data is located at the tail of the entry data; the processing unit is specifically configured to:

[0037] If the first length data, the second length data, and the total length data corresponding to the current entry data are all different, locate the next entry data in the target thesaurus;

[0038] Determine the target word data of the current entry data according to the next entry data;

[0039] Redetermine the syllable length data, word length data, and total length data corresponding to the current entry data according to the target word data.

[0040] Optionally, the processing unit is specifically configured to:

[0041] Accumulate the encoding data corresponding to each character in the word data of the entry data to obtain a first accumulated data;

[0042] Match the first accumulated data according to the word checksum data in the entry data;

[0043] If the word checksum data is different from the first accumulated data, determine that the abnormal detection result corresponding to the entry data includes abnormal word data.

[0044] Optionally, the processing unit is specifically configured to:

[0045] Detect the encoding data corresponding to each character in the word data of the entry data, and divide the characters in the word data into first character data and second character data according to the detection results; wherein, the encoding data corresponding to the first character data conforms to a predetermined encoding feature, and the encoding data corresponding to the second character data does not conform to the predetermined encoding feature;

[0046] If the number of the second character data is 1, accumulate the encoding data corresponding to the first character data to obtain a second accumulated data;

[0047] Calculate the difference between the word checksum data and the second accumulated data to obtain target encoding data;

[0048] Determine target character data according to the target encoding data, and replace the second character data with the target character data.

[0049] Optionally, the processing unit is specifically configured to:

[0050] Query the syllables corresponding to each character in the word data of the entry data to obtain the standard syllable data corresponding to each character;

[0051] Detect whether the syllable data corresponding to each character in the entry data is the same as the standard syllable data;

[0052] If there is any character whose corresponding syllable data is different from the standard syllable data, determine that the abnormal detection result corresponding to the entry data includes abnormal syllable data.

[0053] Optionally, the processing unit is specifically configured to:

[0054] When the anomaly detection result corresponding to the entry data includes abnormal syllable data, update the syllable data different from the standard syllable data to the standard syllable data.

[0055] On the other hand, an embodiment of the present application provides an electronic device, including a processor and a memory;

[0056] The memory is used to store a computer program;

[0057] The processor executes the computer program to implement the foregoing method for processing lexicon data.

[0058] On the other hand, an embodiment of the present application provides a computer-readable storage medium, and the storage medium stores a computer program, and the computer program is executed by a processor to implement the foregoing method for processing lexicon data.

[0059] The embodiments of the present application at least include the following beneficial effects: The present application provides a method, device, equipment and medium for processing lexicon data. In response to an input selection operation of a target object, word data input by the target object is obtained, and then the number of syllables and the number of characters of the word data are detected to generate syllable length data, word length data, and total length data is generated according to the total number of syllable numbers and character numbers. There is a specific mathematical relationship between the syllable length data, the word length data, and the total length data, which can facilitate the implementation of data anomaly detection and repair; query the encoding data corresponding to each character in the word data, and accumulate the encoding data corresponding to all characters in the word data to obtain word checksum data, which is convenient for implementing the verification of character data; then, according to the word data, syllable length data, word length data, total length data, and word checksum data, entry data is established, and the entry data is stored in the target lexicon corresponding to the target object. The technical solution provided by the present application can greatly enrich the information recorded in the entry data, can facilitate the anomaly detection and repair processing of various information in the entry data, is beneficial to accurately record the entry data, and improves the input experience of the object. Description of the Drawings

[0060] The drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solutions of the present application, and do not constitute a limitation to the technical solutions of the present application.

[0061] Figure 1 It is a schematic diagram of the principle of an object lexicon provided in an embodiment of the present application;

[0062] Figure 2 It is a schematic diagram of the process for detecting and repairing lexicon data provided in an embodiment of the present application;

[0063] Figure 3Schematic diagram of the implementation environment of a method for processing lexicon data provided in an embodiment of the present application;

[0064] Figure 4 Schematic flow diagram of a method for processing lexicon data provided in an embodiment of the present application;

[0065] Figure 5 Schematic diagram of an interface for a terminal device to display alternative word data provided in an embodiment of the present application;

[0066] Figure 6 Schematic diagram of a target lexicon provided in an embodiment of the present application;

[0067] Figure 7 Schematic diagram of anomaly detection for entry data provided in an embodiment of the present application;

[0068] Figure 8 Schematic diagram of repair processing for entry data provided in an embodiment of the present application;

[0069] Figure 9 Schematic flow diagram of repair processing for length data provided in an embodiment of the present application;

[0070] Figure 10 Schematic diagram of the structure of a device for processing lexicon data provided in an embodiment of the present application;

[0071] Figure 11 Schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Detailed implementation manners

[0072] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods that are consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0073] It is understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".

[0074] The terms "at least one", "a plurality of", "each", "any one", etc. used in this application, at least one includes one, two, or more than two, a plurality of includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.

[0075] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0076] Before elaborating on the embodiments of this application in detail, some nouns and terms involved in the embodiments of this application are first explained, and the nouns and terms involved in the embodiments of this application are subject to the following explanations.

[0077] 1) Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0078] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various major directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0079] 2) Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Essentially, blockchain is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer. Blockchain can include public blockchains, consortium blockchains, and private blockchains. Among them, a public blockchain refers to a blockchain that anyone can access at any time to read data, send data, or compete for bookkeeping; a consortium blockchain refers to a blockchain jointly managed by several organizations or institutions; a private blockchain refers to a blockchain with certain centralized control. The write right of the ledger of the private blockchain is controlled by a certain organization or institution, and the access and use of data are strictly managed by permissions.

[0080] 3) ASCII (American Standard Code for Information Interchange) is a common character encoding standard. It uses 7-bit binary to represent a character, and a total of 128 characters are included, including English letters, numbers, punctuation marks, and some control characters.

[0081] 4) Unicode is a character encoding standard designed to cover all characters globally. It uses different encoding schemes to support different character sets. The most commonly used encoding schemes are UTF-8 (8-bit Unicode Transformation Format) and UTF-16 (16-bit Unicode Transformation Format).

[0082] 5) MD5 is a commonly used hashing algorithm, full name Message Digest Algorithm 5. It receives an input message of any length and outputs a fixed-length digest (hash value), usually 128 bits (16 bytes).

[0083] Currently, with the rapid development of information technology, related applications have gradually been integrated into people's lives, providing various services for people. For example, in the text input scenario, there are applications of object dictionaries. The object dictionary can record the common word data input by the object, which can facilitate the object to more conveniently and quickly edit the text content and better meet the personalized needs of different objects.

[0084] For example, please refer to Figure 1 , Figure 1A schematic diagram of the principle of an object word library is shown. In some applications, input components for inputting text content may be carried, and corresponding word library functions may be configured for these input components. For an object 110 using the application, a corresponding object word library 120 may be established. In the process of the object 110 using the application to input text content, each word data input by the object 110 may be recorded, and entry data may be established or existing entry data may be updated based on the input word data. For example, for example, the text content input by the object 110 at a certain time includes 3 words, namely "trivia", "express delivery", and "family". In the object word library, entry data corresponding to "express delivery" and "family" have been established, and entry data corresponding to "trivia" has not yet been established. Among them, the input frequency recorded in the entry data corresponding to "express delivery" is 3, and the input frequency recorded in the entry data corresponding to "family" is 12. Then, according to the text content input by object 110 this time, the object vocabulary can be updated, the input frequency recorded in the entry data corresponding to "express delivery" can be updated to 4, the input frequency recorded in the entry data corresponding to "family" can be updated to 13, and new entry data corresponding to "trivia" can be created, recording the input frequency as 1.

[0085] The established object word library 120 can provide the object 110 with common word data with a high frequency of occurrence when the object 110 is editing the text content, so as to facilitate the object 110 to input the text content more conveniently and quickly. Figure 1 For example, it is assumed that the current object 110 uses the pinyin input method when inputting text content, and the input pinyin character is "shishi". In the object word library 120, the entry data corresponding to the pinyin character "shishi" stored include word data such as "implementation", "real time", "try", "fact", and "timely", and the corresponding input frequency of "implementation" is 5, the input frequency of "real time" is 3, the input frequency of "try" is 2, and the input frequencies of "fact" and "timely" are both 1. At this time, the input component can give priority to recommending the word data with higher input frequency corresponding to the input pinyin character "shishi" in the word library to the object 110 for selection, for example, the word data such as "implementation", "real time", and "try" can be recommended. It is understandable that word data such as "implementation", "real time" and "try" are used more frequently when object 110 edits text content. Therefore, when the currently input pinyin character is "shishi", the probability that the word data actually wanted to be input by object 110 is "implementation", "real time" and "try" is also relatively high. Recommending them to object 110 for selection can improve the editing efficiency of the text, meet the personalized needs of the object, and improve the input experience of the object.

[0086] In the related art, generally, entry data is established according to the input frequency of each word data, and all the entry data is integrated to form a word library corresponding to the object. Since the word library often needs to be read and written when it is used, errors or losses may occur during data transmission. It is difficult to perform anomaly detection and data repair on the data content recorded in the current word library, resulting in inaccurate recording of entry data and easily affecting the input experience of the object.

[0087] Exemplarily, please refer to Figure 2 , Figure 2 which shows a schematic flowchart for detecting and repairing word library data provided in an embodiment of the present application. Figure 2 In [the reference], the MD5 value calculated by the hash algorithm is used to detect and repair the word library data. Specifically, each time the word library data is updated, an MD5 value will be calculated according to the updated word library data. When performing anomaly detection on the word library, a new MD5 value will be calculated for the current word library data, and the two calculated MD5 values will be compared. If the two are the same, it is considered that there is no anomaly; if the two are different, it is considered that there is an anomaly. At this time, the word library of the most recent backup will be used to replace the current word library to achieve the repair of the word library. However, this implementation method can only detect the overall anomaly of the word library and cannot be detailed to specific entry data. Moreover, there may also be problems with anomalies in the backup word library, and good repair cannot be achieved. Here, it should be noted that Figure 2 the schematic flowchart shown and the related introduction and description are only used to assist in understanding the problems existing in the related technology and do not mean that it belongs to the publicly disclosed prior art.

[0088] In view of this, an embodiment of the present application provides a method, device, equipment and medium for processing word library data. In response to the input selection operation of the target object, the word data input by the target object is obtained, and then the number of syllables and the number of characters of the word data are detected to generate syllable length data, word length data, and total length data is generated according to the total number of the number of syllables and the number of characters. There is a specific mathematical relationship among the syllable length data, the word length data, and the total length data, which can facilitate the realization of data anomaly detection and repair; query the encoding data corresponding to each character in the word data, and accumulate the encoding data corresponding to all the characters in the word data to obtain word checksum data, which is convenient for realizing the verification of character data; then, according to the word data, the syllable length data, the word length data, the total length data, and the word checksum data, entry data is established and the entry data is stored in the target word library corresponding to the target object. The technical solution provided by the present application can greatly enrich the information recorded in the entry data, can facilitate the anomaly detection and repair processing of various information in the entry data, is beneficial to accurately record the entry data, and improves the input experience of the object.

[0089] The method for processing lexicon data provided in the embodiments of the present application mainly relates to the field of information technology, and may include, for example, but not limited to, various application scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving. Those skilled in the art can understand that the method for processing lexicon data provided in the embodiments of the present application can be executed in various application scenarios. Specifically, the following is an exemplary description:

[0090] Exemplarily, in some embodiments, the method for processing lexicon data in the embodiments of the present application can be applied to the scenario of a chat application. In a chat application, objects often use some specific words for communication, such as common greetings, abbreviations, etc. Using lexicon data can record the word data input by the objects, and can provide functions such as intelligent association and automatic recommendation, which is beneficial to improving the input efficiency and reducing input errors.

[0091] Exemplarily, in some embodiments, the method for processing lexicon data in the embodiments of the present application can be applied to the scenario of a search engine application. For example, in a search engine, objects often input some keywords to conduct searches. Using lexicon data can record the keyword data input by the objects, and according to the search history and preferences of the objects, more accurate search results and personalized recommendations can be provided, thereby improving the search experience of the objects.

[0092] Exemplarily, in some embodiments, the method for processing lexicon data in the embodiments of the present application can be applied to the scenario of a geographical location application. For example, in applications such as map navigation or online food ordering, objects often need to input address information. Using the lexicon can record the address data input by the objects, and provide functions such as intelligent association and automatic completion, simplifying the address input process and improving the object experience.

[0093] Of course, it can be understood that the above application scenarios only serve as examples and do not mean to limit the actual application of the method for processing lexicon data provided in the embodiments of the present application. Those skilled in the art can understand that in different application scenarios, the method for processing lexicon data provided in the embodiments of the present application can be used to execute specified tasks.

[0094] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to the identity or characteristics of an object, such as the information of the object, the behavioral data of the object, the historical data of the object, and the location information of the object, the permission or consent of the object will be obtained first. Moreover, the collection, use, and processing of these data will comply with the relevant laws, regulations, and standards of the relevant countries and regions. In addition, when the embodiments of the present application need to obtain sensitive information of the object, the separate permission or separate consent of the object will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the object, the necessary object-related data for the normal operation of the embodiments of the present application will be obtained.

[0095] Next, with reference to the accompanying drawings, the specific embodiments of the embodiments of the present application will be described in detail. First, a method for processing lexicon data provided in the embodiments of the present application will be described with reference to the accompanying drawings.

[0096] Please refer to Figure 3 , Figure 3 which shows a schematic diagram of the implementation environment of a method for processing lexicon data provided in the embodiments of the present application. In this implementation environment, the main software and hardware entities involved include a terminal device 310 and a background server 320. The terminal device 310 and the background server 320 are communicatively connected.

[0097] Specifically, the method for processing lexicon data provided in the embodiments of the present application can be executed solely on the side of the terminal device 310, solely on the side of the background server 320, or based on the data interaction between the terminal device 310 and the background server 320.

[0098] Exemplarily, taking the method for processing lexicon data provided in the embodiments of the present application being executed based on the data interaction between the terminal device 310 and the background server 320 as an example, in some embodiments, the terminal device 310 can be installed with relevant application programs, and the background server 320 can be a server for processing business data related to the application programs. The user (target object) of the terminal device 310 can interact with the terminal device 310 to perform corresponding operations in the application programs, and these operations can include input operations of text data. The background server 320 can trigger the corresponding application program business logic according to the operations performed by the target object on the terminal device 310 and feedback the results to the terminal device 310.

[0099] In the above scenario, the main process of the method for processing lexicon data provided in the embodiments of the present application by the terminal device 310 and the background server 320 through data interaction includes: the terminal device 310 can, in response to an input selection operation of a target object, obtain word data input by the target object, and then send the word data to the background server 320; after receiving the word data, the background server 320 will detect the number of syllables and the number of characters in the word data, generate syllable length data according to the number of syllables, generate word length data according to the number of characters, and generate total length data according to the total number of syllables and characters; then, the background server 320 queries the encoding data corresponding to each character in the word data, accumulates the encoding data corresponding to all characters in the word data, and obtains word checksum data; then, the background server 320 establishes entry data according to the word data, syllable length data, word length data, total length data, and word checksum data, and stores the entry data in the target lexicon corresponding to the target object. Subsequently, on the side of the terminal device 310, the target lexicon stored on the side of the background server 320 can be called to recommend relevant word data to the target object, thereby improving the text editing efficiency.

[0100] Among them, the terminal device 310 in the above embodiments may include a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc., but is not limited thereto.

[0101] The background server 320 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0102] In addition, the background server 320 may also be a node server in a blockchain network.

[0103] A communication connection can be established between the terminal device 310 and the background server 320 through a wireless network or a wired network. The wireless network or the wired network uses standard communication technologies and / or protocols. The network can be set as the Internet or any other network, such as including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network. Moreover, between the above-mentioned software and hardware entities, the same communication connection method or different communication connection methods can be adopted, and the present application does not make specific restrictions on this.

[0104] Of course, it can be understood that Figure 3 the implementation environment in Figure 3 is only some optional application scenarios of the method for processing the thesaurus data provided in the embodiments of the present application. The actual application is not fixed to the

[0105] software and hardware environment shown. The method provided in the embodiments of the present application can be applied to various technical fields, such as cloud technology, artificial intelligence, intelligent transportation, assisted driving and other fields, and the present application does not make specific restrictions on this.

[0106] As Figure 4 shown, in the embodiments of the present application, a method for processing thesaurus data is provided. The method for processing thesaurus data can be applied to the Figure 3 terminal device 310 or the background server 320 shown. Referring to Figure 4 , the method for processing thesaurus data provided in the embodiments of the present application specifically includes but is not limited to steps 410 to 440:

[0107] Step 410: In response to the input selection operation of the target object, obtain the word data input by the target object; the word data includes at least one character;

[0108] In the embodiments of the present application, when processing the thesaurus data, a thesaurus corresponding to the target object may be established. The target object here refers to the object performing text input, which is generally a person. Of course, it may also include other intelligent devices that can independently perform text input tasks. The present application does not limit this. In the embodiments of the present application, the thesaurus corresponding to the target object is denoted as the target thesaurus. It can be understood that the present application does not limit the time node for establishing the target thesaurus. For example, in some embodiments, the target thesaurus corresponding to the target object may be established when the target object first registers; in some embodiments, the target thesaurus corresponding to the target object may be established when the target object first performs an input operation.

[0109] In this step, when processing the thesaurus data corresponding to the target object, the word data input by the target object may be obtained. Specifically, the word data input by the target object may be obtained in response to the input selection operation of the target object. Here, for the target object, what it generally inputs is syllable data. After the target object inputs the syllable data, the input component will display some alternative word data according to the input syllable data.

[0110] Exemplarily, please refer to Figure 5 , Figure 5 which shows a schematic diagram of an interface for a terminal device to display alternative word data provided in the embodiments of the present application. Figure 5 In, the syllable data input by the target object is "gongyuan". After the target object inputs the syllable data of "gongyuan", at least one alternative word data can be displayed on the terminal device. As Figure 5 shown, within the selected area 510 on the terminal device interface, 5 alternative word data are displayed, namely "park", "AD", "imperial examination hall", "engineering institute", and "palace garden".

[0111] It should be noted that in the embodiments of the present application, the displayed alternative word data may be determined according to the previously established target thesaurus. For example, the number of displayed alternative word data may be preset, denoted as the predetermined number. Then, query which word data the syllable data input by the target object may correspond to, and then obtain the input frequencies corresponding to these word data. Next, arrange them in descending order of input frequency, and select the first predetermined number of word data among them to be displayed to the target object as alternative word data. Of course, it can be understood that the above implementation method for displaying alternative word data is only for illustrative purposes and does not mean to limit the display logic of alternative word data in the present application.

[0112] In the embodiment of the present application, after the candidate word data is displayed on the terminal device interface, the target object can interact with the terminal device to perform an input selection operation, thereby determining the word data to be input. Here, the present application does not limit the specific implementation form of the input selection operation. For example, in some embodiments, the input selection operation can be a touch operation performed by the target object on the terminal device interface, for example, Figure 5 As shown, when the target object wishes to input word data "park", the target object can click the number "1" on the virtual keyboard to select the input word data. In some embodiments, the target object can also directly click the word data in the selected area 510 as the input word data, and this application does not limit this.

[0113] Furthermore, it should be noted that in the embodiment of the present application, the selected area 510 may not be able to display all the optional word data corresponding to the input syllable data at one time, so a change button 520 may be set in the selected area 510. When the target object clicks the change button 520, other optional word data will be updated and displayed.

[0114] In this step, when the target object performs the input selection operation, the word data input by the target object can be obtained. It can be understood that the word data includes at least one character, and may also include two or more characters. The embodiment of the present application does not limit the specific number of characters included in the input word data.

[0115] Step 420, detecting the number of syllables and the number of characters in the word data, generating syllable length data according to the number of syllables, generating word length data according to the number of characters, and generating total length data according to the total number of syllables and characters;

[0116] In this step, after determining the word data input by the target object, the word data can be detected to detect the number of syllables and characters therein. Figure 5 , assuming that the target object selects "park" as the input word data, then correspondingly, the word data includes two syllables, namely "gong" and "yuan", that is, the number of syllables is 2; the word data includes two characters, namely "gong" and "yuan", that is, the number of characters is also 2. Here, for Chinese characters, one pinyin syllable corresponds to one Chinese character, so the number of syllables and the number of characters in the determined word data are the same.

[0117] In this step, after determining the number of syllables and the number of characters in the word data, syllable length data can be generated according to the number of syllables, word length data can be generated according to the number of characters, and total length data can be generated according to the total number of syllables and characters. Figure 5For example, assume that the target object selects "park" as the input word data. Then, the corresponding syllable length data of this word data is 2, the word length data is also 2, and the total length data is 4.

[0118] It can be understood that under normal circumstances, that is, when there is no abnormality, the number of syllables and the number of characters in the word data are the same, and the total length data is twice the number of syllables and the number of characters. This feature can be used to assist in determining whether there is an abnormality in the generated entry data. In the embodiments of the present application, relevant parameters of the word data are determined respectively from the aspects of the number of syllables and the number of characters, which can facilitate subsequent abnormal detection and repair of the word data, and can improve the reliability and usability of the thesaurus data.

[0119] Step 430: Query the encoding data corresponding to each character in the word data, and accumulate the encoding data corresponding to all characters in the word data to obtain word checksum data;

[0120] In this step, for the word data, the encoding data corresponding to each character can also be queried. Here, the encoding data refers to the data obtained by converting a character into a number or a character through a predetermined encoding method. In the embodiments of the present application, there are many encoding methods that can be used for encoding characters, such as ASCII, Unicode, etc. The present application does not limit this, and it can be flexibly selected according to actual needs. Through the predetermined encoding method, a character can be mapped to a unique encoding data, and the encoding data corresponding to different characters is different. The specific data form of the obtained encoding data depends on the encoding method adopted. Exemplarily, when the ASCII encoding method is used, each word data can be mapped to a 7-bit or 8-bit binary value; when the Unicode encoding method is used, such as UTF-8, UTF-16, etc., each word data can be mapped to a binary sequence with a fixed length.

[0121] In this step, after obtaining the encoded data corresponding to each character, the encoded data corresponding to all characters in the word data can be accumulated to obtain an accumulated data, which is denoted as the word checksum data in the embodiments of the present application. The word checksum data can record the integrated information of all characters in the word data. If an error occurs in the record of a certain character in the subsequent word data, the abnormality can be detected through the word checksum data. In the embodiments of the present application, the accumulation of the encoded data corresponding to all characters in the word data depends on the format of the data. For example, in some embodiments, if the encoded data corresponding to a character is in integer form, the numerical addition can be directly used during accumulation; in other embodiments, if the encoded data is in binary form, the binary data can be converted into the corresponding decimal value and then added. The present application does not limit this. Exemplarily, assume that a certain word data includes 3 characters, and the encoded data corresponding to each character is a decimal value, which are 65, 97, and 73 respectively. Then, the accumulation of the encoded data corresponding to these 3 characters results in a total of 235. Therefore, the word checksum data corresponding to this word data is 235.

[0122] It can be understood that in the embodiments of the present application, the word checksum data can be used to verify the integrity of the word data. If an error occurs during data transmission or storage, the recorded word checksum data will not match, indicating that the data is damaged or incorrect, which can facilitate subsequent anomaly detection and repair of the data.

[0123] Step 440: Establish an entry data according to the word data, syllable length data, word length data, total length data, and word checksum data, and store the entry data into the target word library corresponding to the target object.

[0124] In this step, after obtaining the syllable length data, word length data, total length data, and word checksum data corresponding to the word data, an entry data can be established according to the word data, syllable length data, word length data, total length data, and word checksum data. Of course, it can be understood that in the embodiments of the present application, for the convenience of using the entry data, the input frequency of the word data can be recorded therein, and other contents may also be included. The present application does not limit this.

[0125] Exemplarily, please refer to Figure 6 , Figure 6 which shows a schematic diagram of a target word library provided in the embodiments of the present application. Referring to Figure 6 , in the embodiments of the present application, the target word library 610 corresponds to the target object, and it may include multiple entry data, Figure 6The specific data content of the entry data 611 is shown. Specifically, the entry data 611 includes word frequency data, word checksum data, total length data, syllable length data, syllable data, word length data, and word data. The word data corresponding to the entry data 611 is "amusement park", and its word frequency data is 5, indicating that the target object has entered "amusement park" a total of 5 times; the word checksum data is obtained by adding up the encoding data corresponding to each character of "you", "le", and "yuan", specifically "85XF78"; the syllable data in the entry data 611 is "youleyuan", its syllable length data is 3, and the word length data is also 3. Therefore, the corresponding total length data is 6. Of course, it can be understood that in the embodiments of the present application, the entry data is not limited to including Figure 6 the content shown. Those skilled in the art can flexibly add, delete, or replace some of the data according to needs, and the present application does not limit this.

[0126] It should be noted that in the embodiments of the present application, when the target object first enters a certain word data, an entry data corresponding to the word data is established and stored in the target word library. When it is found that the word data entered by the target object already has a corresponding entry data in the target word library, at this time, a new entry data is no longer established, but the word frequency data in the existing entry data is accumulated, that is, the word frequency data in the current entry data is incremented by 1.

[0127] In the embodiments of the present application, there is no limit to the specific storage location of the target word library, and it can be flexibly set according to needs.

[0128] It can be understood that in the method for processing word library data provided in the embodiments of the present application, in response to the input selection operation of the target object, the word data input by the target object is obtained, and then the number of syllables and the number of characters of the word data are detected to generate syllable length data and word length data, and the total length data is generated according to the total number of syllables and characters. There is a specific mathematical relationship between the syllable length data, the word length data, and the total length data, which can facilitate the detection and repair of data anomalies; query the encoding data corresponding to each character in the word data, and accumulate the encoding data corresponding to all characters in the word data to obtain the word checksum data, which is convenient for verifying the character data; then, according to the word data, syllable length data, word length data, total length data, and word checksum data, an entry data is established and the entry data is stored in the target word library corresponding to the target object. The method for processing word library data provided by the present application can greatly enrich the information recorded in the entry data, can facilitate the detection and repair processing of various information in the entry data, is beneficial to accurately record the entry data, and improves the input experience of the object.

[0129] Specifically, in a possible implementation manner, the method provided in the embodiments of the present application further includes:

[0130] In response to the lexicon detection instruction, obtain each entry data in the target lexicon;

[0131] Perform anomaly detection on each entry data to obtain the anomaly detection result corresponding to each entry data; wherein, the anomaly detection result is either there is an anomaly or there is no anomaly;

[0132] When the anomaly detection result corresponding to the entry data is that there is an anomaly, perform repair processing on the entry data.

[0133] In the embodiment of the present application, for the entry data stored in the lexicon, anomaly detection and repair processing can be performed according to requirements. Specifically, in the embodiment of the present application, some mechanisms for triggering anomaly detection can be set. When the relevant conditions are met, a lexicon detection instruction can be triggered. In response to this lexicon detection instruction, anomaly detection can be performed on the target lexicon to obtain the anomaly detection results corresponding to each entry data. Here, the anomaly detection result is specifically used to indicate that there is an anomaly in the entry data or there is no anomaly in the entry data. When it is found that the anomaly detection result corresponding to the entry data in the target lexicon is that there is an anomaly, repair processing can be performed on the entry data.

[0134] Specifically, in the embodiment of the present application, when the lexicon detection instruction is triggered, the target lexicon to be detected can be determined according to this instruction. Then, each entry data can be obtained from the target lexicon. For each entry data, anomaly detection can be performed to determine whether there is an anomaly. The specific method for performing anomaly detection will be introduced in subsequent embodiments and will not be elaborated here. After performing anomaly detection on each entry data, the anomaly detection result corresponding to each entry data can be obtained. The anomaly detection result is either there is an anomaly or there is no anomaly. In the embodiment of the present application, there is no limitation on the specific data form and corresponding meaning of the anomaly detection result. Exemplarily, in some embodiments, the anomaly detection result can be represented by a numerical value. For example, it can be numerical values 0 and 1. When there is an anomaly, the anomaly detection result can be 1, and when there is no anomaly, the anomaly detection result can be 0. Of course, the above is only used to exemplarily illustrate the anomaly detection result in the embodiment of the present application. The actual data form of the anomaly detection result and the corresponding meaning can be flexibly adjusted according to needs, and the present application does not limit this.

[0135] In the embodiment of the present application, if it is found that there is no anomaly in each entry data in the target lexicon, it can be considered that the overall result of this anomaly detection is that there is no anomaly, and no optimization processing needs to be performed on the target lexicon. On the contrary, if it is found that there is an anomaly in a certain or some entry data in the target lexicon, repair processing can be performed on it. Specifically, manual intervention, replacement, correction, etc. can be used for repair to make the entry data return to a normal or reasonable state.

[0136] Specifically, in a possible implementation, the lexicon detection instruction is triggered through the following steps:

[0137] Query the first time period between the current time node and the first time node when the last lexicon detection was performed; when the first time period reaches the preset interval threshold, generate a lexicon detection instruction;

[0138] Alternatively, when the storage location of the target lexicon changes, generate a lexicon detection instruction.

[0139] In the embodiments of the present application, as described above, some mechanisms for triggering anomaly detection can be set, and when the relevant conditions are met, the lexicon detection instruction can be triggered. Specifically, in the embodiments of the present application, some triggering logics for the lexicon detection instruction are provided.

[0140] Exemplarily, in some embodiments, the target lexicon can be detected at regular intervals. At this time, a lexicon detection instruction will be generated every time a certain period of time elapses. Specifically, in the embodiments of the present application, a detection period can be preset, denoted as the preset interval threshold, and its specific duration can be set according to requirements. For example, it can be set to one week, one month, or several days, etc. The present application does not limit this. Then, for each target lexicon, the time node when the last lexicon detection was performed can be queried. In the embodiments of the present application, it is denoted as the first time node. Calculate the time difference between the current time node and the first time node, and denote this difference as the first time period. Then, the size of the first time period and the preset interval threshold can be compared. If the first time period has not reached the preset interval threshold, the accumulation can continue; if the first time period reaches the preset interval threshold, it means that a new round of lexicon detection can be performed, and a lexicon detection instruction can be generated. It can be understood that in the embodiments of the present application, after a new round of lexicon detection is completed, the first time node can be updated to facilitate the determination of the new first time period in the future.

[0141] Exemplarily, in some other embodiments, the conditions for generating the lexicon detection instruction can be set considering factors such as possible data errors or data loss in the target lexicon. It can be understood that generally, the target lexicon often needs to perform read and write operations. For example, it may need to be read from the hard disk to the memory or stored from the memory to the hard disk. In the data storage stage, it is generally difficult to have data errors or data loss, while in the data transfer stage, it is relatively easy to have data errors or data loss due to some signal interferences during transmission. Therefore, in the embodiments of the present application, the storage location of the target lexicon can be monitored in real time. A lexicon detection instruction can be generated when the storage location of the target lexicon changes. In this way, the efficiency and effectiveness of lexicon detection can be improved, and unnecessary data processing processes can be reduced.

[0142] Specifically, in a possible implementation, in response to a thesaurus detection instruction, each entry data in the target thesaurus is obtained, including:

[0143] In response to the thesaurus detection instruction, a hash calculation is performed on the current target thesaurus to obtain a first hash value;

[0144] A second hash value is obtained, and the first hash value is compared with the second hash value; the second hash value is the hash value obtained by performing a hash calculation after the last update of the target thesaurus;

[0145] If the second hash value is different from the first hash value, each entry data in the target thesaurus is obtained.

[0146] In the embodiments of the present application, when detecting the target thesaurus, the whole of it can be detected first to determine whether there is any abnormal situation in the scope of the entire target thesaurus. Specifically, in the embodiments of the present application, when receiving the thesaurus detection instruction, a hash calculation can be performed on the current target thesaurus first to obtain a hash value, denoted as the first hash value. Here, when performing the hash calculation, algorithms such as MD5, SHA-1, and SHA-256 can be used. By processing the target thesaurus through these hash algorithms, a unique hash value can be obtained.

[0147] After obtaining the first hash value, another hash value can be used to verify the first hash value. In the embodiments of the present application, the other hash value used is the second hash value, and the second hash value is the hash value obtained by performing a hash calculation after the last update of the target thesaurus. It can be understood that if the target thesaurus has not been updated and is only stored in a certain location or the storage location is transferred, the content therein should remain unchanged. If the target thesaurus has changed from the time of the last update to the current time node, that is, the first hash value is different from the second hash value, it is very likely that there is a problem during storage or transmission, resulting in some abnormal entry data. At this time, each entry data in the target thesaurus can be obtained, and abnormal detection can be performed on each entry data to determine the entry data with abnormalities. On the contrary, if the target thesaurus has not changed from the time of the last update to the current time node, that is, the first hash value is the same as the second hash value, it means that the target thesaurus has not changed during storage or transmission, and it can be considered that the entry data therein is normal and there is no need to perform one-by-one detection and processing. Of course, in the embodiments of the present application, the conditions for performing abnormal detection on the entry data are not restricted. Even when the first hash value is the same as the second hash value, abnormal detection can be performed on each entry data according to actual requirements.

[0148] It can be understood that in the embodiments of the present application, first, the overall abnormality of the target thesaurus is detected based on the hash value. When a problem is found (that is, the first hash value is different from the second hash value), each entry data in the target thesaurus is obtained for individual abnormality detection, which can improve the efficiency of abnormality detection and reduce unnecessary consumption of computing resources.

[0149] Referring Figure 7 , in the embodiments of the present application, the abnormality detection of the entry data mainly includes three parts, namely, the abnormality detection of the length data, the abnormality detection of the word data, and the abnormality detection of the syllable data. The following will explain these three parts one by one. In the embodiments of the present application, when explaining the abnormality detection, it is introduced in the order of the abnormality detection of the length data, the abnormality detection of the word data, and the abnormality detection of the syllable data. It can be understood that when actually performing the abnormality detection, the order of each abnormality detection is not limited, and it can be determined according to actual needs.

[0150] Specifically, in a possible implementation manner, the abnormality detection of each entry data is performed to obtain the abnormality detection result corresponding to each entry data, including:

[0151] Calculate the product of the syllable length data and 2 to obtain the first length data;

[0152] Calculate the product of the word length data and 2 to obtain the second length data;

[0153] Detect whether the first length data, the second length data, and the total length data are the same;

[0154] If any two of the first length data, the second length data, and the total length data are different, it is determined that the abnormality detection result corresponding to the entry data includes length data abnormality.

[0155] In the embodiments of the present application, when performing anomaly detection on the length data in the entry data, it can be implemented based on the syllable length data, word length data, and total length data. From the previous description, it can be known that under normal circumstances, that is, when no anomaly occurs, the number of syllables and the number of characters in the word data are the same, and the total length data is twice the number of syllables and the number of characters. Therefore, in the embodiments of the present application, the product of the syllable length data in the entry data and 2 can be calculated, and the obtained data is recorded as the first length data, and the product of the word length data and 2 can be calculated, and the obtained data is recorded as the second length data. It can be understood that under normal circumstances, the first length data and the second length data should be the same, and both are consistent with the total length data. If any two of the first length data, the second length data, and the total length data are different, it means that at least one of the recorded syllable length data, word length data, and total length data in the entry data is incorrect. At this time, it can be determined that the anomaly detection result corresponding to the entry data includes length data anomaly.

[0156] In the embodiments of the present application, there is no limitation on the situation where the first length data, the second length data, and the total length data are different. For example, in some embodiments, two of the first length data, the second length data, and the total length data may be the same, and the other is different; in some embodiments, the first length data, the second length data, and the total length data may all be different.

[0157] Specifically, in a possible implementation manner, anomaly detection is performed on each entry data, and the anomaly detection result corresponding to each entry data is obtained, including:

[0158] Accumulate the encoding data corresponding to each character in the word data of the entry data to obtain the first accumulated data;

[0159] Match the first accumulated data according to the word checksum data in the entry data;

[0160] If the word checksum data is different from the first accumulated data, it is determined that the anomaly detection result corresponding to the entry data includes word data anomaly.

[0161] In the embodiments of the present application, anomaly detection can also be performed on the word data in the entry data. Specifically, when performing anomaly detection on the word data, the encoding data corresponding to each character in the word data of the entry data can be accumulated, and the accumulated data is recorded as the first accumulated data. Then, the first accumulated data can be matched using the word checksum data in the entry data. It can be understood that if there is no anomaly in the word data in the entry data, the obtained first accumulated data and the word checksum data should be the same. Therefore, if it is found that the word checksum data and the first accumulated data are the same, it can be considered that the word data has no anomaly. On the contrary, if it is found that the word checksum data and the first accumulated data are different, it can be considered that the word data has an anomaly. At this time, it can be determined that the anomaly detection result corresponding to the entry data includes word data anomaly.

[0162] It can be understood that in the embodiments of the present application, by performing anomaly detection on the word data using the word checksum data, the accuracy of the entry data can be improved, the possibility of errors can be reduced, thereby improving the reliability of the thesaurus data and improving the input experience of the object. Moreover, in the embodiments of the present application, based on the word checksum data to implement the anomaly detection of the word data, only an additional data (i.e., the word checksum data) needs to be stored, which can reduce the storage pressure of the data and improve the utilization rate of the hardware resources.

[0163] Specifically, in a possible implementation manner, the entry data further includes syllable data; performing anomaly detection on each entry data to obtain the anomaly detection result corresponding to each entry data, including:

[0164] Query the syllables corresponding to each character in the word data of the entry data to obtain the standard syllable data corresponding to each character;

[0165] Detect whether the syllable data corresponding to each character in the entry data is the same as the standard syllable data;

[0166] If there is any character whose corresponding syllable data is different from the standard syllable data, determine that the anomaly detection result corresponding to the entry data includes syllable data anomaly.

[0167] In the embodiments of the present application, with reference to Figure 6, within the entry data, syllable data may also be included. The syllable data can record the syllables corresponding to each character in the word data. When performing anomaly detection on the entry data, it is also possible to detect whether there are anomalies in the syllable data therein. Specifically, the syllables corresponding to each character in the word data of the entry data can be queried to obtain the standard syllable data corresponding to each character. Then, it can be detected whether the syllable data stored in the entry data corresponding to each character is the same as the standard syllable data. If the two are consistent, it indicates that there are no anomalies in the syllable data in the entry data; on the contrary, if the two are inconsistent, it indicates that there are anomalies in the syllable data in the entry data. At this time, it can be determined that the anomaly detection result corresponding to the entry data includes syllable data anomalies.

[0168] In the embodiments of the present application, there is no limitation on the situation where there are anomalies in the syllable data stored in the entry data. In some embodiments, it may be that the syllable data corresponding to some characters in the word data of the entry data is abnormal; in some embodiments, it may also be that the syllable data corresponding to all characters in the word data of the entry data is abnormal. The present application does not limit the number of characters corresponding to the specifically abnormal syllable data.

[0169] Similarly, referring to Figure 8 , in the embodiments of the present application, the repair process for the entry data mainly includes three parts, namely, the repair process for the length data, the repair process for the word data, and the repair process for the syllable data. The following will explain these three parts one by one. In the embodiments of the present application, when explaining the repair process, it is introduced in the order of the repair process for the length data, the repair process for the word data, and the repair process for the syllable data. It can be understood that when actually performing the repair process, there is no limitation on the order of each repair process, and it can be determined according to actual needs.

[0170] Specifically, in a possible implementation manner, when the anomaly detection result corresponding to the entry data is abnormal, the repair process for the entry data includes:

[0171] If the first length data and the second length data are the same, and the total length data is different from the first length data, determine the new total length data according to the first length data or the second length data;

[0172] Or, if the first length data and the total length data are the same, and the second length data is different from the first length data, determine the new word length data according to the syllable length data or the total length data;

[0173] Or, if the second length data and the total length data are the same, and the first length data is different from the second length data, determine the new syllable length data according to the word length data or the total length data.

[0174] In the embodiments of the present application, with reference to Figure 9 , Figure 9 FIG. 1 shows a schematic flowchart of a repair process for length data provided in the embodiments of the present application. When repairing the length data in the entry data, targeted processing can be performed according to the existing abnormal situations. Specifically, as described above, when the abnormal detection result includes abnormal length data, it may be that two of the first length data, the second length data, and the total length data are the same, and the other is different. It can be understood that if two of the first length data, the second length data, and the total length data are the same, and the other is different, it is very likely that the different data is abnormal. At this time, the abnormal data can be repaired according to the two identical data.

[0175] In some embodiments, if it is found that the first length data and the second length data are the same, but the total length data is different from the first length data and the second length data, at this time, it is very likely that the total length data is abnormal. The total length data can be repaired according to the first length data and the second length data, that is, the first length data or the second length data can be determined as the new total length data. Exemplarily, for example, both the first length data and the second length data are 6, while the total length data is 4, and the total length data can be repaired to 6.

[0176] In some embodiments, if it is found that the first length data and the total length data are the same, but the second length data is different from the first length data and the total length data, at this time, it is very likely that the second length data is abnormal, that is, the word length data is abnormal. The word length data can be repaired according to the syllable length data and the total length data, that is, the syllable length data can be determined as the new word length data, or half of the total length data can be determined as the new word length data. Exemplarily, for example, both the first length data and the total length data are 6, while the second length data is 4, and the word length data is 2. At this time, the word length data can be repaired to 3.

[0177] In some embodiments, if it is found that the second length data and the total length data are the same, but the first length data is different from the second length data and the total length data, at this time, it is very likely that the second length data is abnormal, that is, the syllable length data is abnormal. The syllable length data can be repaired according to the word length data and the total length data, that is, the word length data can be determined as the new syllable length data, or half of the total length data can be determined as the new syllable length data. Exemplarily, for example, both the second length data and the total length data are 6, while the first length data is 4, and the syllable length data is 2. At this time, the syllable length data can be repaired to 3.

[0178] Specifically, in a possible implementation, the word data is located at the end of the entry data; when the anomaly detection result corresponding to the entry data indicates an anomaly, the entry data is repaired, including:

[0179] If the first length data, second length data, and total length data corresponding to the current entry data are all different, locate the next entry data in the target dictionary;

[0180] Determine the target word data of the current entry data according to the next entry data;

[0181] Redetermine the syllable length data, word length data, and total length data corresponding to the current entry data according to the target word data.

[0182] In the embodiments of the present application, as described above, when the anomaly detection result includes abnormal length data, it is possible that the first length data, second length data, and total length data are all different. At this time, it is necessary to determine the actual length data of the word data, and repair the syllable length data, word length data, and total length data according to the actual length data. Specifically, in the embodiments of the present application, the word data in the entry data can be placed at the end. When it is found that the first length data, second length data, and total length data corresponding to the current entry data are all different, the next entry data can be located in the target dictionary first.

[0183] Specifically, in the embodiments of the present application, the position of the data that conforms to the relationship can be determined according to the relationship between the total length data and the syllable length data, word length data, that is, the next entry data can be located. And the word data in the entry data is at the end, and the starting position of the next entry data is the end position of the word data of the current entry data. Detect forward from the end position of the word data of the current entry data, and judge the characters that conform to the coding data format, and the word data of the current entry data can be determined. In the embodiments of the present application, it is denoted as the target word data. It can be understood that after the target word data is determined, the actual length of the word data of the current entry data can be determined according to the number of characters in the target word data, that is, the number of characters that conform to the coding data format found. Exemplarily, for example, if the determined actual length is 4, the syllable length data corresponding to the current entry data can be redetermined as 4, the word length data as 4, and the total length data as 8.

[0184] Specifically, in a possible implementation, when the anomaly detection result corresponding to the entry data indicates an anomaly, the entry data is repaired, including:

[0185] Detect the encoding data corresponding to each character in the word data of the entry data, and divide the characters in the word data into first character data and second character data according to the detection results; among them, the encoding data corresponding to the first character data conforms to the predetermined encoding characteristics, and the encoding data corresponding to the second character data does not conform to the predetermined encoding characteristics;

[0186] If the number of the second character data is 1, accumulate the encoding data corresponding to the first character data to obtain a second accumulated data;

[0187] Calculate the difference between the word checksum data and the second accumulated data to obtain the target encoding data;

[0188] Determine the target character data according to the target encoding data, and replace the second character data with the target character data.

[0189] In the embodiment of the present application, when an abnormality occurs in the word data of the entry data, generally, an error occurs in the encoding data corresponding to the word data. Therefore, when repairing the word data in the entry data, the encoding data corresponding to each character in the word data of the entry data can be detected. If the encoding data corresponding to a certain character conforms to the predetermined encoding characteristics, then this character can be determined as the first character data; if the encoding data corresponding to a certain character does not conform to the predetermined encoding characteristics, then this character can be determined as the second character data. Here, if the number of the second character data exceeds 1, it means that multiple characters in the word data have abnormalities, and manual intervention is required for repair, or this entry can be discarded. If the number of the second character data is 1, automatic repair can be performed.

[0190] Specifically, when performing automatic repair, the encoding data corresponding to all the first character data can be accumulated, and the obtained accumulated data is recorded as the second accumulated data. Then, the difference between the word checksum data and the second accumulated data can be calculated to obtain the target encoding data. It can be understood that since there is only one second character data, it means that there is probably only one character in the word data that has an abnormality. By taking the difference between the word checksum data and the second accumulated data, the obtained target encoding data is the encoding data corresponding to the actual correct character. Therefore, the target character data can be determined according to the target encoding data, and the second character data can be replaced with the target character data, thereby completing the repair of the entry data.

[0191] Specifically, in a possible implementation manner, when the abnormality detection result corresponding to the entry data indicates an abnormality, the repair process for the entry data includes:

[0192] When the abnormality detection result corresponding to the entry data includes abnormal syllable data, update the syllable data different from the standard syllable data to the standard syllable data.

[0193] In the embodiments of the present application, when abnormal syllable data appears in the entry data, the abnormal syllable data can be replaced with standard syllable data. In this way, the repair process of the syllable data can be efficiently completed.

[0194] Next, in combination with a specific application implementation process, the method for processing the thesaurus data provided in the present application will be introduced and described in detail.

[0195] In the embodiments of the present application, a method for processing thesaurus data is provided. The method mainly includes a process of establishing a target thesaurus and a process of performing abnormal detection and repair processing on the entry data according to the target thesaurus.

[0196] Specifically, in the embodiments of the present application, in response to the input selection operation of the target object, the word data input by the target object is obtained, and then the number of syllables and the number of characters in the word data are detected to generate syllable length data, word length data, and total length data according to the total number of syllables and characters. There is a specific mathematical relationship among the syllable length data, word length data, and total length data, which can facilitate the implementation of abnormal detection and repair of data; query the encoding data corresponding to each character in the word data, and accumulate the encoding data corresponding to all characters in the word data to obtain word checksum data, which is convenient for implementing the verification of character data; then, according to the word data, syllable length data, word length data, total length data, and word checksum data, entry data is established and stored in the target thesaurus corresponding to the target object. The technical solution provided by the present application can greatly enrich the information recorded in the entry data, facilitate the abnormal detection and repair processing of various information in the entry data, is conducive to accurately recording the entry data, and improves the input experience of the object.

[0197] In the embodiments of the present application, when performing abnormal detection on the entry data in the target thesaurus, the abnormal detection of the length data can be performed using the syllable length data, word length data, and total length data. Specifically, the product of the syllable length data and 2 can be calculated to obtain the first length data, and the product of the word length data and 2 can be calculated to obtain the second length data. Then, it is detected whether the first length data, the second length data, and the total length data are the same. If there is a situation where they are not the same, it indicates that there is an abnormality in the length data of the entry data. In the embodiments of the present application, the abnormal detection of the word data in the entry data and the abnormal detection of the syllable data in the entry data can also be performed based on the word checksum data in the entry data. Moreover, according to the syllable length data, word length data, and total length data, the automatic repair processing of the length data can be realized; according to the word checksum data, the automatic repair processing of the word data can be performed.

[0198] It can be understood that the technical solution provided by this application can greatly enrich the information recorded in the entry data, facilitate the detection and repair of various information in the entry data, be conducive to accurately recording the entry data, and improve the input experience of the object.

[0199] Referring to Figure 10 , this embodiment of the application also provides a processing device for lexicon data, and the device includes:

[0200] An acquisition unit 1010, configured to acquire word data input by a target object in response to an input selection operation of the target object; the word data includes at least one character;

[0201] A generation unit 1020, configured to detect the number of syllables and the number of characters in the word data, generate syllable length data according to the number of syllables, generate word length data according to the number of characters, and generate total length data according to the total number of syllables and characters;

[0202] A query unit 1030, configured to query the encoding data corresponding to each character in the word data, and accumulate the encoding data corresponding to all characters in the word data to obtain word checksum data;

[0203] A creation unit 1040, configured to create entry data according to the word data, syllable length data, word length data, total length data, and word checksum data, and store the entry data in a target lexicon corresponding to the target object.

[0204] Optionally, the device further includes a processing unit, and the processing unit is specifically configured to:

[0205] In response to a lexicon detection instruction, acquire each entry data in the target lexicon;

[0206] Perform anomaly detection on each entry data to obtain an anomaly detection result corresponding to each entry data; wherein, the anomaly detection result is that there is an anomaly or there is no anomaly;

[0207] When the anomaly detection result corresponding to the entry data is that there is an anomaly, perform repair processing on the entry data.

[0208] Optionally, the device further includes a generation unit; the generation unit is specifically configured to:

[0209] Query the first time period between the current time node and the first time node when the last lexicon detection was performed; when the first time period reaches a preset interval threshold, generate a lexicon detection instruction;

[0210] Or, when the storage location of the target lexicon changes, generate a lexicon detection instruction.

[0211] Optionally, the processing unit is specifically configured to:

[0212] In response to the lexicon detection instruction, perform a hash calculation on the current target lexicon to obtain a first hash value;

[0213] Obtain a second hash value, and compare the first hash value with the second hash value; the second hash value is the hash value obtained by performing a hash calculation after the last update of the target lexicon;

[0214] If the second hash value is different from the first hash value, obtain each entry data in the target lexicon.

[0215] Optionally, the processing unit is specifically configured to:

[0216] Calculate the product of the syllable length data and 2 to obtain a first length data;

[0217] Calculate the product of the word length data and 2 to obtain a second length data;

[0218] Detect whether the first length data, the second length data, and the total length data are the same;

[0219] If any two of the first length data, the second length data, and the total length data are different, determine that the abnormal detection result corresponding to the entry data includes abnormal length data.

[0220] Optionally, the processing unit is specifically configured to:

[0221] If the first length data and the second length data are the same, and the total length data is different from the first length data, determine a new total length data according to the first length data or the second length data;

[0222] Or, if the first length data and the total length data are the same, and the second length data is different from the first length data, determine a new word length data according to the syllable length data or the total length data;

[0223] Or, if the second length data and the total length data are the same, and the first length data is different from the second length data, determine a new syllable length data according to the word length data or the total length data.

[0224] Optionally, the word data is located at the end of the entry data; the processing unit is specifically configured to:

[0225] If the first length data, the second length data, and the total length data corresponding to the current entry data are all different, locate the next entry data in the target lexicon;

[0226] According to the next entry data, determine the target word data of the current entry data;

[0227] According to the target word data, re-determine the syllable length data, the word length data, and the total length data corresponding to the current entry data.

[0228] Optionally, the processing unit is specifically configured to:

[0229] Accumulate the encoding data corresponding to each character in the word data of the entry data to obtain first accumulated data;

[0230] Match the first accumulated data according to the word checksum data in the entry data;

[0231] If the word checksum data is different from the first accumulated data, determine that the anomaly detection result corresponding to the entry data includes word data anomaly.

[0232] Optionally, the processing unit is specifically configured to:

[0233] Detect the encoding data corresponding to each character in the word data of the entry data, and divide the characters in the word data into first word data and second word data according to the detection results; wherein, the encoding data corresponding to the first word data conforms to the predetermined encoding characteristics, and the encoding data corresponding to the second word data does not conform to the predetermined encoding characteristics;

[0234] If the number of the second word data is 1, accumulate the encoding data corresponding to the first word data to obtain second accumulated data;

[0235] Calculate the difference between the word checksum data and the second accumulated data to obtain target encoding data;

[0236] Determine target word data according to the target encoding data, and replace the second word data with the target word data.

[0237] Optionally, the processing unit is specifically configured to:

[0238] Query the syllables corresponding to each character in the word data of the entry data to obtain the standard syllable data corresponding to each character;

[0239] Detect whether the syllable data corresponding to each character in the entry data is the same as the standard syllable data;

[0240] If there is any character whose corresponding syllable data is different from the standard syllable data, determine that the anomaly detection result corresponding to the entry data includes syllable data anomaly.

[0241] Optionally, the processing unit is specifically configured to:

[0242] When the anomaly detection result corresponding to the entry data includes syllable data anomaly, update the syllable data that is different from the standard syllable data to the standard syllable data.

[0243] It can be understood that, such as Figure 4The content in the embodiments of the method for processing the thesaurus data shown is applicable to the embodiments of the apparatus for processing the thesaurus data of the present application. The functions specifically implemented by the embodiments of the apparatus for processing the thesaurus data of the present application are the same as those in the embodiments of the method for processing the thesaurus data shown in Figure 4 and the beneficial effects achieved are also the same as those in the embodiments of the method for processing the thesaurus data shown in Figure 4 The embodiments of the present application also disclose an electronic device, including:

[0244] At least one processor;

[0245] At least one memory for storing at least one program;

[0246] When the at least one program is executed by the at least one processor, the at least one processor implements the embodiments of the method for processing the thesaurus data shown in

[0247] Figure 4 It can be understood that the content in the embodiments of the method for processing the thesaurus data shown in When the at least one program is executed by the at least one processor, the at least one processor implements the embodiments of the method for processing the thesaurus data shown in

[0248] is applicable to the embodiments of the present electronic device. The functions specifically implemented by the embodiments of the present electronic device are the same as those in the embodiments of the method for processing the thesaurus data shown in Figure 4 and the beneficial effects achieved are also the same as those in the embodiments of the method for processing the thesaurus data shown in Figure 4 Figure 4 The electronic device of the embodiments of the present application may be a terminal device, a computer device, or a server device. The beneficial effects achieved are also the same as those in the embodiments of the method for processing the thesaurus data shown in

[0249] Exemplarily, referring to

[0250] Figure 11 is a schematic structural diagram of an electronic device provided in the embodiments of the present application. Taking the electronic device as a terminal device as an example, Figure 11 in, the electronic device 1100 may include an RF (Radio Frequency) circuit 1110, a memory 1120 including one or more computer-readable storage media, an input unit 1130, a display unit 1140, a sensor 1150, an audio circuit 1160, a short-range wireless transmission module 1170, a processor 1180 including one or more processing cores, and a power supply 1190 and other components. Those skilled in the art can understand that Figure 11 the device structure shown in Figure 11 does not limit the terminal device, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0251] ​The RF circuit 1110 can be used for receiving and transmitting information or signals during a call. Specifically, after receiving the downlink information from the base station, it is handed over to one or more processors 1180 for processing. Additionally, data related to the uplink is sent to the base station. Generally, the RF circuit 1110 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a Subscriber Identity Module (SIM) card, a transceiver, a coupler, a Low Noise Amplifier (LNA), a duplexer, etc. In addition, the RF circuit 1110 can also communicate with the network and other devices via wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), LTE (Long Term Evolution), email, SMS (Short Messaging Service), etc.

[0252] The memory 1120 can be used to store software programs and modules (or units). The processor 1180 executes various functional applications and data processing by running the software programs and modules (or units) stored in the memory 1120. The memory 1120 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, applications required for at least one function (such as a sound playback function, an image playback function), etc.; the data storage area can store data created according to the use of the electronic device 1100 (such as audio data, a phone book), etc. In addition, the memory 1120 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory 1120 can also include a memory controller to provide access to the memory 1120 by the processor 1180 and the input unit 1130. Although Figure 11 the RF circuit 1110 is shown, it can be understood that it does not necessarily constitute a part of the electronic device 1100 and can be omitted entirely within the scope of not changing the essence of the invention according to needs.

[0253] The input unit 1130 can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to object settings and function controls. Specifically, the input unit 1130 can include a touch-sensitive surface 1131 and other input devices 1132. The touch-sensitive surface 1131, also known as a touch display screen or a touchpad, can collect touch operations of an object on or near it (such as operations of the object using any suitable object or accessory such as a finger, a stylus, etc. on or near the touch-sensitive surface 1131), and drive corresponding connection devices according to a pre-set program. Optionally, the touch-sensitive surface 1131 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the object and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1180, and can receive and execute the instructions sent by the processor 1180. In addition, the touch-sensitive surface 1131 can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 1131, the input unit 1130 can also include other input devices 1132. Specifically, the other input devices 1132 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.

[0254] The display unit 1140 can be used to display information input by an object or information provided to the object and control various graphical object interfaces of the electronic device 1100. These graphical object interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 1140 can include a display panel 1141. Optionally, the display panel 1141 can be configured in forms such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode). Further, the touch-sensitive surface 1131 can cover the display panel 1141. After the touch-sensitive surface 1131 detects a touch operation on or near it, it is transmitted to the processor 1180 to determine the type of touch event. Subsequently, the processor 1180 provides a corresponding visual output on the display panel 1141 according to the type of touch event. Although in Figure 11 the touch-sensitive surface 1131 and the display panel 1141 are implemented as two independent components to realize input and input functions, in some embodiments, the touch-sensitive surface 1131 and the display panel 1141 can be integrated to realize input and output functions.

[0255] The electronic device 1100 may further include at least one sensor 1150, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 1141 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1141 or the backlight when the electronic device 1100 is moved to the ear. As a kind of motion sensor, the gravity acceleration sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used in applications for identifying the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors that the electronic device 1100 may also be configured with, they will not be elaborated here.

[0256] The audio circuit 1160, the speaker 1161, and the microphone 1162 can provide an audio interface between the object and the electronic device 1100. The audio circuit 1160 can transmit the electrical signal converted from the received audio data to the speaker 1161, and the speaker 1161 converts it into a sound signal for output; on the other hand, the microphone 1162 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1160 and then converted into audio data. After the audio data is output to the processor 1180 for processing, it is sent to another electronic device through the RF circuit 1110, or the audio data is output to the memory 1120 for further processing. The audio circuit 1160 may also include an earphone jack to provide communication between the peripheral earphone and the electronic device 1100.

[0257] The short-range wireless transmission module 1170 may be a WIFI (wireless fidelity) module, a Bluetooth module, or an infrared module, etc. The electronic device 1100 can transmit information with the wireless transmission module set on other devices through the short-range wireless transmission module 1170.

[0258] The processor 1180 is the control center of the electronic device 1100, connecting various parts of the entire device through various interfaces and lines. By running or executing software programs or modules stored in the memory 1120, and calling data stored in the memory 1120, it executes various functions of the electronic device 1100 and processes data, thereby overall controlling the device. Optionally, the processor 1180 may include one or more processing cores; optionally, the processor 1180 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, the object interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1180 either.

[0259] The electronic device 1100 also includes a power source 1190 (such as a battery) for supplying power to each component. Optionally, the power source 1190 can be logically connected to the processor 1180 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power source 1190 can also include any components such as one or more DC or AC power sources, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0260] Although not shown, the electronic device 1100 may also include a camera, a Bluetooth module, etc., which will not be elaborated here.

[0261] An embodiment of the present application also discloses a computer-readable storage medium, which stores a program executable by a processor. The program executable by the processor, when executed by the processor, is used to implement Figure 4 the method embodiment for processing the thesaurus data as shown.

[0262] It can be understood that Figure 4 the content in the method embodiment for processing the thesaurus data as shown is applicable to this computer-readable storage medium embodiment. The functions specifically implemented by this computer-readable storage medium embodiment are the same as Figure 4 those in the method embodiment for processing the thesaurus data as shown, and the beneficial effects achieved are also the same as Figure 4 those achieved by the method embodiment for processing the thesaurus data as shown.

[0263] An embodiment of the present application also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in the above-mentioned computer-readable storage medium; Figure 11 the processor of the electronic device as shown can read the computer instructions from the above-mentioned computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 4 the method embodiment for processing the thesaurus data as shown.

[0264] It can be understood that Figure 4 the content in the method embodiment for processing the thesaurus data as shown is applicable to this computer program product or computer program embodiment. The functions specifically implemented by this computer program product or computer program embodiment are the same as Figure 4 those in the method embodiment for processing the thesaurus data as shown, and the beneficial effects achieved are also the same as Figure 4 those achieved by the method embodiment for processing the thesaurus data as shown.

[0265] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. Further, the embodiments presented and described in the flowcharts of the present application are provided by way of example for the purpose of providing a more thorough understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are executed independently.

[0266] In addition, although the present application has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present application. Rather, given the attributes, functions and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Thus, those of ordinary skill in the art will be able to implement the present application as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are illustrative only and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.

[0267] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present application, in essence or the part that contributes to the prior art or part of the technical solution, may be embodied in the form of a software product stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0268] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable storage medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0269] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0270] In the above description of this specification, the descriptions referring to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0271] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the claims and their equivalents.

[0272] The above has specifically described the preferred embodiments of the present application, but the present application is not limited to the embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.

Claims

1. A method for processing thesaurus data, characterized in that The method includes: In response to an input selection operation of a target object, obtaining word data input by the target object; the word data includes at least one character. Detecting the number of syllables and the number of characters in the word data, generating syllable length data according to the number of syllables, generating word length data according to the number of characters, and generating total length data according to the total number of syllables and characters. Querying the encoding data corresponding to each character in the word data, and accumulating the encoding data corresponding to all characters in the word data to obtain word checksum data. According to the word data, the syllable length data, the word length data, the total length data, and the word checksum data, establishing entry data and storing the entry data into a target word library corresponding to the target object.

2. The method for processing the thesaurus data according to claim 1, wherein The method further includes: In response to a word library detection instruction, obtaining each entry data in the target word library. Performing anomaly detection on each of the entry data to obtain an anomaly detection result corresponding to each entry data; wherein, the anomaly detection result is that there is an anomaly or there is no anomaly. When the anomaly detection result corresponding to the entry data is that there is an anomaly, performing repair processing on the entry data.

3. The method for processing the thesaurus data according to claim 2, wherein The word library detection instruction is triggered through the following steps: Querying a first time period between the current time node and a first time node when the word library was last detected; when the first time period reaches a preset interval threshold, generating a word library detection instruction. Or, when the storage location of the target word library changes, generating a word library detection instruction.

4. The method for processing the thesaurus data according to claim 2, wherein The obtaining each entry data in the target word library in response to the word library detection instruction includes: In response to the word library detection instruction, performing a hash calculation on the current target word library to obtain a first hash value. Obtaining a second hash value and comparing the first hash value with the second hash value; the second hash value is the hash value obtained by performing a hash calculation after the target word library was last updated. If the second hash value is different from the first hash value, obtaining each entry data in the target word library.

5. The processing method of the lexicon data according to claim 2, wherein The performing anomaly detection on each of the entry data to obtain an anomaly detection result corresponding to each entry data includes: Calculating the product of the syllable length data and 2 to obtain first length data. Calculating the product of the word length data and 2 to obtain second length data. Detecting whether the first length data, the second length data, and the total length data are the same. If any two of the first length data, the second length data, and the total length data are different, determining that the anomaly detection result corresponding to the entry data includes length data anomaly.

6. The method for processing the thesaurus data according to claim 5, wherein The performing repair processing on the entry data when the anomaly detection result corresponding to the entry data is that there is an anomaly includes: If the first length data and the second length data are the same, and the total length data is different from the first length data, determining new total length data according to the first length data or the second length data. Alternatively, if the first length data is the same as the total length data, and the second length data is different from the first length data, determine new word length data according to the syllable length data or the total length data; Alternatively, if the second length data is the same as the total length data, and the first length data is different from the second length data, determine new syllable length data according to the word length data or the total length data.

7. The method for processing the lexicon data according to claim 5, wherein The word data is located at the tail of the entry data; when the anomaly detection result corresponding to the entry data indicates an anomaly, the repair process for the entry data includes: If the first length data, the second length data, and the total length data corresponding to the current entry data are all different, locate the next entry data in the target word library; Determine the target word data of the current entry data according to the next entry data; According to the target word data, re-determine the syllable length data, word length data, and total length data corresponding to the current entry data.

8. The method for processing the thesaurus data according to claim 2, wherein The anomaly detection for each entry data to obtain the anomaly detection result corresponding to each entry data includes: Accumulate the encoding data corresponding to each character in the word data of the entry data to obtain a first accumulated data; Match the first accumulated data according to the word checksum data in the entry data; If the word checksum data is different from the first accumulated data, determine that the anomaly detection result corresponding to the entry data includes word data anomaly.

9. The method for processing the lexicon data according to claim 8, wherein When the anomaly detection result corresponding to the entry data indicates an anomaly, the repair process for the entry data includes: Detect the encoding data corresponding to each character in the word data of the entry data, and divide the characters in the word data into first character data and second character data according to the detection result; wherein, the encoding data corresponding to the first character data conforms to a predetermined encoding feature, and the encoding data corresponding to the second character data does not conform to the predetermined encoding feature; If the number of the second character data is 1, accumulate the encoding data corresponding to the first character data to obtain a second accumulated data; Calculate the difference between the word checksum data and the second accumulated data to obtain target encoding data; Determine target character data according to the target encoding data, and replace the second character data with the target character data.

10. The method for processing the lexicon data according to claim 2, wherein, The entry data further includes syllable data; the anomaly detection for each entry data to obtain the anomaly detection result corresponding to each entry data includes: Query the syllables corresponding to each character in the word data of the entry data to obtain the standard syllable data corresponding to each character; Detect whether the syllable data corresponding to each character in the entry data is the same as the standard syllable data; If there is any character whose corresponding syllable data is different from the standard syllable data, determine that the anomaly detection result corresponding to the entry data includes syllable data anomaly.

11. The method for processing the thesaurus data according to claim 10, wherein When the anomaly detection result corresponding to the entry data indicates an anomaly, the repair process for the entry data includes: When the anomaly detection result corresponding to the entry data includes abnormal syllable data, update the syllable data different from the standard syllable data to the standard syllable data.

12. A processing device for lexicon data, characterized in that, The device includes: An acquisition unit, configured to acquire word data input by the target object in response to an input selection operation of the target object; the word data includes at least one character; A generation unit, configured to detect the number of syllables and the number of characters in the word data, generate syllable length data according to the number of syllables, generate word length data according to the number of characters, and generate total length data according to the total number of the number of syllables and the number of characters; A query unit, configured to query the encoding data corresponding to each character in the word data, and accumulate the encoding data corresponding to all the characters in the word data to obtain word checksum data; A creation unit, configured to create entry data according to the word data, the syllable length data, the word length data, the total length data, and the word checksum data, and store the entry data into a target word library corresponding to the target object.

13. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the method for processing word library data according to any one of claims 1 to 11.

14. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for processing word library data according to any one of claims 1 to 11.

15. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for processing word library data according to any one of claims 1 to 11.