Method and device for managing data items, equipment and storage medium

By dividing the continuously encoded subsequences and determining the mapping characteristics, the problem of inefficient storage and conversion of complex mapping relationships in the prior art is solved, and more efficient data item management and conversion are achieved.

CN120020815APending Publication Date: 2025-05-20BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311550835.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

The prior art takes up a lot of space and is inefficient in storing and converting complex mapping relationships, making it difficult to effectively manage the conversion process between different types of data items.

Method used

By obtaining the sequence of multiple data items, dividing the continuously encoded subsequences, determining the mapping characteristics of the subsequences, including boundaries and encoding differences, and mapping the data items based on these characteristics.

Benefits of technology

It reduces the space occupation of storage mapping relationships and the complexity of the conversion process, and improves the efficiency of data item management and conversion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020815A_ABST
    Figure CN120020815A_ABST
Patent Text Reader

Abstract

According to the implementation mode of the invention, a method and device for managing data items, equipment and a storage medium are provided. The method comprises the following steps: acquiring a first data sequence comprising a plurality of first data items and a second data sequence comprising a plurality of second data items; based on the first data sequence and the second data sequence, the first data sequence is divided into at least one sub-sequence, and codes of a group of first data items in a target sub-sequence in the at least one sub-sequence are continuous; determining mapping features of at least one sub-sequence; and mapping a first data item in the plurality of first data items to a second data item in the plurality of second data items based on the mapping feature of the at least one sub-sequence. In this way, the efficiency can be improved while the accuracy of managing the data items is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various implementations of the present disclosure relate to the field of computers, and in particular, to methods, apparatuses, devices, and computer-readable storage media for managing data items. Background Art

[0002] In a computer system, various types of data items can be stored, and there may be complex mapping relationships between different types of data items. These mapping relationships can be pre-stored, and when data conversion needs to be performed, the mapping relationships are called to convert a data item of one type to another type. However, the mapping relationships are usually complex, which results in a large amount of space being occupied for storing the mapping relationships. Further, the conversion process may involve complex algorithms, which results in low conversion efficiency. At this time, it is desirable to manage data items in a more convenient and efficient manner, and thus implement the conversion process between different types of data items. Summary of the Invention

[0003] In a first aspect of the present disclosure, a method for managing data items is provided. The method includes: obtaining a first data sequence including a plurality of first data items and a second data sequence including a plurality of second data items; based on the first data sequence and the second data sequence, dividing the first data sequence into at least one subsequence, where the encoding of a group of first data items in a target subsequence among the at least one subsequence is continuous; determining the mapping characteristics of the at least one subsequence, where the target mapping characteristics of the target subsequence include: the first boundary and the second boundary of a group of first data items in the target subsequence, and the difference between the encoding of the data items in the target subsequence and the encoding of another data item corresponding to the data item in the second data sequence; and based on the mapping characteristics of the at least one subsequence, mapping the first data items among the plurality of first data items to the second data items among the plurality of second data items.

[0004] In a second aspect of the present disclosure, there is provided an apparatus for managing data items. The apparatus includes: a sequence acquisition module configured to acquire a first data sequence including a plurality of first data items and a second data sequence including a plurality of second data items; a sequence division module configured to divide the first data sequence into at least one subsequence based on the first data sequence and the second data sequence, wherein the encoding of a set of first data items in a target subsequence among the at least one subsequence is continuous; a feature determination module configured to determine mapping features of the at least one subsequence, and the target mapping features of the target subsequence include: the first boundary and the second boundary of a set of first data items in the target subsequence, and the difference between the encoding of the data items in the target subsequence and the encoding of another data item in the second data sequence corresponding to the data item; and a data mapping module configured to map the first data items among the plurality of first data items to the second data items among the plurality of second data items based on the mapping features of the at least one subsequence.

[0005] In a third aspect of the present disclosure, there is provided an electronic device. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to execute the method according to the first aspect of the present disclosure when executed by the at least one processing unit.

[0006] In a fourth aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program, which when executed by a processor causes the processor to implement the method according to the first aspect of the present disclosure.

[0007] It should be understood that the content described in this part is not intended to define the key features or important features of the implementation manners of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In the following, in combination with the drawings and with reference to the following detailed description, the above and other features, advantages and aspects of the various implementation manners of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0009] Figure 1 A schematic diagram showing an example environment in which the implementation manners of the present disclosure can be implemented;

[0010] Figure 2A Showing the Unicode corresponding to the upper and lower case characters of English;

[0011] Figure 2B Showing the character set corresponding to the upper and lower case characters of English;

[0012] Figure 2C A hash table for storing upper and lower case characters in English is shown;

[0013] Figure 3 A flowchart of a method for managing data items according to some implementations of the present disclosure is shown;

[0014] Figure 4 A schematic diagram showing an example of determining at least one subsequence according to some implementations of the present disclosure is shown;

[0015] Figure 5 A schematic diagram showing an example of determining target mapping features according to some implementations of the present disclosure is shown;

[0016] Figure 6 A block diagram schematically showing the process of updating an index according to some implementations of the present disclosure is shown;

[0017] Figure 7 A block diagram of a device for managing data items according to some implementations of the present disclosure is shown; and

[0018] Figure 8 A block diagram of a device capable of implementing multiple implementations of the present disclosure is shown. Detailed Implementations

[0019] Implementations of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain implementations of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. On the contrary, these implementations are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and implementations of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.

[0020] In the description of the implementations of the present disclosure, the term "including" and its like should be understood as an open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions below.

[0021] It can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.

[0022] It is understandable that before using the technical solutions disclosed in various implementations of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to users and the authorization of users should be obtained through appropriate means in accordance with relevant laws and regulations.

[0023] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.

[0024] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may, for example, be in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0025] It is understandable that the above process of notifying and obtaining user authorization is only illustrative and does not constitute a limitation on the implementation manners of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manners of the present disclosure.

[0026] The term "in response to" used herein represents a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the execution timing of the subsequent actions executed in response to the event or condition and the time when the event occurs or the condition is established are not necessarily strongly correlated. For example, in some cases, the subsequent actions can be executed immediately when the event occurs or the condition is established; while in other cases, the subsequent actions can be executed after a period of time after the event occurs or the condition is established.

[0027] Example environment

[0028] Figure 1 FIG. shows a schematic diagram of an exemplary environment 100 in which the implementations of the present disclosure can be implemented. In this exemplary environment 100, an application 120 is installed in a terminal device 110. A user 140 can interact with the application 120 via the terminal device 110 and / or an attached device of the terminal device 110. The application 120 can be any type of application, and it can receive any appropriate type of content (such as text content, audio content, etc.) input by the user 140.

[0029] In Figure 1In an environment 100, if an application 120 is active, a terminal device 110 can present a page 150 of the application 120. The page 150 can include various types of pages that the application 120 can provide, such as content presentation pages, content creation pages, content publishing pages, message pages, personal home pages, and so on. The application 120 can, for example, present received user input in the page 150. For example, the application 120 can present an input box in the page 150. The application 120 can receive text content input by a user 140 via the input box and present the text content in the page 150. Also for example, the application 120 can also convert received content of other types (i.e., types other than text type), such as audio content, into text content and present the text content in the page 150.

[0030] According to an example implementation of the present disclosure, the terminal device 110 communicates with a server 130 to implement the supply of services for the application 120. The terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. According to an example implementation of the present disclosure, the terminal device 110 can also support any type of user interface (such as a "wearable" circuit, etc.).

[0031] The server 130 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. The server 130 can be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in a cloud environment, and so on.

[0032] It should be understood that the structures and functions of the various elements in the environment 100 are described only for exemplary purposes and do not imply any limitation on the scope of the present disclosure.

[0033] As discussed above, applications and websites can provide data processing services. Data processing can, for example, include data management, data reception, data conversion, and so on. Taking an input method application as an example, it can receive and present the text content input by the user. The input method application also supports data processing functions such as case conversion of characters, language switching, text matching, etc.

[0034] An application typically performs conversions between different types of data items. In the context of the present disclosure, the specific details of managing data items will be described by taking the conversion between uppercase and lowercase letters in a multilingual environment as an example. For example, an application can provide operation controls associated with data processing, such as an operation control for switching the case of characters. It can respond to the triggering operation received for the operation control for switching the case of characters and switch the case of the next data sequence to be received. When the user inputs multiple data sequences, the operation control may need to be triggered multiple times to switch the case, which will affect the efficiency of the user input and the efficiency of the application performing case conversion.

[0035] Taking the execution of case switching as an example, in order to improve the efficiency of executing case switching, an application can, for example, provide extended content matching the user input based on the user input. The extended content can include the user input after case transformation. For example, when the user input is a proper noun, the application can perform case extension on the received user input, compare the extended content with a thesaurus for query, and in the case of a query result, determine the query result as the basic candidate. The application can also perform case conversion on the basic candidate to determine the supplementary candidate. The application can sort all candidates (including the basic candidate and the supplementary candidate) according to certain rules and provide the sorting result to the user. For example, if the user input is the text "http", the application can perform the above processing on the user input to obtain candidates such as "Http", "HTTP", "HTTPS", etc. The application can provide the user input "http" and the obtained multiple candidates "Http", "HTTP", "HTTPS", etc. to the user together.

[0036] Data can be encoded so that each data item has its corresponding encoding. The encoding here can, for example, include general character encodings such as Unicode, ASCII code, scan code, etc., and the present disclosure does not limit the specific encoding. For example, for any language (including but not limited to Chinese, English, Russian, French, etc., and the present disclosure does not limit the specific language), the encoding of its character set can be determined. Taking English as an example, Figure 2A shows the Unicode 200A of English. As Figure 2AAs shown, the Unicode values of the capital letters A - Z are from U+0041 to U+005A (the values are consecutive and increase by 1 in sequence), and the Unicode values of the lowercase letters a - z are from U+0061 to U+007A (the values are consecutive and increase by 1 in sequence). The difference between the Unicode value of a capital letter and the Unicode value of the corresponding lowercase letter is 32.

[0037] The case - sensitive character sets usually include multiple 1:1 mapping pairs (i.e., multiple capital - letter - lowercase - letter mapping pairs). For most hash tables, programming languages provide data structures for storing and querying mapping pairs, such as hash tables, red - black trees, etc. Traditionally, a character set 200B as shown in Figure 2B can be obtained. The character set 200B can indicate the correspondence between capital letters and lowercase letters. For example, if the user presses the key 'A', it can be determined that the corresponding capital letter of this key is 'A' and the corresponding lowercase letter is 'a'. A certain amount of continuously - addressed physical space (also called a bucket) can be pre - allocated. Each bucket is implemented as a linked - list structure. By traversing the character set 200B, the mapping pairs (pair, i.e., capital - letter - lowercase - letter) of individual characters can be obtained in sequence. For the lowercase character (lower) in the mapping pair, an index (index = Hash(lower)) can be calculated via a hash function. Based on this index, the index - th bucket can be randomly accessed. If the linked list in the bucket is empty, the mapping pair is inserted into the linked list of this bucket (bucket[index]), and at this time, the linked list has only one element, which is the mapping pair. If the linked list in the bucket is not empty, the mapping pair can be inserted at the end of the linked list of the bucket. A hash table as shown in Figure 2C can be obtained.

[0038] Taking the query of the corresponding capital letter based on a lowercase letter as an example, if the character entered by the user is the lowercase letter 'e', its corresponding index can be determined as M based on the hash algorithm. Then, bucket[M] can be accessed, and it is determined that the linked list in bucket[M] is not empty. Traverse this linked list and compare the lowercase characters in sequence. It can be determined that the second element in the linked list matches the lowercase character 'e', and the corresponding capital letter 'E' is returned, and the query ends. The above describes the situation where the corresponding linked list is not empty and there are matching elements in it. The following describes several situations where no result can be returned. For example, if the character entered by the user is the lowercase letter's', and its corresponding index is determined as 2 based on the hash algorithm. Access bucket[2], and it is determined that the linked list in bucket[2] is empty, and the query ends. Another example is that if the character entered by the user is the lowercase letter 'k', and its corresponding index is determined as 1 based on the hash algorithm. Access bucket[1], and it is determined that the linked list in bucket[1] is not empty. Traverse this linked list and compare the lowercase characters in sequence. It can be determined that there is no element in the linked list that matches the lowercase character 'k', and the query ends.

[0039] In summary, in traditional case conversion, all character mapping pairs need to be stored, and additional data structures (such as indexes) need to be introduced. Index calculation is required, and then the uppercase / lowercase character matching the input character is queried. Thus, a large amount of storage space is traditionally consumed, resulting in low storage efficiency. Traditionally, additional index calculations are also required, and the linked list is traversed based on the calculated index, resulting in low query efficiency as well.

[0040] Process for managing data items

[0041] To at least partially address the deficiencies in the prior art, according to some example implementations of the present disclosure, a method for managing data items is proposed. Generally speaking, a first data sequence including a plurality of first data items and a second data sequence including a plurality of second data items can be obtained. Based on the first data sequence and the second data sequence, the first data sequence can be divided into at least one subsequence. The encoding of a set of first data items in the target subsequence among the at least one subsequence is continuous. The mapping characteristics of the at least one subsequence can be determined. The target mapping characteristics of the target subsequence include the first boundary and the second boundary of a set of first data items in the target subsequence, and the difference between the encoding of the data items in the target subsequence and the encoding of another data item in the second data sequence corresponding to the data item. Based on the mapping characteristics of the at least one subsequence, the first data item among the plurality of first data items can be mapped to the second data item among the plurality of second data items.

[0042] See Figure 3 Describing an exemplary implementation according to the present disclosure, Figure 3 FIG. 300 shows a flowchart of a process 300 for managing data items according to some implementations of the present disclosure. The process 300 can be implemented at the terminal device 110 and / or the server 130. For the sake of description, hereinafter, the terminal device 110 and the server 130 are collectively referred to as an electronic device, and the process 300 will be described with the electronic device as the execution subject. For ease of discussion, the process 300 will be described with reference to Figure 1 the environment 100. It should be noted that the operations performed by the electronic device can specifically be performed by relevant applications installed on the electronic device.

[0043] In block 310, the electronic device obtains a first data sequence including a plurality of first data items and a second data sequence including a plurality of second data items.

[0044] The data items here can be, for example, characters, text, images, or any other appropriate type of data items. For example, multiple first data items and multiple second data items can include multiple characters in a character set in a multilingual environment. The first data items and the second data items have a mapping relationship. Such a mapping relationship can be any appropriate predefined mapping relationship. The mapping relationship can be a one-way mapping relationship or a two-way mapping relationship. For example, multiple first data items include any one of uppercase characters and lowercase characters, and multiple second data items include the other one of uppercase characters and lowercase characters. For the convenience of description, the following takes a one-way mapping relationship (that is, determining the corresponding second data item based on the first data item) as an example for exemplary description.

[0045] Exemplarily, the first data item can be a lowercase character in a certain language, and the second data item can be the uppercase character in that language. The first data item can also represent a certain key on the keyboard (such as shift, ctrl, and alt, etc.), and the second data item can represent the unified code of these keys, etc. The following takes the first data item as a lowercase character in a certain language, the second data item as the corresponding uppercase character in that language, and the mapping relationship that the lowercase character can be mapped to the uppercase character as an example for exemplary description.

[0046] The electronic device can obtain the first data sequence of multiple first data items and the second data sequence including multiple second data items in any appropriate manner. It can be understood that the number of multiple first data items and multiple second data items can be the same, and each first data item can be mapped to the corresponding second data item. Taking the first data item and the second data item as uppercase and lowercase characters respectively, multiple first data items can include multiple lowercase characters, and multiple second data items can include multiple uppercase characters. It should be noted that multiple first data items and multiple second data items can include lowercase characters and uppercase characters corresponding to different languages.

[0047] The first data sequence of multiple first data items can be, for example, a sequence of multiple encodings corresponding to multiple first data items. Similarly, the second data sequence of multiple second data items can also be a sequence of multiple encodings corresponding to multiple second data items. The following takes the encoding as the unified code as an example for exemplary description. Taking English as an example, multiple first data items and multiple second data items respectively include lowercase characters (that is, a - z) and uppercase characters (that is, A - Z) of English. The unified codes of multiple first data items are U+0061 to U+007A (the values increase by 1 in sequence), and the unified codes of multiple second data items are U+0041 to U+005A (the values increase by 1 in sequence). It can be understood that multiple first data items and multiple second data can also respectively include lowercase characters and uppercase characters of other languages.

[0048] At block 320, based on the first data sequence and the second data sequence, the electronic device divides the first data sequence into at least one subsequence, and the encoding of a group of first data items in the target subsequence among the at least one subsequence is continuous. According to an example implementation of the present disclosure, in order to distinguish data items corresponding to a given language and the corresponding data sequences, the electronic device 110 may, for example, divide the first data sequence into at least one subsequence based on the encodings corresponding to multiple characters included in the multiple first data items. Thus, the encoding of the data items in each subsequence obtained by the electronic device is continuous, that is, the electronic device 110 can obtain a corresponding plurality of subsequences (each subsequence only includes lowercase characters).

[0049] According to an example implementation of the present disclosure, the electronic device may determine a pairing sequence based on the first data sequence and the second data sequence. The pairing sequence includes a plurality of pairs, and each pair in the plurality of pairs includes a first data item among the plurality of first data items and a second data item among the plurality of second data items. Exemplarily, if the first data item is a lowercase character and the second data item is an uppercase character, each pair includes a lowercase character and the corresponding uppercase character (such as a and A in English). The obtained plurality of pairs may include aA, bB, cC, etc.

[0050] The electronic device may then determine at least one subsequence based on the pairing sequence. Specifically, the electronic device may identify the head pair at the head of the pairing sequence as the current pair. The head pair may be the first pair among the plurality of pairs. The plurality of pairs in the pairing sequence may be sorted based on a predetermined rule. Exemplarily, the plurality of first data items and the plurality of second data items may be arranged in the order of the encodings of the plurality of first data items and the plurality of second data items. That is, in the case where the data items only include the corresponding uppercase and lowercase characters in English, the pair aA may be the first pair (i.e., the head pair), and the pair zZ may be the last pair.

[0051] The electronic device may process the current pair and, in the case of processing the current pair, set the current pair to the subsequent pair after the current pair in the pairing sequence. That is, in the case of processing the first pair, the electronic device may determine the second pair as the current pair and process the second pair. Similarly, the electronic device may, in the case of processing the previous pair, sequentially process each subsequent pair until all the pairs included in the pairing sequence are processed. In this way, each pair can be traversed step by step, thereby realizing the mapping between all data items.

[0052] According to an exemplary implementation of the present disclosure, corresponding mapping features can be constructed for each subsequence. The mapping features can be represented by (a first boundary, a second edit, a difference), where the first edit represents, for example, the upper boundary of the subsequence, the second boundary represents, for example, the lower boundary), and the difference can represent the difference between the encodings of the corresponding data items. Regarding the specific manner of processing the current pair, the electronic device can determine the current subsequence in at least one subsequence based on the current pair.

[0053] In the initial stage, the electronic device can set the first boundary and the second boundary in the current mapping feature of the current subsequence to the encodings of the first data item in the current pair, respectively. For example, if the current pair is aA, the electronic device 110 can set the first boundary and the second boundary to the lowercase character a in the pair, and its encoding can be, for example, Unicode U+0061. Here, the first boundary includes any one of the upper boundary and the lower boundary, and the second boundary includes the other one of the upper boundary (which can also be referred to as the maximum boundary) and the lower boundary (which can also be referred to as the minimum boundary). Hereinafter, an exemplary description will be given with the first boundary being the maximum boundary and the second boundary being the minimum boundary. Alternatively and / or additionally, the meanings of the first boundary and the second boundary can be exchanged, which will not be elaborated herein.

[0054] The electronic device can also set the difference in the current mapping feature of the current subsequence to the difference between the encoding of the first data item in the current pair and the encoding of the second data item in the current pair. For example, if the current pair is aA, the electronic device can set the difference to the difference between the encoding corresponding to the lowercase character a in the pair and the encoding corresponding to the uppercase character A, and the encoding can be, for example, the difference between Unicode U+0061 and Unicode U+0041. Since Unicode is a hexadecimal encoding, the difference is -32 (decimal).

[0055] After the electronic device determines the first boundary, the second boundary, and the difference based on the current pair, it can determine that the processing of the current pair is completed. The electronic device then sets the current pair to the subsequent pair after the current pair in the pair sequence to process the next pair. For example, the electronic device can process the next pair bB in response to the completion of the processing of the pair aA.

[0056] According to an example implementation of the present disclosure, an electronic device may determine whether a difference between an encoding of a first data item in a current pairing and an encoding of a second data item in the current pairing matches a difference in mapping features of a current subsequence. For example, the electronic device may determine whether a difference between an encoding of a first data item in the current pairing and an encoding of a second data item in the current pairing is the same as a difference in mapping features of the current subsequence. If they are the same, the electronic device may determine that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing matches the difference in mapping features of the current subsequence. If they are different, the electronic device may determine that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing does not match the difference in mapping features of the current subsequence.

[0057] If it is determined that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing does not match the difference in mapping features of the current subsequence, the electronic device may determine that the current pairing in the pairing sequence and the previous pairing of the current pairing should be divided into different subsequences. In this case, the electronic device may determine the current pairing in the pairing sequence and the part after the current pairing to create another subsequence. That is, the electronic device may divide the first data sequence based on the previous pairing and the current pairing.

[0058] If it is determined that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing matches the difference in mapping features of the current subsequence, the electronic device may determine that the current pairing and the previous pairing of the current pairing can be divided into the same subsequence. According to an example implementation of the present disclosure, the electronic device may directly update the first boundary of the current mapping to the encoding of the first data item in the current pairing. For example, if the current pairing is bB, the previous pairing is aA, the first boundary and the second boundary both currently correspond to the Unicode U+0061, and the difference is -32, the electronic device may update the first boundary in the current mapping features of the current subsequence based on the Unicode U+0062 of the lowercase character b in the current pairing bB. The updated first boundary, second boundary, and difference may be U+0062, U+0061, and -32.

[0059] According to an example implementation of the present disclosure, the electronic device also needs to determine whether the first boundary of the current mapping features is continuous with the encoding of the first data item in the current pairing. If it is not continuous, the electronic device may determine that the encodings of the current pairing and the previous pairing of the current pairing are not continuous. In this case, the electronic device may create another subsequence based on the current pairing in the pairing sequence and the part after the current pairing.

[0060] It should be understood that in the English language environment, the encodings of lowercase and uppercase characters are both continuous, and the difference between the corresponding pairs is -32. At this time, only one subsequence is generated. Specifically, if they are continuous (such as U+0061 and U+0062 shown above), the electronic device can update the first boundary based on the encoding of the first data item in the current pair. The electronic device can determine whether the current pair is the last pair in the pair sequence, and if it is determined that the current pair is not the last pair in the pair sequence, set the current pair to the subsequent pair after the current pair in the pair sequence. That is, the electronic device can continue to process subsequent pairs based on the above method until all pairs in the pair sequence are processed. The electronic device can then determine at least one subsequence based on at least one mapping feature. Thus, the electronic device can divide the first data sequence into at least one subsequence.

[0061] It can be understood that for the same language, it can correspond to at least one subsequence. For example, if the encodings of the lowercase characters of a certain language are continuous encodings, then the language corresponds to only 1 subsequence. If the encodings of the lowercase characters of the language are discontinuous encodings or the encoding differences of the characters in the pair are different, then the language can correspond to multiple subsequences.

[0062] Exemplarily, for Danish, the encodings of lowercase characters a-z (U+0061 - U+007A) are continuous, the encodings of uppercase characters A-Z (U+0041 - U+005A) are continuous, and the difference between lowercase characters and uppercase characters is -32. The encodings of lowercase characters and are Unicode U+00E5, U+00E6, and U+00F8 respectively. The uppercase characters corresponding to these 3 lowercase characters and are Unicode U+00C5, U+00C6, and Unicode U+00D8 respectively. The difference between the encoding of each lowercase character and the encoding of the uppercase character corresponding to the lowercase character is also -32. At this time, in Danish, a-z will be divided into one subsequence, will be divided into one subsequence, and will be divided into another subsequence. At this time, there are three subsequences, and each subsequence has its own mapping feature.

[0063] Continuing to refer to Figure 3 , in block 330, the electronic device determines the mapping features of at least one subsequence. The target mapping feature of the target subsequence includes: the first boundary and the second boundary of a group of first data items in the target subsequence, and the difference between the encoding of the data item in the target subsequence and the encoding of another data item corresponding to the data item in the second data sequence.

[0064] The encoding of the data item in the target subsequence is the encoding of the first data item.

[0065] Specifically, the electronic device can determine a first boundary and a second boundary based on the first data item, and can determine the difference of the subsequence where it is located based on the difference between the first data item and the second data item in the corresponding pairing. Therefore, after the electronic device obtains the target subsequence (which can be any one of at least one subsequence), it can determine the first boundary (i.e., the maximum value in the continuous encoding corresponding to the continuous first data items), the second boundary (i.e., the minimum value in the continuous encoding corresponding to the continuous first data items), and the difference (the difference between each first data item in the sequence and the second data item in the corresponding pairing, that is, the difference between the encoding of the data item in the target subsequence and the encoding of another data item corresponding to the data item in the second data sequence) of the target subsequence.

[0066] The following combines Figure 4 to describe the process of dividing the first data sequence into at least one subsequence. Figure 4 FIG. 400 is a schematic diagram showing an example of determining at least one subsequence according to some implementations of the present disclosure.

[0067] In block 402, the electronic device can determine a pairing sequence (CharMap). The pairing sequence includes a plurality of pairings. For example, CharMap = {pair_1, pair_2,..., pair_n}. Exemplarily, if the plurality of pairings include pairings composed of corresponding upper and lower case characters in English (for example, [lower case character, upper case character]), then CharMap includes [a, A], [b, B],..., [z, Z] these 26 pairings.

[0068] In block 404, the electronic device can initialize the mapping feature set (CharBounds, which can include at least one mapping feature) to be used for storing the mappings corresponding to the plurality of pairings to make it empty.

[0069] In block 406, the electronic device can sort the plurality of pairings in the pairing sequence. The electronic device can, for example, sort the plurality of pairings according to the Unicode value corresponding to the lower case characters.

[0070] In block 408, the electronic device can, according to the method described above, access each pairing in the pairing sequence one by one and process the accessed pairing in turn.

[0071] During the process of processing each pairing, during the processing of the current pairing, in block 410, the electronic device can determine whether the mapping feature set is empty. That is, it can determine whether there are already determined mapping features.

[0072] When the mapping feature set is empty, the electronic device can determine that the current pairing is the first sequence in the pairing sequence (i.e., the head pairing). At block 412, the electronic device can initialize a new mapping feature. At this time, CharBound is set based on the first pair [a, A].

[0073] At block 414, the electronic device can determine the minimum boundary (CharBound.min) of the mapping feature as the code (pair.lower) corresponding to the lowercase character (e.g., a) in the current pairing. That is, CharBound.min = pair.lower.

[0074] At block 416, the electronic device can determine the maximum boundary (CharBound.max) of the mapping feature as the code corresponding to the lowercase character (e.g., a) in the current pairing. That is, CharBound.max = pair.lower.

[0075] At block 418, the electronic device can determine the difference between the code of the lowercase character in the current pairing and the code of the uppercase character in the current pairing (pair.upper) as the difference (CharBound.diff) in the current mapping feature. That is, CharBound.diff = pair.upper - pair.lower = -32.

[0076] At block 420, the electronic device can insert the newly determined mapping feature behind the mapping feature set. In the case where there is no mapping feature previously, the electronic device can determine the currently determined mapping feature as the first mapping feature in the mapping feature set (i.e., at this time, the mapping feature is (U+0061, U+0061, -32)).

[0077] After processing the current pairing, the electronic device can update the next pairing of the current device to the current pairing, that is, the electronic device can continue to process the next pairing. At block 422, the electronic device can determine whether the updated current pairing reaches the end of the pairing sequence, that is, it can determine whether the updated current pairing is the last sequence in the pairing sequence.

[0078] If the updated current pairing does not reach the end of the pairing sequence, the electronic device can return to execute the process shown in block 408. That is, the electronic device can continue to process the updated current pairing based on the above method.

[0079] In the case where the current pairing is not the first sequence in the pairing sequence, the electronic device can determine that the mapping feature set is not empty. At block 424, the electronic device can take out the mapping feature at the end of the mapping feature set. That is, the electronic device can take out the mapping feature corresponding to the previous pairing of the current pairing.

[0080] At block 426, the electronic device may calculate the difference between the code of the lowercase character and the code of the uppercase character in the current pairing. That is, the difference diff of pair = pair.upper - pair.lower.

[0081] At block 428, the electronic device may determine whether the difference of the current pairing is the same as the difference of the retrieved mapping feature. That is, determine whether diff is equal to CharBound.diff.

[0082] If they are the same, at block 430, the electronic device may further determine whether the code corresponding to the lowercase character in the current pairing is consecutive with the code corresponding to the lowercase character in the previous pairing. As described above, the maximum boundary in the mapping feature is the code corresponding to the lowercase character in the previous pairing, and the electronic device may determine the continuity by determining the difference between the maximum boundary and the code corresponding to the lowercase character in the current pairing. That is, determine whether CharBound.max is equal to pair.lower of the current pairing minus 1.

[0083] In the case where the differences are different (that is, diff is not equal to CharBound.diff) and / or not consecutive (that is, CharBound.max is not equal to pair.lower - 1), the electronic device may determine that the lowercase character in the current pairing and the lowercase character in the previous pairing will be divided into different subsequences and thus a new CharBound needs to be created. The electronic device may perform the steps described in block 412, that is, the electronic device may newly determine a mapping feature. The electronic device may then determine the maximum boundary, minimum boundary, difference, etc. of the new mapping feature based on the current pairing and the subsequent pairings of the current pairing.

[0084] In the case where the differences are the same and consecutive, at block 432, the electronic device may update the maximum boundary of the currently retrieved mapping feature to the code corresponding to the lowercase character of the current pairing. The electronic device may perform the above process until the current pairing is the last sequence in the pairing sequence (that is, the electronic device has completed the processing of each pairing in the pairing sequence). In this case, at block 434, the electronic device may determine that the mapping feature set is constructed. After traversing all lowercase English characters, at this time the mapping feature is (U+007A, U+0061, -32), and the above mapping feature may be used to convert all lowercase characters to the corresponding uppercase characters.

[0085] Return Figure 3, at block 340, the electronic device maps a first data item among a plurality of first data items to a second data item among a plurality of second data items based on mapping features of at least one subsequence. According to an example implementation of the present disclosure, if a processing request for specifying a target data item to be mapped is received, the electronic device may determine a target mapping feature corresponding to the target data item in the mapping feature set. Exemplarily, the electronic device may determine that a processing request for the target data item indicated by the user input is received in response to receiving the user input. For example, the electronic device may determine that a processing request for 'a' is received in response to receiving the user input of the lowercase character 'a'.

[0086] The encoding of the target data item should be included between the first boundary and the second boundary in the target mapping feature. It is assumed that the size relationship among the first boundary, the second boundary, and the encoding of the target data item should be the first boundary ≥ the encoding of the target data item ≥ the second boundary. In the case where there are multiple subsequences, the electronic device may determine multiple mapping features corresponding to these multiple subsequences. The electronic device may determine the target mapping feature from these multiple mapping features based on any suitable method. For example, the electronic device may traverse all the data before the two boundaries to search for the encoding of the target data item. Alternatively and / or additionally, the electronic device may determine the target mapping feature in the mapping feature set based on the dichotomy method.

[0087] The following combines Figure 5 to describe the process of determining the target mapping feature. Figure 5 FIG. shows a schematic diagram of an example 500 for determining a target mapping feature according to some implementations of the present disclosure.

[0088] At block 502, the electronic device may obtain the data item input by the user (i.e., the target data item). For example, the electronic device may receive the corresponding data item in response to receiving the user's trigger of the corresponding key. For example, the electronic device may determine that the user input received is the lowercase character 'c' in English in response to receiving the trigger of the key 'c' on the keyboard.

[0089] At block 504, the electronic device may obtain a given mapping feature set. This mapping feature set may be previously determined and stored locally by the electronic device itself, or may be determined and provided to the electronic device by other electronic devices. In the scenario of converting a lowercase English character to an uppercase character, the mapping feature is (U+007A, U+0061, -32).

[0090] At block 506, the electronic device determines whether the mapping feature set is empty. That is, the electronic device may determine whether the mapping feature set includes mapping features.

[0091] If the mapping feature set is empty, at block 508, the electronic device may determine that there is no mapping feature corresponding to the data item input by the user. The electronic device may determine that the data item input by the user is not a data item among the multiple first data items corresponding to this mapping feature set. According to an example implementation of the present disclosure, the electronic device may also provide a corresponding indication in response to determining that there is no mapping feature corresponding to the data item input by the user. For example, the electronic device may provide an indication that the data item does not belong to the mapping feature set, or provide an indication that the data item does not belong to the first subsequence (which may be any subsequence among the at least one subsequence) of the at least one subsequence corresponding to the multiple first data items. The electronic device may provide the indication in any suitable manner such as by providing text, audio, video, etc. Thereby, the electronic device can prompt the user that it is impossible to determine another data item corresponding to the data item input by the user, and can provide a user experience.

[0092] If the mapping feature set is not empty, the electronic device may determine the number of at least one mapping feature included in the mapping feature set (CharBounds.size(), and the number is 1 at this time). At block 510, the electronic device may determine the upper boundary index (index_max) for determining the target mapping feature. The electronic device may determine the index as the number of at least one mapping feature minus 1. That is, index = CharBounds.size() – 1.

[0093] At block 512, the electronic device may determine both the lower boundary index (index_min) and the middle index (index_middle) of the target mapping feature as 0. That is, index_min = 0, index_middle = 0.

[0094] At block 514, the electronic device may determine whether the value of the lower boundary index is less than or equal to the value of the upper boundary index. If the value of the lower boundary index is greater than the value of the upper boundary index, the electronic device may return to execute block 508, that is, the electronic device may determine that there is no mapping feature corresponding to the data item input by the user.

[0095] If the value of the lower boundary index is less than or equal to the value of the upper boundary index, the electronic device may, for example, determine the target mapping feature in the mapping feature set based on the dichotomy method. Specifically, at block 516, the electronic device may update the value of the middle index to the mean of the lower boundary index and the upper boundary index. That is, index_middle = (index_max + index_min) / 2.

[0096] At block 518, the electronic device may determine the mapping feature corresponding to the middle index in the mapping feature set as the target mapping feature. That is, CharBound = CharBounds[index_middle].

[0097] The electronic device can also determine the encoding corresponding to the data item input by the user. For example, the electronic device can determine that the Unicode corresponding to the lowercase character c is U+0063. At block 520, the electronic device can determine whether this Unicode is an encoding between the encoding corresponding to the minimum boundary of the target mapping feature and the encoding corresponding to the maximum boundary. For example, the electronic device can determine whether the encoding of the character c is between CharBound.max and CharBound.min.

[0098] If the encoding corresponding to the data item input by the user is an encoding between the encoding corresponding to the minimum boundary of the target mapping feature and the encoding corresponding to the maximum boundary, block 522 is executed. At block 522, the electronic device can determine that the data item is a data item among at least one data item corresponding to the target mapping feature. For example, the electronic device can determine that the input character c is a lowercase character.

[0099] At block 524, the electronic device can determine another data item corresponding to the input item input by the user based on the difference of the target mapping feature. For example, if the difference is -32, the electronic device can determine, based on the encoding of the lowercase character c and this difference, that the encoding of the corresponding data item (i.e., the corresponding uppercase character C) is CharBound.min+(-32)=U+0043.

[0100] At block 526, the electronic device can query the corresponding data item based on the determined encoding and provide it to the user. For example, the electronic device can query based on the determined encoding U+0043 that the uppercase character corresponding to the lowercase character c is the uppercase character C and return the uppercase character C.

[0101] If the encoding corresponding to the data item input by the user is not an encoding between the encoding corresponding to the minimum boundary of the target mapping feature and the encoding corresponding to the maximum boundary, the electronic device can further determine whether the encoding corresponding to the input data item is an encoding less than the encoding corresponding to the minimum boundary or an encoding greater than the encoding corresponding to the maximum boundary.

[0102] According to an example implementation of the present disclosure, at block 528, the electronic device can determine whether the encoding corresponding to the input data item is less than the encoding of the minimum boundary of the target mapping feature. The electronic device can update the upper boundary index or the lower boundary index of the target mapping feature based on the determination result.

[0103] If the code corresponding to the input data item is less than the code of the minimum boundary of the target mapping feature, at block 530, the electronic device may update the upper boundary index of the target mapping feature to the current middle index minus 1. That is, index_max = index_middle – 1. If the code corresponding to the input data item is greater than or equal to the code of the minimum boundary of the target mapping feature, at block 532, the electronic device may update the lower boundary index of the target mapping feature to the current middle index plus 1. That is, index_min = index_middle + 1.

[0104] After the upper boundary index / lower boundary index of the target mapping feature is updated, the electronic device may perform the steps shown in block 514. That is, the electronic device may update the middle index based on the updated upper boundary index / lower boundary index. The electronic device may then update the target mapping feature based on the updated middle index.

[0105] It should be understood that the mapping process is described above only by taking the conversion of lowercase English characters to uppercase characters as an example. More details about the case conversion of Danish will be provided below. The mapping features can be determined in the manner described above Figure 4 as shown in Table 1 below.

[0106] Table 1 Mapping Features

[0107]

[0108] According to an example implementation of the present disclosure, the technical solution described above may be performed in an input method environment. Refer to Figure 6 for more details. The Figure 6 block diagram 600 schematically shows the process of updating indexes according to some implementations of the present disclosure. As Figure 6 shown, certain special terms have specific representations, and users may not pay attention to the case representation when inputting, and spelling mistakes may occur at this time. For example, the user may input the lowercase character "http" in the input box 610, and at this time, the input method may pop up multiple candidate words in the candidate box 620 for the user to select.

[0109] Specifically, the mapping feature (U+007A, U+0061, -32) may be constructed based on the method described above, and this mapping feature may be used to convert each lowercase character to the corresponding uppercase character: h->H, t->T, and p->P. At this time, only the above mapping feature needs to be stored at the terminal device 110 to implement the above conversion process. Compared with using a hash table as Figure 2C shown, the required storage space is greatly reduced. Further, during the conversion process, it is not necessary to traverse the hash table, but it can be based on as Figure 5Use the process shown to search for the encodings of h, t, t, and p of lowercase characters between the encodings of U+007A and U+0061 by dichotomy, and then sum the found encodings with "-32".

[0110] In summary, according to the implementation manner of the present disclosure, while ensuring the accuracy of managing data items, the storage space occupied by the storage mapping relationship and the complexity of the corresponding mapping process can be reduced, thereby improving efficiency.

[0111] Example process

[0112] According to an exemplary implementation manner of the present disclosure, a method for managing data items is provided. The method includes: obtaining a first data sequence including a plurality of first data items and a second data sequence including a plurality of second data items; based on the first data sequence and the second data sequence, dividing the first data sequence into at least one subsequence, and the encodings of a group of first data items in the target subsequence among the at least one subsequence are continuous; determining the mapping characteristics of the at least one subsequence, and the target mapping characteristics of the target subsequence include: the first boundary and the second boundary of a group of first data items in the target subsequence, and the difference between the encoding of the data item in the target subsequence and the encoding of another data item in the second data sequence corresponding to the data item; and based on the mapping characteristics of the at least one subsequence, mapping the first data item among the plurality of first data items to the second data item among the plurality of second data items.

[0113] According to an exemplary implementation manner of the present disclosure, dividing the first data sequence into at least one subsequence includes: based on the first data sequence and the second data sequence, determining a pairing sequence, the pairing sequence including a plurality of pairings, and the pairing in the plurality of pairings includes the pairing of the first data item among the plurality of first data items and the second data item among the plurality of second data items; and determining at least one subsequence based on the pairing sequence.

[0114] According to an exemplary implementation manner of the present disclosure, determining at least one subsequence based on the pairing sequence includes: identifying the head pairing at the head of the pairing sequence as the current pairing; and determining the current subsequence among the at least one subsequence based on the current pairing.

[0115] According to an exemplary implementation manner of the present disclosure, determining the current subsequence includes: respectively setting the first boundary and the second boundary in the current mapping characteristics of the current subsequence as the encoding of the first data item in the current pairing; and setting the difference in the current mapping characteristics of the current subsequence as the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing.

[0116] According to an example implementation of the present disclosure, the method further includes: setting the current pairing to a subsequent pairing after the current pairing in the pairing sequence; and in response to determining that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing does not match the difference in the mapping feature of the current subsequence, creating another subsequence based on the current pairing in the pairing sequence and the portion after the current pairing.

[0117] According to an example implementation of the present disclosure, the method further includes: in response to determining that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing matches the difference in the mapping feature of the current subsequence, determining whether the first boundary of the current mapping feature is continuous with the encoding of the first data item in the current pairing; and in response to determining that the first boundary of the current mapping feature is discontinuous with the encoding of the first data item in the current pairing, creating another subsequence based on the current pairing in the pairing sequence and the portion after the current pairing.

[0118] According to an example implementation of the present disclosure, the method further includes: in response to determining that the first boundary of the current mapping feature is continuous with the encoding of the first data item in the current pairing, updating the first boundary of the current mapping to the encoding of the first data item in the current pairing.

[0119] According to an example implementation of the present disclosure, setting the current pairing to a subsequent pairing after the current pairing in the pairing sequence includes: in response to determining that the current pairing is not the last pairing in the pairing sequence, setting the current pairing to a subsequent pairing after the current pairing in the pairing sequence.

[0120] According to an example implementation of the present disclosure, mapping a first data item among a plurality of first data items to a second data item among a plurality of second data items based on the mapping features of at least one subsequence includes: in response to receiving a processing request for specifying a target data item to be mapped, determining a target mapping feature corresponding to the target data item in at least one mapping feature, where the encoding of the target data item is between the first boundary and the second boundary in the target mapping feature; and mapping the target data item to the second data item among the plurality of second data items based on the encoding of the target data item and the difference in the target mapping feature.

[0121] According to an example implementation of the present disclosure, determining a target mapping feature corresponding to the target data item in at least one mapping feature includes: determining the target mapping feature in at least one mapping feature based on the dichotomy method.

[0122] According to an example implementation of the present disclosure, the method further includes: in response to determining that there is no mapping feature corresponding to the target data item, providing an indication that the target data item does not belong to the first subsequence.

[0123] According to an exemplary implementation of the present disclosure, the multiple first data items and the multiple second data items include multiple characters in a character set in a multilingual environment, wherein the encoding includes a universal character encoding.

[0124] According to an exemplary implementation of the present disclosure, the multiple first data items include either uppercase characters or lowercase characters, and the multiple second data items include the other of uppercase characters and lowercase characters.

[0125] According to an exemplary implementation of the present disclosure, the multiple first data items and the multiple second data items are arranged in the order of the encoding of the multiple first data items and the multiple second data items. The first boundary includes either the upper boundary or the lower boundary, and the second boundary includes the other of the upper boundary and the lower boundary.

[0126] Example apparatus and device

[0127] Specific details of the method for managing data items have been described above. According to an exemplary implementation of the present disclosure, a device for managing data items is provided. Figure 7 FIG. shows a schematic structural block diagram of a device 700 for managing data items according to certain implementations of the present disclosure. The device 700 may be implemented as or included in a server 130. Each module / component in the device 700 may be implemented by hardware, software, firmware, or any combination thereof.

[0128] As shown in the figure, the device 700 includes a sequence acquisition module 710 configured to acquire a first data sequence including multiple first data items and a second data sequence including multiple second data items. The device 700 further includes a sequence division module 720 configured to divide the first data sequence into at least one subsequence based on the first data sequence and the second data sequence, and the encoding of a group of first data items in the target subsequence among the at least one subsequence is continuous. The device 700 further includes a feature determination module 730 configured to determine the mapping features of the at least one subsequence. The target mapping features of the target subsequence include: the first boundary and the second boundary of a group of first data items in the target subsequence, and the difference between the encoding of the data items in the target subsequence and the encoding of the other data items in the second data sequence corresponding to the data items. The device 700 further includes a data mapping module 740 configured to map the first data items among the multiple first data items to the second data items among the multiple second data items based on the mapping features of the at least one subsequence.

[0129] According to an example implementation of the present disclosure, the sequence partitioning module 720 is further configured to: determine a pairing sequence based on a first data sequence and a second data sequence, the pairing sequence including a plurality of pairings, where a pairing in the plurality of pairings includes a pairing of a first data item among a plurality of first data items and a second data item among a plurality of second data items; and determine at least one subsequence based on the pairing sequence.

[0130] According to an example implementation of the present disclosure, the sequence partitioning module 720 is further configured to: identify a head pairing at the head of the pairing sequence as the current pairing; and determine a current subsequence among at least one subsequence based on the current pairing.

[0131] According to an example implementation of the present disclosure, the sequence partitioning module 720 is further configured to: respectively set a first boundary and a second boundary in the current mapping feature of the current subsequence as the encoding of the first data item in the current pairing; and set the difference in the current mapping feature of the current subsequence as the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing.

[0132] According to an example implementation of the present disclosure, the sequence partitioning module 720 is further configured to: set the current pairing as the subsequent pairing after the current pairing in the pairing sequence; and in response to determining that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing does not match the difference in the mapping feature of the current subsequence, create another subsequence based on the current pairing in the pairing sequence and the portion after the current pairing.

[0133] According to an example implementation of the present disclosure, the sequence partitioning module 720 is further configured to: in response to determining that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing matches the difference in the mapping feature of the current subsequence, determine whether the first boundary of the current mapping feature is continuous with the encoding of the first data item in the current pairing; and in response to determining that the first boundary of the current mapping feature is discontinuous with the encoding of the first data item in the current pairing, create another subsequence based on the current pairing in the pairing sequence and the portion after the current pairing.

[0134] According to an example implementation of the present disclosure, the sequence partitioning module 720 is further configured to: in response to determining that the first boundary of the current mapping feature is continuous with the encoding of the first data item in the current pairing, update the first boundary of the current mapping to the encoding of the first data item in the current pairing.

[0135] According to an exemplary implementation of the present disclosure, the sequence partitioning module 720 is further configured to: in response to determining that the current pairing is not the last pairing in the pairing sequence, set the current pairing to the subsequent pairing after the current pairing in the pairing sequence.

[0136] According to an exemplary implementation of the present disclosure, the data mapping module 740 is further configured to: in response to receiving a processing request for specifying a target data item to be mapped, determine a target mapping feature corresponding to the target data item in at least one mapping feature, where the encoding of the target data item is between a first boundary and a second boundary in the target mapping feature; and map the target data item to a second data item among a plurality of second data items based on the encoding of the target data item and the differences in the target mapping feature.

[0137] According to an exemplary implementation of the present disclosure, the data mapping module 740 is further configured to: determine the target mapping feature in at least one mapping feature based on the dichotomy method.

[0138] According to an exemplary implementation of the present disclosure, the data mapping module 740 is further configured to: in response to determining that there is no mapping feature corresponding to the target data item, provide an indication that the target data item does not belong to the first subsequence.

[0139] According to an exemplary implementation of the present disclosure, the plurality of first data items and the plurality of second data items include a plurality of characters in a character set in a multilingual environment, where the encoding includes a universal character encoding.

[0140] According to an exemplary implementation of the present disclosure, the plurality of first data items include either uppercase characters or lowercase characters, and the plurality of second data items include the other of uppercase characters and lowercase characters.

[0141] According to an exemplary implementation of the present disclosure, the plurality of first data items and the plurality of second data items are arranged in the order of the encodings of the plurality of first data items and the plurality of second data items, the first boundary includes either an upper boundary or a lower boundary, and the second boundary includes the other of the upper boundary and the lower boundary.

[0142] Figure 8 The block diagram of an electronic device 800 capable of implementing multiple implementations of the present disclosure is shown. It should be understood that Figure 8 The illustrated electronic device 800 is merely exemplary and should not constitute any limitation to the functions and scope of the implementations described herein. Figure 8 The illustrated electronic device 800 can be used to implement the methods described above.

[0143] As Figure 8As shown, the electronic device 800 is in the form of a general-purpose computing device. The components of the electronic device 800 may include, but are not limited to, one or more processors or processing units 810, a memory 820, a storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. The processing unit 810 may be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 800.

[0144] The electronic device 800 generally includes multiple computer storage media. Such media can be any available media accessible to the electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 may be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 may be removable or non-removable media and may include machine-readable media, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 800.

[0145] The electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 8 , a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 820 may include a computer program product 825 having one or more program modules that are configured to perform the various methods or actions of the various implementations of the present disclosure.

[0146] The communication unit 840 enables communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 800 may be implemented in a single computing cluster or multiple computer machines that are capable of communicating through a communication connection. Thus, the electronic device 800 may operate in a networked environment using a logical connection with one or more other servers, network personal computers (PCs), or another network node.

[0147] The input device 850 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 860 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 800 can also communicate with one or more external devices (not shown) as needed through the communication unit 840. The external devices such as a storage device, a display device, etc., communicate with one or more devices that enable a user to interact with the electronic device 800, or communicate with any device that enables the electronic device 800 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0148] According to some implementations of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, where the computer-executable instructions are executed by a processor to implement the method described above. According to some implementations of the present disclosure, a computer program product is also provided. The computer program product is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above. According to some implementations of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.

[0149] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0150] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams.

[0151] Computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0152] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0153] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skilled persons in the art to understand the various implementations disclosed herein.

Claims

1. A method for managing data items, comprising: Acquire a first data sequence including a plurality of first data items and a second data sequence including a plurality of second data items; Based on the first data sequence and the second data sequence, the first data sequence is divided into at least one subsequence, and the encoding of a group of first data items in a target subsequence in the at least one subsequence is continuous; determining a mapping feature of the at least one subsequence, the target mapping feature of the target subsequence comprising: a first boundary and a second boundary of a group of first data items in the target subsequence, and a difference between an encoding of a data item in the target subsequence and an encoding of another data item in the second data sequence corresponding to the data item; as well as Based on the mapping feature of the at least one subsequence, a first data item of the plurality of first data items is mapped to a second data item of the plurality of second data items.

2. The method according to claim 1, wherein dividing the first data sequence into the at least one subsequence comprises: Determine a pairing sequence based on the first data sequence and the second data sequence, the pairing sequence comprising a plurality of pairings, wherein a pair in the plurality of pairings comprises a pairing of a first data item in the plurality of first data items and a second data item in the plurality of second data items; as well as The at least one subsequence is determined based on the paired sequence.

3. The method of claim 2, wherein determining the at least one subsequence based on the paired sequence comprises: Marking the head pairing at the head of the pairing sequence as the current pairing; as well as A current subsequence of the at least one subsequence is determined based on the current pairing.

4. The method of claim 3, wherein determining the current subsequence comprises: Setting a first boundary and a second boundary in a current mapping feature of the current subsequence as the encoding of a first data item in the current pairing respectively; as well as The difference in the current mapped features of the current subsequence is set to the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing.

5. The method according to claim 4, further comprising: Setting the current pairing as a subsequent pairing after the current pairing in the pairing sequence; as well as In response to determining that a difference between an encoding of a first data item in the current pairing and an encoding of a second data item in the current pairing does not match a difference in a mapping feature of the current subsequence, creating another subsequence based on the current pairing and a portion of the pairing sequence after the current pairing.

6. The method according to claim 5, further comprising: In response to determining that a difference between an encoding of a first data item in the current pairing and an encoding of a second data item in the current pairing matches a difference in a mapping feature of the current subsequence, determining whether a first boundary of the current mapping feature is continuous with the encoding of the first data item in the current pairing; and In response to determining that a first boundary of the current mapping feature is discontinuous with an encoding of a first data item in the current pairing, the another subsequence is created based on the current pairing and a portion after the current pairing in the pairing sequence.

7. The method according to claim 5, further comprising: In response to determining that the first boundary of the current mapping feature is continuous with the encoding of the first data item in the current pairing, the first boundary of the current mapping is updated to the encoding of the first data item in the current pairing.

8. The method of claim 5, wherein setting the current pairing as a subsequent pairing after the current pairing in the pairing sequence comprises: In response to determining that the current pairing is not the last pairing in the pairing sequence, the current pairing is set as a subsequent pairing after the current pairing in the pairing sequence.

9. The method of claim 1, wherein mapping a first data item of the plurality of first data items to a second data item of the plurality of second data items based on the mapping feature of the at least one subsequence comprises: In response to receiving a processing request for specifying a target data item to be mapped, determining a target mapping feature corresponding to the target data item in the at least one mapping feature, the encoding of the target data item being located between a first boundary and a second boundary in the target mapping feature; and The target data item is mapped to a second data item of the plurality of second data items based on the encoding of the target data item and the difference in the target mapping characteristic.

10. The method according to claim 9, wherein determining a target mapping feature corresponding to the target data item in the at least one mapping feature comprises: The target mapping feature is determined among the at least one mapping feature based on a dichotomy method.

11. The method according to claim 9, further comprising: In response to determining that there is no mapping feature corresponding to the target data item, an indication is provided that the target data item does not belong to the first subsequence.

12. The method according to claim 1, wherein the plurality of first data items and the plurality of second data items comprise a plurality of characters in a character set in a multi-language environment, and wherein the encoding comprises a universal character encoding.

13. The method of claim 12, wherein the plurality of first data items include any one of uppercase characters and lowercase characters, and the plurality of second data items include the other of the uppercase characters and the lowercase characters.

14. The method according to claim 1, wherein the plurality of first data items and the plurality of second data items are arranged in an order of encoding of the plurality of first data items and the plurality of second data items, the first boundary includes any one of an upper boundary and a lower boundary, and the second boundary includes the other of the upper boundary and the lower boundary.

15. An apparatus for managing data items, comprising: A sequence acquisition module, configured to acquire a first data sequence including a plurality of first data items and a second data sequence including a plurality of second data items; a sequence division module, configured to divide the first data sequence into at least one subsequence based on the first data sequence and the second data sequence, wherein the encodings of a group of first data items in a target subsequence in the at least one subsequence are continuous; a feature determination module configured to determine a mapping feature of the at least one subsequence, wherein the target mapping feature of the target subsequence comprises: a first boundary and a second boundary of a group of first data items in the target subsequence, and a difference between an encoding of a data item in the target subsequence and an encoding of another data item corresponding to the data item in the second data sequence; as well as The data mapping module is configured to map a first data item among the plurality of first data items to a second data item among the plurality of second data items based on a mapping feature of the at least one subsequence.

16. An electronic device, comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 14 when executed by the at least one processing unit.

17. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to any one of claims 1 to 14.