Method and apparatus for managing data items, and device and storage medium

By dividing the continuous encoding subsequence and determining the mapping characteristics, the problems of large storage occupancy and low conversion efficiency caused by complex mapping relationships in the prior art are solved, and more efficient data item management and conversion are achieved.

WO2025108282A1PCT designated stage expired Publication Date: 2025-05-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/133064
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2024-11-19
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When managing and converting different types of data items, the complex mapping relationship leads to large storage usage and low conversion efficiency.

Method used

By obtaining the sequence of multiple data items, dividing the continuously encoded subsequences, determining the mapping characteristics of the subsequences, and mapping the data items based on these characteristics.

Benefits of technology

It reduces the space occupation of storage mapping relationships, improves the efficiency of data conversion, and simplifies the data item management process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024133064_30052025_PF_FP_ABST
    Figure CN2024133064_30052025_PF_FP_ABST
Patent Text Reader

Abstract

According to the implementations of the present disclosure, provided are a method and apparatus for managing data items, and a device and a storage medium. The method comprises: acquiring a first data sequence comprising a plurality of first data items and a second data sequence comprising a plurality of second data items; on the basis of the first data sequence and the second data sequence, dividing the first data sequence into at least one sub-sequence, wherein the encoding of a group of first data items in a target sub-sequence among the at least one sub-sequence is continuous; determining mapping features of the at least one sub-sequence; and on the basis of the mapping features of the at least one sub-sequence, mapping the first data items among the plurality of first data items to the second data items among the plurality of second data items. In this way, while the accuracy of data item management can be ensured, the efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device and storage medium for managing data items

[0001] This application claims priority to the Chinese invention patent application entitled “Methods, apparatus, devices and storage media for managing data items” filed on November 20, 2023, with application number 202311550835.X. The entire contents of that application are incorporated by reference into this application. Technical Field

[0002] Implementations of the present disclosure relate to the field of computers, and in particular, to methods, devices, apparatuses, and computer-readable storage media for managing data items. Background Art

[0003] In a computer system, multiple types of data items can be stored, and complex mapping relationships may exist between different types of data items. These mapping relationships can be stored in advance, and when data conversion needs to be performed, the mapping relationship is called so that the data item of one type is converted to another type. However, mapping relationships are usually more complicated, which causes the storage mapping relationship to take up a large amount of space. Further, the conversion process may involve complex algorithms, which causes conversion efficiency to be low. At this point, it is expected that the data items can be managed in a more convenient and effective manner, thereby realizing the conversion process between different types of data items. Summary of the Invention

[0004] In a first aspect of the present disclosure, a method for managing data items is provided. The method includes: obtaining a first data sequence including a plurality of first data items and a second data sequence including a plurality of second data items; dividing the first data sequence into at least one subsequence based on the first data sequence and the second data sequence, wherein the encodings of a group of first data items in a target subsequence in the at least one subsequence are continuous; determining a mapping feature of the at least one subsequence, wherein the target mapping feature of the target subsequence includes: a first boundary and a second boundary of a group of first data items in the target subsequence, and a difference between the encoding of the data item in the target subsequence and the encoding of another data item corresponding to the data item in the second data sequence; and mapping a first data item in the plurality of first data items to a second data item in the plurality of second data items based on the mapping feature of the at least one subsequence.

[0005] In a second aspect of the present disclosure, a device for managing data items is provided. The device includes: a sequence acquisition module configured to acquire a first data sequence including a plurality of first data items and a second data sequence including a plurality of second data items; a sequence division module configured to divide the first data sequence into at least one subsequence based on the first data sequence and the second data sequence, wherein the encoding of a group of first data items in a target subsequence in the at least one subsequence is continuous; a feature determination module configured to determine a mapping feature of the at least one subsequence, wherein the target mapping feature of the target subsequence includes: a first boundary and a second boundary of a group of first data items in the target subsequence, and a difference between the encoding of the data item in the target subsequence and the encoding of another data item corresponding to the data item in the second data sequence; and a data mapping module configured to map a first data item in the plurality of first data items to a second data item in the plurality of second data items based on the mapping feature of the at least one subsequence.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the present disclosure.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor implements the method according to the first aspect of the present disclosure.

[0008] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the implementation of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other features, advantages and aspects of various implementations of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0010] FIG1 illustrates a schematic diagram of an example environment in which implementations of the present disclosure can be implemented;

[0011] FIG2A shows the Unicode corresponding to uppercase and lowercase English characters;

[0012] FIG2B shows a character set corresponding to uppercase and lowercase English characters;

[0013] FIG2C shows a hash table for storing uppercase and lowercase characters in English;

[0014] FIG3 shows a flowchart of a method for managing data items according to some implementations of the present disclosure;

[0015] FIG4 is a schematic diagram showing an example of determining at least one subsequence according to some implementations of the present disclosure;

[0016] FIG5 is a schematic diagram illustrating an example of determining target mapping features according to some implementations of the present disclosure;

[0017] FIG6 schematically shows a block diagram of a process for updating an index according to some implementations of the present disclosure;

[0018] FIG7 shows a block diagram of an apparatus for managing data items according to some implementations of the present disclosure; and

[0019] FIG8 illustrates a block diagram of a device capable of implementing various implementations of the present disclosure. DETAILED DESCRIPTION

[0020] The following describes implementations of the present disclosure in more detail with reference to the accompanying drawings. Although certain implementations of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the implementations described herein. Rather, these implementations are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and implementations of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0021] In the description of the implementations of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "an implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". The terms "first", "second", etc. can refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0022] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0023] It is understandable that before using the technical solutions disclosed in each implementation of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0024] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0025] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0026] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0027] As used herein, the term "in response to" refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of executing a subsequent action executed in response to the event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is satisfied. For example, in some cases, a subsequent action may be executed immediately upon the occurrence of the event or the satisfaction of the condition; in other cases, the subsequent action may be executed some time after the occurrence of the event or the satisfaction of the condition.

[0028] Sample Environment

[0029] FIG1 illustrates a schematic diagram of an example environment 100 in which implementations of the present disclosure can be implemented. In this example environment 100, an application 120 is installed on a terminal device 110. A user 140 can interact with the application 120 via the terminal device 110 and / or devices attached to the terminal device 110. The application 120 can be any type of application that can receive any appropriate type of content (e.g., text content, audio content, etc.) input by the user 140.

[0030] In the environment 100 of FIG1 , if the application 120 is active, the terminal device 110 may present a page 150 of the application 120. The page 150 may include various types of pages that the application 120 can provide, such as a content presentation page, a content creation page, a content publishing page, a message page, a personal homepage, and the like. The application 120 may, for example, present the received user input in the page 150. For example, the application 120 may present an input box in the page 150. The application 120 may receive text content input by the user 140 via the input box and present the text content in the page 150. For another example, the application 120 may also convert received content of other types (i.e., types other than text) (e.g., audio content) into text content and present the text content in the page 150.

[0031] According to an example implementation of the present disclosure, the terminal device 110 communicates with the server 130 to implement the provision of services for the application 120. The terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. According to an example implementation of the present disclosure, the terminal device 110 can also support any type of interface for the user (such as "wearable" circuit, etc.).

[0032] Server 130 may be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. Server 130 may be any type of computing system / server capable of providing computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, and the like.

[0033] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.

[0034] As discussed above, applications and websites can provide data processing services. Data processing can include, for example, data management, data reception, and data conversion. For example, an input method application can receive and present text input from the user. Input method applications also support data processing functions such as character case switching, language switching, and text matching.

[0035] Applications typically perform conversions between different types of data items. In the context of this disclosure, the specific details of managing data items will be described using the conversion between uppercase and lowercase letters in a multilingual environment as an example. For example, an application may provide an operation control associated with data processing, such as an operation control for switching character case. In response to receiving a trigger operation on the operation control for switching character case, the case of the next data sequence to be received may be switched. When a user enters multiple data sequences, the user may need to trigger the operation control multiple times to switch case, which affects the efficiency of the user's input and the efficiency of the application in performing case conversion.

[0036] Taking case switching as an example, in order to improve the efficiency of case switching, the application can, for example, provide the user with extended content that matches the user input based on the user input. The extended content may include the user input that has been case-converted. For example, when the user input is a proper noun, the application can case-expand the received user input, perform a vocabulary comparison query on the expanded content, and if a result is found, determine the query result as the basic candidate. The application can also case-convert the basic candidate to determine the supplementary candidate. The application can sort all candidates (including basic candidates and supplementary candidates) according to certain rules and provide the sorted results to the user. For example, if the user input is the text "http", the application can perform the above processing on the user input to obtain candidates "Http", "HTTP", "HTTPS", etc. The application can provide the user with the user input "http" and the multiple obtained candidates "Http", "HTTP", "HTTPS", etc.

[0037] Data can be encoded so that each data item has its corresponding encoding. The encoding here can include, for example, universal character encoding, such as Unicode, ASCII code, scan code, etc., and the present disclosure does not limit the specific encoding. For example, for any language (including but not limited to Chinese, English, Russian, French, etc., the present disclosure does not limit the specific language), the encoding of its character set can be determined. Taking English as an example, Figure 2A shows the Unicode 200A of English. As shown in Figure 2A, the Unicode 210 of its uppercase characters AZ is U+0041 to U+005A (the numerical values ​​are continuous and increase by 1 in sequence), and the Unicode 220 of the lowercase characters az is U+0061 to U+007A (the numerical values ​​are continuous and increase by 1 in sequence), and the difference between the Unicode of the uppercase characters and the Unicode of the corresponding lowercase characters is 32.

[0038] The uppercase and lowercase character sets usually include multiple 1:1 mapping pairs (that is, multiple uppercase character-lowercase character mapping pairs). Most programming languages ​​provide data structures for storing and querying mapping pairs, such as hash tables, red-black trees, etc. Traditionally, a character set 200B as shown in Figure 2B can be obtained. The character set 200B can indicate the correspondence between uppercase characters and lowercase characters. For example, if the user presses the key "A", it can be determined that the uppercase character corresponding to the key is "A" and the corresponding lowercase character is "a". A certain amount of continuous physical address space (also called a bucket) can be applied in advance, and each bucket is implemented as a linked list structure. By traversing the character set 200B, the mapping pairs (that is, uppercase characters-lowercase characters) of single characters can be obtained in turn. For the lowercase character (lower) in the mapping pair, the index (index = Hash (lower)) can be calculated via a hash function. Based on this index (index), we can randomly access the index-th bucket. If the bucket's linked list is empty, we insert the mapping pair into the bucket's linked list (bucket[index]). At this point, the linked list contains only one element: the mapping pair. If the bucket's linked list is not empty, we insert the mapping pair at the end of the bucket's linked list. This results in the hash table shown in Figure 2C.

[0039] Taking the query based on the uppercase corresponding to the lowercase character as an example, if the character entered by the user is the lowercase character "e", its corresponding index can be determined to be M based on the hash algorithm. Then, bucket [M] can be accessed to determine that the linked list of bucket [M] is not empty. By traversing the linked list and comparing the lowercase characters one by one, it can be determined that the second element of the linked list matches the lowercase character "e", and the corresponding uppercase character "E" is returned, and the query ends. The above describes the case where the corresponding linked list is not empty and there are matching elements. The following describes several cases where no results can be returned. For example, if the character entered by the user is the lowercase character "s", and its corresponding index is determined to be 2 based on the hash algorithm. Access bucket [2] and determine that the linked list of bucket [2] is empty, and the query ends. For another example, if the character entered by the user is the lowercase character "k", and its corresponding index is determined to be 1 based on the hash algorithm. Access bucket [1] and determine that the linked list of bucket [1] is not empty. By traversing the linked list and comparing the lowercase characters one by one, it can be determined that there is no element in the linked list that matches the lowercase character "k", and the query ends.

[0040] In summary, traditional case conversion requires storing all character mapping pairs and introducing additional data structures (such as indexes). The index needs to be calculated and then the uppercase / lowercase characters that match the input characters are searched. This traditionally consumes a large amount of storage space, which results in low storage efficiency. Traditionally, additional index calculations are required, and the linked list is traversed based on the calculated index, which also results in low query efficiency.

[0041] Processes for managing data items

[0042] In order to at least partially address the deficiencies in the prior art, according to some example implementations of the present disclosure, a method for managing data items is proposed. In summary, a first data sequence including a plurality of first data items and a second data sequence including a plurality of second data items can be obtained. The first data sequence can be divided into at least one subsequence based on the first data sequence and the second data sequence. The encoding of a group of first data items in a target subsequence in at least one subsequence is continuous. The mapping characteristics of the at least one subsequence can be determined. The target mapping characteristics of the target subsequence include the first boundary and the second boundary of a group of first data items in the target subsequence, and the difference between the encoding of the data item in the target subsequence and the encoding of another data item corresponding to the data item in the second data sequence. Based on the mapping characteristics of the at least one subsequence, a first data item in the plurality of first data items can be mapped to a second data item in the plurality of second data items.

[0043] Referring to FIG3 , an exemplary implementation according to the present disclosure is described. FIG3 shows a flowchart of a process 300 for managing data items according to some implementations of the present disclosure. Process 300 can be implemented at a terminal device 110 and / or a server 130. For ease of description, the terminal device 110 and the server 130 are collectively referred to as electronic devices below, that is, the process 300 is described with the electronic device as the execution subject. For ease of discussion, the process 300 will be described with reference to the environment 100 of FIG1 . It should be noted that the operations performed by the electronic device may specifically be performed by relevant applications installed on the electronic device.

[0044] In block 310 , the electronic device acquires a first data sequence including a plurality of first data items and a second data sequence including a plurality of second data items.

[0045] The data items here can be, for example, characters, texts, images or data items of any other appropriate type. For example, multiple first data items and multiple second data items can include multiple characters in a character set in a multilingual environment. The first data item and the second data item have a mapping relationship. Such a mapping relationship can be any appropriate mapping relationship that is predefined. The mapping relationship can be a unidirectional mapping relationship or a bidirectional mapping relationship. For example, multiple first data items include any one of uppercase characters and lowercase characters, and multiple second data items include the other one of uppercase characters and lowercase characters. For the convenience of description, the following illustrative description is given by taking a unidirectional mapping relationship (i.e., determining the corresponding second data item based on the first data item) as an example.

[0046] For example, the first data item may be a lowercase character of a certain language, and the second data item may be an uppercase character of the language. The first data item may also represent a key on a keyboard (e.g., shift, ctrl, and alt), and the second data item may represent the Unicode of these keys, etc. The following is an example of a mapping relationship in which the first data item is a lowercase character of a certain language, the second data item is an uppercase character of the corresponding language, and the lowercase characters can be mapped to the uppercase characters.

[0047] The electronic device can obtain a first data sequence of multiple first data items and a second data sequence including multiple second data items in any appropriate manner. It is understandable that the number of multiple first data items and the number of multiple second data items can be the same, and each first data item can be mapped to a corresponding second data item. Taking the example that the first data item and the second data item are uppercase and lowercase characters respectively, the multiple first data items can include multiple lowercase characters, and the multiple second data items can include multiple uppercase characters. It should be noted that the multiple first data items and the multiple second data items can include lowercase characters and uppercase characters corresponding to different languages.

[0048] The first data sequence of the plurality of first data items can be, for example, a plurality of coded sequences corresponding to the plurality of first data items. Similarly, the second data sequence of the plurality of second data items can also be a plurality of coded sequences corresponding to the plurality of second data items. The following is an exemplary description using encoding as a Unicode as an example. Taking English as an example, the plurality of first data items and the plurality of second data items respectively include lowercase characters (i.e., az) and uppercase characters (i.e., AZ) of English. The Unicodes of the plurality of first data items are U+0061 to U+007A (the values ​​are incremented by 1 in sequence), and the Unicodes of the plurality of second data items are U+0041 to U+005A (the values ​​are incremented by 1 in sequence). It can be understood that the plurality of first data items and the plurality of second data items can also respectively include lowercase characters and uppercase characters of other languages.

[0049] In box 320, the electronic device divides the first data sequence into at least one subsequence based on the first data sequence and the second data sequence, and the encoding of a group of first data items in the target subsequence in the at least one subsequence is continuous. According to an example implementation of the present disclosure, in order to distinguish data items and corresponding data sequences corresponding to a given language, the electronic device 110 can, for example, divide the first data sequence into at least one subsequence based on the encoding corresponding to multiple characters included in the multiple first data items. As a result, the encoding of the data items in each subsequence obtained by the electronic device is continuous, that is, the electronic device 110 can obtain corresponding multiple subsequences (each subsequence only includes lowercase characters).

[0050] According to an example implementation of the present disclosure, an electronic device may determine a pairing sequence based on a first data sequence and a second data sequence. The pairing sequence includes multiple pairs, each of which includes a first data item from a plurality of first data items and a second data item from a plurality of second data items. For example, if the first data item is a lowercase character and the second data item is an uppercase character, each pair includes the lowercase character and the corresponding uppercase character (e.g., a and A in English). The resulting multiple pairs may include aA, bB, cC, and so on.

[0051] The electronic device can then determine at least one subsequence based on the pairing sequence. Specifically, the electronic device can identify the head pairing at the head of the pairing sequence as the current pairing. The head pairing can be the first pairing among multiple pairings. Multiple pairings in the pairing sequence can be sorted based on predetermined rules. Exemplarily, multiple first data items and multiple second data items can be arranged in the order of the encoding of the multiple first data items and the multiple second data items. That is, in the case where the data items only include uppercase and lowercase characters corresponding to English, the pairing aA can be the first pairing (that is, the head pairing), and the pairing zZ can be the last pairing.

[0052] The electronic device can process the current pairing and, upon completion, set the current pairing as the subsequent pairing in the pairing sequence. That is, upon completion of processing the first pairing, the electronic device can determine the second pairing as the current pairing and process the second pairing. Similarly, upon completion of processing the previous pairing, the electronic device can sequentially process each subsequent pairing until all pairs included in the pairing sequence are processed. In this way, each pairing can be traversed step by step, thereby achieving a mapping between all data items.

[0053] According to an example implementation of the present disclosure, a corresponding mapping feature can be constructed for each subsequence. The mapping feature can be represented by (first boundary, second edit, difference), where the first edit represents, for example, the upper boundary of the subsequence and the second boundary represents, for example, the lower boundary. The difference can represent the difference between the encodings of the corresponding data items. Regarding the specific manner of processing the current pairing, the electronic device can determine the current subsequence in the at least one subsequence based on the current pairing.

[0054] In the initial stage, the electronic device may set the first boundary and the second boundary in the current mapping feature of the current subsequence to the encoding of the first data item in the current pairing. For example, if the current pairing is aA, the electronic device 110 may set the first boundary and the second boundary to the lowercase character a in the pairing, and its encoding may be, for example, Unicode U+0061. The first boundary here includes any one of the upper boundary and the lower boundary, and the second boundary includes the other one of the upper boundary (which may also be referred to as the maximum boundary) and the lower boundary (which may also be referred to as the minimum boundary). The following is an exemplary description taking the first boundary as the maximum boundary and the second boundary as the minimum boundary as an example. Alternatively and / or additionally, the meanings of the first boundary and the second boundary can be interchanged, which will not be repeated herein.

[0055] The electronic device may also set the difference in the current mapping feature of the current subsequence to the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing. For example, if the current pairing is aA, the electronic device may set the difference to the difference between the encoding corresponding to the lowercase character a and the encoding corresponding to the uppercase character A in the pairing. The encoding may be, for example, the difference between Unicode U+0061 and Unicode U+0041. Since Unicode is a hexadecimal encoding, the difference is -32 (decimal).

[0056] After determining the first boundary, the second boundary, and the difference based on the current pairing, the electronic device may determine that the current pairing process is complete. The electronic device then sets the current pairing as the subsequent pairing after the current pairing in the pairing sequence to process the next pairing. For example, in response to pairing aA being completed, the electronic device may proceed to the next pairing bB.

[0057] According to an example implementation of the present disclosure, an electronic device may determine whether the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing matches the difference in the mapping feature of the current subsequence. For example, the electronic device may determine whether the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing is the same as the difference in the mapping feature of the current subsequence. If they are the same, the electronic device may determine that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing matches the difference in the mapping feature of the current subsequence. If they are different, the electronic device may determine that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing does not match the difference in the mapping feature of the current subsequence.

[0058] If it is determined that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing does not match the difference in the mapping characteristics of the current subsequence, the electronic device may determine that the current pairing and the previous pairing in the pairing sequence should be divided into different subsequences. In this case, the electronic device may determine the current pairing and the portion after the current pairing in the pairing sequence to create another subsequence. In other words, the electronic device may divide the first data sequence based on the previous pairing and the current pairing.

[0059] If it is determined that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing matches the difference in the mapping feature of the current subsequence, the electronic device can determine that the current pairing and the previous pairing of the current pairing can be divided into the same subsequence. According to an example implementation of the present disclosure, the electronic device can directly update the first boundary of the current mapping to the encoding of the first data item in the current pairing. For example, if the current pairing is bB and the previous pairing is aA, the first boundary and the second boundary currently both correspond to the Unicode U+0061, and the difference is -32, then the electronic device can update the first boundary in the current mapping feature of the current subsequence based on the Unicode U+0062 of the lowercase character b in the current pairing bB. The updated first boundary, second boundary and difference can be U+0062, U+0061 and -32.

[0060] According to an exemplary implementation of the present disclosure, the electronic device may also determine whether the encoding of the first boundary of the currently mapped feature is continuous with the encoding of the first data item in the current pairing. If not, the electronic device may determine that the encoding of the current pairing and the pairing preceding the current pairing are not continuous. In this case, the electronic device may create another subsequence based on the current pairing and the portion following the current pairing in the pairing sequence.

[0061] It should be understood that in an English environment, the encoding of lowercase characters and uppercase characters are continuous, and the difference between the corresponding pairs is -32. Only one subsequence is generated at this time. Specifically, if it is continuous (such as U+0061 and U+0062 shown above), the electronic device can update the first boundary based on the encoding of the first data item in the current pairing. The electronic device can determine whether the current pairing is the last pairing in the pairing sequence, and if it is determined that the current pairing is not the last pairing in the pairing sequence, set the current pairing to the subsequent pairing after the current pairing in the pairing sequence. That is, the electronic device can continue to process subsequent pairings based on the above method until all pairings in the pairing sequence are processed. The electronic device can then determine at least one subsequence based on at least one mapping feature. Thus, the electronic device can divide the first data sequence into at least one subsequence.

[0062] It is understood that the same language may correspond to at least one subsequence. For example, if the lowercase characters of a language correspond to continuous codes, then the language corresponds to only one subsequence. If the lowercase characters of the language correspond to discontinuous codes or the codes of the paired characters differ, then the language may correspond to multiple subsequences.

[0063] For example, for Danish, the encodings of lowercase characters az (U+0061-U+007A) are continuous, the encodings of uppercase characters AZ (U+0041-U+005A) are continuous, and the difference between the lowercase characters and the uppercase characters is -32. and The encodings are Unicode U+00E5, U+00E6 and U+00F8, and the uppercase characters corresponding to these three lowercase characters are and The codes for are Unicode U+00C5, U+00C6, and Unicode U+00D8. The difference between the code of each lowercase character and the code of the corresponding uppercase character is also -32. At this time, in Danish, az will be divided into a subsequence, will be divided into a subsequence, and will be divided into another subsequence. Now there are three subsequences, and each subsequence has its own mapping features.

[0064] Continuing with reference to FIG3 , in box 330, the electronic device determines mapping characteristics of at least one subsequence, the target mapping characteristics of the target subsequence including: first boundaries and second boundaries of a group of first data items in the target subsequence, and a difference between an encoding of a data item in the target subsequence and an encoding of another data item corresponding to the data item in the second data sequence.

[0065] The encoding of the data item in the target subsequence is the encoding of the first data item.

[0066] Specifically, the electronic device can determine the first boundary and the second boundary based on the first data item, and can determine the difference in the subsequence based on the difference between the first data item and the second data item in the corresponding pair. Therefore, after the electronic device obtains the target subsequence (which can be any subsequence in at least one subsequence), it can determine the first boundary (i.e., the maximum value in the continuous codes corresponding to the consecutive first data items), the second boundary (i.e., the minimum value in the continuous codes corresponding to the consecutive first data items), and the difference (the difference between each first data item in the sequence and the second data item in the corresponding pair, i.e., the difference between the code of the data item in the target subsequence and the code of another data item corresponding to the data item in the second data sequence) of the target subsequence.

[0067] The process of dividing the first data sequence into at least one subsequence is described below in conjunction with Figure 4. Figure 4 shows a schematic diagram of an example 400 of determining at least one subsequence according to some implementations of the present disclosure.

[0068] At block 402, the electronic device may determine a pairing sequence (CharMap). The pairing sequence includes multiple pairs. For example, CharMap = {pair_1, pair_2, ..., pair_n}. For example, if the multiple pairs include pairs consisting of uppercase and lowercase characters corresponding to English (e.g., [lowercase character, uppercase character]), then the CharMap includes 26 pairs: [a, A], [b, B], ..., [z, Z].

[0069] In block 404 , the electronic device may initialize a mapping feature set (CharBounds, which may include at least one mapping feature) to be used for storing mapping features corresponding to a plurality of pairings, so that the mapping feature set is empty.

[0070] At block 406, the electronic device may sort the multiple pairs in the pairing sequence. For example, the electronic device may sort the multiple pairs according to the Unicode values ​​corresponding to lowercase characters.

[0071] In block 408 , the electronic device may access each pairing in the pairing sequence one by one according to the method described above, and process the accessed pairings in sequence.

[0072] In the process of processing each pairing, during processing the current pairing, at block 410 , the electronic device may determine whether the mapping feature set is empty, that is, whether there are already determined mapping features.

[0073] When the mapping feature set is empty, the electronic device may determine that the current pairing is the first sequence in the pairing sequence (ie, the head pairing). In block 412, the electronic device may initialize a new mapping feature, in which case the CharBound is set based on the first pair [a, A].

[0074] In block 414 , the electronic device may determine the minimum boundary (CharBound.min) of the mapping feature as the code (pair.lower) corresponding to the lowercase character (eg, a) in the current pair, that is, CharBound.min=pair.lower.

[0075] In block 416 , the electronic device may determine the maximum value boundary (CharBound.max) of the mapping feature as the code corresponding to the lowercase character (eg, a) in the current pair, that is, CharBound.max=pair.lower.

[0076] In block 418 , the electronic device may determine the difference between the encoding of the lowercase character in the current pair and the encoding of the uppercase character in the current pair (pair.upper) as the difference in the current mapping feature (CharBound.diff), that is, CharBound.diff=pair.upper−pair.lower=−32.

[0077] At block 420, the electronic device may insert the newly determined mapping feature after the mapping feature set. If no mapping feature previously existed, the electronic device may determine the currently determined mapping feature as the first mapping feature in the mapping feature set (i.e., the mapping feature is (U+0061, U+0061, -32)).

[0078] After processing the current pairing, the electronic device may update the subsequent pairing of the current device as the current pairing, meaning that the electronic device may continue processing the next pairing. In block 422, the electronic device may determine whether the updated current pairing has reached the end of the pairing sequence, meaning whether the updated current pairing is the last pair in the pairing sequence.

[0079] If the updated current pairing has not reached the end of the pairing sequence, the electronic device may return to executing the process shown in block 408. That is, the electronic device may continue to process the updated current pairing based on the above manner.

[0080] If the current pairing is not the first pairing in the pairing sequence, the electronic device may determine that the mapping feature set is not empty. In block 424, the electronic device may retrieve the mapping feature at the end of the mapping feature set. In other words, the electronic device may retrieve the mapping feature corresponding to the pairing preceding the current pairing.

[0081] In block 426 , the electronic device may calculate the difference between the encoding of the lowercase character and the encoding of the uppercase character in the current pair, that is, the difference of the pair diff = pair.upper−pair.lower.

[0082] In block 428 , the electronic device may determine whether the difference value of the current pairing is the same as the difference value of the retrieved mapping feature, that is, determine whether the diff is equal to CharBound.diff.

[0083] If they are the same, at block 430, the electronic device can further determine whether the codes corresponding to the lowercase characters in the current pair are continuous with the codes corresponding to the lowercase characters in the previous pair. As previously described, the maximum value boundary in the mapping feature is the code corresponding to the lowercase characters in the previous pair. The electronic device can determine continuity by determining the difference between this maximum value boundary and the code corresponding to the lowercase characters in the current pair. Specifically, it determines whether CharBound.max is equal to the current pair's pair.lower minus 1.

[0084] If the difference is different (i.e., diff is not equal to CharBound.diff) and / or discontinuous (i.e., CharBound.max is not equal to pair.lower-1), the electronic device may determine that the lowercase characters in the current pair and the lowercase characters in the previous pair are divided into different subsequences and therefore need to create a new CharBound. The electronic device may perform the steps described in block 412, i.e., the electronic device may determine a new mapping feature. The electronic device may then determine the maximum value boundary, minimum value boundary, and difference of the new mapping feature based on the current pair and subsequent pairs of the current pair.

[0085] If the difference values ​​are the same and continuous, in box 432, the electronic device can update the maximum value boundary of the currently retrieved mapping feature to the code corresponding to the lowercase character of the current pairing. The electronic device can perform the above process until the current pairing is the last sequence in the pairing sequence (that is, the electronic device has completed the processing of each pairing in the pairing sequence). In this case, in box 434, the electronic device can determine that the mapping feature set is complete. After traversing all lowercase English characters, the mapping feature is now (U+007A, U+0061, -32), and the above mapping feature can be used to convert all lowercase characters to corresponding uppercase characters.

[0086] Returning to Figure 3, at box 340, the electronic device maps a first data item in the plurality of first data items to a second data item in the plurality of second data items based on the mapping features of at least one subsequence. According to an example implementation of the present disclosure, if a processing request for specifying a target data item to be mapped is received, the electronic device may determine a target mapping feature corresponding to the target data item in the mapping feature set. Exemplarily, the electronic device may determine that a processing request for the target data item indicated by the user input is received in response to receiving user input. For example, the electronic device may determine that a processing request for a is received in response to receiving user input of a lowercase character a.

[0087] The encoding of the target data item should be included between the first boundary and the second boundary in the target mapping feature. It is assumed that the size relationship between the first boundary, the second boundary and the encoding of the target data item should be first boundary ≥ encoding of the target data item ≥ second boundary. In the case where there are multiple subsequences, the electronic device can determine multiple mapping features corresponding to these multiple subsequences. The electronic device can determine the target mapping feature from these multiple mapping features based on any appropriate method. For example, the electronic device can traverse all the data before the two boundaries, and then look for the encoding of the target data item. Alternatively and / or additionally, the electronic device can determine the target mapping feature in the mapping feature set based on dichotomy.

[0088] The process of determining target mapping features is described below in conjunction with Figure 5. Figure 5 shows a schematic diagram of an example 500 of determining target mapping features according to some implementations of the present disclosure.

[0089] At block 502, the electronic device may obtain a data item input by the user (i.e., a target data item). For example, the electronic device may receive the corresponding data item in response to receiving a triggering of a corresponding key by the user. For example, the electronic device may determine that the received user input is the lowercase English character "c" in response to receiving a triggering of the key "c" on a keyboard.

[0090] At block 504, the electronic device may obtain a given mapping feature set. This mapping feature set may have been previously determined and stored locally by the electronic device, or may have been determined and provided to the electronic device by another electronic device. In the scenario of converting lowercase English characters to uppercase, the mapping feature set is (U+007A, U+0061, -32).

[0091] In block 506 , the electronic device determines whether the mapping feature set is empty, that is, the electronic device may determine whether the mapping feature set includes a mapping feature.

[0092] If the mapping feature set is empty, in box 508, the electronic device may determine that there is no mapping feature corresponding to the data item input by the user. The electronic device may determine that the data item input by the user is not a data item in the multiple first data items corresponding to this mapping feature set. According to an example implementation of the present disclosure, the electronic device may also provide a corresponding indication in response to determining that there is no mapping feature corresponding to the data item input by the user. For example, the electronic device may provide an indication indicating that the data item does not belong to the mapping feature set, or provide an indication indicating that the data item does not belong to the first subsequence (which may be any subsequence in the at least one subsequence) in at least one subsequence corresponding to multiple first data items. The electronic device may provide an indication, for example, by providing text, audio, video, or any other appropriate method. Thus, the electronic device may prompt the user that it is unable to determine another data item corresponding to the data item input by the user, and may provide a user experience.

[0093] If the mapping feature set is not empty, the electronic device may determine the number of at least one mapping feature included in the mapping feature set (CharBounds.size(), which is 1 in this case). At block 510, the electronic device may determine the upper bound index (index_max) for determining the target mapping feature. The electronic device may determine the index as the number of at least one mapping feature minus 1. That is, index = CharBounds.size() – 1.

[0094] In block 512 , the electronic device may determine both the lower boundary index (index_min) and the middle index (index_middle) of the target mapping feature to be 0. That is, index_min=0, index_middle=0.

[0095] At block 514, the electronic device may determine whether the value of the lower boundary index is less than or equal to the value of the upper boundary index. If the value of the lower boundary index is greater than the value of the upper boundary index, the electronic device may return to block 508, i.e., the electronic device may determine that there is no mapping feature corresponding to the data item input by the user.

[0096] If the value of the lower boundary index is less than or equal to the value of the upper boundary index, the electronic device may determine the target mapping feature from the mapping feature set based on, for example, a binary search. Specifically, at block 516, the electronic device may update the value of the middle index to the average of the lower boundary index and the upper boundary index, i.e., index_middle = (index_max + index_min) / 2.

[0097] In block 518 , the electronic device may determine the mapping feature corresponding to the middle index in the mapping feature set as the target mapping feature, that is, CharBound=CharBounds[index_middle].

[0098] The electronic device may also determine the encoding corresponding to the data item input by the user. For example, the electronic device may determine that the Unicode code corresponding to the lowercase character c is U+0063. At block 520, the electronic device may determine whether the Unicode code is between the encoding corresponding to the minimum boundary and the encoding corresponding to the maximum boundary of the target mapping feature. For example, the electronic device may determine whether the encoding of the character c is between CharBound.max and CharBound.min.

[0099] If the code corresponding to the data item input by the user is between the code corresponding to the minimum boundary and the code corresponding to the maximum boundary of the target mapping feature, block 522 is executed. At block 522, the electronic device may determine that the data item is a data item in the at least one data item corresponding to the target mapping feature. For example, the electronic device may determine that the input character c is a lowercase character.

[0100] At block 524, the electronic device may determine another data item corresponding to the input item entered by the user based on the difference in the target mapping feature. For example, if the difference is -32, the electronic device may determine that the encoding of the corresponding data item (i.e., the corresponding uppercase character C) is CharBound.min+(-32)=U+0043 based on the encoding of the lowercase character c and the difference.

[0101] In block 526, the electronic device may query the corresponding data item based on the determined encoding and provide it to the user. For example, the electronic device may query the uppercase character corresponding to the lowercase character c based on the determined encoding U+0043 and return the uppercase character C.

[0102] If the code corresponding to the data item input by the user is not a code between the code corresponding to the minimum boundary and the code corresponding to the maximum boundary of the target mapping feature, the electronic device can further determine whether the code corresponding to the input data item is a code smaller than the code corresponding to the minimum boundary or a code greater than the code corresponding to the maximum boundary.

[0103] According to an exemplary implementation of the present disclosure, at block 528, the electronic device may determine whether the code corresponding to the input data item is less than the code of the minimum boundary of the target mapping feature. The electronic device may update the upper boundary index or the lower boundary index of the target mapping feature based on the determination result.

[0104] If the code corresponding to the input data item is less than the code of the minimum boundary of the target mapping feature, at block 530, the electronic device may update the upper boundary index of the target mapping feature to the current middle index minus 1. That is, index_max = index_middle – 1. If the code corresponding to the input data item is greater than or equal to the code of the minimum boundary of the target mapping feature, at block 532, the electronic device may update the lower boundary index of the target mapping feature to the current middle index plus 1. That is, index_min = index_middle + 1.

[0105] After the upper boundary index / lower boundary index of the target mapping feature is updated, the electronic device may execute the steps shown in block 514. That is, the electronic device may update the intermediate index based on the updated upper boundary index / lower boundary index. The electronic device may then update the target mapping feature based on the updated intermediate index.

[0106] It should be understood that the above description only uses the example of converting lowercase English characters to uppercase characters to describe the mapping process. Further details on Danish case conversion will be provided below. The mapping features can be determined in the manner described in FIG. 4 above, as shown in Table 1 below.

[0107] Table 1 Mapping characteristics

[0108] According to an example implementation of the present disclosure, the technical solution described above can be executed in an input method environment. See Figure 6 for more details, which schematically shows a block diagram 600 of the process of updating the index according to some implementations of the present disclosure. As shown in Figure 6, some specialized terms have specific representations, and users may not pay attention to uppercase and lowercase representations when entering, which may result in spelling errors. For example, a user can enter the lowercase characters "http" in input box 610, and the input method can pop up multiple candidate words in candidate box 620 for the user to select.

[0109] Specifically, a mapping feature (U+007A, U+0061, -32) can be constructed based on the method described above, and the mapping feature can be used to convert each lowercase character to the corresponding uppercase character: h->H, t->T, and p->P. At this time, the terminal device 110 only needs to store the above mapping feature to implement the above conversion process. Compared with using the hash table shown in Figure 2C, the required storage space is greatly reduced. Furthermore, during the conversion process, it is not necessary to traverse the hash table, but based on the process shown in Figure 5, the lowercase character h, t, t, p codes can be searched separately between the codes U+007A and U+0061 using a binary search method, and then the found codes can be summed with "-32".

[0110] In summary, according to the implementation of the present disclosure, the storage space occupied by the storage mapping relationship and the complexity of the corresponding mapping process can be reduced while ensuring the accuracy of the management data items, thereby improving efficiency.

[0111] Example Process

[0112] According to an example implementation of the present disclosure, a method for managing data items is provided. The method includes: obtaining a first data sequence including a plurality of first data items and a second data sequence including a plurality of second data items; dividing the first data sequence into at least one subsequence based on the first data sequence and the second data sequence, wherein the encodings of a group of first data items in a target subsequence in the at least one subsequence are continuous; determining a mapping feature of the at least one subsequence, wherein the target mapping feature of the target subsequence includes: a first boundary and a second boundary of a group of first data items in the target subsequence, and a difference between the encoding of the data item in the target subsequence and the encoding of another data item corresponding to the data item in the second data sequence; and mapping a first data item in the plurality of first data items to a second data item in the plurality of second data items based on the mapping feature of the at least one subsequence.

[0113] According to an example implementation of the present disclosure, dividing a first data sequence into at least one subsequence includes: determining a pairing sequence based on the first data sequence and the second data sequence, the pairing sequence including multiple pairings, the pairings in the multiple pairings including pairings of a first data item from multiple first data items and a second data item from multiple second data items; and determining at least one subsequence based on the pairing sequence.

[0114] According to an example implementation of the present disclosure, determining at least one subsequence based on the paired sequence includes: identifying a head pairing at a head in the paired sequence as a current pairing; and determining a current subsequence in the at least one subsequence based on the current pairing.

[0115] According to an example implementation of the present disclosure, determining the current subsequence includes: setting the first boundary and the second boundary in the current mapping feature of the current subsequence to the encoding of the first data item in the current pairing, respectively; and setting the difference in the current mapping feature of the current subsequence to the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing.

[0116] According to an example implementation of the present disclosure, the method further includes: setting the current pairing as a subsequent pairing after the current pairing in the pairing sequence; and in response to determining that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing does not match the difference in the mapping feature of the current subsequence, creating another subsequence based on the current pairing in the pairing sequence and the portion after the current pairing.

[0117] According to an example implementation of the present disclosure, the method further includes: in response to determining that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing matches the difference in the mapping feature of the current subsequence, determining whether the first boundary of the current mapping feature and the encoding of the first data item in the current pairing are continuous; and in response to determining that the first boundary of the current mapping feature and the encoding of the first data item in the current pairing are discontinuous, creating another subsequence based on the current pairing in the pairing sequence and the portion after the current pairing.

[0118] According to an example implementation of the present disclosure, the method further includes: in response to determining that the first boundary of the current mapping feature is continuous with the encoding of the first data item in the current pairing, updating the first boundary of the current mapping to the encoding of the first data item in the current pairing.

[0119] According to an example implementation of the present disclosure, setting the current pairing as a subsequent pairing after the current pairing in the pairing sequence includes: in response to determining that the current pairing is not the last pairing in the pairing sequence, setting the current pairing as a subsequent pairing after the current pairing in the pairing sequence.

[0120] According to an example implementation of the present disclosure, mapping a first data item among a plurality of first data items to a second data item among a plurality of second data items based on a mapping feature of at least one subsequence includes: in response to receiving a processing request for specifying a target data item to be mapped, determining a target mapping feature corresponding to the target data item in at least one mapping feature, an encoding of the target data item being located between a first boundary and a second boundary in the target mapping feature; and mapping the target data item to a second data item among the plurality of second data items based on the encoding of the target data item and a difference in the target mapping feature.

[0121] According to an example implementation of the present disclosure, determining a target mapping feature corresponding to the target data item in the at least one mapping feature includes: determining the target mapping feature in the at least one mapping feature based on a dichotomy method.

[0122] According to an example implementation of the present disclosure, the method further includes: in response to determining that there is no mapping feature corresponding to the target data item, providing an indication that the target data item does not belong to the first subsequence.

[0123] According to an exemplary implementation of the present disclosure, the plurality of first data items and the plurality of second data items include a plurality of characters in a character set in a multi-language environment, wherein the encoding includes a universal character encoding.

[0124] According to an example implementation of the present disclosure, the plurality of first data items include any one of uppercase characters and lowercase characters, and the plurality of second data items include the other one of the uppercase characters and lowercase characters.

[0125] According to an example implementation of the present disclosure, multiple first data items and multiple second data items are arranged in an encoding order of the multiple first data items and the multiple second data items, the first boundary includes any one of the upper boundary and the lower boundary, and the second boundary includes the other one of the upper boundary and the lower boundary.

[0126] Example devices and equipment

[0127] The specific details of the method for managing data items have been described above. According to an exemplary implementation of the present disclosure, an apparatus for managing data items is provided. FIG7 shows a schematic structural block diagram of an apparatus 700 for managing data items according to certain implementations of the present disclosure. Apparatus 700 may be implemented as or included in server 130. Each module / component in apparatus 700 may be implemented by hardware, software, firmware, or any combination thereof.

[0128] As shown, the apparatus 700 includes a sequence acquisition module 710 configured to acquire a first data sequence comprising a plurality of first data items and a second data sequence comprising a plurality of second data items. The apparatus 700 also includes a sequence division module 720 configured to divide the first data sequence into at least one subsequence based on the first data sequence and the second data sequence, wherein the encodings of a group of first data items in a target subsequence in the at least one subsequence are continuous. The apparatus 700 also includes a feature determination module 730 configured to determine a mapping feature of the at least one subsequence, wherein the target mapping feature of the target subsequence includes: a first boundary and a second boundary of a group of first data items in the target subsequence, and a difference between the encoding of a data item in the target subsequence and the encoding of another data item corresponding to the data item in the second data sequence. The apparatus 700 also includes a data mapping module 740 configured to map a first data item in the plurality of first data items to a second data item in the plurality of second data items based on the mapping feature of the at least one subsequence.

[0129] According to an example implementation of the present disclosure, the sequence partitioning module 720 is further configured to: determine a pairing sequence based on the first data sequence and the second data sequence, the pairing sequence including multiple pairings, the pairings in the multiple pairings including a pairing of a first data item in the multiple first data items and a second data item in the multiple second data items; and determine at least one subsequence based on the pairing sequence.

[0130] According to an example implementation of the present disclosure, the sequence partitioning module 720 is further configured to: identify a head pairing at the head of the pairing sequence as a current pairing; and determine a current subsequence in the at least one subsequence based on the current pairing.

[0131] According to an example implementation of the present disclosure, the sequence partitioning module 720 is further configured to: set the first boundary and the second boundary in the current mapping feature of the current subsequence to the encoding of the first data item in the current pairing, respectively; and set the difference in the current mapping feature of the current subsequence to the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing.

[0132] According to an example implementation of the present disclosure, the sequence partitioning module 720 is further configured to: set the current pairing as a subsequent pairing after the current pairing in the pairing sequence; and in response to determining that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing does not match the difference in the mapping features of the current subsequence, create another subsequence based on the current pairing in the pairing sequence and the portion after the current pairing.

[0133] According to an example implementation of the present disclosure, the sequence partitioning module 720 is further configured to: in response to determining that the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing matches the difference in the mapping feature of the current subsequence, determine whether the first boundary of the current mapping feature and the encoding of the first data item in the current pairing are continuous; and in response to determining that the first boundary of the current mapping feature and the encoding of the first data item in the current pairing are discontinuous, create another subsequence based on the current pairing in the pairing sequence and the portion after the current pairing.

[0134] According to an example implementation of the present disclosure, the sequence partitioning module 720 is further configured to: in response to determining that the first boundary of the current mapping feature is continuous with the encoding of the first data item in the current pairing, update the first boundary of the current mapping to the encoding of the first data item in the current pairing.

[0135] According to an example implementation of the present disclosure, the sequence partitioning module 720 is further configured to: in response to determining that the current pairing is not the last pairing in the pairing sequence, set the current pairing as a subsequent pairing after the current pairing in the pairing sequence.

[0136] According to an example implementation of the present disclosure, the data mapping module 740 is further configured to: in response to receiving a processing request for specifying a target data item to be mapped, determine a target mapping feature corresponding to the target data item in at least one mapping feature, the encoding of the target data item being located between a first boundary and a second boundary in the target mapping feature; and map the target data item to a second data item among a plurality of second data items based on the encoding of the target data item and a difference in the target mapping feature.

[0137] According to an example implementation of the present disclosure, the data mapping module 740 is further configured to determine a target mapping feature among the at least one mapping feature based on a dichotomy method.

[0138] According to an example implementation of the present disclosure, the data mapping module 740 is further configured to: in response to determining that there is no mapping feature corresponding to the target data item, provide an indication that the target data item does not belong to the first subsequence.

[0139] According to an exemplary implementation of the present disclosure, the plurality of first data items and the plurality of second data items include a plurality of characters in a character set in a multi-language environment, wherein the encoding includes a universal character encoding.

[0140] According to an example implementation of the present disclosure, the plurality of first data items include any one of uppercase characters and lowercase characters, and the plurality of second data items include the other one of the uppercase characters and lowercase characters.

[0141] According to an example implementation of the present disclosure, multiple first data items and multiple second data items are arranged in an encoding order of the multiple first data items and the multiple second data items, the first boundary includes any one of the upper boundary and the lower boundary, and the second boundary includes the other one of the upper boundary and the lower boundary.

[0142] FIG8 shows a block diagram of an electronic device 800 capable of implementing various implementations of the present disclosure. It should be understood that the electronic device 800 shown in FIG8 is merely exemplary and should not be construed as limiting the functionality and scope of the implementations described herein. The electronic device 800 shown in FIG8 can be used to implement the methods described above.

[0143] As shown in FIG8 , electronic device 800 is in the form of a general-purpose computing device. Components of electronic device 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. Processing unit 810 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 800.

[0144] The electronic device 800 typically includes a plurality of computer storage media. Such media can be any available media accessible to the electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 800.

[0145] The electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG8 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 820 may include a computer program product 825 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.

[0146] The communication unit 840 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 800 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 800 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0147] The input device 850 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 860 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 800 may also communicate with one or more external devices (not shown) via the communication unit 840 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 800, or with any device that allows the electronic device 800 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0148] According to some implementations of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to some implementations of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above. According to some implementations of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.

[0149] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0150] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0151] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0152] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0153] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for managing data items, comprising: Acquire a first data sequence including a plurality of first data items and a second data sequence including a plurality of second data items; Based on the first data sequence and the second data sequence, the first data sequence is divided into at least one subsequence, and the encoding of a group of first data items in a target subsequence in the at least one subsequence is continuous; determining a mapping feature of the at least one subsequence, the target mapping feature of the target subsequence comprising: a first boundary and a second boundary of a group of first data items in the target subsequence, and a difference between an encoding of a data item in the target subsequence and an encoding of another data item in the second data sequence corresponding to the data item; as well as Based on the mapping feature of the at least one subsequence, a first data item of the plurality of first data items is mapped to a second data item of the plurality of second data items.

2. The method according to claim 1, wherein dividing the first data sequence into the at least one subsequence comprises: Determine a pairing sequence based on the first data sequence and the second data sequence, the pairing sequence comprising a plurality of pairings, wherein a pair in the plurality of pairings comprises a pairing of a first data item in the plurality of first data items and a second data item in the plurality of second data items; as well as The at least one subsequence is determined based on the paired sequence.

3. The method of claim 2, wherein determining the at least one subsequence based on the paired sequence comprises: Marking the head pairing at the head of the pairing sequence as the current pairing; as well as A current subsequence of the at least one subsequence is determined based on the current pairing.

4. The method of claim 3, wherein determining the current subsequence comprises: Setting a first boundary and a second boundary in a current mapping feature of the current subsequence as the encoding of a first data item in the current pairing respectively; as well as The difference in the current mapped features of the current subsequence is set to the difference between the encoding of the first data item in the current pairing and the encoding of the second data item in the current pairing.

5. The method according to claim 4, further comprising: Setting the current pairing as a subsequent pairing after the current pairing in the pairing sequence; as well as In response to determining that a difference between an encoding of a first data item in the current pairing and an encoding of a second data item in the current pairing does not match a difference in a mapping feature of the current subsequence, creating another subsequence based on the current pairing and a portion of the pairing sequence after the current pairing.

6. The method according to claim 5, further comprising: In response to determining that a difference between an encoding of a first data item in the current pairing and an encoding of a second data item in the current pairing matches a difference in a mapping feature of the current subsequence, determining whether a first boundary of the current mapping feature is continuous with the encoding of the first data item in the current pairing; and In response to determining that a first boundary of the current mapping feature is discontinuous with an encoding of a first data item in the current pairing, the another subsequence is created based on the current pairing and a portion after the current pairing in the pairing sequence.

7. The method according to claim 5, further comprising: In response to determining that the first boundary of the current mapping feature is continuous with the encoding of the first data item in the current pairing, the first boundary of the current mapping is updated to the encoding of the first data item in the current pairing.

8. The method of claim 5, wherein setting the current pairing as a subsequent pairing after the current pairing in the pairing sequence comprises: In response to determining that the current pairing is not the last pairing in the pairing sequence, the current pairing is set as a subsequent pairing after the current pairing in the pairing sequence.

9. The method of claim 1, wherein mapping a first data item of the plurality of first data items to a second data item of the plurality of second data items based on the mapping feature of the at least one subsequence comprises: In response to receiving a processing request for specifying a target data item to be mapped, determining a target mapping feature corresponding to the target data item in the at least one mapping feature, the encoding of the target data item being located between a first boundary and a second boundary in the target mapping feature; and The target data item is mapped to a second data item of the plurality of second data items based on the encoding of the target data item and the difference in the target mapping characteristic.

10. The method according to claim 9, wherein determining a target mapping feature corresponding to the target data item in the at least one mapping feature comprises: The target mapping feature is determined among the at least one mapping feature based on a dichotomy method.

11. The method according to claim 9, further comprising: In response to determining that there is no mapping feature corresponding to the target data item, an indication is provided that the target data item does not belong to the first subsequence.

12. The method according to claim 1, wherein the plurality of first data items and the plurality of second data items comprise a plurality of characters in a character set in a multi-language environment, and wherein the encoding comprises a universal character encoding.

13. The method of claim 12, wherein the plurality of first data items include any one of uppercase characters and lowercase characters, and the plurality of second data items include the other of the uppercase characters and the lowercase characters.

14. The method according to claim 1, wherein the plurality of first data items and the plurality of second data items are arranged in an order of encoding of the plurality of first data items and the plurality of second data items, the first boundary includes any one of an upper boundary and a lower boundary, and the second boundary includes the other of the upper boundary and the lower boundary.

15. An apparatus for managing data items, comprising: A sequence acquisition module, configured to acquire a first data sequence including a plurality of first data items and a second data sequence including a plurality of second data items; a sequence division module, configured to divide the first data sequence into at least one subsequence based on the first data sequence and the second data sequence, wherein the encodings of a group of first data items in a target subsequence in the at least one subsequence are continuous; a feature determination module configured to determine a mapping feature of the at least one subsequence, wherein the target mapping feature of the target subsequence comprises: a first boundary and a second boundary of a group of first data items in the target subsequence, and a difference between an encoding of a data item in the target subsequence and an encoding of another data item corresponding to the data item in the second data sequence; as well as The data mapping module is configured to map a first data item among the plurality of first data items to a second data item among the plurality of second data items based on a mapping feature of the at least one subsequence.

16. An electronic device, comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 14 when executed by the at least one processing unit.

17. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Event stream data processing method and device, electronic equipment and storage medium

    CN113810478A

  • Efficient encoding and decoding of mixed data strings in RFID tags and other media

    US20100302078A1

  • Hardware-Accelerated Lossless Data Compression

    US20110307659A1

  • Flexible partitioning of data

    US20140337392A1

  • Systems and methods for sequence encoding, storage, and compression

    US20180089369A1